CLARIN Tool Portal

698 record(s) found

Search results

The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1

2 resources

The model for lemmatisation of non-standard Serbian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1794), using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~94.92. The difference to the previous version of the model is that this version is trained on a combination of two corpora (SETimes.SR, ReLDI-NormTagNER-sr).

Use "The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1"
Tests for Word Embeddings

4 resources

Evaluation tools (WBST, HWBST, EWBST) for word embedding models used to assess and compare the usefulness of different word embeddings

Use "Tests for Word Embeddings"
Czech PDT-C 1.0 Model for UDPipe 2 (2023-11-16)

2 resources

Tokenizer, POS Tagger, Lemmatizer, and Parser model based on the PDT-C 1.0 treebank (https://hdl.handle.net/11234/1-3185). The model documentation including performance can be found at https://ufal.mff.cuni.cz/udpipe/2/models#czech_pdtc1.0_model . To use these models, you need UDPipe version 2.1, which you can download from https://ufal.mff.cuni.cz/udpipe/2 .

Use "Czech PDT-C 1.0 Model for UDPipe 2 (2023-11-16)"
Linguistic digital repository based on DSpace 5.2

2 resources

One of the goals of LINDAT/CLARIN Centre for Language Research Infrastructure is to provide technical background to institutions or researchers who wants to share their tools and data used for research in linguistics or related research fields. The digital repository is built on a highly customised DSpace platform.

Use "Linguistic digital repository based on DSpace 5.2"
The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0

3 resources

This model for morphosyntactic annotation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1210), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~95.11.

Use "The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0"
Liner2.5

2 resources

Generic framework for information extraction tasks, including recognition of named entities, temporal expressions, spatial expressions and events.

Use "Liner2.5"
Translation Models (en-ru) (v1.0)

2 resources

En-Ru translation models, exported via TensorFlow Serving, available in the Lindat translation service (https://lindat.mff.cuni.cz/services/translation/). Models are compatible with Tensor2tensor version 1.6.6. For details about the model training (data, model hyper-parameters), please contact the archive maintainer. Evaluation on newstest2020 (BLEU): en->ru: 18.0 ru->en: 30.4 (Evaluated using multeval: https://github.com/jhclark/multeval)

Use "Translation Models (en-ru) (v1.0)"
The Model latinpipe-evalatin24-240520 for LatinPipe 2024

2 resources

The latinpipe-evalatin24-240520 is a PhilBerta-based model for LatinPipe 2024 <https://github.com/ufal/evalatin2024-latinpipe>, performing tagging, lemmatization, and dependency parsing of Latin, based on the winning entry to the EvaLatin 2024 <https://circse.github.io/LT4HALA/2024/EvaLatin> shared task. It is released under the CC BY-NC-SA 4.0 license.

Use "The Model latinpipe-evalatin24-240520 for LatinPipe 2024"
Universal Dependencies 1.2 Models for Parsito

2 resources

Parsing models for all Universal Depenencies 1.2 Treebanks, created solely using UD 1.2 data (http://hdl.handle.net/11234/1-1548). To use these models, you need Parsito binary, which you can download from http://hdl.handle.net/11234/1-1584.

Use "Universal Dependencies 1.2 Models for Parsito"
CUBBITT Translation Models (en-pl) (v1.0)

3 resources

CUBBITT En-Pl translation models, exported via TensorFlow Serving, available in the Lindat translation service (https://lindat.mff.cuni.cz/services/translation/). Models are compatible with Tensor2tensor version 1.6.6. For details about the model training (data, model hyper-parameters), please contact the archive maintainer. Evaluation on newstest2020 (BLEU): en->pl: 12.3 pl->en: 20.0 (Evaluated using multeval: https://github.com/jhclark/multeval)

Use "CUBBITT Translation Models (en-pl) (v1.0)"

Result filters

Metadata provider

Language

Resource type

Type of tool

Tool task

Field of study

Availability

Organisation

Project

Keywords

Search results

The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1

Tests for Word Embeddings

Czech PDT-C 1.0 Model for UDPipe 2 (2023-11-16)

Linguistic digital repository based on DSpace 5.2

The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0

Liner2.5

Translation Models (en-ru) (v1.0)

The Model latinpipe-evalatin24-240520 for LatinPipe 2024

Universal Dependencies 1.2 Models for Parsito

CUBBITT Translation Models (en-pl) (v1.0)