Finland Swedish Online is a platform offering online courses for learners of Finland Swedish. The service is provided by the University of Helsinki. The service is based on Icelandic Online provided by the University of Iceland. The courses are offered at different levels. They are learner centered with interactive visual and listening exercises organized around themes relevant to life in Finland. The courses are supported by glossaries, grammars and dictionaries.
Access Finland Swedish Online
Try out the related service for Icelandic, Iclandic Online
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2024112801
INCEpTION is a certified open-source web annotation service that has been developed by the Faculty of Computer Science of Technische Universität Darmstadt and is available to all registered users of the CLARIN:EL Research Infrastructure.
INCEpTION offers a generic multi-user annotation environment aiming
INCEpTION service is hosted at Kielipankki’s CLARIN partners at CLARIN:EL in Greece. (Click here to view their Privacy Policy.)
To start using the INCEpTION service Click ”Use Service” > ”Log in to access” > ”CLARIN Service Provider Federation login” and select your home organization.
For more information see the INCEpTION User Documentation.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2024081601
NTS on monikielinen monitorikorpus, joka sisältää maantieteellisesti paikannettuja twiittejä ja niihin liittyviä metatietoja Pohjoismaista. Kaikkiaan se sisältää lähes 74 miljoonaa viestiä sadoilta tuhansilta käyttäjätileiltä Tanskasta, Suomesta, Islannista, Norjasta ja Ruotsista. NTS-tiedot kattavat ajanjakson tammikuun 2013 ja toukokuun 2023 välillä, ja ne kerättiin Twitter Academic API:n avulla, joka on nyt suljettu.
NTS:n tarkoituksena on helpottaa SSH:n perustutkimusta. NTS:ssä on helppokäyttöinen graafinen käyttöliittymä, joka tukee nopeaa tiedonsaantia, jotta tutkijat voivat keskittyä tietojen analysointiin. Tietoaineisto mahdollistaa erityyppiset tutkimukset. Esimerkiksi on mahdollista tutkia julkista keskustelua ja tunteita lähihistorian tapahtumista (esim. COVID-19-pandemia, Nato-jäsenyysprosessi jne.). Tietokokonaisuus on myös resurssi sosiolingvistiselle tutkimukselle ja monikielisyyden tutkijoille.
Tutustu verkkosivustoon.
Jos käytät NTS-käyttöliittymää ja hyödynnät tuloksia julkaisuissasi, mainitse hiljattain julkaistu artikkeli, joka on saatavilla verkossa:
[1] Laitinen, Mikko, Jonas Lundberg, Magnus Levin & Rafael Martins. 2018. The Nordic Tweet Stream: A Dynamic Real-Time Monitor Corpus of Big and Rich Language Data, Proc. of Digital Humanities in the Nordic Countries 3rd Conference, Helsinki, Finland, March 7-9, 2018, CEUR-WS.org, online CEUR-WS.org/Vol-2084/short10.pdf.
Tämän sivun pysyvä tunniste: http://urn.fi/urn:nbn:fi:lb-2024041502
The NTS is a multilingual monitor corpus of geolocated tweets and associated metadata from the Nordic region. Altogether, it contains nearly 74 million messages from hundreds of thousands of user accounts from Denmark, Finland, Iceland, Norway, and Sweden. The NTS data cover the period between January 2013 and May 2023 and were collected using the Twitter Academic API, which is now closed.
The purpose of the NTS is to facilitate fundamental research in SSH. The NTS comes with an easy-to-use graphic interface that supports quick data access so that researchers can focus on data analysis. The dataset enables various types of research. For instance, it is possible to study public discourses and sentiment concerning events in recent history (e.g., the COVID-19 pandemic, the NATO membership process, etc.). The dataset is also a resource for sociolinguistic research and for scholars of multilingualism.
Please visit the website.
If you use the NTS interface and use the findings in your publications, please cite the recent paper, which is available online:
[1] Laitinen, Mikko, Jonas Lundberg, Magnus Levin & Rafael Martins. 2018. The Nordic Tweet Stream: A Dynamic Real-Time Monitor Corpus of Big and Rich Language Data, Proc. of Digital Humanities in the Nordic Countries 3rd Conference, Helsinki, Finland, March 7-9, 2018, CEUR-WS.org, online CEUR-WS.org/Vol-2084/short10.pdf.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2024041501
UDPipe is a trainable pipeline for tokenization, tagging, lemmatization and dependency parsing of CoNLL-U files. UDPipe is language-agnostic and can be trained given annotated data in CoNLL-U format. Trained models are provided for nearly all UD treebanks. UDPipe is available as a binary for Linux/Windows/OS X, as a library for C++, Python, Perl, Java, C#, and as a web service. Third-party R CRAN package also exists.
UDPipe is a free software distributed under the Mozilla Public License 2.0 and the linguistic models are free for non-commercial use and distributed under the CC BY-NC-SA license, although for some models the original data used to create the model may impose additional licensing conditions. UDPipe is versioned using Semantic Versioning.
Copyright 2017 by the Institute of Formal and Applied Linguistics, Faculty of Mathematics and Physics, Charles University, Czech Republic.
Kielipankki version: | |
UDPipe Kielipankki version Metadata and license |
Access to Puhti |
Source version: | |
UDPipe Metadata and license |
Access to GitHub |
Look for all versions of this tool in META-SHARE |
For more information on this tool have a look at the UDPipe User’s manual
More information on the Kielipankki version:
Using UDPipe on CSC’s servers requires a CSC user account: https://research.csc.fi/accounts-and-projects
UDPipe is installed in CSC’s computing environment (invoke with: module load udpipe) in the following configuration:
Software: UDPipe 1.2.0
Models: 2.3-181115
UDPipe was compiled and installed from Source without local modifications. Please refer to the user’s manual.
The tool was installed using Ansible scripts that can be found here: https://github.com/CSCfi/Kielipankki-palvelut/tree/Dec2018/commandline/roles/udpipe
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2024021901
Kielipankki version: | |
Turku Dependency Parser Pipeline, Kielipankki version (TDPP-LBF) Metadata and license |
Access to GitHub |
TurkuNLP Finnish Dependency Parser: | |
Finnish dependency parser developed by TurkuNLP (TDPP) Metadata and license |
Access to GitHub |
Look for all versions of this tool in META-SHARE |
The Turku Dependency Parser Pipeline, Kielipankki version (TDPP-LBF) is a version of the open source dependency parsing pipeline developed by the University of Turku NLP group for analyzing Finnish text, adapted by Kielipankki – the Language Bank of Finland.
For further information on the source version please visit the project’s website.
On Kielipankki’s GitHub repository you can find VRT tools adapted from the original pipeline (vrt-tdp-…):
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2024021503
Transkribus is a comprehensive platform for the digitisation, AI-powered text recognition, transcription and searching of historical documents.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021110305
A tool developed for analyzing the semantic similarity of words.
The demo is based on word embeddings induced using the word2vec method, trained on 4.5B words of Finnish from the Finnish Internet Parsebank project and over 2B words of Finnish from Suomi24. On the Parsebank project page you can also download the vectors in binary form. The software behind the demo is open-source, available on GitHub. The demo is maintained by the Turku NLP group.
Demo: | |
TurkuNLP word embedding demo Metadata and license |
Try out the demo |
Tool: | |
word2vec Metadata and license |
Tool project page |
Search for all versions of this resource in META-SHARE |
For word embeddings trained with word2vec and available in Kielipankki – The Language Bank of Finland please visit the wordvec resource group page.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021110304
The Language Bank’s Webanno instance was shutdown 15.8.2024
Existing and new users are encouraged to start using the much newer INCEpTION service hosted at our CLARIN partners at CLARIN:EL in Greece. (Click here to view their Privacy Policy.)
To start using the INCEpTION service Click ”Use Service” > ”Log in to access” > ”CLARIN Service Provider Federation login” and select your home organization.
For more information see the INCEpTION User Documentation. If you have questions contact us at kielipankki (ät) csc.fi.
For reference: Historical documentation about WebAnno is available in Github.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021110303
Due to very low usage, the Mylly service was shut down. If you still have data in Mylly or in case you wish to utilise the Mylly tool scripts on other services, read the instructions here.
Mylly is a versatile data analysis platform with interactive visualizations and workflows. It can be used to build workflows with a variety of tools, including morphosyntactic parsing, character set conversion and speech recognition.
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021110302
Sparv, Språkbanken’s text analysis tool, is a multilingual toolkit provided by the Swedish Språkbanken for parsing and annotating text in various languages.
Latest version: | |
Sparv Metadata and license |
Access |
Look for all versions of this tool in META-SHARE |
Latest Sparv release on GitHub
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021110301
The Turku Neural Parser Pipeline is a neural parsing pipeline for segmentation, morphological tagging, dependency parsing and lemmatization with pre-trained models for more than 50 languages.
The pipeline is installed in CSC’s computing environment as a Singularity container for the languages Finnish, Swedish and English.
Kielipankki version: | |
Turku Neural Parser Pipeline, Kielipankki version (TNPP-LBF) Metadata and license |
Access to Puhti |
TurkuNLP Finnish Neural Parser: | |
Turku Neural Parser Pipeline (TNPP) Metadata and license |
Access to GitHub |
Look for all versions of this tool in META-SHARE |
Kielipankki – the Language Bank of Finland has adapted the parser for its VRT format ( CWB-VRT):
Source for the Kielipankki version on GitHub
On Puhti you can see a list of all installed versions and languages using:
module use /appl/soft/ai/singularity/modulefiles/
module spider turku-neural-parser
For more information on this tool have a look at the following links:
Parser Demo
Turku-neural-parser-pipeline manual TNPP no longer maintained by TurkuNLP, see the note from May 2024!
TurkuNLP DockerHub
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021101102
This software package provides finnish-postag, a part-of-speech and morphology tagger for Finnish, and finnish-nertag, a named entity recogniser for Finnish.
This software is also installed in CSC’s computing environment (module load finnish-tagtools).
Both tools take running text from standard input and produce tabular output (one token per line) to standard output. See –help messages for more details.
An installer is provided in the form of a Makefile. More information can be found in the README file in the download folder.
Latest version: | |
Finnish Tagtools 1.6 Metadata and license |
Download the resource |
Look for all versions of this tool in META-SHARE |
This resource group page has a Persistent Identifier: http://urn.fi/urn:nbn:fi:lb-2021101101
Last modified on 2024-02-15