<?xml version="1.0" encoding="utf-8" standalone="no"?>
<?xml-stylesheet type='text/xsl' href='/oai-pmh/oai2.xsl'?>
<OAI-PMH xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns="http://www.openarchives.org/OAI/2.0/">
  <responseDate>2026-10-10</responseDate>
  <request verb="GetRecord" identifier="oai:researchdata.se:doi-10-23695-ygw3-gf17/0" metadataPrefix="oai_dc">https://api.researchdata.se/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:researchdata.se:doi-10-23695-ygw3-gf17/0</identifier>
        <datestamp>2020-12-09</datestamp>
        <setSpec>subject:ssif:10208</setSpec>
        <setSpec>subject:ssif:102</setSpec>
        <setSpec>subject:ssif:1</setSpec>
        <setSpec>principal:slug:university-of-gothenburg</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:type>info:eu-repo/semantics/other</dc:type>
          <dc:type>http://purl.org/dc/dcmitype/Dataset</dc:type>
          <dc:identifier>https://doi.org/10.23695/YGW3-GF17</dc:identifier>
          <dc:title xml:lang="en">POS-tagging model: Stanza</dc:title>
          <dc:title xml:lang="sv">Ordklasstaggningsmodell: Stanza</dc:title>
          <dc:creator>https://ror.org/05qhvy459</dc:creator>
          <dc:subject xml:lang="en">Natural Language Processing</dc:subject>
          <dc:subject xml:lang="sv">Språkbehandling och datorlingvistik</dc:subject>
          <dc:description xml:lang="en">Models
Stanza is currently the default annotation tool used by Sparv. We provide two Stanza POS-tagging models.
stanza_eval is trained on SUC3 with Talbanken_SBX_dev as dev set. The advantage of this model is that it can be evaluated, using Talbanken_SBX_test or SIC2. The evaluation results are reported in the table below.

Test set
Exact match
POS
MSD

Talbanken_SBX_test
0.973
0.983
0.988

SIC2
0.918
0.932
0.957

 Read more about the evaluation here.
stanza_full is trained on SUC3 + Talbanken_SBX_test + SIC2 with Talbanken_SBX_dev as dev set. We cannot evaluate the performance of this model, but we expect it to perform better than stanza_eval, or at least not worse. This is the model used by Sparv.
 We updated the "pretrain" file in spring 2025. This was a minor format change.
Using the models on your own
Unzip the model you want to use and the "pretrain" file (which contains word2vec embeddings encoded in a format required by Stanza). Follow the instructions provided by Stanza</dc:description>
          <dc:description xml:lang="sv">Models
Stanza is currently the default annotation tool used by Sparv. We provide two Stanza POS-tagging models.
stanza_eval is trained on SUC3 with Talbanken_SBX_dev as dev set. The advantage of this model is that it can be evaluated, using Talbanken_SBX_test or SIC2. The evaluation results are reported in the table below.

Test set
Exact match
POS
MSD

Talbanken_SBX_test
0.973
0.983
0.988

SIC2
0.918
0.932
0.957

 Read more about the evaluation here.
stanza_full is trained on SUC3 + Talbanken_SBX_test + SIC2 with Talbanken_SBX_dev as dev set. We cannot evaluate the performance of this model, but we expect it to perform better than stanza_eval, or at least not worse. This is the model used by Sparv.
 We updated the "pretrain" file in spring 2025. This was a minor format change.
Using the models on your own
Unzip the model you want to use and the "pretrain" file (which contains word2vec embeddings encoded in a format required by Stanza). Follow the instructions provided by Stanza</dc:description>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
          <dc:publisher xml:lang="en">University of Gothenburg</dc:publisher>
          <dc:publisher xml:lang="sv">Göteborgs universitet</dc:publisher>
          <dc:date>2024-01-01T00:00:00Z</dc:date>
          <dc:language>swe</dc:language>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>