<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//TaxonX//DTD Taxonomic Treatment Publishing DTD v0 20100105//EN" "../../nlm/tax-treatment-NS0.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:tp="http://www.plazi.org/taxpub" article-type="research-article" dtd-version="3.0" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">109</journal-id>
      <journal-id journal-id-type="index">urn:lsid:arphahub.com:pub:3dc5f44e-8666-58db-bc76-a455210e8891</journal-id>
      <journal-title-group>
        <journal-title xml:lang="en">JUCS - Journal of Universal Computer Science</journal-title>
        <abbrev-journal-title xml:lang="en">jucs</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="ppub">0948-695X</issn>
      <issn pub-type="epub">0948-6968</issn>
      <publisher>
        <publisher-name>Journal of Universal Computer Science</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.3217/jucs-014-11-1911</article-id>
      <article-id pub-id-type="publisher-id">29105</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="scientific_subject">
          <subject>H.3.3 - Information Search and Retrieval</subject>
          <subject>H.3.4 - Systems and Software</subject>
          <subject>H.5.4 - Hypertext/Hypermedia</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Exploring Information Extraction Resilience</article-title>
      </title-group>
      <contrib-group content-type="authors">
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Gregg</surname>
            <given-names>Dawn G.</given-names>
          </name>
          <email xlink:type="simple">dawn.gregg@cudenver.edu</email>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="A1">
        <label>1</label>
        <addr-line content-type="verbatim">University of Colorado, Denver, United States of America</addr-line>
        <institution>University of Colorado</institution>
        <addr-line content-type="city">Denver</addr-line>
        <country>United States of America</country>
      </aff>
      <author-notes>
        <fn fn-type="corresp">
          <p>Corresponding author: Dawn G. Gregg (<email xlink:type="simple">dawn.gregg@cudenver.edu</email>).</p>
        </fn>
        <fn fn-type="edited-by">
          <p>Academic editor: </p>
        </fn>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2008</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>06</month>
        <year>2008</year>
      </pub-date>
      <volume>14</volume>
      <issue>11</issue>
      <fpage>1911</fpage>
      <lpage>1920</lpage>
      <uri content-type="arpha" xlink:href="http://openbiodiv.net/F5039B91-EB9E-5F36-AD07-A3FDA4DDE2B2">F5039B91-EB9E-5F36-AD07-A3FDA4DDE2B2</uri>
      <uri content-type="zenodo_dep_id" xlink:href="https://zenodo.org/record/7000354">7000354</uri>
      <permissions>
        <copyright-statement>Dawn G. Gregg</copyright-statement>
        <license license-type="creative-commons-attribution" xlink:href="" xlink:type="simple">
          <license-p>This article is freely available under the J.UCS Open Content License.</license-p>
        </license>
      </permissions>
      <abstract>
        <label>Abstract</label>
        <p>There are many challenges developers face when attempting to reliably extract data from the Web. One of these challenges is the resilience of the extraction system to changes in the web pages information is being extracted from. This article compares the resilience of information extraction systems that use position based extraction with an ontology based extraction system and a system that combines position based extraction with ontology based extraction. The findings demonstrate the advantages of using a system that combines multiple extraction techniques, especially in environments where web sites change frequently and where data collection is conducted over an extended period of time.</p>
      </abstract>
    </article-meta>
  </front>
</article>
