<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//TaxonX//DTD Taxonomic Treatment Publishing DTD v0 20100105//EN" "../../nlm/tax-treatment-NS0.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:tp="http://www.plazi.org/taxpub" article-type="research-article" dtd-version="3.0" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">109</journal-id>
      <journal-id journal-id-type="index">urn:lsid:arphahub.com:pub:3dc5f44e-8666-58db-bc76-a455210e8891</journal-id>
      <journal-title-group>
        <journal-title xml:lang="en">JUCS - Journal of Universal Computer Science</journal-title>
        <abbrev-journal-title xml:lang="en">jucs</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="ppub">0948-695X</issn>
      <issn pub-type="epub">0948-6968</issn>
      <publisher>
        <publisher-name>Journal of Universal Computer Science</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.3217/jucs-020-04-0488</article-id>
      <article-id pub-id-type="publisher-id">23103</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="scientific_subject">
          <subject>I.4.10 - Image Representation</subject>
          <subject>I.4.5 - Reconstruction</subject>
          <subject>I.4.7 - Feature Measurement</subject>
          <subject>I.4 - IMAGE PROCESSING AND COMPUTER VISION</subject>
          <subject>I.5.3 - Clustering</subject>
          <subject>I.5 - PATTERN RECOGNITION</subject>
          <subject>I.7.2 - Document Preparation</subject>
          <subject>I.7 - DOCUMENT AND TEXT PROCESSING</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>An Approach to Skew Detection of Printed Documents</article-title>
      </title-group>
      <contrib-group content-type="authors">
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Brodić</surname>
            <given-names>Darko</given-names>
          </name>
          <email xlink:type="simple">dbrodic@tf.bor.ac.rs</email>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Mello</surname>
            <given-names>Carlos A. B.</given-names>
          </name>
          <xref ref-type="aff" rid="A2">2</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Maluckov</surname>
            <given-names>Čedomir A.</given-names>
          </name>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Milivojevic</surname>
            <given-names>Zoran N.</given-names>
          </name>
          <xref ref-type="aff" rid="A3">3</xref>
        </contrib>
      </contrib-group>
      <aff id="A1">
        <label>1</label>
        <addr-line content-type="verbatim">University of Belgrade, Bor, Serbia</addr-line>
        <institution>University of Belgrade</institution>
        <addr-line content-type="city">Bor</addr-line>
        <country>Serbia</country>
      </aff>
      <aff id="A2">
        <label>2</label>
        <addr-line content-type="verbatim">Universidade Federal de Pernambuco, Recife, Brazil</addr-line>
        <institution>Universidade Federal de Pernambuco</institution>
        <addr-line content-type="city">Recife</addr-line>
        <country>Brazil</country>
      </aff>
      <aff id="A3">
        <label>3</label>
        <addr-line content-type="verbatim">College of Applied Technical Science Niš, Niš, Serbia</addr-line>
        <institution>College of Applied Technical Science Niš</institution>
        <addr-line content-type="city">Niš</addr-line>
        <country>Serbia</country>
      </aff>
      <author-notes>
        <fn fn-type="corresp">
          <p>Corresponding author: Darko Brodić (<email xlink:type="simple">dbrodic@tf.bor.ac.rs</email>).</p>
        </fn>
        <fn fn-type="edited-by">
          <p>Academic editor: </p>
        </fn>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2014</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>04</month>
        <year>2014</year>
      </pub-date>
      <volume>20</volume>
      <issue>4</issue>
      <fpage>488</fpage>
      <lpage>506</lpage>
      <uri content-type="arpha" xlink:href="http://openbiodiv.net/C21BB285-1BA2-55B8-8526-69E912605668">C21BB285-1BA2-55B8-8526-69E912605668</uri>
      <uri content-type="zenodo_dep_id" xlink:href="https://zenodo.org/record/5505007">5505007</uri>
      <history>
        <date date-type="received">
          <day>27</day>
          <month>05</month>
          <year>2013</year>
        </date>
        <date date-type="accepted">
          <day>14</day>
          <month>02</month>
          <year>2014</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Darko Brodić, Carlos A. B. Mello, Čedomir A. Maluckov, Zoran N. Milivojevic</copyright-statement>
        <license license-type="creative-commons-attribution" xlink:href="" xlink:type="simple">
          <license-p>This article is freely available under the J.UCS Open Content License.</license-p>
        </license>
      </permissions>
      <abstract>
        <label>Abstract</label>
        <p>In this paper, we propose an approach to estimate the text skew for printed documents. This is an important step to prevent errors in further stages of an automatic document processing system (as text segmentation). Our approach is based on the statistical analysis of the height of the connected components. In a nutshell, our algorithm is comprised of four steps: (i) removal of redundant data; (ii) establishment of the connected components, which represent filled convex hulls around each text element; (iii) enlargement of these components using morphological erosion; (iv) removal of the largest connected component to identify the first estimation of text skew. According to it, the connected components are enlarged by oriented morphological erosion and the longest of them is extracted. Statistical moments are applied to this longest component to evaluate its orientation and the global text skew of the document is identified. At the end of this process, the original document is rotated back based on the calculated angle. The performance of the proposed algorithm is examined by testing on a custom dataset. The results support the robustness of our approach.</p>
      </abstract>
    </article-meta>
  </front>
</article>
