<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//TaxonX//DTD Taxonomic Treatment Publishing DTD v0 20100105//EN" "../../nlm/tax-treatment-NS0.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:tp="http://www.plazi.org/taxpub" article-type="research-article" dtd-version="3.0" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">109</journal-id>
      <journal-id journal-id-type="index">urn:lsid:arphahub.com:pub:3dc5f44e-8666-58db-bc76-a455210e8891</journal-id>
      <journal-title-group>
        <journal-title xml:lang="en">JUCS - Journal of Universal Computer Science</journal-title>
        <abbrev-journal-title xml:lang="en">jucs</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="ppub">0948-695X</issn>
      <issn pub-type="epub">0948-6968</issn>
      <publisher>
        <publisher-name>Journal of Universal Computer Science</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.3217/jucs-021-13-1726</article-id>
      <article-id pub-id-type="publisher-id">23825</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="scientific_subject">
          <subject>H.3.3 - Information Search and Retrieval</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Learning to Choose the Best System Configuration in Information Retrieval: the Case of Repeated Queries</article-title>
      </title-group>
      <contrib-group content-type="authors">
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Bigot</surname>
            <given-names>Anthony</given-names>
          </name>
          <email xlink:type="simple">anthony.bigot@irit.fr</email>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Déjean</surname>
            <given-names>Sébastien</given-names>
          </name>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Mothe</surname>
            <given-names>Josiane</given-names>
          </name>
          <xref ref-type="aff" rid="A2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="A1">
        <label>1</label>
        <addr-line content-type="verbatim">Université de Toulouse, Toulouse, France</addr-line>
        <institution>Université de Toulouse</institution>
        <addr-line content-type="city">Toulouse</addr-line>
        <country>France</country>
      </aff>
      <aff id="A2">
        <label>2</label>
        <addr-line content-type="verbatim">Universite de Toulouse, Toulouse,, France</addr-line>
        <institution>Universite de Toulouse</institution>
        <addr-line content-type="city">Toulouse,</addr-line>
        <country>France</country>
      </aff>
      <author-notes>
        <fn fn-type="corresp">
          <p>Corresponding author: Anthony Bigot (<email xlink:type="simple">anthony.bigot@irit.fr</email>).</p>
        </fn>
        <fn fn-type="edited-by">
          <p>Academic editor: </p>
        </fn>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2015</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>28</day>
        <month>12</month>
        <year>2015</year>
      </pub-date>
      <volume>21</volume>
      <issue>13</issue>
      <fpage>1726</fpage>
      <lpage>1745</lpage>
      <uri content-type="arpha" xlink:href="http://openbiodiv.net/42695883-DB82-5570-8BDF-0E26AF485035">42695883-DB82-5570-8BDF-0E26AF485035</uri>
      <uri content-type="zenodo_dep_id" xlink:href="https://zenodo.org/record/5505969">5505969</uri>
      <history>
        <date date-type="received">
          <day>04</day>
          <month>05</month>
          <year>2015</year>
        </date>
        <date date-type="accepted">
          <day>06</day>
          <month>07</month>
          <year>2015</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Anthony Bigot, Sébastien Déjean, Josiane Mothe</copyright-statement>
        <license license-type="creative-commons-attribution" xlink:href="" xlink:type="simple">
          <license-p>This article is freely available under the J.UCS Open Content License.</license-p>
        </license>
      </permissions>
      <abstract>
        <label>Abstract</label>
        <p>This paper presents a method that automatically decides which system configuration should be used to process a query. This method is developed for the case of repeated queries and implements a new kind of meta-system. It is based on a training process: the meta-system learns the best system configuration to use on a per query basis. After training, the meta-search system knows which configuration should treat a given query. The Learning to Choose method we developed selects the best configurations among many. This selective process rests on data analytics applied to system parameter values and their link with system effectiveness. Moreover, we optimize the parameters on a per-query basis. The training phase uses a limited amount of document relevance judgment. When the query is repeated or when an equal-query is submitted to the system, the meta-system automatically knows which parameters it should use to treat the query. This method fits the case of changing collections since what is learned is the relationship between a query and the best parameters to use to process it, rather than the relationship between a query and documents to retrieve. In this paper, we describe how data analysis can help to select among various configurations the ones that will be useful. The "Learning to choose" method is presented and evaluated using simulated data from TREC campaigns. We show that system performance highly increases in terms of precision, specifically for the queries that are difficult or medium difficult to answer. The other parameters of the method are also studied.</p>
      </abstract>
    </article-meta>
  </front>
</article>
