Learning domain taxonomies: the TaxoLine approach

DOIhttps://doi.org/10.1108/IJWIS-04-2017-0024
Pages281-301
Date21 August 2017
Published date21 August 2017
AuthorOmar El Idrissi Esserhrouchni,Bouchra Frikh,Brahim Ouhbi,Ismail Khalil Ibrahim
Subject MatterInformation & knowledge management,Information & communications technology,Information systems,Library & information science,Information behaviour & retrieval,Metadata,Internet
Learning domain taxonomies:
the TaxoLine approach
Omar El Idrissi Esserhrouchni
Université Moulay Ismail Ecole Nationale Supérieure d’Arts et Metiers
(ENSAM), Morocco
Bouchra Frikh
Université Sidi Mohamed Ben Abdallah Ecole Supérieur de Technologie
(ESTF)–LTTI Lab Fès, Morocco
Brahim Ouhbi
Université Moulay Ismail Ecole Nationale Supérieure d’Arts et Metiers
(ENSAM), Meknès, Morocco, and
Ismail Khalil Ibrahim
Department of Telecooperation, Johannes Kepler University Linz,
Linz, Austria
Abstract
Purpose The aim of this paper is to present an online framework for building a domain taxonomy, called
TaxoLine, from Web documents automatically.
Design/methodology/approach TaxoLine proposes an innovative methodology that combines
frequency and conditional mutual information to improve the quality of the domain taxonomy. The system
also includes a set of mechanisms that improve the execution time needed to build the ontology.
Findings The performance of the TaxoLine framework was applied to nine different nancial corpora.
The generated taxonomies are evaluated against a gold-standard ontology and are compared to
state-of-the-art ontology learning methods.
Originality/value The experimental results show that TaxoLine produces high precision and recall for
both concept and relation extraction than well-known ontology learning algorithms. Furthermore, it also
shows promising results in terms of execution time needed to build the domain taxonomy.
Keywords Chir-statistic, Conditional mutual information, Domain taxonomy, Financial ontology,
Financial taxonomy, Ontology learning
Paper type Research paper
1. Introduction and motivation
The proliferation of the internet and Web content-producing services have allowed the
creation of a huge amount of unstructured data on the Web. These data represent an
impressive source of information for knowledge extraction. However, to infer knowledge and
extract semantics from this information, it is necessary to represent them in a structured way
to make them machine-processable. This could be performed by generating domain
taxonomies from Web documents corpora (Maedche et al., 2003).
A taxonomy is dened as a specic form of an ontology, which is a formal and an explicit
specication of a shared conceptualization (Guber, 1993). Ontologies are more complex than
taxonomies. They consist of multiple types of semantic relations between concepts,
including part-of and other domain-specic relations. A taxonomy is a concept hierarchy
that includes only the hierarchical parent-child relations. Taxonomies can be represented in
The current issue and full text archive of this journal is available on Emerald Insight at:
www.emeraldinsight.com/1744-0084.htm
Learning
domain
taxonomies
281
Received 24 April 2017
Revised 24 April 2017
Accepted 1 May 2017
InternationalJournal of Web
InformationSystems
Vol.13 No. 3, 2017
pp.281-301
©Emerald Publishing Limited
1744-0084
DOI 10.1108/IJWIS-04-2017-0024
a declarative form by using semantic markup languages like RDF and OWL. They have been
used in various elds such as e-commerce (Domingue et al., 2003), medicine (Wen et al., 2014)
and biomedical informatics (Paul and Maji, 2014).
However, manual construction of a domain taxonomy remains a costly task and
time-consuming process (Dhingra and Bhatia, 2012). It requires an extended knowledge of
the studied domain (expert). Therefore, to reduce the effort of building domain taxonomies,
several approaches that target building domain taxonomies, automatically or
semi-automatically, have been developed using different techniques (Cimiano et al., 2004;
Hearst, 1992;Meijer et al., 2014;Novalija et al., 2011).
Nevertheless, according to Buitelaar et al. (2005), the process of learning a domain
taxonomy is based on many sub-tasks and issues that have to be addressed. One of the main
issues is the data sparseness: usually, it is difcult to obtain enough data that cover the
domain of interest thoroughly. However, learning concept hierarchy requires a large amount
of data to build an appropriate domain taxonomy that covers the most important concepts.
To overcome this problem, we use the Web as a source of knowledge.
The next issue is whether the use of frequency analysis only should select the most
relevant concept and should identify precisely their taxonomic relationships. We
hypothesize that a domain taxonomy which is constructed with an algorithm that combines
statistic frequency and semantic similarity measure will extract the taxonomy more
precisely than an algorithm based on frequency only.
Another issue that, to the best of our knowledge, was never discussed in earlier works and
that needs to be addressed, is the execution time constraint. The execution time of an
algorithm is a very important parameter to evaluate its applicability. We believe that an
effective algorithm is the one which extracts concept hierarchy and builds the taxonomy
accurately and promptly.
Frikh et al. (2011) presented HCHIRSIM, an algorithm to build domain taxonomy
automatically from Web documents. The approach combines both a frequency statistic and
a similarity measure based on mutual information to build a domain taxonomy. The
algorithm was evaluated in the cancer domain against the well-known MESH ontology and
showed promising results. However, the approach still has some limitations in nding
hyponyms between concepts. Indeed, the mutual information used to measure the similarity
between two concepts does not take into consideration the context of their parent. For
instance, the term “Bank” means a nancial institution in the nancial context and a ight
maneuver in the aviation context. Nevertheless, both taxonomic relations will be extracted
by the algorithm and will be integrated in the result taxonomy, even if the studied domain is
nance or aviation. Thus, the accuracy of the result taxonomy is decreased.
In this work, we present a novel system, called TaxoLine[1] which overcomes this
limitation by introducing a new measure based on conditional mutual information.
Therefore, to identify the dependency between two concepts, the proposed algorithm
measures the similarity between them on the basis of the presence of their parent concept.
On the basis of the tests carried out, TaxoLine is more efcient and fully optimized
because it introduces a new measure based on the conditional mutual information (CMI), and
compared to benchmark algorithms, it perfectly improves the execution time of the
taxonomy learning process by using indexation method, optimized data structures and
threads in the implementation of the algorithm. To the best of our knowledge, earlier
research has rarely considered the execution time (time needed for the algorithm to construct
the taxonomy) as a property in the evaluation process. The purpose of our new algorithm is
to make the system fast enough to be hosted on the Web and used online.
IJWIS
13,3
282

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex

Unlock full access with a free 7-day trial

Transform your legal research with vLex

  • Complete access to the largest collection of common law case law on one platform

  • Generate AI case summaries that instantly highlight key legal issues

  • Advanced search capabilities with precise filtering and sorting options

  • Comprehensive legal content with documents across 100+ jurisdictions

  • Trusted by 2 million professionals including top global firms

  • Access AI-Powered Research with Vincent AI: Natural language queries with verified citations

vLex