A scalable approach for efficiently generating structured dataset topic profiles

Besnik Fetahu, Stefan Dietze, Bernardo Pereira Nunes, Marco Antonio Casanova, Davide Taibi, Wolfgang Nejdl

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

29 Citations (Scopus)

Abstract

The increasing adoption of Linked Data principles has led to an abundance of datasets on the Web. However, take-up and reuse is hindered by the lack of descriptive information about the nature of the data, such as their topic coverage, dynamics or evolution. To address this issue, we propose an approach for creating linked dataset profiles. A profile consists of structured dataset metadata describing topics and their relevance. Profiles are generated through the configuration of techniques for resource sampling from datasets, topic extraction from reference datasets and their ranking based on graphical models. To enable a good trade-off between scalability and accuracy of generated profiles, appropriate parameters are determined experimentally. Our evaluation considers topic profiles for all accessible datasets from the Linked Open Data cloud. The results show that our approach generates accurate profiles even with comparably small sample sizes (10%) and outperforms established topic modelling approaches.

Original languageEnglish
Title of host publicationThe Semantic Web
Subtitle of host publicationTrends and Challenges - 11th International Conference, ESWC 2014, Proceedings
PublisherSpringer Verlag
Pages519-534
Number of pages16
ISBN (Print)9783319074429
DOIs
Publication statusPublished - 2014
Externally publishedYes
Event11th International Conference on Semantic Web: Trends and Challenges, ESWC 2014 - Anissaras, Crete, Greece
Duration: 25 May 201429 May 2014

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume8465 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference11th International Conference on Semantic Web: Trends and Challenges, ESWC 2014
Country/TerritoryGreece
CityAnissaras, Crete
Period25/05/1429/05/14

Fingerprint

Dive into the research topics of 'A scalable approach for efficiently generating structured dataset topic profiles'. Together they form a unique fingerprint.

Cite this