Development of Browser Extension for HTML Web Page Content Extraction

Karabulut, Murat; Mayda, Islam

Development of Browser Extension for HTML Web Page Content Extraction

Tarih

2020

Yazarlar

Karabulut, Murat

Mayda, Islam

Yayıncı

IEEE

Erişim Hakkı

info:eu-repo/semantics/closedAccess

Özet

As the amount of content on the websites increases, automatic content extraction from Web pages becomes more important. Although many studies have been done in the literature on this subject, a method that fully solves the problem has not been revealed due to the flexible structure of HTML. The performances of the methods that show success at certain rates also decrease over time with the changing and developing Web structure. In this study, a browser extension was developed to automatically download text content on Web pages. This developed extension provides an output with 100% recall rate by cleaning the text content on the Web page from all tags and codes with a parser that utilizes the Document Object Model (DOM) structure. This browser extension that operates independently from the language has been tested on different types of popular Web sites in Turkey and has been shown to work successfully.

Açıklama

2nd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA) -- JUN 26-27, 2020 -- TURKEY

Anahtar Kelimeler

web content extraction; web data extraction; web scraping

Kaynak

2nd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (Hora 2020)

WoS Q Değeri

N/A

Scopus Q Değeri

N/A

Bağlantı

https://doi.org/10.1109/hora49412.2020.9152891
https://hdl.handle.net/20.500.14704/960

Koleksiyon

WoS İndeksli Yayınlar Koleksiyonu
Mühendislik ve Mimarlık Fakültesi Koleksiyonu
Scopus İndeksli Yayınlar Koleksiyonu

Detaylı Öğe Kaydı

Development of Browser Extension for HTML Web Page Content Extraction

Tarih

Yazarlar

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Erişim Hakkı

Özet

Açıklama

Anahtar Kelimeler

Kaynak

WoS Q Değeri

Scopus Q Değeri

Cilt

Sayı

Künye

Bağlantı

Koleksiyon