SEN-CTD: semantic enhancement network with content-title discrepancy for fake news detection
| Date | 04 November 2024 |
| Pages | 603-620 |
| DOI | https://doi.org/10.1108/IJWIS-04-2024-0116 |
| Published date | 04 November 2024 |
| Subject Matter | Information & knowledge management,Information & communications technology,Information systems,Library & information science,Information behaviour & retrieval,Metadata,Internet |
| Author | Jiaqi Fang,Kun Ma,Yanfang Qiu,Ke Ji,Zhenxiang Chen,Bo Yang |
SEN-CTD: semantic enhancement
network with content-title discrepancy
for fake news detection
Jiaqi Fang,Kun Ma,Yan fa ng Qiu,Ke Ji and Zhenxiang Chen
School of Information Science and Engineering, University of Jinan,
Jinan, China, and
Bo Yang
School of Information Science and Engineering, University of Jinan, Jinan, China
and Quancheng Laboratory, Jinan, China
Abstract
Purpose –The discrepancy between the contentof an article and its title is a key characteristic of fake news.
Current methods for detectingfake news often ignore the significant difference in length between the content
and its title. In addition, relying solely on textual discrepancies between the title and content to distinguish
between real and fake news has proven ineffective. The purpose of this paper is todevelop a new approach
called semantic enhancement network with content–title discrepancy (SEN–CTD), which enhances the
accuracyof fake news detection.
Design/methodology/approach –The SEN–CTD framework is composed of two primary modules: the
SEN and the content–title comparison network(CTCN). The SEN is designed to enrich the representation of
news titles by integratingexternal information and position informationto capture the context. Meanwhile, the
CTCN focuses on assessingthe consistency between the content of news articles and their corresponding titles
examiningboth emotional tones and semantic attributes.
Findings –The SEN–CTD model performs well on the GossipCop, PolitiFact and RealNews data sets,
achieving accuraciesof 80.28%, 86.88% and 84.96%, respectively.These results highlight its effectivenessin
accuratelydetecting fake news across different types of content.
Originality/value –The SEN is specifically designedto improve the representation of extremely short texts,
enhancing the depth and accuracy of analyses for brief content. The CTCN is tailored to examine the
consistency betweennews titles and their corresponding content, ensuring a thorough comparativeevaluation
of both emotionaland semantic discrepancies.
Keywords Fake news detection, Heterogeneousinformation network, Discrepancy between content and title,
Attention mechanism
Paper type Research paper
1. Introduction
Fake news, always fabricated by making some minor changes to the correct statement, is highly
deceptive and indistinguishable. The widespread of fake news in diverse domains, such as politics
(Allcott and Gentzkow, 2017) and public health (Salman Bin Naeem and Bhatti, 2020), has posed
This work was supported by the Natural Science Foundation of Shandong Province
(ZR2022LZH016), the National Natural Science Foundation of China (72471103), the Shandong
Provincial Key R&D Program of China (2021CXGC010103) and the Shandong Provincial Teaching
Research Project of Graduate Education (SDYAL2022102and SDYJG21034).
Declarations: The authors declare that they have no known competing financial interests or
personal relationships that could have appeared to influence the work reported in this paper.
International
Journal of Web
Information
Systems
603
Received18 April2024
Revised 22 June 2024
24 July 2024
29 August2024
Accepted29 August 2024
InternationalJournal of Web
InformationSystems
Vol.20 No. 6, 2024
pp. 603-620
© Emerald Publishing Limited
1744-0084
DOI 10.1108/IJWIS-04-2024-0116
The current issue and full text archive of this journal is available on Emerald Insight at:
https://www.emerald.com/insight/1744-0084.htm
a huge threat to Web security and human society. The spread of fake news is largely driven by
misleading article titles, which can be attributed to two main factors (Shrestha and Spezzano,
2021). First, many Internet users only read titles and ignore the full content. Second, some
We-mediaeditors create sensational titles to attract attention and increase click-through rates.
Some existing fake news detection methods, such as bidirectional long short-term memory
(Bi-LSTM), and convolutional neural networks (CNN) ( Jamal Abdul et al.,2021;Huosong
et al.,2023), rely heavily on sequential features. Although these methods can capture local word
sequences well, they struggle with the problem of long-distance and nonconsecutive word
interactions (Zhang et al.,2020). More recent methods based on transformer models (Ben et al.,
2021;Shaina and Chen, 2022;Zhiwei et al.,2023) excel at capturing sequential context within
local word sequences, which is crucial for text classification (Ding et al.,2020). However, they
fail in modeling the distinct interaction patterns between real and fake news sentences (Vaibhav
et al.,2019). Multimodal learning methods have also integrated features from both images and
text (Qichao et al.,2023;Chuanming et al.,2022). Yet, these approaches often fail to extract the
high-order complementary information from multimodal context (Qian et al.,2021).
The relationship between article content and titles is another key factor in fake news
detection. For example, Peter et al. (Ting et al., 2019) analyzed title–content alignment to
identify fake news. Similarly, Guo et al. (2023) incorporated text similarity, multimodal
information and authorsentiment analysis in their approach.
However, these methods face two main challenges. First, news titles are typically fewer
than 10 words long (Gu et al., 2021), briefly to capture less comprehensive semantic
information than longer texts. Second, the significant length disparity between news titles
and content (Milojević, 2017) leads to textual misalignments, making it difficult to rely
solely on textual differencesto distinguish real news from fake news.
To address these challenges, we propose a semantic enhancement network with content–title
discrepancy (SEN–CTD). Our model uses two differenttypes of information. One of them is the
textual information extracted from the news it self. Another type is the discrepancy information
captured from news titles and content. Specifically, the contribution of SEN–CTD can be divided
into two parts: (1) SEN, which is used to extract rich semantic information from news headlines
and content and (2) content–title comparison network (CTCN), which is used to capture the
discrepancy of news titles and content. Our contributions are summarized as follows:
Semantic enhancement network:The SEN involves two designs: supplementing text with
external information and fully using position information. By leveraging heterogeneous
information networks(HIN) to integrate multisource data, the deficiencies in text content can
be addressed. In addition, the introduction of word position information through
bidirectional encoders enhances the contextual representation of the text. By combining
these two designs, we find that the incorporation of position information mitigates the
challenge of syntactic loss in graph-basedtext modeling.
Content–title comparison network: The CTCN is proposed to evaluate the alignment
between news content and titles. By examiningtheir emotional and semantic features, CTCN
assesses the consistency from both discrepancy and similarity perspectives. It is found that
the incorporation of emotional comparison enhances the ability to discern between real and
fake news articles. In addition, the simultaneous consideration of both discrepancy and
similarity facilitatesa more comprehensive understanding of their consistency.
2. Related work
2.1 Fake news detection
Fake news detection methods can be divided into content-based and social-context-based
methods (Shu et al.,2017). Social-context-based methods analyze how news spreads on
IJWIS
20,6
604
Get this document and AI-powered insights with a free trial of vLex and Vincent AI
Get Started for FreeUnlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations