Evaluating credibility of interest reflection on Twitter
| Pages | 343-362 |
| Date | 11 November 2014 |
| Published date | 11 November 2014 |
| DOI | https://doi.org/10.1108/IJWIS-04-2014-0019 |
| Author | Hao Han,Hidekazu Nakawatase,Keizo Oyama |
| Subject Matter | Information & knowledge management,Information & communications technology |
Evaluating credibility of interest
reection on Twitter
Hao Han
Department of Information Sciences, Kanagawa University, Hiratsuka,
Japan, and
Hidekazu Nakawatase and Keizo Oyama
Digital Content and Media Sciences Research Division,
National Institute of Informatics, Chiyoda, Japan
Abstract
Purpose – The purpose of this article was to conrm whether users’ interests are reected by tweeted
Web pages, and to evaluate the credibility of interest reection of tweeted Web pages.
Design/methodology/approach – Interest reection of Twitter is investigated based on the context
of sharing behavior. A context-oriented approach is proposed to evaluate the interest reection of
tweeted Web pages based on machine learning. Some different distribution models of similarity are
present, and infer whether tweeted Web pages reect respective users’ interests by analyzing user
access proles.
Findings – The analysis of browsing behaviors nds that many users partially hide their own
concerns, hobbies and interests, and emphasize the concerns about social phenomenon. The extensive
experimental results showed the context-oriented approach is effective on real net view data.
Originality/value – As the rst-of-its-kind study on evaluating the credibility of interest reection on
Twitter, extensive experiments have been conducted on the data sets containing real net view data. For
higher accuracy and less subjectivity, various features are generated from user’s Web view and Twitter
submission background with some different context factors.
Keywords Twitter, Credibility, Browsing behavior, Integrated analysis, Net view data,
Social network
Paper type Research paper
1. Introduction
Twitter (www.twitter.com/) is the most popular micro-blogging service and plays an
important role in social network. The service has rapidly gained worldwide popularity,
with over 500 million registered users (Twitter Turns Six, 2012) and generating over 400
million tweets daily (Celebrating Twitter7, 2013). It attracts more and more users to
share the interesting Web pages such as news (Kwak et al., 2010). The Web pages are
effectively spread and shared as a form of shortened URLs (http://t.co/) in the posted
messages (tweet) by reproduction (retweet and reply) through the users subscription
(follower).
We gratefully acknowledge the data provided by Yamana Laboratory (Department of Computer
Science, Waseda University, Japan). We also acknowledge the technical support provided by Feng
Xiao (Department of Computer Science, Tokyo Institute of Technology, Japan). This work was
supported by a Grant-in-Aid for Scientic Research A (No.22240007) from the Japan Society for
the Promotion of Science (JSPS).
The current issue and full text archive of this journal is available at
www.emeraldinsight.com/1744-0084.htm
Evaluating
credibility of
interest reection
343
International Journal of Web
Information Systems
Vol. 10 No. 4, 2014
pp. 343-362
© Emerald Group Publishing Limited
1744-0084
DOI 10.1108/IJWIS-04-2014-0019
The rapid development of Twitter stimulates social media-oriented researches, and
personalized community recommendation such as Web page recommendation becomes
a popular issue and a hot research topic. Personalized recommendation is based on the
analysis of users’ concerns and interests. There is a comprehensible but unproven
assumption:
Users access interested webpages that attract their curiosities and attentions, and share them
on Twitter. If a user shared ten sports webpages and one politics webpage, this user should
have more interest in sports than politics. So, it is natural to nd out users’ interests and
concerned topics by analyzing their shared webpages.
This assumption is widely reected in many recommendation approaches as common
sense.
However, an important difference between Twitter and other services, especially
traditional blogs, is ignored. As a quick and convenient way of sharing Web pages, most
tweets for sharing Web pages contain few or even no original opinions and comments of
the users. They copy and paste URLs into tweets or just click embedded Tweet Buttons
to share the Web pages. A tweet can be posted within an extremely short period of time
such as 5 seconds, which is far inadequate for posting a traditional blog article. This
quick pace of Web life brings many uncertain elements in sharing behaviors. For
example, some users may share Web pages just for attracting more retweets, replies and
followers on Twitter (Halpern) without seriously considering whether these Web pages
bring concerns or interests to themselves since sharing costs few seconds and brings
nothing harmful to their Web proles. Thus, “big data has big noise”, as users may tend
to post many “not worth reading” tweets (Andre et al., 2012), we cannot fully believe that
the tweeted Web pages reect the users’ interests.
In this paper, we present an in-depth analysis of the user browsing behaviors on
Twitter by using net view data, Twitter data, and Web pages as data sets. Our analysis
emphasis is laid on Web pages provided by external Web sites instead of retweets, posted
photos or chatters. We focus on two aspects:
(1) conrming whether users’ interests are reected by tweeted Web pages; and
(2) evaluating the credibility of interest reection of tweeted Web pages.
Regarding point (1), we investigate the context of sharing behavior and explain the
investigation results. As for point (2), we propose a novel approach to judge the
credibility of interest reection of tweeted Web pages based on clustering and support
vector machines (SVMs). Extensive experiments have been performed on real net view
data to show our approach is exible and effective.
The contributions of this paper are summarized as follows:
• Thedifferent sources of data are integrated and the analysis range is extended to
the general accessed Web pages (not limited to tweeted Web pages). To realize the
effective and comprehensive comparison of the difference between the accessed
Web pages and shared Web pages for respective users, we present some different
distribution models of similarity between those Web pages during a short time
period (e.g. 1 hour), and infer whether tweeted Web pages reect respective users’
interests by analyzing user access proles (all accessed Web pages) during a long
time period (e.g. one month).
IJWIS
10,4
344
Get this document and AI-powered insights with a free trial of vLex and Vincent AI
Get Started for FreeUnlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations
Unlock full access with a free 7-day trial
Transform your legal research with vLex
-
Complete access to the largest collection of common law case law on one platform
-
Generate AI case summaries that instantly highlight key legal issues
-
Advanced search capabilities with precise filtering and sorting options
-
Comprehensive legal content with documents across 100+ jurisdictions
-
Trusted by 2 million professionals including top global firms
-
Access AI-Powered Research with Vincent AI: Natural language queries with verified citations