SlideShare a Scribd company logo
Personal Web Usage Mining
Mining client side Web Usage Data
Web Usage Mining
The discovery of patterns in the browsing and navigation
data of Web users.
Web usage mining has been an important technology for
understanding user’s behaviors on the Web.
Currently, most Web usage mining research has been
focusing on the Web server side.
The main purpose of research is to improve a Web site’s
service and the server’s performance.
Data sources for Web usage mining are primarily Web
server logs. Although it is very important and interesting
to investigate server side issues, we argue that an
equally important and potentially fruitful aspect of Web
usage mining is the mining of client side usage data.
Web Usage Mining on Server Side
Currently, Web usage mining finds patterns in Web server
logs. The logs are preprocessed to group requests from
the same user into sessions.
A session contains the requests from a single visit of a
user to the Web site. During the preprocessing, irrelevant
information for Web usage mining such as background
images and unsuccessful requests is ignored. The users
are identified by the IP addresses in the log and all
requests from the same IP address within a certain time-
window are put into a session.
Different heuristics have been developed to deal with the
inaccuracy due to caching, IP sharing or blocking, and
network congestion.
Web Usage Mining on Server Side
Some common characteristics
• Their goal is to improve Web services and performance
Through the improvement of Web sites, including their
contents, structure, presentation, and delivery.
• They focus on the mining of server side data. Their data
sources are almost exclusively server logs, sometimes with
site structure and/or page contents.
• They target groups of users instead of individual users. It is
overwhelming for a Web site to deal with users on an
individual basis.
Personal Web Usage Mining
Individual’s Web usage, rather than group behaviors.
By looking into a user’s Web usage data, we hope to
understand the user’s interests, behaviors, and
preferences.
In other words, we are building the user’s Web profile.
We call this personal Web usage mining since it focuses
on personal Web usage.
Personal Web Usage Mining
Some of the reasons we advocate personal Web usage
mining are as follows.
• The goal of personal Web usage mining is to help and
enhance individual users Web use. It intends to make the
Web easier to use from a single user’s point of view.
• Client side data provide a more accurate and complete
picture of a user’s Web activities.
• We can achieve true individualism and personalization.
● Users have full control of what, when, and how their
data can be used for mining.
● Personal Web usage has increased significantly
recently,
Personal Web Usage Mining
Some researchers are building intelligent agents or
Internet agents that will help individuals use the Web. For
example, many agents were built for information filtering
and gathering on the Web.
WARREN is a multi-agent system for compiling financial
information.
WEBMATE edits a personal newpaper.
WebSifter is a meta-search agent which uses taxonomy
to improve search on the Web.
Other examples include home page finder , user
interface learning agent, and Web browsing assistant.
Although some aspects and pieces of personal Web
usage mining may be around in various areas such as
intelligent agent, Web warehousing, and Web usage
mining,
Personal Web Usage Mining
Two kinds of user Web activities are recorded for
analysis:
● The remote activities
include requests sent by a user to a Web server. Such
kind of click stream data includes the URLs of pages as
well as any keywords, queries, forms, and cookies sent
with the URL.
The remote activities can be captured by almost all Web
browsers. Besides, the browsers also cache the Web pages
in most cases.
Personal Web Usage Mining
Two kinds of user Web activities are recorded for
analysis:
● The local activities
include actions the user can take at his or her desktop
without the knowledge of Web servers. They include,
but are not limited to, the following.
Save a page, Print a page, Click Back on browser, Click Forward on
browser, Click Reload on browser, Click Stop on browser, Email a
link/page, Add a bookmark, Minimize/maximize/close window, Change
visual settings such as font size.
The local activities can be recorded by an activity recorder, which is a
client side program running on top of the browser.
Personal Web Usage Mining
These two kinds of activities are put together into an
activity log. Each entry in the activity log will contain a
timestamp and an activity. Some will contain extra
information such as URL, cache address,keyword,
cookie, email address, and font size.
The schema of the log looks like this:
(timestamp, activity, [URL], [cache address],[keyword],
[cookie], [email address], [other optional fields])
Personal Web Usage Mining
There are four major modules in the framework:
● Logging
● Data Warehousing
● Data mining
● Tool/Application.
Personal Web Usage Mining
In the logging module, user Web activities are stored
into the activity, as well as the cached pages.
In the data warehousing module, the logs and cached
pages are cleansed, extracted, transformed, aggregated,
and stored in a data warehouse. The data warehouse will
facilitate search, query, and OLAP operations, in the
mean time providing data sources for mining.
In the data mining module, various data mining
algorithms are applied to the data in the data warehouse,
whose findings will be used by the tools and applications
in the tool/application module.
Personal web usage mining

More Related Content

PPTX
Share, Follow, and Sync: How SharePoint 2013 uses Personal MySites for Social...
PPT
Enterprise Collaboration and Employee Engagement with Microsoft SharePoint My...
PPTX
Leveraging User Profiles and MySites
PPTX
Users, Profiles, and MySites: Managing a Changing SharePoint User population
PPTX
SPConnections - Search Administration in SharePoint 2013
PPTX
Bulding anextraneto365
PPTX
Introduction to the sharepoint 2013 userprofile service By Quontra
PPTX
SharePoint 2010 Pages
Share, Follow, and Sync: How SharePoint 2013 uses Personal MySites for Social...
Enterprise Collaboration and Employee Engagement with Microsoft SharePoint My...
Leveraging User Profiles and MySites
Users, Profiles, and MySites: Managing a Changing SharePoint User population
SPConnections - Search Administration in SharePoint 2013
Bulding anextraneto365
Introduction to the sharepoint 2013 userprofile service By Quontra
SharePoint 2010 Pages

What's hot (6)

PPTX
SharePoint And WCM
PPT
Ofc216 Shah German Webcms
PPTX
What’s new in share point 2013
PPT
Web 2.0 and Depository Web Sites: A Winning Combination (FDLP Version)
PPTX
SPCA2013 - Best Practices Document Management in SharePoint (Online) 2013
PPTX
Magnolia Innovation Spotlight - DX Summit 2018 - Agile Content Delivery
SharePoint And WCM
Ofc216 Shah German Webcms
What’s new in share point 2013
Web 2.0 and Depository Web Sites: A Winning Combination (FDLP Version)
SPCA2013 - Best Practices Document Management in SharePoint (Online) 2013
Magnolia Innovation Spotlight - DX Summit 2018 - Agile Content Delivery
Ad

Similar to Personal web usage mining (20)

PDF
Pxc3893553
PDF
C017231726
PDF
Implementation of Intelligent Web Server Monitoring
PDF
COMPARISON ANALYSIS OF WEB USAGE MINING USING PATTERN RECOGNITION TECHNIQUES
PDF
Identifying the Number of Visitors to improve Website Usability from Educatio...
PPTX
Web mining
PDF
Bb31269380
PPTX
WEB MINING.
PDF
Automatic recommendation for online users using web usage mining
PDF
Automatic Recommendation for Online Users Using Web Usage Mining
PDF
Web Data mining-A Research area in Web usage mining
PDF
Web mining .pdf module 6 dwm third year ce
PDF
Web Page Recommendation Using Web Mining
PDF
Research Paper
PPTX
Web Analytics Primer
PDF
RESEARCH ISSUES IN WEB MINING
PDF
RESEARCH ISSUES IN WEB MINING
PDF
RESEARCH ISSUES IN WEB MINING
PDF
RESEARCH ISSUES IN WEB MINING
PDF
RESEARCH ISSUES IN WEB MINING
Pxc3893553
C017231726
Implementation of Intelligent Web Server Monitoring
COMPARISON ANALYSIS OF WEB USAGE MINING USING PATTERN RECOGNITION TECHNIQUES
Identifying the Number of Visitors to improve Website Usability from Educatio...
Web mining
Bb31269380
WEB MINING.
Automatic recommendation for online users using web usage mining
Automatic Recommendation for Online Users Using Web Usage Mining
Web Data mining-A Research area in Web usage mining
Web mining .pdf module 6 dwm third year ce
Web Page Recommendation Using Web Mining
Research Paper
Web Analytics Primer
RESEARCH ISSUES IN WEB MINING
RESEARCH ISSUES IN WEB MINING
RESEARCH ISSUES IN WEB MINING
RESEARCH ISSUES IN WEB MINING
RESEARCH ISSUES IN WEB MINING
Ad

More from Daminda Herath (10)

ODP
Data mining
ODP
Data mining
ODP
Web mining
ODP
Web content mining
ODP
Personal Web Usage Mining
PPT
Social Aspect of the Internet
ODP
Web Content Mining
PPT
JavaScript Libraries
PPT
1. Overview of Distributed Systems
Data mining
Data mining
Web mining
Web content mining
Personal Web Usage Mining
Social Aspect of the Internet
Web Content Mining
JavaScript Libraries
1. Overview of Distributed Systems

Recently uploaded (20)

PDF
Classroom Observation Tools for Teachers
PDF
3rd Neelam Sanjeevareddy Memorial Lecture.pdf
PDF
102 student loan defaulters named and shamed – Is someone you know on the list?
PPTX
PPH.pptx obstetrics and gynecology in nursing
PDF
01-Introduction-to-Information-Management.pdf
PDF
FourierSeries-QuestionsWithAnswers(Part-A).pdf
PDF
Microbial disease of the cardiovascular and lymphatic systems
PPTX
Final Presentation General Medicine 03-08-2024.pptx
PDF
Chapter 2 Heredity, Prenatal Development, and Birth.pdf
PPTX
Cell Structure & Organelles in detailed.
PDF
TR - Agricultural Crops Production NC III.pdf
PPTX
school management -TNTEU- B.Ed., Semester II Unit 1.pptx
PPTX
PPT- ENG7_QUARTER1_LESSON1_WEEK1. IMAGERY -DESCRIPTIONS pptx.pptx
PDF
STATICS OF THE RIGID BODIES Hibbelers.pdf
PDF
Complications of Minimal Access Surgery at WLH
PDF
Module 4: Burden of Disease Tutorial Slides S2 2025
PPTX
Institutional Correction lecture only . . .
PPTX
Introduction_to_Human_Anatomy_and_Physiology_for_B.Pharm.pptx
PDF
Anesthesia in Laparoscopic Surgery in India
PDF
O5-L3 Freight Transport Ops (International) V1.pdf
Classroom Observation Tools for Teachers
3rd Neelam Sanjeevareddy Memorial Lecture.pdf
102 student loan defaulters named and shamed – Is someone you know on the list?
PPH.pptx obstetrics and gynecology in nursing
01-Introduction-to-Information-Management.pdf
FourierSeries-QuestionsWithAnswers(Part-A).pdf
Microbial disease of the cardiovascular and lymphatic systems
Final Presentation General Medicine 03-08-2024.pptx
Chapter 2 Heredity, Prenatal Development, and Birth.pdf
Cell Structure & Organelles in detailed.
TR - Agricultural Crops Production NC III.pdf
school management -TNTEU- B.Ed., Semester II Unit 1.pptx
PPT- ENG7_QUARTER1_LESSON1_WEEK1. IMAGERY -DESCRIPTIONS pptx.pptx
STATICS OF THE RIGID BODIES Hibbelers.pdf
Complications of Minimal Access Surgery at WLH
Module 4: Burden of Disease Tutorial Slides S2 2025
Institutional Correction lecture only . . .
Introduction_to_Human_Anatomy_and_Physiology_for_B.Pharm.pptx
Anesthesia in Laparoscopic Surgery in India
O5-L3 Freight Transport Ops (International) V1.pdf

Personal web usage mining

  • 1. Personal Web Usage Mining Mining client side Web Usage Data
  • 2. Web Usage Mining The discovery of patterns in the browsing and navigation data of Web users. Web usage mining has been an important technology for understanding user’s behaviors on the Web. Currently, most Web usage mining research has been focusing on the Web server side. The main purpose of research is to improve a Web site’s service and the server’s performance. Data sources for Web usage mining are primarily Web server logs. Although it is very important and interesting to investigate server side issues, we argue that an equally important and potentially fruitful aspect of Web usage mining is the mining of client side usage data.
  • 3. Web Usage Mining on Server Side Currently, Web usage mining finds patterns in Web server logs. The logs are preprocessed to group requests from the same user into sessions. A session contains the requests from a single visit of a user to the Web site. During the preprocessing, irrelevant information for Web usage mining such as background images and unsuccessful requests is ignored. The users are identified by the IP addresses in the log and all requests from the same IP address within a certain time- window are put into a session. Different heuristics have been developed to deal with the inaccuracy due to caching, IP sharing or blocking, and network congestion.
  • 4. Web Usage Mining on Server Side Some common characteristics • Their goal is to improve Web services and performance Through the improvement of Web sites, including their contents, structure, presentation, and delivery. • They focus on the mining of server side data. Their data sources are almost exclusively server logs, sometimes with site structure and/or page contents. • They target groups of users instead of individual users. It is overwhelming for a Web site to deal with users on an individual basis.
  • 5. Personal Web Usage Mining Individual’s Web usage, rather than group behaviors. By looking into a user’s Web usage data, we hope to understand the user’s interests, behaviors, and preferences. In other words, we are building the user’s Web profile. We call this personal Web usage mining since it focuses on personal Web usage.
  • 6. Personal Web Usage Mining Some of the reasons we advocate personal Web usage mining are as follows. • The goal of personal Web usage mining is to help and enhance individual users Web use. It intends to make the Web easier to use from a single user’s point of view. • Client side data provide a more accurate and complete picture of a user’s Web activities. • We can achieve true individualism and personalization. ● Users have full control of what, when, and how their data can be used for mining. ● Personal Web usage has increased significantly recently,
  • 7. Personal Web Usage Mining Some researchers are building intelligent agents or Internet agents that will help individuals use the Web. For example, many agents were built for information filtering and gathering on the Web. WARREN is a multi-agent system for compiling financial information. WEBMATE edits a personal newpaper. WebSifter is a meta-search agent which uses taxonomy to improve search on the Web. Other examples include home page finder , user interface learning agent, and Web browsing assistant. Although some aspects and pieces of personal Web usage mining may be around in various areas such as intelligent agent, Web warehousing, and Web usage mining,
  • 8. Personal Web Usage Mining Two kinds of user Web activities are recorded for analysis: ● The remote activities include requests sent by a user to a Web server. Such kind of click stream data includes the URLs of pages as well as any keywords, queries, forms, and cookies sent with the URL. The remote activities can be captured by almost all Web browsers. Besides, the browsers also cache the Web pages in most cases.
  • 9. Personal Web Usage Mining Two kinds of user Web activities are recorded for analysis: ● The local activities include actions the user can take at his or her desktop without the knowledge of Web servers. They include, but are not limited to, the following. Save a page, Print a page, Click Back on browser, Click Forward on browser, Click Reload on browser, Click Stop on browser, Email a link/page, Add a bookmark, Minimize/maximize/close window, Change visual settings such as font size. The local activities can be recorded by an activity recorder, which is a client side program running on top of the browser.
  • 10. Personal Web Usage Mining These two kinds of activities are put together into an activity log. Each entry in the activity log will contain a timestamp and an activity. Some will contain extra information such as URL, cache address,keyword, cookie, email address, and font size. The schema of the log looks like this: (timestamp, activity, [URL], [cache address],[keyword], [cookie], [email address], [other optional fields])
  • 11. Personal Web Usage Mining There are four major modules in the framework: ● Logging ● Data Warehousing ● Data mining ● Tool/Application.
  • 12. Personal Web Usage Mining In the logging module, user Web activities are stored into the activity, as well as the cached pages. In the data warehousing module, the logs and cached pages are cleansed, extracted, transformed, aggregated, and stored in a data warehouse. The data warehouse will facilitate search, query, and OLAP operations, in the mean time providing data sources for mining. In the data mining module, various data mining algorithms are applied to the data in the data warehouse, whose findings will be used by the tools and applications in the tool/application module.