Text classification method of Chinese web page based on steam clustering
A text and webpage technology, which is applied in the clustering field of massive webpage texts, can solve the problems of uncertain data unit dimensions and difficult analysis, and achieve the effects of high operating speed, high processing efficiency, and wide coverage
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Examples
Embodiment Construction
[0029] A kind of Chinese web page text classification method and embodiment based on flow clustering that the present invention proposes are described in detail as follows:
[0030] First, define a single text structure consisting of the title vector, label vector, text vector, author vector, related link vector and publication time of the text;
[0031] The text class is a set of publication time T coming at a certain time t 1 ,T 2 ,... T n (in days) the corresponding text P 1 , P 2 ,...P 3 A collection of , the class structure is composed of multiple feature vectors and class weights and update time, expressed as ( , ω, t), where Respectively, the weighted linear sum of the title vector, label vector, text vector, author vector, and related blog post link vector of all texts in this class; Represents the weight of this class, f(t)=2 -λt is the decay function (λ is recommended to take 0.1, that is, 10 days as the half-life), t is the publication date of the text cl...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com