Personalized and accurate classification system for Chinese web pages based on big data
By designing a personalized and accurate classification system for Chinese web pages based on big data, using filtering algorithms for useless HTML tags, improved word segmentation algorithms and weight calculation methods, the problem of massive web page information classification is solved, and the rapid and accurate automatic classification of web pages is achieved, which meets personalized needs and is applied in multiple fields.
Patent Information
- Application Number
- CN202410710621.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-06-04
AI Technical Summary
The existing technology is difficult to effectively manage and classify massive web page information, resulting in a long time for users to obtain the required information, and lack of suitable web page feature dimensionality reduction methods and personalized Chinese word segmentation models, which affects the accuracy and efficiency of automatic web page classification.
A personalized and accurate classification system for Chinese web pages based on big data is designed, and a filtering algorithm for HTML useless tags, an improved sequence optimal matching word segmentation algorithm and TF*IDF*CHI weight calculation method are used. Combined with the distribution uncertainty of the CHI calculation quantity calculation feature items, various modules of the Chinese web page automatic classification model are constructed, including massive web page data collection, pre-processing, feature screening and precise classification.
It realizes the rapid accuracy of automatic web page classification, improves the ambiguity recognition ability in word segmentation process, reduces user search time, and reaches 96.3%, meets the needs of personalized web page classification, and is widely used in digital libraries, news classification and search engines.
Smart Images

Figure CN118839047B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a personalized and accurate classification system for Chinese web pages, and in particular to a personalized and accurate classification system for Chinese web pages based on big data, which belongs to the technical field of automatic classification of web pages. Background Art
[0002] With the rapid development and widespread application of the Internet, mankind has entered the era of information explosion. Mobile Internet has penetrated into various fields such as news, colleges and universities, government, sports, education, finance, etc., and network applications have also spread to thousands of households. Faced with these huge and complex online resources, it is very difficult for people to quickly and effectively find information that meets user needs from hundreds of millions of web pages. Therefore, how to effectively manage and classify the massive amount of information on the Internet and reduce the time it takes for users to obtain the information they need has become a huge challenge.
[0003] Traditionally, this problem is solved manually. Arranging personnel from different professional fields to analyze web page content and classify web pages into one or several categories according to different web page contents will consume a lot of manpower and financial resources, which is not worth promoting. In the current massive growth of web page data, manual classification is also unrealistic because the speed of manual classification cannot keep up with the growth rate of web pages. In this case, the automatic web page classification technology is used to automatically classify and manage web pages, classify web page sets according to different web page content features, and classify similar web pages into one or more categories, so as to help users quickly and accurately find the information they need, and to a large extent solve the problem of chaotic network resource information.
[0004] The personalized classification technology of Chinese web pages has great practical significance and application value. First, the application object of Chinese web page classification technology is a large number of web pages on the Internet. It automatically classifies Chinese web pages, improves the speed and efficiency of web page text processing, and allows users to clearly face the messy information on the Internet. Secondly, the automatic classification of Chinese web pages can greatly reduce the user's search time and facilitate users to directly search for the required web page categories. When users enter keywords on the search engine, they will get a large number of web page search results, but most of the web page search results obtained cannot meet the needs of users. Through the personalized classification of Chinese web pages, web pages are divided into different groups. Users can accurately locate the group most relevant to the target web page, so as to get more search results more quickly. Third, web page classification has a wide range of applications in many fields, such as digital libraries, search engines, information filtering, etc. The resources of digital libraries are huge, and automatic classification counting is often required to digitally manage digital resources, which has high application significance and value.
[0005] The problems that need to be solved by the Chinese webpage personalized classification system in the prior art and the key technical difficulties of this application include:
[0006] (1) Faced with the vast amount of online resources, it is very difficult to quickly and effectively find information that meets user needs from hundreds of millions of web pages. Currently, it is impossible to effectively manage and classify the massive amount of information on the Internet. It takes a long time for users to obtain the required information. The existing technology cannot solve this problem manually. Arranging personnel from different professional fields to analyze the content of web pages and classifying web pages into one or several categories according to different web page contents will consume a lot of manpower and financial resources. The manual classification is highly subjective and is not worth promoting. In the current massive growth of web page data, it is not realistic to use manual classification because the speed of manual classification cannot keep up with the growth rate of web pages. In this case, there is an urgent need for an automatic web page classification technology to automatically classify and manage web pages, classify web page sets according to different web page content features, classify similar web pages into one or more categories, improve the ability to identify ambiguities in the word segmentation process, help users quickly and accurately find the information they need, and solve the problem of chaotic network resource information to a large extent.
[0007] (2) Web page text classification is more complicated. Web pages have their own rich structural features, such as HTML tags, advertising information, copyright statements, etc. The existing technology does not combine text automatic classification with web page structural information when automatically classifying web pages. It lacks a filtering algorithm for useless HTML tags and is not suitable for automatic classification of web page text. Moreover, Chinese is more complex in structure than English. Compared with English, which directly separates words and phrases with spaces, Chinese has no obvious division in structure and is only separated by punctuation marks at the end of sentences. The existing technology lacks a personalized Chinese word segmentation model for web page classification and cannot find a Chinese word segmentation method suitable for web page classification to perform word segmentation on web page text. The existing web page feature weight representation method is TF-IDF and Boolean weight, which are mostly suitable for the representation of ordinary plain text. If it is directly used for the representation of web page text, the structural features of the web page itself will be ignored. The HTML language used by Web pages is structured, and each tag represents different semantic information. Web page structure information and text information features need to be combined for representation. The existing technology lacks an efficient web page feature weight representation method.
[0008] (3) The existing technology lacks a suitable web page feature dimensionality reduction method. The complexity of the feature items of web page text makes the feature space have high dimensionality. How to remove features that are irrelevant to classification from the feature items and perform feature screening is a research difficulty in the classification of Chinese text on the Web. The existing technology does not have a high-accuracy precise classification system that is applicable to corpora in different fields. When the number of web page text sets is huge, the expressiveness of most precise classification systems drops sharply. Precise classification systems with good expressiveness are mostly for data in a specific field, which means that there is a great fit between the precise classification system and the data corpus. The existing technology lacks research on classification algorithms for restricted corpora and lacks a classification algorithm based on restricted corpora. It is impossible to build an efficient Chinese automatic classification model. The accuracy and efficiency of the web page classification system are low, and it cannot meet personalized needs and current application needs. Summary of the invention
[0009] Aiming at the extraction of web page data, this application designs a filtering algorithm for useless HTML tags to obtain web page text content with higher value. In the maximum sequence matching word segmentation algorithm, three-word intersection type ambiguous field processing is adopted to improve the ability of ambiguous recognition in the word segmentation process. The weight calculation method based on TF*IDF is improved, and the weight is calculated in combination with the CHI calculation amount, which comprehensively considers the number of times the feature item appears in a certain type of text and all texts, the influence of category information on the weight, and the influence of the feature appearance position on the weight. The Chinese automatic classification model is implemented, and the construction method of each module of the automatic classification of Chinese web pages is designed to effectively organize and process the massive information on the Internet, so that people can better search for the resources they want. The automatic classification of web pages in this application is an important technology for realizing rapid information retrieval. It has been widely used in digital libraries, news classification, search engines and other fields.
[0010] In order to achieve the above technical effects, the technical solutions adopted in this application are as follows:
[0011] The personalized and precise classification system for Chinese web pages based on big data directly classifies the massive web pages on the Internet. The web page set is captured from the Internet according to a certain strategy, and then the web page data is pre-processed, and the features of the pre-processed text information are screened, and finally classified using a precise classification system:
[0012] P1: Based on the different importance of feature items of different tags in web pages for classification, the tag structure characteristics of web pages are analyzed, and a filtering algorithm for useless HTML tags is constructed. The corresponding weights are assigned to high-value tag sets, and the titles, keywords and body texts that have a greater impact on web page classification are extracted;
[0013] P2: In the text pre-processing process, the sequential optimal matching word segmentation algorithm is improved, combined with the characteristics of Chinese text, and a processing framework of three-character intersection ambiguous fields is adopted to enhance the algorithm's ambiguous recognition ability;
[0014] P3: In the feature screening stage, based on the distribution of feature items between classes and in each category, the CHI calculation amount is incorporated to calculate the distribution uncertainty of feature items, and the TF*IDF*CHI weight calculation method is used to comprehensively consider the number of times the feature items appear in a certain category and all texts, the impact of category information on feature weights, and the location of feature appearance;
[0015] P4: Build the modules of the automatic web page classification model, including:
[0016] Module 1: Massive web page data collection module: crawl and collect web page URLs and build a web page text set;
[0017] Module 2: Web page data pre-processing module: extract high-value text content from web pages, and perform word segmentation and denoising;
[0018] Module 3: Feature extraction module: screening and extracting features;
[0019] Module 4: Accurate classification module: Construct an accurate classification system to classify the final constructed text vector.
[0020] Preferably, the massive web page data collection module: the web page automatic and accurate classification system does not directly classify the URL, but first uses a web crawler to crawl the web page text content corresponding to the URL, and then classifies the web page text content. The following web page collection module is designed:
[0021] (1) URLs and web page content are stored in a database: MySQL database is used as the URL pool, and automatic deduplication is performed based on the MySQL database. The corresponding field URL in the table is set to be unique to automatically deduplicate the URL;
[0022] (2) Adopting a depth-first crawler strategy: crawling web pages using a depth-first crawler approach to collect as much web page content as possible;
[0023] (3) Using multi-threading to collect web pages: Using multi-threading to collect web pages can reduce the time for database operations and waiting. At the same time, using multi-threading to collect web pages can ensure that the program runs for a long time without errors. Threads use independent running units. When a thread ends, the requested memory unit is released, but the number of threads cannot be increased indefinitely.
[0024] (4) Non-recursive programming.
[0025] Preferably, the algorithm flow of the web page collector is as follows:
[0026] ① Create a database table tbl_URL and save the initial URL input into the table tbl_URL. At this time, the table only contains the URL field information, and the rest of the fields are default values. The default value of the field state is 0, indicating that the corresponding URL has not yet been crawled. The field state indicates the current state of the web page;
[0027] ② Read the URL from the database to ensure that there is only one thread accessing the database to read the URL at the same time. Use a mutex semaphore to implement mutually exclusive access to the database. When a thread reads the URL, the remaining threads are blocked until the thread ends accessing the database. The database command for reading the URL is: select top 20*from tbl_URL where state=0order by id desc. A maximum of 20 URLs are extracted at a time. The captured value is represented by num. The variable is configured through the configuration file. Before the capture, state=0 indicates that the webpage URL has not been captured. After the operation is executed, the state field is set to 1, indicating that the webpage is being captured. After the operation is completed, the blocking ends and the memory is released to the next queued thread.
[0028] ③ When crawling a page, there are three situations: First, when parsing the URL, if the port and server information corresponding to the URL cannot be obtained, the URL is wrong and the state value of the URL is set to 5; second, if the port and server information corresponding to the URL can be correctly read, the corresponding page content will continue to be read. After reading, the state value of the URL is set to 2, indicating that the web page has been crawled, and the web page content is saved to the field File of the table tbl_URL. At the same time, other hyperlinks in the web page are searched, and other hyperlinks are extracted and saved to the database; third, if the port and server information corresponding to the URL can be correctly read, but the connection times out when continuing to read the corresponding web page content, it means that the web page corresponding to the URL no longer exists. At this time, the corresponding unAccessible value is increased by 1, and the critical value of unAccessible is set to 10. When the value reaches the critical value, the state field is set to 5;
[0029] ④ After one crawl is completed, the thread ends and repeats the above operation to continue the next crawl until there is no record with a state value of 0 in the database. The program ends and the web page collection is completed.
[0030] Preferably, the web page data pre-processing module: extracts information from the web page content and performs text pre-processing on the obtained text information, including Chinese word segmentation and text de-noising. The process module is divided into web page mass data extraction and text pre-processing.
[0031] Preferably, massive data extraction of web pages: extract the text information of the main content of the web page and the related content in the high-value tags to obtain the corpus source of the original features. The effective information of the web page is in the title, hyperlink <ahref>, meta keywords and description, in the body, useless tags include script tags <script>、注释标签<!——>、表单及相关标签<form><option>input>、样式标签<style>、对象标签<object>、applet标签<applet>、格式标签<hr>,将这七类标签作为无用标签处理,归于集合T,在网页前置处理模块对这些无用标签进行剔除;
[0032] 基于本申请定义的高价值标签和无用标签的代表含义,保留高价值标签中的文字内容,去除无用标签中的噪声信息,设计网页数据提取算法如下:
[0033] 第一步:输入html网页;
[0034] 第二步:读取html网页,查找<head>和< / head>标签,解析<meta>标签并记录,查找<title>和< / title>,输出html网页标题;
[0035] 第三步:循环读入html网页文本内容:
[0036] 经过以上算法对网页进行内容提取,去掉无用标签,保留高价值的网页标签有:tag={title,meta,B,I,U,H1,H2,H3,H4,H5,a},信息提取方法由extract实现,输入为网页的html字符串,输出为标题、正文等字符串。
[0037] 优选地,文本前置处理:将网页表示成向量模式,对第一步网页抽取后的文本信息进行中文分词、去停用词处理,首先要去掉标点符号,得到完全的文字信息;
[0038] 中文分词采用基于字符串的方法,处理未登录词识别和歧义识别,对顺序最优匹配法进行改进,MaxLen代表初始最优匹配长度,L代表待分词的句子长度,P指向待分词的句子,Len表示实际取出的词条长度,每个汉字是两个字节,因此如果匹配失败需要减掉一个汉字重新匹配,则取出的字符串长度需要减去2,即Len=Len-2;
[0039] 基于分词过程中最容易有歧义的词是三字长交集型词,本申请针对三字长交集型的词将最优匹配分词法作如下改进:设置两个缓冲区Smian[p,q]及Ssub[i,j],Smian[p,q]表示从字符串的第p个字符开始取q个字符,则初始时p=0,q=L,L是整个字符串的长度,表示为Smian[O,L],即从字符串第一个字符开始取L个字符;Ssub[i,j]表示从字符的第i个字符开始取j个字符,代表从整个待分词的字符串中取出的字符串,在分词过程中,是Smian[p,q]的一个子集;
[0040] 本申请将MaxLen的值设为4,带分词的字符串长度为L,将待分词字符串放入缓冲区Smian[p,q]中,最长匹配字符为Len,其初始值为MaxLen,根据L和MaxLen的长度不同,算法分为两种情况:
[0041] 情况1:L≥MaxLen:从Smian[p,q]中的第一个字开始取长度为Len的字串放入Ssub[i,j]中,i=1,j=Len,将Ssub[i,j]中的字串和词表中的词语逐一进行匹配,如果匹配失败,则从第2个字开始再取长度为Len的字串放入Ssub[i,j],i=2,j=Len,将Ssub[i,j]中的字串再次和词表中的词逐一匹配,如果匹配失败,则按照上述步骤重复,从第3、4、5、…、L-Len个字开始取长度为Len的字串放入Ssub[i,j]中进行匹配,如果上述过程中所有匹配均失败,则表明字串中没有长Len的词,减去一个字进行匹配,将Len减2,重复上述步骤,直到匹配出成功的词,当有词匹配成功时,将该词切分出来,将该词左右两边的字串分开作为新的字串递归调用以上过程;
[0042] 情况2:L<MaxLen:此时待分词的字串的长度最大为3,首先对整个字串进行匹配,如果匹配不成功,则为了避免三字长交集型分词歧义字段从第二个字开始匹配,如果匹配成功则将该词切分出来,分词成功;否则取前两个字进行匹配,如果匹配成功则将这两个字作为词切分,分词成功;否则匹配不成功,则将该待分词字串看做是每个字是单字词,分词结束;
[0043] 在对文本进行中文分词后,保存下来许多出现较为频繁但和分类关系不大的词,即停用词,对停用词进行剔除处理,进行文本去噪。
[0044] 优选地,特征提取模块:
[0045] 第1步:输入经过文本前置处理后的初始特征词集;
[0046] 输出:经过特征筛选后的最终特征词集
[0047] 第2步:对于初始特征集中的每个词,利用式1:
[0048]
[0049] 对每个词和每个类别计算特征值,A表示特征项t和cj类文本同时出现的次数,B表示特征项t不出现在cj类文本中的次数,C表示cj类文本出现但t不出现的次数,D表示特征项t不出现又不属于cj类的次数,N表示训练集中所有的文本数,x2衡量类别ci和特征项t之间的关系,对每一对类别ci和特征t都计算x2(t,ci)值,然后按照由高到低排序,剔除值较低的特征;
[0050] 第3步:对于每个类别的所有特征词计算按特征值由低到高进行排序;
[0051] 第4步:取前K个词作为该类别的特征项,首先设置K初始值为1000,然后据实对K值不断进行调整;
[0052] 第5步:计算所有类别的特征项,对特征空间进行统一;
[0053] 第6步:输出经过特征筛选后的最终特征词集;
[0054] 对特征项计算对应的权重,权重影响因子设置2个:一是一个特征项在某类文本中出现越多表明它对该类别越重要,区分能力越强;二是一个特征项在所有的文本中出现的越多,表明该词对类别的区分能力越弱,从两个方面对TF*IDF进行改进;
[0055] 在TFIDF的权重计算中融入类别信息对特征权重的影响因子,采用x2计算量来表示特征词和类别的相关性,改进后的公式如式2:
[0056] Wij=tfij×idfij×CHI(ti) 式2
[0057] 其中,Wij表示特征词的权重,TF和IDF的意义和TF*IDF公式中的相同,CHI表示特征项的特征值计算量;
[0058] 本申请在网页海量数据抽取模块中抽取到的标签Tag={Title,Meta,B,I,U,H1,H2,H3}中的文本和正文文本,定义标签权重集合为W={wt|t∈Tag},wt表示标签t的权重,并具体赋值;
[0059] 最终,修改权重计算公式为式3所示:
[0060]
[0061] 经过特征筛选后,得到网页文本的最终特征项集,通过以上公式计算特征向量的权重,用向量对网页文本进行表示,至此特征筛选模块完成,即可进入网页分类模块。
[0062] 优选地,精准分类模块:网页分类模块的核心是训练方法和分类算法,训练阶段中,输入所有训练集样本,调用精准分类系统的训练算法先进行独立于具体精准分类系统,然后特定于精准分类系统的处理;分类阶段中,精准分类系统将待分类文本表达成向量后,用训练阶段生成的精准分类系统分类。
[0063] 优选地,本申请在构建分类算法时,采用最邻近分类模型对文本进行分类,网页分类模块的工作步骤如下:
[0064] 步骤一:读取训练集中所有文本的空间向量化后的数据;
[0065] 步骤二:筛选k值,指定k个最邻近文本作为匹配数量;
[0066] 步骤三:计算训练集中每个文本和测试文本的相似度,计算公式为式4所示:
[0067]
[0068] 其中,x表示测试文本的特征向量,d代表训练文本的特征向量,n代表特征向量有n维,Wk表示向量的第k维;
[0069] 步骤四:将步骤三中计算出来的文本相似度由低到高排序,选出距离最相近的k个文本;
[0070] 步骤五:在k个相似的训练文本中,对其中每个类别的权重进行计算,计算式为:
[0071] p(x,cj)=∑sim(x,di)×y(di,cj)-T 式5
[0072] 其中,T为临界值,如果di属于cj类,y(di,cj)为1,否则为0;
[0073] 步骤六:比较步骤五中计算出的各个类别的权重,将文档分到数值最大的类中。
[0074] 与现有技术相比,本申请的创新点和优势在于:
[0075] (1)本申请针对网页数据的抽取,基于网页中不同标签的特征项对于分类的重要程度不同,针对网页的标签结构特征,设计了对HTML无用标签的过滤算法,并对高价值的标签集合赋予了对应的权值,提取得到对网页分类影响较大的标题、关键词及正文文本等。在文本前置处理过程中,对顺序最优匹配分词算法进行了改进,结合中文文本的特征,采用了三字长交集型歧义字段的处理思想,提高了算法的歧义识别能力。在特征筛选阶段,考虑到TF*IDF权重计算方法没有考虑到特征项在类间分布情况和每个类别中的分布情况,引入了CHI计算量计算特征项的分布不确定性,采用了TF*IDF*CHI权重计算方法,综合考虑了特征项在某一类和所有文本中出现的次数、类别信息对特征权重的影响及特征出现位置,较传统的TF*IDF权重计算方法更有合理性。提出了网页自动分类模型各模块的设计和实现方法,测试结果表明,分类准确率达到96.3%,满足个性化网页分类需求。
[0076] (2)本申请针对网页数据的提取,设计了对HTML无用标签的过滤算法,得到较高价值的网页文本内容。在最大顺序匹配分词算法上,采用三字长交集型歧义字段处理,提高了分词过程中的歧义识别能力。改进了基于TF*IDF的权重计算方法,结合CHI计算量计算权重,综合考虑了特征项在某类文本和所有文本中出现次数、类别信息对权重的影响和特征出现位置对权重的影响。实现了中文自动分类模型,设计了中文网页自动分类各个模块的构建方法,有效组织和处理网络上的海量信息,让人们更好的搜索到自己想要的资源,本申请网页自动分类是实现快速信息检索的重要技术,目前已在数字图书馆、新闻分类和搜索引擎等领域得到了广泛应用。
[0077] (3)本申请中文网页个性化分类技术有着巨大的技术优势和应用价值。首先,本申请的应用对象是互联网上的大量网页,自动对中文网页进行分类,提高网页文本处理的速度和效率,让用户能够清晰了然的面对网上杂乱的信息。其次,中文网页自动分类可以大幅减少用户的搜索时间,方便用户直接搜索所需要的网页类别,当用户在搜索引擎上输入关键词后,会得到大量的网页检索结果,但是得出的大部分网页检索结果却无法满足用户的需求,通过中文网页个性化分类,将网页分到各个不同的组,用户可以准确的定位到与目标网页最相关的组,从而更加快速的得到更多的搜索结果。第三,本申请能在多个领域有广泛的应用,如数字图书馆、搜索引擎、信息过滤等,对数字资源进行数字化管理,有较高的应用意义和价值。附图说明
[0078] 图1是本申请建立数据库表tbl_URL字段信息示意图。
[0079] 图2是字段state网页当前状态赋值表示示意图。
[0080] 图3是网页收集器的算法流程图。
[0081] 图4是网页的信息抽取模块的实现流程图。
[0082] 图5是本申请中用到的部分停用词示意图。
[0083] 图6是本申请定义标签权重具体赋值示意图。
[0084] 图7是网页精准分类模块工作流程图。具体实施方式
[0085] 下面结合附图,对本申请提供的基于大数据的中文网页个性化精准分类系统的技术方案进行进一步的描述,使本领域的技术人员能够更好的理解本申请并能够予以实施。
[0086] 如何有效组织和处理网络上的海量信息,让人们更好的搜索到自己想要的资源,是信息处理领域的重要任务。网页自动分类是实现快速信息检索的重要技术,目前已在数字图书馆、新闻分类和搜索引擎等领域得到了广泛应用。
[0087] 1.针对网页数据的抽取,由于网页中不同标签的特征项对于分类的重要程度不同,针对网页的标签结构特征,对网页特性进行了详细分析,设计了对HTML无用标签的过滤算法,并对高价值的标签集合赋予了对应的权值,提取得到对网页分类影响较大的标题、关键词及正文文本等。
[0088] 2.在文本前置处理过程中,对顺序最优匹配分词算法进行了改进,结合中文文本的特征,采用了三字长交集型歧义字段的处理思想,提高了算法的歧义识别能力。
[0089] 3.在特征筛选阶段,考虑到TF*IDF权重计算方法没有考虑到特征项在类间分布情况和每个类别中的分布情况,引入了CHI计算量计算特征项的分布不确定性,采用了TF*IDF*CHI权重计算方法,综合考虑了特征项在某一类和所有文本中出现的次数、类别信息对特征权重的影响及特征出现位置,较传统的TF*IDF权重计算方法更有合理性。
[0090] 4.在网页自动分类技术的研究基础上,提出了网页自动分类模型各模块的设计和实现方法,其中包括了网页抽取模块、网页前置处理模块、特征筛选模块、网页分类模块,并对模型进行了实验和测试,测试结果表明,分类模型具有较高的准确率,能达到设计要求:
[0091] 模块1:海量网页数据收集模块:对网页url进行爬取采集,构建网页文本集;
[0092] 模块2:网页数据前置处理模块:抽取网页高价值文本内容,并进行分词和去噪处理;
[0093] 模块3:特征提取模块:对特征进行筛选和提取;
[0094] 模块4:精准分类模块:构造精准分类系统对最终构建的文本向量分类。
[0095] 一、海量网页数据收集模块
[0096] 网页自动精准分类系统不直接对URL进行分类,而是先用网络爬虫爬取URL相对应的网页文本内容,然后对网页文本内容分类,本申请根据系统需求,设计如下网页收集模块:
[0097] (1)URL及网页内容用数据库保存:采用mysql数据库做URL池,基于mysql数据库自动去重,将表的对应字段url设置为唯一来对url自动去重;
[0098] (2)采用深度优先爬虫策略:采取深度优先爬虫方式爬取网页,让采集得到的网页内容尽可能多;
[0099] (3)采用多线程收集网页:采用多线程收集网页,降低数据库操作和等待的时间,同时,采用多线程收集网页,,提高系统的稳定性,保证程序长时间运行且不出错,线程采用独立运行的单元,在一个线程结束时释放申请的内存单元,但线程不能无限增加;
[0100] (4)非递归程序设计:递归的程序设计占用内存资源较多,难以控制,且不能长时间运行,因此采用非递归的程序设计。
[0101] 网页收集器的算法流程如下:
[0102] ①建立数据库表tbl_URL,字段信息如图1所示,将初始URL输入保存到表tbl_URL中,此时表中只有URL的字段信息,其余字段均为默认值,字段state的默认值为0,表示还没有对对应的url抓取网页;其中,字段state表示网页的当前状态,赋值的表示情况如图2所示;
[0103] ②从数据库中读取url,保证同一时间访问数据库读取url的线程只有一个,采用互斥信号量实现数据库的互斥访问,当一个线程读取url时,阻塞其余线程,直到该线程结束对数据库的访问,读取url的数据库命令为:select top 20*from tbl_URL where state=0order by id desc,一次最多提取20条url,抓取的数值用num表示,通过配置文件对该变量进行配置,抓取前state=0表示未抓取的网页url,执行操作后将state字段置为1,表示网页正在被抓取,操作结束后,结束阻塞,释放内存给下一个排队的线程;
[0104] ③抓取页面时,包括三种情况:一是对url进行解析时,无法获得该url相对应的端口和服务器等相关信息,则该url出错,将该url的state值置为5;二是能正确读取到该url相对应的端口和服务器等信息,则继续读取对应的页面内容,读取后将此url的state值置为2,表示该网页已抓取完毕,并将该网页内容保存到表tbl_URL的字段File中,同时查找该网页中的其他超链接,将其他超链接提取出来保存到数据库;三是能正确读取到该url对应的端口和服务器等信息,但继续读取相对应的网页内容时连接超时,则表示此url相对应的网页已不存在,此时把对应的unAccessible的数值加1,unAccessible的临界值设置为10,当该值达到临界值时,将state字段置为5;
[0105] ④一次的抓取完成后,则该线程结束,重复以上操作,继续下一次的抓取,直到数据库中没有state值为0的记录,程序结束,网页收集完成。
[0106] 网页收集器的算法流程图如图3所示。
[0107] 二、网页数据前置处理模块
[0108] 对网页内容进行信息抽取以及对得到的文本信息进行文本前置处理,包括中文分词和文本去躁,过程模块分为网页海量数据抽取和文本前置处理。
[0109] (一)网页海量数据抽取
[0110] 网页具有半结构化的特征,既有网页文本又有HTML,并且网页文本里不仅有正文内容,还有大量的文本噪声,抽取的内容是网页的正文内容和高价值标签中的相关内容,所以第一步就是对这些内容的文字信息进行抽取得到原始特征的语料源,网页的有效信息存在于标题title、超链接<ahref>、关键字meta keywords和描述description、正文body中,无用标签则包括有脚本标签<script>、注释标签<!——>、表单及相关标签<form><option>input>、样式标签<style>、对象标签<object>、applet标签<applet>、格式标签<hr>,将这七类标签作为无用标签处理,归于集合T,在网页前置处理模块对这些无用标签进行剔除。
[0111] 基于本申请定义的高价值标签和无用标签的代表含义,保留高价值标签中的文字内容,去除无用标签中的噪声信息,设计网页数据提取算法如下:
[0112] 第一步:输入html网页;
[0113] 第二步:读取html网页,查找<head>和< / head>标签,解析<meta>标签并记录,查找<title>和< / title>,输出html网页标题;
[0114] 第三步:循环读入html网页文本内容:
[0115] if读取到起始标签"<”,则继续读一直到读到">”,将该标签压入栈内,转第三步;
[0116] if读取到除"〉”外的其他字符,接着读取,一直到读取到起始标签"<”,将中间的内容当作正文保留,转第三步;
[0117] if读取到结束标签"< / ”,直接读取到出现">”,转第四步;
[0118] if读到超链接标签"”,接着读到相对应的结束标签"”,记录超链接标签之间的内容,根据锚文本判断链接是否和本页面主题相关,如果相关则插入到正文,转第三步;
[0119] if读取到结束符"\0”,转第六步;
[0120] 第四步:if读到无用标签集T中的结束标签,依次弹出位于栈顶的元素,一直到弹出与无用标签相对应的起始标签后,删除结束标签和起始标签当中的全部内容,转第六步;
[0121] if结束标签不是集合T中的标签;
[0122] if栈顶起始标签与结束标签相对应,弹出栈,转第六步;else转第五步;
[0123] 第五步:由栈顶起始标签来判断结束标签进行补充,弹出栈,转第六步;
[0124] 第六步:if文本未读取完,转第三步;
[0125] Else if栈为空,结束;
[0126] Else栈非空栈,转第五步;
[0127] 输出:正文信息;
[0128] 经过以上算法对网页进行内容提取,去掉无用标签,保留高价值的网页标签有:tag={title,meta,B,I,U,H1,H2,H3,H4,H5,a},信息提取方法由extract实现,输入为网页的html字符串,输出为标题、正文等字符串,网页的信息抽取模块的实现流程图如图4所示。
[0129] (二)文本前置处理
[0130] 将网页表示成向量模式,对第一步网页抽取后的文本信息进行中文分词、去停用词处理,首先要去掉标点符号,得到完全的文字信息。
[0131] 中文分词采用基于字符串的方法,处理未登录词识别和歧义识别,对顺序最优匹配法进行改进;
[0132] MaxLen代表初始最优匹配长度,L代表待分词的句子长度,P指向待分词的句子,Len表示实际取出的词条长度,每个汉字是两个字节,因此如果匹配失败需要减掉一个汉字重新匹配,则取出的字符串长度需要减去2,即Len=Len-2;
[0133] 但是,由于顺序最优匹配法将词条逐词匹配,因此需要进行无数次的字符匹配,影响了算法的速度,且由于顺序最优匹配法采用了稳定的词典,将词表作为分词的唯一标准,没有考虑到词中含词的情况,容易出现分词错误和歧义,因此在此方法的基础上,本申请对顺序最优匹配法进行了改进。
[0134] 基于分词过程中最容易有歧义的词是三字长交集型词,本申请针对三字长交集型的词将最优匹配分词法作如下改进:设置两个缓冲区Smian[p,q]及Ssub[i,j],Smian[p,q]表示从字符串的第p个字符开始取q个字符,则初始时p=0,q=L,L是整个字符串的长度,表示为Smian[O,L],即从字符串第一个字符开始取L个字符;Ssub[i,j]表示从字符的第i个字符开始取j个字符,代表从整个待分词的字符串中取出的字符串,在分词过程中,是Smian[p,q]的一个子集。
[0135] 在中文分词中最易产生分词歧义的类型是二子长父差型词语,因此本申请将MaxLen的值设为4,带分词的字符串长度为L,将待分词字符串放入缓冲区Smian[p,q]中,最长匹配字符为Len,其初始值为MaxLen,根据L和MaxLen的长度不同,算法分为两种情况:
[0136] 情况1:L≥MaxLen:从Smian[p,q]中的第一个字开始取长度为Len的字串放入Ssub[i,j]中,i=1,j=Len,将Ssub[i,j]中的字串和词表中的词语逐一进行匹配,如果匹配失败,则从第2个字开始再取长度为Len的字串放入Ssub[i,j],i=2,j=Len,将Ssub[i,j]中的字串再次和词表中的词逐一匹配,如果匹配失败,则按照上述步骤重复,从第3、4、5、…、L-Len个字开始取长度为Len的字串放入Ssub[i,j]中进行匹配,如果上述过程中所有匹配均失败,则表明字串中没有长Len的词,减去一个字进行匹配,将Len减2,重复上述步骤,直到匹配出成功的词。当有词匹配成功时,将该词切分出来,将该词左右两边的字串分开作为新的字串递归调用以上过程;
[0137] 情况2:L<MaxLen:此时待分词的字串的长度最大为3,首先对整个字串进行匹配,如果匹配不成功,则为了避免三字长交集型分词歧义字段从第二个字开始匹配,如果匹配成功则将该词切分出来,分词成功;否则取前两个字进行匹配,如果匹配成功则将这两个字作为词切分,分词成功;否则匹配不成功,则将该待分词字串看做是每个字是单字词,分词结束;
[0138] 在对文本进行中文分词后,保存下来许多出现较为频繁但和分类关系不大的词,即停用词,对停用词进行剔除处理,进行文本去噪。现有技术已经有一些采用较广泛的停用词表,但是表的构造与所选取的语料密切相关,很难通用于不同的领域,所以本申请对停用词表进行完善和补充。本申请中用到的部分停用词如图5所示。
[0139] 对停用词进行剔除,减少特征词的个数,降低文本中的噪声信息,还能减小存储空间,方便之后的分类操作,去停用词后,大量对分类无用的词如语气词、代词等被去掉了,特征词的质量得到了提高,可以降低向量的维数,提高特征词的辨识度。
[0140] 三、特征提取模块
[0141] 网页经过网页内容抽取和文本前置处理后,得到的词为特征项,特征项的个数为特征项的维数,特征项的维数过高会导致算法复杂度的提高以及分类效率的下降,因为高维度的特征项无法直接精确的提取最有效的信息。因此,下一步需要进行特征筛选。
[0142] 本申请特征筛选算法构造如下:
[0143] 第1步:输入经过文本前置处理后的初始特征词集;
[0144] 输出:经过特征筛选后的最终特征词集
[0145] 第2步:对于初始特征集中的每个词,利用式1:
[0146]
[0147] 对每个词和每个类别计算特征值,A表示特征项t和cj类文本同时出现的次数,B表示特征项t不出现在cj类文本中的次数,C表示cj类文本出现但t不出现的次数,D表示特征项t不出现又不属于cj类的次数,N表示训练集中所有的文本数,x2衡量类别ci和特征项t之间的关系,对每一对类别ci和特征t都计算x2(t,ci)值,然后按照由高到低排序,剔除值较低的特征;
[0148] 第3步:对于每个类别的所有特征词计算按特征值由低到高进行排序;
[0149] 第4步:取前K个词作为该类别的特征项,首先设置K初始值为1000,然后据实对K值不断进行调整;
[0150] 第5步:计算所有类别的特征项,对特征空间进行统一;
[0151] 第6步:输出经过特征筛选后的最终特征词集;
[0152] 经过以上特征筛选后,筛选出的特征项是最能代表文本类别的特征词,但这些词对于文本类别的判断的贡献度不同,即它们的权重不同,因此,对特征项计算对应的权重,权重影响因子设置2个:一是一个特征项在某类文本中出现越多表明它对该类别越重要,区分能力越强;二是一个特征项在所有的文本中出现的越多,表明该词对类别的区分能力越弱。本申请中采用TF*IDF算法对权重进行计算,但是由于网页的半结构性特征,该算法忽略了类型信息和特征出现的位置对权重的影响,因此,需要从两个方面对TF*IDF进行改进。
[0153] 特征词在某个类别的文本中出现越多,在文本间分布的越集中越重要,因此在TFIDF的权重计算中融入类别信息对特征权重的影响因子,本申请采用x2计算量来表示特征词和类别的相关性,改进后的公式如式2:
[0154] Wij=tfij×idfij×CHI(ti) 式2
[0155] 其中,Wij表示特征词的权重,TF和IDF的意义和TF*IDF公式中的相同,CHI表示特征项的特征值计算量。
[0156] 由于网页文本中,不同位置标签的文本会对权重有着不同的影响,因此本申请在网页海量数据抽取模块中抽取到的标签Tag={Title,Meta,B,I,U,H1,H2,H3}中的文本和正文文本,定义标签权重集合为W={wt|t∈Tag},wt表示标签t的权重,并具体赋值如图6所示;
[0157] 最终,修改权重计算公式为式3所示:
[0158]
[0159] 经过特征筛选后,得到网页文本的最终特征项集,通过以上公式计算特征向量的权重,用向量对网页文本进行表示,至此特征筛选模块完成,即可进入网页分类模块。
[0160] 四、精准分类模块
[0161] 网页分类模块的核心是训练方法和分类算法,训练阶段中,输入所有训练集样本,调用精准分类系统的训练算法先进行独立于具体精准分类系统,然后特定于精准分类系统的处理;分类阶段中,精准分类系统将待分类文本表达成向量后,用训练阶段生成的精准分类系统分类。本申请在构建分类算法时,采用最邻近分类模型对文本进行分类,网页分类模块的工作步骤如下:
[0162] 步骤一:读取训练集中所有文本的空间向量化后的数据;
[0163] 步骤二:筛选k值,指定k个最邻近文本作为匹配数量;
[0164] 步骤三:计算训练集中每个文本和测试文本的相似度,计算公式为式4所示:
[0165]
[0166] 其中,x表示测试文本的特征向量,d代表训练文本的特征向量,n代表特征向量有n维,Wk表示向量的第k维;
[0167] 步骤四:将步骤三中计算出来的文本相似度由低到高排序,选出距离最相近的k个文本;
[0168] 步骤五:在k个相似的训练文本中,对其中每个类别的权重进行计算,计算式为:
[0169] p(x,cj)=∑sim(x,di)×y(di,cj)-T 式5
[0170] 其中,T为临界值,如果di属于cj类,y(di,cj)为1,否则为0;
[0171] 步骤六:比较步骤五中计算出的各个类别的权重,将文档分到数值最大的类中。
[0172] 其工作流程如图7所示。
[0173] 五、实验及分析
[0174] 首先,对精准分类系统进行模型的训练,在命令行中输入java-jarClassifier.jar-train C:\class\语料库\train;
[0175] 接下来,通过模型进行测试,必须要先有训练才能运行测试,输入java-jarClassifier.jar-predict C:\class\语料库\test,模型测试实验中,采用准确率(Precision)对分类效果进行考察,衡量实际分类后得到的结果中,真正符合分类类别的文档比率,体现了精准分类系统的性能准确程度。计算公式为:
[0176]
[0177] 其中,a是正确归到该类别的文档数,b是错误归到该类别的文档数,c是本该归到该类别但错误归到别的类别的文档数。输入java-jar Classifier.jar-test C:\class\语料库\train测试准确率,由实验结果可知,分类的准确率达到96.3%,精准分类系统达到了较好的效果。< / script> < / ahref>
Claims
1. A personalized and accurate classification system for Chinese web pages based on big data, characterized by: Directly classify the massive web pages on the Internet. The web page set is captured from the Internet according to a certain strategy, and then the web page data is pre-processed, and the features of the pre-processed text information are screened, and finally classified using a precise classification system: P1: Based on the different importance of feature items of different tags in web pages for classification, the tag structure characteristics of web pages are analyzed, and a filtering algorithm for useless HTML tags is constructed. The corresponding weights are assigned to high-value tag sets, and the titles, keywords and body texts that have a greater impact on web page classification are extracted; P2: In the text pre-processing process, the sequential optimal matching word segmentation algorithm is improved, combined with the characteristics of Chinese text, and a processing framework of three-character intersection ambiguous fields is adopted to enhance the algorithm's ambiguous recognition ability; P3: In the feature screening stage, based on the distribution of feature items between classes and in each category, the CHI calculation amount is incorporated to calculate the distribution uncertainty of feature items, and the TF*IDF*CHI weight calculation method is used to comprehensively consider the number of times the feature items appear in a certain category and all texts, the impact of category information on feature weights, and the location of feature appearance; P4: Build the modules of the automatic web page classification model, including: Module 1: Massive web page data collection module: crawl and collect web page URLs and build a web page text set; Module 2: Web page data pre-processing module: extract high-value text content from web pages, and perform word segmentation and denoising; Module 3: Feature extraction module: screening and extracting features; Module 4: Accurate classification module: construct an accurate classification system to classify the final constructed text vector; Feature extraction module: Step 1: Input the initial feature word set after text pre-processing; Output: The final feature word set after feature screening Step 2: For each word in the initial feature set, use Formula 1: Formula 1 Calculate the feature value for each word and each category, A represents the feature items t and c i The number of times the class text appears at the same time, B means that feature item t does not appear in c i The number of times in the class text, C represents c i The number of times the class text appears but t does not appear, D means that the feature item t does not appear and does not belong to c i The number of classes, N represents the number of texts in the training set, x 2 Measurement Categoryc i The relationship between the feature item t, for each pair of categories c i and feature t both calculate x 2 (t,c i ) values, and then sort them from high to low, eliminating features with lower values; Step 3: Calculate the feature words for each category , sort by eigenvalue from low to high; Step 4: Take the first K words as the feature items of this category. First, set the initial value of K to 1000, and then adjust the value of K according to the actual situation. Step 5: Calculate the feature items of all categories and unify the feature space; Step 6: Output the final feature word set after feature screening; The corresponding weights are calculated for the feature items, and two weight influencing factors are set: first, the more a feature item appears in a certain type of text, the more important it is to the category and the stronger its distinguishing ability is; second, the more a feature item appears in all texts, the weaker its ability to distinguish categories is. TF*IDF is improved from two aspects; The influence factor of category information on feature weight is incorporated into the weight calculation of TFIDF, using x 2 The calculation amount is used to express the correlation between feature words and categories. The improved formula is as follows: Formula 2 Among them, W ij Indicates the weight of the feature word. The meanings of TF and IDF are the same as those in the TF*IDF formula. CHI indicates the calculated characteristic value of the feature item. The text and body text in the tag Tag={title,meta,B,I,U,H1,H2,H3} extracted in the webpage massive data extraction module are defined as the tag weight set: , w t Represents the weight of label t and assigns a specific value; Finally, the modified weight calculation formula is shown in Formula 3: Formula 3 After feature screening, the final feature item set of the web page text is obtained. The weight of the feature vector is calculated by the above formula, and the web page text is represented by the vector. At this point, the feature screening module is completed and you can enter the web page classification module.
2. According to the big data-based personalized Chinese web page accurate classification system of claim 1, it is characterized by: Massive web page data collection module: The web page automatic and accurate classification system does not directly classify the URL, but first uses a web crawler to crawl the web page text content corresponding to the URL, and then classifies the web page text content. The following web page collection module is designed: (1) URLs and web page content are stored in a database: MySQL database is used as the URL pool, and automatic deduplication is performed based on the MySQL database. The corresponding field URL in the table is set to unique to automatically deduplicate the URL; (2) Adopting a depth-first crawler strategy: crawling web pages using a depth-first crawler approach to collect as much web page content as possible; (3) Use multi-threading to collect web pages: Use multi-threading to collect web pages to reduce database operation and waiting time. At the same time, use multi-threading to collect web pages to ensure that the program runs for a long time without errors. Threads use independent running units. When a thread ends, the requested memory unit is released, but the number of threads cannot be increased indefinitely. (4) Non-recursive programming.
3. According to the Chinese webpage personalized accurate classification system based on big data as claimed in claim 1, it is characterized in that: The algorithm flow of the web page collector is as follows: ① Create a database table tbl_URL and save the initial URL input into the table tbl_URL. At this time, the table only contains the URL field information, and the rest of the fields are default values. The default value of the field state is 0, indicating that the corresponding URL has not yet been crawled. The field state indicates the current state of the web page; ② Read the URL from the database to ensure that there is only one thread accessing the database to read the URL at the same time. Use a mutex semaphore to implement mutually exclusive access to the database. When a thread reads the URL, the remaining threads are blocked until the thread ends accessing the database. The database command for reading the URL is: select top 20 * from tbl_URL where state=0order by id desc. A maximum of 20 URLs can be extracted at a time. The captured value is represented by num. The variable is configured through the configuration file. Before the capture, state=0 indicates that the webpage URL has not been captured. After the operation is executed, the state field is set to 1, indicating that the webpage is being captured. After the operation is completed, the blocking ends and the memory is released to the next queued thread. ③ When crawling a page, there are three situations: First, when parsing the URL, if the port and server-related information corresponding to the URL cannot be obtained, the URL is wrong and the state value of the URL is set to 5; second, if the port and server information corresponding to the URL can be correctly read, the corresponding page content will continue to be read. After reading, the state value of the URL is set to 2, indicating that the web page has been crawled, and the web page content is saved to the field File of the table tbl_URL. At the same time, other hyperlinks in the web page are searched, and other hyperlinks are extracted and saved to the database; third, if the port and server information corresponding to the URL can be correctly read, but the connection times out when continuing to read the corresponding web page content, it means that the web page corresponding to the URL no longer exists. At this time, the corresponding unAccessible value is increased by 1, and the critical value of unAccessible is set to 10. When the value reaches the critical value, the state field is set to 5; ④ After one crawl is completed, the thread ends and repeats the above operation to continue the next crawl until there is no record with a state value of 0 in the database. The program ends and the web page collection is completed.
4. According to the Chinese webpage personalized accurate classification system based on big data as claimed in claim 1, it is characterized in that: Web page data pre-processing module: extract information from web page content and perform text pre-processing on the obtained text information, including Chinese word segmentation and text denoising. The process module is divided into web page massive data extraction and text pre-processing.
5. According to the Chinese webpage personalized accurate classification system based on big data as claimed in claim 1, it is characterized in that: Extracting massive data from web pages: Extracting text information from the main content of web pages and related content in high-value tags to obtain the original feature corpus source. The effective information of web pages is contained in the title, hyperlinks, <ahref>, keywords, descriptions, and body text. Useless tags include script tags. <script>、注释标签<!——>、表单及相关标签<form><option>input>、样式标签<style>、对象标签<object>、applet标签<applet>、格式标签<hr>,将这七类标签作为无用标签处理,归于集合T,在网页前置处理模块对这些无用标签进行剔除;基于定义的高价值标签和无用标签的代表含义,保留高价值标签中的文字内容,去除无用标签中的噪声信息,设计网页数据提取算法如下:第一步:输入html网页;第二步:读取html网页,查找<head>和< / head>标签,解析<meta>标签并记录,查找<title>和< / title>,输出html网页标题;第三步:循环读入html网页文本内容:经过以上算法对网页进行内容提取,去掉无用标签,保留高价值的网页标签有:tag={title,meta,B,I,U,H1,H2,H3,H4,H5,a},信息提取方法由extract实现,输入为网页的html字符串,输出为标题、正文字符串。6.根据权利要求1所述基于大数据的中文网页个性化精准分类系统,其特征在于,文本前置处理:将网页表示成向量模式,对第一步网页抽取后的文本信息进行中文分词、去停用词处理,首先要去掉标点符号,得到完全的文字信息;中文分词采用基于字符串的方法,处理未登录词识别和歧义识别,对顺序最优匹配法进行改进, MaxLen代表初始最优匹配长度,L代表待分词的句子长度,P指向待分词的句子,Len表示实际取出的词条长度,每个汉字是两个字节,因此如果匹配失败需要减掉一个汉字重新匹配,则取出的字符串长度需要减去2,即Len=Len-2;基于分词过程中最容易有歧义的词是三字长交集型词,针对三字长交集型的词将最优匹配分词法作如下改进:设置两个缓冲区Smian[p,q]及Ssub[i,j],Smian[p,q]表示从字符串的第p个字符开始取q个字符,则初始时p=0,q=L,L是整个字符串的长度,表示为Smian[O,L],即从字符串第一个字符开始取L个字符;Ssub [i,j]表示从字符的第i个字符开始取j个字符,代表从整个待分词的字符串中取出的字符串,在分词过程中,是Smian[p,q]的一个子集;将MaxLen的值设为4,带分词的字符串长度为L,将待分词字符串放入缓冲区Smian[p,q]中,最长匹配字符为Len,其初始值为MaxLen,根据L和MaxLen的长度不同,算法分为两种情况:情况1:L≥MaxLen:从Smian[p,q]中的第一个字开始取长度为Len的字串放入Ssub[i,j]中,i=1,j=Len,将Ssub[i,j]中的字串和词表中的词语逐一进行匹配,如果匹配失败,则从第2个字开始再取长度为Len的字串放入Ssub[i,j],i=2,j=Len,将Ssub[i,j]中的字串再次和词表中的词逐一匹配,如果匹配失败,则按照上述步骤重复,从第3、4、5、…、L-Len个字开始取长度为Len的字串放入Ssub[i,j]中进行匹配,如果上述过程中所有匹配均失败,则表明字串中没有长Len的词,减去一个字进行匹配,将Len减2,重复上述步骤,直到匹配出成功的词,当有词匹配成功时,将该词切分出来,将该词左右两边的字串分开作为新的字串递归调用以上过程;情况2:L<MaxLen:此时待分词的字串的长度最大为3,首先对整个字串进行匹配,如果匹配不成功,则为了避免三字长交集型分词歧义字段从第二个字开始匹配,如果匹配成功则将该词切分出来,分词成功;否则取前两个字进行匹配,如果匹配成功则将这两个字作为词切分,分词成功;否则匹配不成功,则将该待分词字串看做是每个字是单字词,分词结束;在对文本进行中文分词后,保存下来许多出现较为频繁但和分类关系不大的词,即停用词,对停用词进行剔除处理,进行文本去噪。7.根据权利要求1所述基于大数据的中文网页个性化精准分类系统,其特征在于,精准分类模块:网页分类模块的核心是训练方法和分类算法,训练阶段中,输入所有训练集样本,调用精准分类系统的训练算法先进行独立于具体精准分类系统,然后特定于精准分类系统的处理;分类阶段中,精准分类系统将待分类文本表达成向量后,用训练阶段生成的精准分类系统分类。8.根据权利要求7所述基于大数据的中文网页个性化精准分类系统,其特征在于,在构建分类算法时,采用最邻近分类模型对文本进行分类,网页分类模块的工作步骤如下:步骤一:读取训练集中所有文本的空间向量化后的数据;步骤二:筛选k值,指定k个最邻近文本作为匹配数量;步骤三:计算训练集中每个文本和测试文本的相似度,计算公式为式4所示: 式4其中,x表示测试文本的特征向量,d代表训练文本的特征向量,n代表特征向量有n维,Wk表示向量的第k维;步骤四:将步骤三中计算出来的文本相似度由低到高排序,选出距离最相近的k个文本;步骤五:在k个相似的训练文本中,对其中每个类别的权重进行计算,计算式为: 式5其中,T为临界值,如果di属于cj类,y(di,cj)为1,否则为0;步骤六:比较步骤五中计算出的各个类别的权重,将文档分到数值最大的类中。< / script> < / ahref>
Citation Information
Patent Citations
CBL feature extraction and denoising webpage accurate classification method
CN113516202A