Webpage English text similarity calculation method

By dividing text and extracting keywords on the English web browsing page, calculating feedback weights and adjusting word weights, the problem of insufficient application of the existing TF-IDF algorithm in complex scenarios is solved, and the quality improvement of the English web similarity model and the effectiveness of the recommendation strategy is enhanced.

CN120181072AInactive Publication Date: 2025-06-20LIAONING QIDIAN EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510661346.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing TF-IDF algorithm is not highly applicable in complex scenarios and cannot fully express the key information in the English web browsing page, resulting in low quality of the English web page similarity model, affecting the effectiveness of the recommendation strategy.

Method used

By dividing the text of the English web browsing page, keywords with advertising attributes, subject attributes and feedback attributes are extracted, the feedback relationship between the directory keywords and the feedback keywords is calculated, the feedback weight is obtained, the word rights of the program keywords are adjusted, and the word rights are further modified by verifying the credibility of the feedback keywords, and finally the text similarity between English web pages is calculated.

Benefits of technology

Deeply explore the core attributes of catalog keywords, improve the efficiency and application value of English web pages, obtain high-quality English web page clustering results, and enhance the effectiveness of recommendation strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181072A_ABST
    Figure CN120181072A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a webpage English text similarity calculation method, which comprises the following steps of: obtaining a feedback weight of each directory keyword according to each directory keyword and all feedback keywords in each target English webpage browsing page; and according to all the unit texts, obtaining the commonly-used property of each unit text, obtaining the credibility of each type of feedback keywords according to the commonly-used property of each unit text, and obtaining the final word right of each directory keyword according to the feedback weight of each directory keyword and the credibility of each type of feedback keywords. Through the optimized word right extraction method, the clustering result of the English webpage browsing page contains key information of the English webpage core attribute, and the English webpage classification efficiency and the application value are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method for calculating the similarity of English texts on web pages. Background Art

[0002] The existing TF-IDF algorithm in the field of natural language processing is used for information retrieval and text mining. The role of the TF-IDF algorithm is to evaluate the importance of words in a text, and it has been widely used in many fields, including process automation approval, literature retrieval, data management, user portrait description, and so on. However, the applicability of the traditional TF-IDF algorithm in complex scenarios is not high because its logic for extracting word weights is too simple and lacks the relevance with other words, and it cannot fully reflect the importance of a word in some scenarios.

[0003] For example, taking an English website that provides product search for users as an example, the website often adopts various types of recommendation strategies for users, and each recommendation strategy requires a large number of similar recommendable English web pages to support. Therefore, the similarity calculation model between English web pages is crucial for the generation of recommendation strategies. However, there are often many complex text blocks on the English web page browsing page. Taking the English product web page as an example, these text blocks include product parameters, promotional activities, user reviews, advertisements, etc. These text blocks all belong to the attached text information of the product. The existing TF-IDF word weight extraction algorithm relies too much on word frequency features and cannot fully express the key information in the English web page browsing page, which in turn limits the quality of the English web page similarity model, and this may lead to the ineffectiveness of the recommendation strategy. Summary of the Invention

[0004] The present invention provides a method for calculating the similarity of English texts on web pages to solve the problem that the existing TF-IDF word weight extraction algorithm relies too much on word frequency features, limits the quality of the English web page similarity model, and may lead to the ineffectiveness of the recommendation strategy.

[0005] A method for calculating the similarity of English texts on web pages according to the present invention adopts the following technical solution: An embodiment of the present invention provides a method for calculating the similarity of English texts on web pages, and the method includes the following steps: Obtain the browsing pages of all English web pages, take each English web page browsing page as the target English web page browsing page, and divide the target English web page browsing page into three types of texts, including advertisement attributes, main body attributes, and feedback attributes, to obtain all the keywords of the three types of texts in the target English web page browsing page; All directory keywords and feedback keywords in the target English web page browsing page are obtained based on all the keywords of the three types of texts in the target English web page browsing page. All the feedback keywords in the target English web page browsing page are divided into several unit texts. The feedback weight of each directory keyword is obtained based on each directory keyword and all the feedback keywords in each target English web page browsing page. The commonality of any one unit text containing each type of feedback keyword is obtained based on all the unit texts. The same feedback keywords are taken as one type of feedback keyword. The credibility of each type of feedback keyword is obtained based on the commonality of any one unit text containing each type of feedback keyword. The credibility of each type of feedback keyword represents the credibility of the feedback information expressed by each type of feedback keyword for any one or more directory keywords. The final word weight of each directory keyword is obtained based on the feedback weight of each directory keyword and the credibility of each type of feedback keyword. The English web page similarity weight of each directory keyword is obtained based on the final word weight of each directory keyword, and the clustering result of all English web pages is obtained based on the English web page similarity weight of each directory keyword.

[0006] Furthermore, the step of obtaining all directory keywords and feedback keywords in the target English web page browsing page based on all the keywords of the three types of texts in the target English web page browsing page includes the following specific steps: All the keywords of the advertisement attribute and the main body attribute text are used as directory keywords; All the keywords of the feedback attribute text are used as feedback keywords.

[0007] Furthermore, the step of obtaining the feedback weight of each directory keyword based on each directory keyword and all the feedback keywords in each target English web page browsing page includes the following specific steps: All the directory keywords and feedback keywords in each target English web page browsing page are converted into word vectors; Based on the word vectors of the directory keywords and feedback keywords in each target English web page browsing page, the cosine similarity between each directory keyword and each type of feedback keyword is obtained, and the type of feedback keyword with the largest cosine similarity to each directory keyword is obtained; Furthermore, the feedback weight of each directory keyword is obtained: where k represents the k-th directory keyword, represents the feedback weight of the k-th directory keyword, represents the word vector of the k-th directory keyword, e represents that the category corresponding to the feedback keyword with the largest cosine similarity to the k-th directory keyword is the e-th category, represents the word vector of the e-th type of feedback keyword, represents the number of unit texts containing the e-th type of feedback keywords among all unit texts in each target English web page browsing page; represents the word vector of the k-th directory keyword and the word vector of the e-th type of feedback keyword of the cosine similarity.

[0008] Furthermore, obtaining the commonality of any unit text containing each type of feedback keyword based on all unit texts includes the following specific steps: Combining any two feedback keywords in each unit text pairwise to obtain a number of feedback keyword combinations; Calculating the commonality of any unit text containing each type of feedback keyword: where n represents the n-th type of feedback keyword, b represents the b-th feedback keyword combination in any unit text containing the n-th type of feedback keyword, represents the number of all feedback keyword combinations in any unit text containing the n-th type of feedback keyword, represents the cosine similarity of the b-th feedback keyword combination in any unit text containing the n-th type of feedback keyword, represents the average value of the cosine similarities of all feedback keyword combinations in any unit text containing the n-th type of feedback keyword; r represents the r-th feedback keyword in any unit text containing the n-th type of feedback keyword, represents the number of all feedback keywords in any unit text containing the n-th type of feedback keyword, represents the probability of the r-th feedback keyword in any unit text containing the n-th type of feedback keyword in all unit texts; represents the commonality of any unit text containing the n-th type of feedback keyword.

[0009] Furthermore, obtaining the credibility of each type of feedback keyword based on the commonality of any unit text containing each type of feedback keyword includes the following specific steps: where n represents the n-th type of feedback keyword, z represents the z-th unit text containing the n-th type of feedback keyword, represents the commonality of the z-th unit text containing the n-th type of feedback keyword, represents the number of unit texts containing the n-th type of keyword among all unit texts in each target English web page browsing page, represents the credibility of the n-th type of feedback keyword.

[0010] Further, obtaining the final word weight of each directory keyword based on the feedback weight of each directory keyword and the credibility of each type of feedback keyword includes the following specific steps: Among them, k represents the kth directory keyword, e represents that the category corresponding to the feedback keyword with the largest cosine similarity to the kth directory keyword is the eth category, represents the feedback weight of the kth directory keyword, represents the credibility of the eth type of feedback keyword, represents the final word weight of the kth directory keyword.

[0011] Further, obtaining the English web page similarity weight of each directory keyword based on the final word weight of each directory keyword includes the following specific steps: Obtain the TF-IDF weights of all directory keywords; Among them, represents the final word weight of the kth directory keyword, represents the TF-IDF weight of the kth directory keyword, represents the English web page similarity weight of the kth directory keyword.

[0012] Further, obtaining the clustering result of all English web pages based on the English web page similarity weight of each directory keyword includes the following specific steps: When taking any English web page browsing page as the target English web page browsing page, other English web page browsing pages are non-target English web page browsing pages; Among them, a represents the ath target English web page browsing page, b represents the bth non-target English web page browsing page, u represents the uth directory keyword of the ath target English web page browsing page, represents the number of all directory keywords of the ath target English web page browsing page, represents the English web page similarity weight of the uth directory keyword of the ath target English web page browsing page, represents the word vector of the uth directory keyword of the ath target English web page browsing page, c represents that the directory keyword in the bth non-target English web page browsing page with the largest cosine similarity to the uth directory keyword of the ath target English web page browsing page is the cth directory keyword, represents the word vector of the cth directory keyword in the bth non-target English web page browsing page, represents the probability of the \(u\)-th directory keyword in the \(a\)-th target English web page browsing page, represents the probability of the \(c\)-th directory keyword in the \(b\)-th non-target English web page browsing page, represents the text similarity between the \(a\)-th target English web page browsing page and the \(b\)-th non-target English web page browsing page, represents a natural number, represents the cosine similarity between the \(u\)-th directory keyword in the \(a\)-th target English web page browsing page and the \(c\)-th directory keyword in the \(b\)-th non-target English web page browsing page; Cluster all English web pages according to the text similarity between each target English web page browsing page and other non-target English web page browsing pages to obtain a clustering result.

[0013] Further, the step of clustering all English web pages according to the text similarity between each target English web page browsing page and other non-target English web page browsing pages to obtain a clustering result includes the following specific steps: Obtain the average value of the text similarities between each English web page browsing page as the target English web page browsing page and all other English web page browsing pages, sort all English web page browsing pages in ascending order according to the average value of the text similarities between each English web page browsing page and all other English web page browsing pages to obtain an English web page browsing page sequence, and each English web page browsing page is a point on the English web page browsing page sequence; Use DBSCAN to cluster all English web page browsing pages to obtain a clustering result of all English web page browsing pages.

[0014] Further, the step of using DBSCAN to cluster all English web page browsing pages to obtain a clustering result of all English web page browsing pages includes the following specific steps: Preset the neighborhood radius of DBSCAN on the English web page browsing page sequence, obtain the average value of the text similarities of all English web page browsing pages within the neighborhood centered on each English web page browsing page, and use the average value of the text similarities of all English web page browsing pages within the neighborhood centered on each English web page browsing page as the density parameter of this English web page browsing page; Preset a density parameter threshold, regard the English web page browsing pages with density parameters greater than or equal to the preset density parameter threshold as core points, and regard the English web page browsing pages with density parameters less than the preset density parameter threshold as boundary points until all English web page browsing pages are visited. Obtain a clustering result of all English web page browsing pages according to all core points and boundary points on the English web page browsing page sequence, and the clustering result is several non-adjacent clusters on the English web page browsing page sequence.

[0015] The beneficial effects of the technical solution of the present invention are as follows: By dividing the text of the English web browsing page, the present invention obtains the feedback weight according to the feedback relationship between the directory keywords and the feedback keywords of the English web browsing page, adjusts the word weight of the directory keywords according to the feedback weight, deeply excavates the core attributes of the directory keywords, and then further corrects the word weight of the directory keywords by verifying the credibility of the feedback keywords. By using the finally corrected word weight to obtain the text similarity between English web browsing pages, a high-quality English web page clustering result can be obtained for the purpose of classifying the core attributes of English web pages, greatly improving the English web page classification efficiency and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the steps of a method for calculating the similarity of English text of web pages according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, will describe in detail the specific implementation manner, structure, features and effects of a method for calculating the similarity of English text of web pages according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0020] The following will specifically describe the specific solution of a method for calculating the similarity of English text of web pages provided by the present invention with reference to the accompanying drawings.

[0021] Taking an English website that provides product search for users as an example, after the user searches for the target product, the English website provides multiple English web browsing pages. Please refer to Figure 1 , which shows a flowchart of the steps of a method for calculating the similarity of English text of web pages according to an embodiment of the present invention. The method includes the following steps: Step S001: Obtain all English web page browsing pages. Take each English web page browsing page as the target English web page browsing page, and split the target English web page browsing page into three types of texts, including advertisement attributes, main body attributes, and feedback attributes, to obtain all the keywords of the three types of texts in the target English web page browsing page.

[0022] Obtain all English web pages to be classified; When taking each English web page as the target English web page, obtain the browsing page of each target English web page. The text information of the browsing page of the target English web page is mainly divided into advertisement attributes, main body attributes, and feedback attributes; Advertisement attributes are all the information with advertisement and publicity nature such as links, copywriting, pictures, etc. about the target product that appear in the target English web page browsing page. The main function of the advertisement is to infect users and belongs to the appealing type of text; Main body attributes are about the parameter details, services, etc. of the target product in the target English web page browsing page and belong to the display type of text; Feedback attributes are the comments, bullet screens, messages, etc. of users about the target product in the target English web page browsing page and belong to the feedback type of text; The above text information with different attributes has its own different HTML structures. Use a web crawler, such as Beautiful Soup, to capture the HTML content in the English web page browsing page, and then use CSS selectors to locate the text information of the three different attributes in the target English web page browsing page, and split all the text information in each English web page browsing page into three types of texts; Use the Text Rank tool to split all the keywords of the three types of texts in the target English web page browsing page. Use a data cleaning tool to remove punctuation marks and special characters from the three types of texts respectively, and perform part-of-speech tagging on all the split keywords in the three types of texts. Remove the keywords with the part-of-speech of prepositions, conjunctions, and pronouns from the text through the stop word list to reduce the data dimension and noise interference of the text.

[0023] Step S002: Obtain all the directory keywords and feedback keywords in the target English web page browsing page according to all the keywords of the three types of texts in the target English web page browsing page. Divide all the feedback keywords in the target English web page browsing page into several unit texts. Obtain the feedback weight of each directory keyword according to each directory keyword and all the feedback keywords in each target English web page browsing page; Obtain the commonness of each unit text according to all the unit texts, and obtain the credibility of each type of feedback keyword according to the commonness of each unit text; Obtain the final word weight of each directory keyword according to the feedback weight of each directory keyword and the credibility of each type of feedback keyword.

[0024] When calculating the similarity of English texts on web pages, it is necessary to construct a similarity model based on the keyword weights in the texts, and only construct the similarity model based on the independent attributes of the English web pages themselves. In the target English web page browsing page, the advertisement attribute and the main body attribute belong to the independent attributes of the English web page itself, while the feedback attribute is a non-independent attribute containing user information; the text information of the advertisement attribute and the main body attribute describes all the core attributes of the target English web page. For example, the appealing texts such as "slimming, fair complexion, sexy, cute" in the advertisement attribute, and the display texts such as "color, size, price" in the main body attribute; the feedback attribute describes the feedback information of the user on some core attributes in the advertisement attribute and the main body attribute, that is, the comments, bullet screens, messages, etc. on the target product. For example, in the user comment "This dress is really slimming and fair-complexioned, and the price is not expensive", the keywords such as "very slimming, fair complexion, price, not expensive" are the feedback keywords for the user to give feedback on the core attributes such as "slimming, fair complexion" in the product advertisement attribute and "price" in the main body attribute. When calculating the similarity of English texts on web pages, it is necessary to consider which of the keyword in the advertisement attribute and the main body attribute are the core attribute keywords that are most attractive to users. This requires in-depth mining in combination with the feedback attribute to obtain accurate keyword weights.

[0025] Traditional TF-IDF only assigns different weights to keywords based on statistical ideas, and its feasibility is poor when actually applied to the calculation of the similarity of English texts on web pages. It cannot accurately extract the core keywords with the most core attributes and attractiveness in the target English web page browsing page, which may lead to anomalies in the subsequent classification of English web pages, and the classification results of English web pages have no application value.

[0026] 1. First, the English web pages have various attribute entries in the advertisement attribute and the main body attribute. All the attribute entries form the attribute directory of the English web page. Therefore, all the keywords in the text information of the advertisement attribute and the main body attribute in each target English web page browsing page are used as directory keywords; All the keywords in the text information of the feedback attribute in each target English web page browsing page are used as feedback keywords. The same feedback keywords are grouped as one type of feedback keywords, and all the feedback keywords in each target English web page browsing page are divided into several unit texts according to the user ID; Then, according to the feedback attribute, obtain the feedback weight of each directory keyword, specifically: Use the Word2Vec technology in the field of natural language processing to convert all the directory keywords and feedback keywords in each target English web page browsing page into word vectors; The same feedback keywords are grouped as one type of feedback keywords; According to the word vectors of the directory keywords and feedback keywords in each target English web page browsing page, obtain the cosine similarity between each directory keyword and all types of feedback keywords, and obtain the type of feedback keywords with the largest cosine similarity to each directory keyword; Among them, k represents the k-th directory keyword, represents the feedback weight of the k-th directory keyword, represents the word vector of the k-th directory keyword, e represents that the category corresponding to the feedback keyword with the largest cosine similarity to the k-th directory keyword is the e-th category, represents the word vector of the e-th type of feedback keyword, represents the number of unit texts containing the e-th type of feedback keyword in all unit texts of each target English web page browsing page; represents the word vector of the k-th directory keyword and the word vector of the e-th type of feedback keyword of the cosine similarity; It should be noted that: Since the text of the feedback attribute is generated based on the target keywords including the core attributes of the English web page in the call-type advertising attribute and the display-type main attribute, and the directory keyword is the feedback source of the feedback keyword, that is, the feedback information of whether the user approves or disapproves of each directory keyword. Therefore, the higher the similarity between the directory keyword in the advertising attribute and main attribute of the target English web page browsing page and the feedback keyword in the feedback attribute, the higher the probability that the directory keyword belongs to the source of the feedback information; and the greater the number of occurrences of the feedback keyword in all unit texts, the more feedback times it represents. Therefore, the higher the maximum cosine similarity between the directory keyword and the feedback keyword, and the more times the feedback keyword with the largest cosine similarity to the directory keyword appears in the unit text, the higher the feedback weight of the directory keyword.

[0027] Through the feedback weight of the feedback keyword to any one or more directory keywords in the advertising attribute and main attribute, the deeper core attribute of the directory keyword can be reflected, and the English web pages with similar core attributes can be better classified in the subsequent clustering model.

[0028] 2. However, since there will also be some commonly used comment words in the directory keywords, it is necessary to verify the credibility of the text information of the feedback attribute to avoid the commonly used words having a high weight. Specifically: Combine any two feedback keywords in each unit text in pairs to obtain a number of feedback keyword combinations; Calculate the commonality of any one unit text containing each type of feedback keyword: Among them, n represents the nth type of feedback keyword, and b represents the bth feedback keyword combination in any unit text containing the nth type of feedback keyword. represents the total number of feedback keyword combinations in any unit text containing the nth type of feedback keyword. represents the cosine similarity of the bth feedback keyword combination in any unit text containing the nth type of feedback keyword. represents the average value of the cosine similarities of all feedback keyword combinations in any unit text containing the nth type of feedback keyword; r represents the rth feedback keyword in any unit text containing the nth type of feedback keyword. represents the total number of feedback keywords in any unit text containing the nth type of feedback keyword. represents the probability of the rth feedback keyword in any unit text containing the nth type of feedback keyword among all unit texts. represents the commonness of any unit text containing the nth type of feedback keyword. It should be noted that: since both the cosine similarity between feedback keywords and the probability of feedback keywords take values between 0 and 1, therefore the value range of is also between 0 and 1. The higher the commonness of the unit text, the less credible the feedback keywords that often appear in such unit texts. Therefore, according to the commonness of any unit text containing each type of feedback keyword, the credibility of each type of feedback keyword is obtained: Among them, n represents the nth type of feedback keyword, and z represents the zth unit text containing the nth type of feedback keyword. represents the commonness of the zth unit text containing the nth type of feedback keyword. represents the number of unit texts containing the nth type of keyword among all unit texts in each target English web page browsing page. represents the credibility of the nth type of feedback keyword. represents the average value of the commonness of all unit texts containing the nth type of keyword in the target English web page browsing page. The lower the average value, the lower the commonness of the unit texts to which the nth type of feedback keyword is applied. Then the nth type of feedback keyword is not a common word. Therefore, the nth type of feedback keyword is more credible for the feedback information expressed by any one or more directory keywords. Since The value range of is between 0 and 1. Therefore, subtract the average commonness of all unit texts containing the nth type of keyword in the target English web browsing page from the constant 1. The purpose is to correct the logical relationship and obtain the credibility of the nth type of feedback keyword. 。

[0029] The higher the credibility of the feedback keyword, the greater the impact on the word weight of the directory keyword.

[0030] 3. Further, according to the feedback weight of each directory keyword and the credibility of each type of feedback keyword, obtain the final word weight of each directory keyword: Among them, k represents the kth directory keyword, e represents that the category corresponding to the feedback keyword with the largest cosine similarity to the kth directory keyword is the eth category. represents the feedback weight of the kth directory keyword. represents the credibility of the eth type of feedback keyword. represents the final word weight of the kth directory keyword; represents multiplying the credibility of the eth type of feedback keyword by the feedback weight of the kth directory keyword to obtain the final word weight. 。

[0031] Obtain the final word weights of all directory keywords. The final word weights not only contain the core attributes of the directory keywords but also verify the true word weights of the directory keywords through credibility, and a subsequent better clustering model can be obtained.

[0032] Step S003: Obtain the English web page similarity weight of each directory keyword according to the final word weight of each directory keyword, and obtain the clustering result of all English web pages according to the English web page similarity weight of each directory keyword.

[0033] Obtain the TF-IDF weights of all directory keywords; Use the final word weights of all directory keywords to optimize the TF-IDF weights of all directory keywords, and adjust the weights of some keywords containing the core attributes of English web pages when calculating the English web page similarity.

[0034] represents the final word weight of the kth directory keyword. represents the TF-IDF weight of the kth directory keyword. represents the English web page similarity weight of the kth directory keyword; Multiply the final word weight of the k-th directory keyword by the TF-IDF weight of the k-th directory keyword to obtain the adjusted TF-IDF weight of the k-th directory keyword, and use this adjusted TF-IDF weight as the English web page similarity weight of the k-th directory keyword; Then obtain all the English web page browsing pages to be classified; calculate the text similarity between the browsing page of each English web page as the target English web page and the browsing pages of each other English web page, specifically: Take any English web page browsing page as the target English web page browsing page, and obtain all the directory keywords of the target English web page browsing page, the word vector of each directory keyword, the final word weight of each directory keyword, and the probability of each directory keyword appearing in the target English web page browsing page; Call all the English web page browsing pages other than the target English web page browsing page non-target English web page browsing pages, and obtain all the directory keywords of each non-target English web page browsing page, the word vector of each directory keyword, and the probability of each directory keyword appearing in each non-target English web page browsing page; According to the word vectors of all the directory keywords of the target English web page browsing page and the word vectors of the directory keywords of each non-target English web page browsing page, obtain the directory keyword with the highest cosine similarity between each directory keyword in the target English web page browsing page and each non-target English web page browsing page; Furthermore, calculate the text similarity between each target English web page browsing page and other non-target English web page browsing pages: Among them, a represents the a-th target English web page browsing page, b represents the b-th non-target English web page browsing page, u represents the u-th directory keyword of the a-th target English web page browsing page, represents the number of all directory keywords of the a-th target English web page browsing page, represents the English web page similarity weight of the u-th directory keyword of the a-th target English web page browsing page, represents the word vector of the u-th directory keyword of the a-th target English web page browsing page, c represents the c-th directory keyword in the b-th non-target English web page browsing page that has the maximum cosine similarity with the u-th directory keyword of the a-th target English web page browsing page, represents the word vector of the c-th directory keyword in the b-th non-target English web page browsing page, represents the probability of the u-th directory keyword of the a-th target English web page browsing page, represents the probability of the c-th directory keyword in the b-th non-target English web page browsing page, represents the text similarity between the a-th target English web page browsing page and the b-th non-target English web page browsing page, represents a natural number, represents the cosine similarity between the u-th directory keyword of the a-th target English web page browsing page and the c-th directory keyword in the b-th non-target English web page browsing page; represents the absolute value of the difference between the probability of the u-th directory keyword of the a-th target English web page browsing page and the probability of the c-th directory keyword in the b-th non-target English web page browsing page; represents the core attribute convergence between the u-th directory keyword of the a-th target English web page browsing page and the c-th directory keyword in the b-th non-target English web page browsing page, and the cosine similarity between the u-th directory keyword of the a-th target English web page browsing page and the c-th directory keyword in the b-th non-target English web page browsing page The larger it is, and the absolute value of the difference between the probabilities of the u-th directory keyword and the c-th directory keyword is smaller, the larger it is, which also means that the u-th directory keyword of the a-th target English web page browsing page and the c-th directory keyword in the b-th non-target English web page browsing page have more identical core attributes. Adding to the denominator is to avoid the denominator being 0; represents multiplying the English web page similarity weight of each directory keyword in the target English web page browsing page by the core attribute convergence between each directory keyword in the target English web page browsing page and each directory keyword in each other non-target English web page browsing page, represents using the English web page similarity weights of all directory keywords in the target English web page browsing page to perform weighted averaging on the core attribute convergence between all directory keywords in the target English web page browsing page and all directory keywords in each other non-target English web page browsing page, to obtain the text similarity between the a-th target English web page browsing page and the b-th non-target English web page browsing page ; Since in this embodiment, when measuring the text similarity between English web page browsing pages, the text similarity is calculated based on the differences between other non-target English web page browsing pages and each target English web page browsing page. When the target English web page browsing page and the non-target English web page browsing page are interchanged, the text similarity will change with the target English web page browsing page. Therefore, when constructing a clustering model using the text similarity later, the density center clustering method needs to be adopted, and this embodiment adopts the DBSCAN clustering model; First, assume that each English web page browsing page is regarded as a target English web page browsing page. Then, calculate the average value of the text similarity between each English web page browsing page and all other English web page browsing pages when it is used as the target English web page browsing page. Sort all English web page browsing pages in ascending order according to the average value of the text similarity between each target English web page browsing page and all other English web page browsing pages to obtain an English web page browsing page sequence, and each English web page browsing page is a point on the sequence. Then, preset the neighborhood radius of DBSCAN on the English web page browsing page sequence as R = 15, and obtain the average value of the text similarity of all English web page browsing pages within the neighborhood centered on each English web page browsing page. Take the average value of the text similarity of all English web page browsing pages within the neighborhood centered on each English web page browsing page as the density parameter of this English web page browsing page. Set the density parameter threshold to 0.7. Regard the English web page browsing pages with density parameters greater than or equal to this threshold as core points, and regard the English web page browsing pages with density parameters less than this threshold as boundary points. Starting from any core point, recursively examine all English web page browsing pages within its neighborhood. If the English web page browsing pages within the neighborhood are also core points, then merge them into one cluster. This process will continue until there are no more English web page browsing pages that can be added to the cluster. For those boundary points that are not core points but are within the neighborhood of core points, also add them to the cluster where the core point is located. Then, mark all English web page browsing pages that have not been assigned to a cluster as noise points. Repeat the above steps until all English web page browsing pages have been visited, and obtain several non-adjacent clusters on the English web page browsing page sequence. These clusters are the clustering results of the English web page browsing pages.

[0035] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for calculating the similarity of English web page texts, characterized in that, The method includes the following steps: Obtain the browsing pages of all English web pages. Take each English web page browsing page as the target English web page browsing page, and divide the target English web page browsing page into three types of texts, including advertisement attributes, main body attributes, and feedback attributes, to obtain all the keywords of the three types of texts in the target English web page browsing page; Obtain all the directory keywords and feedback keywords in the target English web page browsing page according to all the keywords of the three types of texts in the target English web page browsing page. Divide all the feedback keywords in the target English web page browsing page into several unit texts. Obtain the feedback weight of each directory keyword according to each directory keyword and all the feedback keywords in each target English web page browsing page; Obtain the commonness of any one unit text containing each type of feedback keyword according to all the unit texts. Take the same feedback keywords as one type of feedback keyword, and obtain the credibility of each type of feedback keyword according to the commonness of any one unit text containing each type of feedback keyword. The credibility of each type of feedback keyword represents the credibility of the feedback information expressed by each type of feedback keyword for any one or more directory keywords; Obtain the final word weight of each directory keyword according to the feedback weight of each directory keyword and the credibility of each type of feedback keyword; Obtain the English web page similarity weight of each directory keyword according to the final word weight of each directory keyword, and obtain the clustering result of all English web pages according to the English web page similarity weight of each directory keyword.

2. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, The specific steps included in obtaining all the directory keywords and feedback keywords in the target English web page browsing page according to all the keywords of the three types of texts in the target English web page browsing page are as follows: Take all the keywords of the advertisement attributes and main body attribute texts as directory keywords; Take all the keywords of the feedback attribute text as feedback keywords.

3. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, The specific steps included in obtaining the feedback weight of each directory keyword according to each directory keyword and all the feedback keywords in each target English web page browsing page are as follows: Convert all the directory keywords and feedback keywords in each target English web page browsing page into word vectors; According to the word vectors of the directory keywords and feedback keywords in each target English web page browsing page, obtain the cosine similarity between each directory keyword and each type of feedback keyword, and obtain the type of feedback keyword with the largest cosine similarity to each directory keyword; Furthermore, obtain the feedback weight of each directory keyword: Among them, k represents the k-th directory keyword, indicating the feedback weight of the k-th directory keyword, representing the word vector of the k-th directory keyword, where e represents the category of the feedback keyword with the largest cosine similarity to the k-th directory keyword being the e-th category, representing the word vector of the e-th category of feedback keywords, representing the number of unit texts containing the e-th category of feedback keywords among all unit texts in each target English web page browsing page; representing the word vector of the k-th directory keyword and the word vector of the e-th category of feedback keywords is the cosine similarity.

4. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, The specific steps included in obtaining the commonness of any one unit text containing each type of feedback keyword according to all the unit texts are as follows: Combine any two feedback keywords in each unit text in pairs to obtain several feedback keyword combinations; Calculate the commonness of any one unit text containing each type of feedback keyword: Among them, n represents the nth type of feedback keyword, and b represents the bth feedback keyword combination in any unit text containing the nth type of feedback keyword. represents the total number of feedback keyword combinations in any unit text containing the nth type of feedback keyword. represents the cosine similarity of the bth feedback keyword combination in any unit text containing the nth type of feedback keyword. represents the average value of the cosine similarities of all feedback keyword combinations in any unit text containing the nth type of feedback keyword; r represents the rth feedback keyword in any unit text containing the nth type of feedback keyword. represents the total number of feedback keywords in any unit text containing the nth type of feedback keyword. represents the probability of the rth feedback keyword in any unit text containing the nth type of feedback keyword among all unit texts. represents the commonality of any unit text containing the nth type of feedback keyword.

5. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, The specific steps included in obtaining the credibility of each type of feedback keyword according to the commonness of any one unit text containing each type of feedback keyword are as follows: Among them, n represents the nth type of feedback keyword, and z represents the zth unit text containing the nth type of feedback keyword. represents the commonness of the zth unit text containing the nth type of feedback keyword. represents the number of unit texts containing the nth keyword among all unit texts in each target English web page browsing page. represents the credibility of the nth type of feedback keyword.

6. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, The specific steps included in obtaining the final word weight of each directory keyword according to the feedback weight of each directory keyword and the credibility of each type of feedback keyword are as follows: Among them, k represents the k-th directory keyword, and e represents that the category corresponding to the feedback keyword with the largest cosine similarity to the k-th directory keyword is the e-th category. Indicates the feedback weight of the k-th directory keyword. Represents the credibility of the feedback keywords in the e-th category. Represents the final word weight of the k-th directory keyword.

7. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, Obtaining the English web page similarity weight for each directory keyword according to the final word weight of each directory keyword includes the following specific steps: Obtaining the TF-IDF weights of all directory keywords; Among them, represents the final word weight of the k-th directory keyword, represents the TF-IDF weight of the k-th directory keyword, represents the English web page similarity weight of the k-th directory keyword.

8. The method for calculating the similarity of English web page texts according to claim 1, characterized in that, Obtaining the clustering result of all English web pages according to the English web page similarity weight of each directory keyword includes the following specific steps: When taking any English web page browsing page as the target English web page browsing page, other English web page browsing pages are non-target English web page browsing pages; Among them, a represents the a-th target English web page browsing page, b represents the b-th non-target English web page browsing page, u represents the u-th directory keyword of the a-th target English web page browsing page, represents the total number of directory keywords of the a-th target English web page browsing page, represents the English web page similarity weight of the u-th directory keyword of the a-th target English web page browsing page, represents the word vector of the u-th directory keyword of the a-th target English web page browsing page. c represents that the directory keyword in the b-th non-target English web page browsing page with the maximum cosine similarity to the u-th directory keyword of the a-th target English web page browsing page is the c-th directory keyword, represents the word vector of the c-th directory keyword in the b-th non-target English web page browsing page, represents the probability of the u-th directory keyword of the a-th target English web page browsing page, represents the probability of the c-th directory keyword in the b-th non-target English web page browsing page, represents the text similarity between the a-th target English web page browsing page and the b-th non-target English web page browsing page, represents a natural number, represents the cosine similarity between the u-th directory keyword of the a-th target English web page browsing page and the c-th directory keyword in the b-th non-target English web page browsing page; Clustering all English web pages according to the text similarity between each target English web page browsing page and other non-target English web page browsing pages to obtain the clustering result.

9. The method for calculating the similarity of web page English texts according to claim 8, wherein, Clustering all English web pages according to the text similarity between each target English web page browsing page and other non-target English web page browsing pages to obtain the clustering result includes the following specific steps: Obtaining the average value of the text similarity between each English web page browsing page as the target English web page browsing page and all other English web page browsing pages, sorting all English web page browsing pages in ascending order according to the average value of the text similarity between each English web page browsing page and all other English web page browsing pages to obtain the English web page browsing page sequence, and each English web page browsing page is a point on the English web page browsing page sequence; Using DBSCAN to cluster all English web page browsing pages to obtain the clustering result of all English web page browsing pages.

10. The method for calculating the similarity of web page English texts according to claim 9, wherein, Using DBSCAN to cluster all English web page browsing pages to obtain the clustering result of all English web page browsing pages includes the following specific steps: Presetting the neighborhood radius of DBSCAN on the English web page browsing page sequence, obtaining the average value of the text similarity of all English web page browsing pages within the neighborhood centered on each English web page browsing page, and taking the average value of the text similarity of all English web page browsing pages within the neighborhood centered on each English web page browsing page as the density parameter of this English web page browsing page; Presetting the density parameter threshold, taking the English web page browsing pages with density parameters greater than or equal to the preset density parameter threshold as core points, taking the English web page browsing pages with density parameters less than the preset density parameter threshold as boundary points until all English web page browsing pages are all visited, and obtaining the clustering result of all English web page browsing pages according to all core points and boundary points on the English web page browsing page sequence. The clustering result is several non-adjacent clusters on the English web page browsing page sequence.