Cheating Web Page Detection Method and Device Integrating a Propagation Model and an Ensemble Classification Model
By integrating the propagation model and the ensemble classification model, forward trust propagation and reverse non-trust propagation, combined with the punishment operation of trust value and non-trust value, the problem of low detection accuracy of cheating web pages caused by insufficient labeling samples is solved, and more efficient cheating web page recognition is achieved.
Patent Information
- Application Number
- CN202510578106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing cheating web page detection methods are insufficient in the labeling sample, resulting in low detection accuracy and cannot effectively identify and distinguish cheating web pages.
The method of fusion propagation model and ensemble classification model is adopted to determine the credibility and cheating of unmarked sample web pages through forward trust propagation and reverse non-trust propagation, combined with the penalty operation of trust value and non-trust value, and the integrated classification model is iteratively updated to improve detection accuracy.
The accuracy of cheating web page detection is improved, and the performance of the dissemination model and classification model is enhanced by using a small number of labeled samples, and the quality of search results and information retrieval efficiency.
Smart Images

Figure CN120105314B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of web page detection, and in particular to a cheating web page detection method and device that integrates a propagation model and an integrated classification model. Background Art
[0002] Some people use various means to deceive search engine ranking algorithms, so that some web pages are ranked higher than they actually deserve. All these deceptive behaviors that attempt to improve web page rankings are usually called web spam. If spam web pages appear at the front of the search engine result list, it will worsen the search engine ranking results, affect search performance, bring great obstacles to users in the process of obtaining information, and reduce user experience and user trust in search engines. From the perspective of search engines, even if these spam web pages are not ranked high enough to disturb users, crawling, indexing and storing these web pages also increases the space overhead during indexing and the time overhead during retrieval. In addition, many spam web pages contain viruses or malware that can harm search engines and users' computers. Therefore, detecting spam has become one of the major challenges of modern search engines.
[0003] In the existing technology, cheating web page detection can be regarded as a special binary classification problem (cheating or normal), so classification is one of the effective technologies for detecting web page cheating. However, the existing classification method faces the problem of insufficient labeled samples, resulting in low detection accuracy of cheating web pages. Summary of the invention
[0004] Based on this, it is necessary to provide a cheating web page detection method and device that integrates a propagation model and an integrated classification model to address the above technical problems, which can improve the accuracy of cheating web page detection.
[0005] The present invention adopts the following technical solutions:
[0006] The present invention provides a cheating web page detection method that integrates a propagation model and an integrated classification model, comprising:
[0007] The labeled sample web page set, the unlabeled sample web page set and the web page link structure diagram are input into a propagation model based on three-valued subjective logic or double-end penalty, and forward trust propagation and reverse distrust propagation are performed through the labeled sample web page set and the web page link structure diagram to obtain the trust value and distrust value of the unlabeled sample web page in the unlabeled sample web page set; the double-end penalty includes forward propagation of trust value along the outbound link and reverse propagation of distrust value along the inbound link, while performing penalty operations on the trust value and distrust value at the source and target ends of the propagation;
[0008] Integrate and classify unlabeled sample web pages through an integrated classification model to determine the integrated classification results of the unlabeled sample web pages; the integrated classification model is trained based on a set of labeled sample web pages;
[0009] Determine the trustworthy web pages as the unlabeled sample web pages whose trust values are among the top first preset rankings and whose integrated classification results are trustworthy web pages, and determine the cheating web pages as the unlabeled sample web pages whose non-trust values are among the top second preset rankings and whose integrated classification results are untrustworthy web pages;
[0010] Add the trustworthy web pages and cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively update the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web pages to be detected.
[0011] Preferably, the propagation model is a propagation model based on three-valued subjective logic, and reverse non-trust propagation is performed through the labeled sample web page set and the web page link structure diagram to obtain the non-trust values of the unlabeled sample web pages in the unlabeled sample web page set, including:
[0012] Determine the initial cheating views of each unlabeled sample web page according to the link relationships between each unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram;
[0013] Determine the view matrix according to the initial cheating views of each unlabeled sample web page, the initial cheating views of the sample web pages pointed to by each unlabeled sample web page, and the web page link relationships;
[0014] Obtain the cheating sample web pages in the labeled sample web page set and the web page link structure diagram, and determine the updated cheating views of each unlabeled sample web page through reverse non-trust propagation and combination operations according to the view matrix and the cheating sample web pages;
[0015] Quantify the updated cheating views of each unlabeled sample web page to obtain the non-trust value of each unlabeled sample web page.
[0016] Preferably, determining the initial cheating views of each unlabeled sample web page according to the link relationships between each unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram includes:
[0017] For any unlabeled sample web page, obtain the in-chain sample web pages along the in-chain direction of the unlabeled sample web page and the out-chain sample web pages along the out-chain direction of the unlabeled sample web page according to the link relationships between the unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram;
[0018] If the total number of web pages of the in-link sample web pages and the out-link sample web pages is 0, then determine that the unlabeled sample web page is an uncertain opinion;
[0019] If the total number of web pages of the in-link sample web pages and the out-link sample web pages is not 0, then determine the initial cheating opinion of the unlabeled sample web page according to the number of trusted sample web pages, cheating sample web pages and uncertain category sample web pages in the in-link sample web pages and the out-link sample web pages.
[0020] Preferably, according to the opinion matrix and the cheating sample web pages, determine the updated cheating opinion of each unlabeled sample web page through reverse untrusted propagation and combination operations, including:
[0021] For any unlabeled sample web page, obtain the out-link sample web pages along the out-link direction of the unlabeled sample web page;
[0022] If the number of out-link sample web pages is 1, obtain the cheating opinion of the out-link sample web page and the cheating opinion in the opinion matrix, and use the propagation operation of the cheating opinion to determine the updated cheating opinion of the unlabeled sample web page;
[0023] If the number of out-link sample web pages is multiple, combine the updated cheating opinions of the multiple out-link sample web pages by the propagation operation to obtain the updated cheating opinion of the unlabeled sample web page.
[0024] Preferably, the propagation model is a propagation model based on double-end penalty. Through the labeled sample web page set and the web page link structure diagram, perform forward trusted propagation and reverse untrusted propagation to obtain the trust value and untrust value of the unlabeled sample web pages in the unlabeled sample web page set, including:
[0025] For any unlabeled sample web page, obtain the in-link sample web pages along the in-link direction of the unlabeled sample web page, and the out-link sample web pages along the out-link direction of the unlabeled sample web page from the labeled sample web page set and the web page link structure diagram;
[0026] Perform source-end penalty on the unlabeled sample web page through the trust value and untrust value of the out-link sample web page to obtain the first credible value of the unlabeled sample web page;
[0027] Perform source-end penalty on the unlabeled sample web page through the trust value and untrust value of the in-link sample web page to obtain the first non-credible value of the unlabeled sample web page;
[0028] According to the credible value of the out-link sample web page, and the first credible value and the first non-credible value of the unlabeled sample, perform target-end penalty on the unlabeled sample web page to obtain the trust value of the unlabeled sample;
[0029] Based on the distrust value of the in-link sample web pages, as well as the first trust value and the first distrust value of the unlabeled samples, perform target-end punishment on the unlabeled sample web pages to obtain the distrust value of the unlabeled samples.
[0030] The present invention provides a cheating web page detection device integrating a propagation model and an integrated classification model, including:
[0031] A propagation module, configured to input a labeled sample web page set, an unlabeled sample web page set, and a web page link structure diagram into a propagation model based on three-valued subjective logic or double-end punishment, and perform forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust value and the distrust value of the unlabeled sample web pages in the unlabeled sample web page set; the double-end punishment includes performing punishment operations on the trust value and the distrust value at the source end and the target end of the propagation while propagating the trust value forward along the out-links and propagating the distrust value backward along the in-links;
[0032] A classification module, configured to perform integrated classification on the unlabeled sample web pages through an integrated classification model to determine the integrated classification result of the unlabeled sample web pages; the integrated classification model is trained according to the labeled sample web page set;
[0033] A determination module, configured to determine the unlabeled sample web pages with the trust value in the top first preset ranking and the integrated classification result being a trusted web page as trusted web pages, and determine the unlabeled sample web pages with the distrust value in the top second preset ranking and the integrated classification result being an untrusted web page as cheating web pages;
[0034] An update module, configured to add the trusted web pages and the cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and perform iterative updates on the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web pages to be detected.
[0035] The present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned cheating web page detection method integrating a propagation model and an integrated classification model is implemented.
[0036] The present invention provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above-mentioned cheating web page detection method integrating a propagation model and an integrated classification model is implemented.
[0037] The above-mentioned at least one technical solution adopted by the present invention can achieve the following beneficial effects:
[0038] In the present invention, the results of the propagation model based on three-valued subjective logic or two-sided penalty and the integrated classification model are fused together to accurately determine whether an unlabeled sample web page is a trustworthy web page or a cheating web page. In this way, by using the integrated classification model of a small number of labeled sample web pages and combining the results of forward trust propagation and reverse distrust propagation, more trustworthy and cheating web page samples are obtained and added to the labeled sample web page set. Then, the integrated classification model and the propagation model are iteratively updated through the labeled sample web page set, further enhancing the classification performance of the integrated classification model and the propagation ability of the propagation algorithm, thereby improving the accuracy of cheating web page detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0040] Figure 1 is a schematic flow chart of a method for detecting cheating web pages by fusing a propagation model and an integrated classification model provided by the present invention;
[0041] Figure 2 is a web page link graph provided by the present invention;
[0042] Figure 3 is a schematic flow chart of a method for detecting cheating web pages by fusing a propagation model based on three-valued subjective logic and an integrated classification model provided by the present invention;
[0043] Figure 4 is a schematic diagram of a computer device for implementing a method for detecting cheating web pages by fusing a propagation model and an integrated classification model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] According to the "55th Statistical Report on the Development of China's Internet Network", as of December 2024, the scale of Internet users in China reached 1.108 billion, and the Internet penetration rate was 78.6%. Among them, the scale of search engine users reached 878 million, accounting for 79.2% of the overall Internet users. For a given query, usually search engines can return hundreds of thousands of results. However, research has found that 85% of search engine users only view the content on the first page, that is, the top ten web pages, and about 60% of users only visit the first 5 results on the first page. Therefore, the ranking position in the search engine's returned results has become a concern for network service providers. Especially for some commercial sites, moving up in the search engine rankings can translate into increased sales, revenue, and profits.
[0046] Gyöngyi et al. estimated that approximately 10% to 15% of the content on web pages is cheating content; Teachers Liu Yiqun, Ma Shaoping, etc. conducted a sampling analysis of approximately 800 million Chinese web pages and concluded that approximately 15% of the Chinese web resources are cheating web pages.
[0047] Currently, the battle between cheating and anti-cheating is like an "arms race". Once a cheating technique is detected and prohibited, new cheating techniques will emerge. Therefore, although researchers have proposed many anti-cheating techniques, there are still many cheating web pages in the search engine's returned result list, and anti-cheating will continue.
[0048] Since web page cheating detection can be regarded as a special binary classification problem (cheating or normal), cheating web page detection that combines the propagation model and the integrated classification model is one of the effective techniques for detecting web page cheating. However, the classification method faces problems such as insufficient labeled samples and the need to improve classification accuracy. In addition, since the link relationship between web pages can be modeled as a directed graph, the trust / non-trust values of web pages can be calculated using the trust and non-trust propagation models starting from the labeled normal web page seeds and cheating web page seeds respectively, so as to detect web page cheating. However, due to the non-connectivity of the hyperlink structure graph between web pages, the limited number of labeled seeds, and the fact that representing the link relationship between web pages with a single numerical value is not precise enough, the propagation ability of this method is limited. Therefore, researching how to improve the propagation ability of the trust / non-trust propagation model, and organically integrating the credible web pages obtained from its forward trust propagation and the cheating web pages obtained from its reverse non-trust propagation with the results of the integrated classification technology in web page cheating detection, and using a small number of labeled samples to improve the detection rate of cheating web pages, enhance the quality of search results and the efficiency of people retrieving information, has important theoretical significance and practical application value.
[0049] Based on this, the present invention provides a method for detecting cheating web pages by integrating a propagation model and an integrated classification model. This method first trains an integrated classification model and a trust / non-trust propagation model with limited propagation ability using a small number of labeled sample web pages, and then combines the results of the trust / non-trust propagation model and the integrated classification model in the detection of cheating web pages. By adding more samples to the set of labeled sample web pages, a propagation model with stronger propagation ability and an integrated classification model with better classification performance can be obtained, reducing the cost of labeled samples, improving the learning performance, and enhancing the detection rate of web page cheating.
[0050] The method for detecting cheating web pages by integrating a propagation model and an integrated classification model in the present invention can be applied to a server. The server can be a server set up in a business platform or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention. The following will be described with the server as the execution entity.
[0051] The following will detail the technical solutions provided by each embodiment of the present invention in conjunction with the accompanying drawings.
[0052] Figure 1 FIG. is a schematic flow chart of a method for detecting cheating web pages by integrating a propagation model and an integrated classification model in the present invention, which specifically includes the following steps:
[0053] S101, input the set of labeled sample web pages, the set of unlabeled sample web pages, and the web page link structure diagram into a propagation model based on ternary subjective logic or double-end penalty. Through positive trust propagation and negative non-trust propagation using the set of labeled sample web pages and the web page link structure diagram, obtain the trust values and non-trust values of the unlabeled sample web pages in the set of unlabeled sample web pages; the double-end penalty includes performing penalty operations on the trust values and non-trust values at the source end and the target end of the propagation while propagating the trust values forward along the out-links and the non-trust values backward along the in-links.
[0054] Among them, the set of labeled sample web pages includes multiple labeled sample web pages. The multiple labeled sample web pages can include multiple trustworthy sample web pages, multiple cheating sample web pages, and multiple sample web pages of uncertain categories; the trustworthy sample web pages, cheating sample web pages, and sample web pages of uncertain categories in the set of labeled sample web pages can be determined after observing the sample web pages; the set of unlabeled sample web pages can include multiple unlabeled sample web pages, and the unlabeled sample web pages are sample web pages that are uncertain whether they are cheating web pages or trustworthy web pages. The web page link structure diagram is determined according to the link relationships between web pages; the web page link structure diagram can include the link relationships between each sample web page in the set of labeled sample web pages and the set of unlabeled sample web pages.
[0055] In an exemplary embodiment, the propagation model is a propagation model based on three-valued subjective logic. Positive trust propagation and reverse distrust propagation are performed through a labeled sample web page set and a web page link structure diagram to obtain the trust value and distrust value of unlabeled sample web pages in the unlabeled sample web page set, including: performing positive trust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust value of unlabeled sample web pages in the unlabeled sample web page set; performing reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the distrust value of unlabeled sample web pages in the unlabeled sample web page set.
[0056] Specifically, in positive trust propagation, based on the trustworthy sample web pages of the labeled sample web page set and the web page link structure diagram, positive trust learning is performed on the unlabeled sample web pages to obtain the trust value of the unlabeled sample web pages. In reverse distrust propagation, based on the cheating sample web pages of the labeled sample web page set and the web page link structure diagram, reverse distrust learning is performed on the unlabeled sample web pages to obtain the distrust value of the unlabeled sample web pages.
[0057] In an exemplary embodiment, performing reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the distrust value of unlabeled sample web pages in the unlabeled sample web page set includes: determining the initial cheating view of each unlabeled sample web page according to the link relationship between each unlabeled sample web page and its in-neighbor and out-neighbor web pages in the web page link structure diagram; determining the view matrix according to the initial cheating view of each unlabeled sample web page, the initial cheating view of the sample web page pointed to by each unlabeled sample web page, and the web page link relationship; obtaining the cheating sample web pages in the labeled sample web page set and the web page link structure diagram, and determining the updated cheating view of each unlabeled sample web page through reverse distrust propagation and combination operations according to the view matrix and the cheating sample web pages; quantifying the updated cheating view of each unlabeled sample web page to obtain the distrust value of each unlabeled sample web page.
[0058] Among them, according to the link relationship between each unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram, the initial cheating view of each unlabeled sample web page is determined, including: for any unlabeled sample web page, according to the link relationship between the unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram, obtain the in-link sample web pages along the in-link direction of the unlabeled sample web page, and the out-link sample web pages along the out-link direction of the unlabeled sample web page; if the total number of web pages of the in-link sample web pages and the out-link sample web pages is 0, then determine that the unlabeled sample web page has an uncertain view; if the total number of web pages of the in-link sample web pages and the out-link sample web pages is not 0, then according to the number of trusted sample web pages, cheating sample web pages and uncertain category sample web pages in the in-link sample web pages and the out-link sample web pages, determine the initial cheating view of the unlabeled sample web page.
[0059] Specifically, construct a web page link structure diagram with views as edge weights, and the link relationship between web pages can be modeled as a directed graph , where the vertices represent web pages, and the edges represent hyperlinks from web page to , represents the weight of edge . The in-neighbor and out-neighbor web pages of a web page are either trusted, or untrusted / cheating, or of an uncertain category. Therefore, the degree to which a web page is a trusted web page or a cheating web page can be determined by the number and trust degree of the in-neighbor and out-neighbor web pages of a web page.
[0060] For the detection of cheating web pages, usually a group of web pages are manually labeled as normal / trusted or cheating web pages, and those that are not labeled cannot have their categories determined. As Figure 2 shows, given a web page , the web pages it points to and the web pages that point to it include trusted web pages, cheating web pages and web pages of an uncertain category. Assume that the number of cheating web pages, trusted web pages and web pages of an uncertain category among the in-and-out neighbors of the unlabeled sample web page is 0, then determine that the unlabeled sample web page has an uncertain view, the unlabeled sample web page has no in-neighbors and out-neighbors, and the category of the unlabeled sample web page cannot be determined.
[0061] If the total number of web pages of the in-link sample web pages and the out-link sample web pages is not 0, then assume that the numbers of cheating web pages, trusted web pages and web pages of an uncertain category among the in-and-out neighbors of web page j are and respectively, and define the initial cheating view of the web page based on the three-valued subjective logic model as , where and respectively represent web pages are the probabilities of a cheating web page and a trusted web page, respectively represent web pages the posterior uncertainty and prior uncertainty probabilities of whether a web page is cheating. These probabilities are calculated as shown in formula (1). Since there are three categories of web pages: trusted, cheating, and uncertain, a constant 3 is set in the prior uncertainty.
[0062] (1).
[0063] Therefore, if the total number of web pages in the in - link sample web pages and out - link sample web pages is not zero, the number of trusted sample web pages, cheating sample web pages, and uncertain - category sample web pages in the in - link sample web pages and out - link sample web pages can be substituted into formula (1) to calculate the initial cheating view of the unlabeled sample web pages.
[0064] Based on the three - valued subjective logic model, the weight of an edge in the web page hyperlink structure diagram is represented as an opinion , where is the opinion of web page on the cheating situation of web page . Given the web page hyperlink structure diagram G and the labeled sample set , the relationship between web pages is represented using the opinion matrix given by formula (2), where represents the number of all sample web pages. N represents the number of all sample web pages.
[0065] (2).
[0066] Therefore, according to the initial cheating view of each unlabeled sample web page, the initial cheating view of the sample web page pointed to by each unlabeled sample web page, and the web page link relationship, an opinion matrix is determined; each element in the opinion matrix represents the cheating view between every two sample web pages. For example, if has no hyperlink, that is , then the opinion of web page on the cheating situation of web page is an uncertain opinion . For , if is a cheating sample web page seed, it is assigned a definite opinion , indicating full belief that web page is a cheating web page; otherwise, it is an uncertain opinion ; where and are cheating sample web pages or unlabeled sample web pages.
[0067] If , then , which is calculated by formula (1). If , then . For If and the web page is a cheating web page, then .
[0068] The opinion matrix records the relationships between mutually linked web pages. This matrix will be used in non-trust propagation and combinatorial calculations to iteratively calculate the untrustworthiness of all web pages.
[0069] In an exemplary embodiment, according to the opinion matrix and cheating sample web pages, the updated cheating opinions of each unlabeled sample web page are determined through reverse non-trust propagation and combinatorial operations, including: for any unlabeled sample web page, obtaining the out-link sample web pages along the out-link direction of the unlabeled sample web page; if the number of out-link sample web pages is 1, obtaining the cheating opinions of the out-link sample web page and the cheating opinions in the opinion matrix, and using the propagation operation of the cheating opinions to determine the updated cheating opinions of the unlabeled sample web page; if the number of out-link sample web pages is multiple, combining the updated cheating opinions of the multiple out-link sample web pages obtained by the propagation operation to obtain the updated cheating opinions of the unlabeled sample web page.
[0070] Specifically, using the cheating sample web pages as seeds, their non-trust opinions are assigned as =(1, 0, 0, 0). Assuming that after the non-trust propagation algorithm iteratively executes k times, the cheating opinions of all web pages are stored in the vector , where represents the cheating opinion of the k -th web page after iterating j times. The non-trust propagation algorithm first uses formula (3) to initialize the opinion of web page j . Then, using propagation and combinatorial operations, it iteratively calculates the cheating opinions of all web pages along the in-links, so that the opinion vector is updated to , ..., .
[0071] (3).
[0072] If the unlabeled sample web page has a link pointing to the cheating sample web page , then the cheating opinion of the cheating sample web page will be passed to the unlabeled sample web page through the edge using the propagation operation given by formula (4) . Use the propagation and combination operations to obtain the cheating view of the unlabeled sample web pages at the th iteration. .
[0073] (4).
[0074] For in formula (4), if , then , otherwise, .
[0075] Among them, if the number of out-link sample web pages pointed to by the unlabeled sample web page link to the cheating sample web page is 1, the cheating view after the update of the unlabeled sample can be calculated according to formula (4).
[0076] If the number of out-link sample web pages pointed to by the unlabeled sample web page link to the cheating sample web page is multiple, the cheating views after the update of the multiple out-link sample web pages by the propagation operation are combined to obtain the cheating view after the update of the unlabeled sample web page.
[0077] For example, if the unlabeled sample web page points to two cheating sample web pages, such as and , then and respectively use the propagation operation to pass a view to the web page , denoted as and , and the reverse non-trust propagation combines the two views using the combination operation given by formula (5) .
[0078] (5).
[0079] If the unlabeled sample web page points to cheating sample web pages, respectively marked as , the combination operation can be performed on the views propagated from the web pages to the web page , as shown in formula (6).
[0080] (6).
[0081] Similarly, the cheating views of all other web pages will be updated to obtain , as shown in formula (7).
[0082] (7);
[0083] Among them, and are the combination and propagation operations respectively, represents the opinion vectors of all web pages after iterations, represents the opinion of web page after iterations, represents the opinion vectors of all web pages after iterations,
[0084] For the opinion , it contains four values and cannot be directly used for web page ranking. It needs to be converted into a single distrust value. Formula (8) is used to calculate the probability that web page is a cheating web page, where x and y are the coefficients of posterior uncertainty and prior uncertainty, respectively representing how much of the posterior uncertainty and prior uncertainty probabilities indicate that the web page is untrustworthy.
[0085] (8);
[0086] Among them, represents the probability that web page
[0087] In this embodiment, the reverse distrust propagation starts from multiple cheating seeds with opinion vectors of (1, 0, 0, 0), and iteratively updates the cheating opinions of other web pages through the propagation and combination operations of the cheating opinion matrix and the opinion vectors.
[0088] In an exemplary embodiment, the propagation model is a propagation model based on double - end penalty. Positive trust propagation and reverse distrust propagation are performed through a labeled sample web page set and a web page link structure diagram to obtain the trust value and distrust value of unlabeled sample web pages in the unlabeled sample web page set, including: for any unlabeled sample web page, obtaining the in - link sample web pages along the in - link direction of the unlabeled sample web page and the out - link sample web pages along the out - link direction of the unlabeled sample web page from the labeled sample web page set and the web page link structure diagram; performing source - end penalty on the unlabeled sample web page through the trust value and distrust value of the out - link sample web pages to obtain the first credible value of the unlabeled sample web page; performing source - end penalty on the unlabeled sample web page through the trust value and distrust value of the in - link sample web pages to obtain the first non - credible value of the unlabeled sample web page; performing target - end penalty on the unlabeled sample web page according to the credible value of the out - link sample web pages, and the first credible value and the first non - credible value of the unlabeled sample to obtain the trust value of the unlabeled sample; performing target - end penalty on the unlabeled sample web page according to the distrust value of the in - link sample web pages, and the first credible value and the first non - credible value of the unlabeled sample to obtain the distrust value of the unlabeled sample.
[0089] Positive trust propagation and reverse distrust propagation can be performed on unlabeled sample web pages through the double - end penalty algorithm to obtain the trust value and distrust value of unlabeled sample web pages. Among them, the double - end penalty algorithm includes: when the link relationship between web pages is q pointing to p ... When calculating the trust value of p ..., the double - end penalty algorithm will perform two calculations. The first calculation is to punish the trust value according to the distrust value of the q node in the previous iteration to obtain a trust value, and then use the calculated trust value as the p current trust value, and use the p current distrust value to punish the trust value to obtain the final trust value. The calculation of the distrust value is similar. Iteratively calculate and traverse all nodes in the network diagram until convergence.
[0090] This algorithm will select a trustworthy seed set and a cheating web page seed set, and give all web pages a d- t 1 and d-t 2 as the credible score, a d-d 1 and d-d 2 as the non - credible score. In each iterative calculation process, first calculate the d-t 1 score and the d-d 1 score respectively according to the principle of source - end penalty, and then use the obtained d-t 1 score as the trust value, d-d 1 score as the non - trust value, and perform d-t2 and d-d The calculation of the 2 - score is achieved by iterative calculation until the algorithm converges. Through the process of trust and distrust propagation and combination, the operation of imposing penalties is simultaneously carried out on the source - end web page and the target - end web page. Finally, the trust values and distrust values of all web pages are determined, the ranking of trustworthy web pages is improved, and the cheating web pages are downgraded.
[0091] As shown in formulas (9) - (12). The source - end penalty for the unlabeled sample web page is performed using the trust value and distrust value of the out - link sample web page. The first trust value of the unlabeled sample web page can be calculated by formula (9). The source - end penalty for the unlabeled sample web page is performed using the trust value and distrust value of the in - link sample web page. The first distrust value of the unlabeled sample web page is calculated by formula (10). According to the trust value of the out - link sample web page, and the first trust value and the first distrust value of the unlabeled sample, the target - end penalty for the unlabeled sample web page is carried out, and the trust value of the unlabeled sample can be calculated by formula (11); According to the distrust value of the in - link sample web page, and the first trust value and the first distrust value of the unlabeled sample, the target - end penalty for the unlabeled sample web page is carried out, and the distrust value of the unlabeled sample can be calculated by formula (12).
[0092] (9);
[0093] (10);
[0094] (11);
[0095] (12);
[0096] Where, represents the first trust value of and respectively represent the trust value of and represents the first distrust value of and respectively represent the distrust value of and and are attenuation factors, is a trust vector, is a distrust vector; In formulas (9) and (11), the web - page link relationship between and is -> , represents the web page The out-degree. In formulas (10) and (12), and The web page link relationship with -> , Indicates the in-degree of web page .
[0097] S102. Integrate and classify unlabeled sample web pages through an integrated classification model to determine the integrated classification results of the unlabeled sample web pages.
[0098] Among them, the integrated classification model is trained based on a set of labeled sample web pages. The integrated classification model can include multiple base classifiers. The multiple base classifiers can be constructed by the same or different machine learning algorithms and / or different features. That is, one feature can be used to construct multiple base classifiers by multiple different machine learning algorithms. For example, if there are 2 types of features and 3 machine learning algorithms, then the number of base classifiers in the integrated classification model can be 6.
[0099] The construction process of the integrated classification model can include: obtaining various features of each sample web page in the set of labeled sample web pages; and for any one feature, constructing multiple different base classifiers according to the features of each sample web page in the set of labeled sample web pages and multiple different classification algorithms; determining the integrated classification model according to the multiple different base classifiers constructed for each feature. The features of the sample web pages in the set of labeled sample web pages can include content features and link features. The classification algorithms can include: Decision Tree C4.5, K-Nearest Neighbor, Logistic Regression, Backpropagation Neural Network, etc.
[0100] After obtaining the integrated classification model, the unlabeled sample web pages can be input into the integrated classification model, and the integrated classification results of the unlabeled sample web pages are output through the integrated classification model. The integrated classification results are cheating web pages or trustworthy web pages.
[0101] It should be noted that in the integrated classification model, after obtaining the classification results of multiple base classifiers for the unlabeled sample web pages, the classification results of the multiple base classifiers for the unlabeled sample web pages can be input into a pre-constructed integration function to determine the integration function value of the unlabeled sample web pages, and the integrated classification results of the unlabeled web pages are determined according to the integration function value.
[0102] For example, if the integration function value is greater than 0, it is determined that the integrated classification result is a trustworthy web page; if the integration function value is less than or equal to 0, it is determined that the integrated classification result is a cheating web page.
[0103] S103. Determine the web pages with unlabeled samples that have trust values in the top first preset ranking and whose integrated classification results are trusted web pages as trusted web pages, and determine the web pages with unlabeled samples that do not have trust values in the top second preset ranking and whose integrated classification results are untrusted web pages as cheating web pages.
[0104] For example, if the first preset ranking is g1 and the second preset ranking is g2, determine the web pages with unlabeled samples that have the top g1 trust values with the highest positive trust propagation and whose integrated classification results are trusted web pages as trusted web pages, and determine the web pages with unlabeled samples that have the top g2 non-trust values with the highest negative non-trust propagation and whose integrated classification results are cheating web pages as cheating web pages.
[0105] S104. Add the trusted web pages and cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively update the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web page to be detected.
[0106] Add the trusted web pages and cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively update the integrated classification model and the propagation model through the labeled sample web page set to obtain the updated integrated classification model and propagation model. The updated integrated classification model can be used to detect cheating web pages for the web page to be detected and determine the cheating web page detection result of the web page to be detected.
[0107] Optionally, the labeled sample web page set obtained by adding the trusted web pages and cheating web pages in the unlabeled sample web page can be passed through the propagation model for positive trust propagation and negative non-trust propagation again. In this way, more labeled samples can be provided for the propagation model.
[0108] In an exemplary embodiment, based on the need for theoretical and application research on trust / non-trust propagation models and integrated classification techniques rooted in three-valued subjective logic or two-sided penalty in web page cheating detection and taking this as the starting point, with the basic goal of effectively fusing the results of the propagation model and integrated classification and using a small number of labeled samples to improve cheating web page detection, systematically study the core process and corresponding key issues of the propagation model based on three-valued subjective logic or two-sided penalty to obtain more trusted and cheating web pages using a small number of labeled trusted and cheating web page seeds and fuse its results with the results of integrated classification. As Figure 3 shown, Figure 3 is a schematic flowchart of a method for detecting cheating web pages by fusing a propagation model and an integrated classification model provided by the present invention. According to Figure 3The internal connection between the main research contents shown in the figure, this method mainly implements various research works from the following aspects: 1) According to the link structure diagram between websites, a website trust / cheating evaluation model is established, and combined with labeled samples, a trust and cheating seed set is provided for the trust and distrust propagation model based on three-valued subjective logic or double-end punishment; 2) Based on the propagation model of three-valued subjective logic or double-end punishment, according to the link structure diagram between websites and the labeled sample set, a website link structure diagram with opinions as edge weights is constructed to provide basic data for trust and distrust propagation; 3) Based on the trust / distrust transmission and combination operations in the propagation model, propagation starts from the trust / cheating seeds to obtain more trust / cheating websites; 4) According to the existing labeled sample set, feature selection and integration are performed to obtain a new training set, and it is used for different classification algorithms to obtain multiple different classifiers, providing an integrated classification model for the classification of unlabeled sample sets; 5) The classification results of multiple classifiers on the unlabeled sample set and the results obtained by the trust / distrust model are integrated to provide more labeled samples for the training sample set.
[0109] Specifically, it includes: firstly, determining the seed sample web page set through the website link structure diagram and the PageRank algorithm and the inverse PageRank algorithm. The seed sample web page set includes trusted seed sample web pages and untrusted seed sample web pages. The trusted sample web pages in the labeled sample set and the sample web pages ranked high by the trust propagation algorithm are used as trust seeds of the forward trust propagation model. The cheating sample web pages in the labeled sample set and the sample web pages ranked high by the untrust propagation algorithm are used as cheating seeds of the reverse untrusted propagation model.
[0110] The labeled sample set is subjected to feature selection and generation to obtain the features of the labeled sample set. The features of the labeled sample set are used as training sets. Different classification algorithms are trained through the training sets to obtain multiple different classifiers. Figure 3 In the example, c classifiers are used.
[0111] The unlabeled web page sample set is integratedly classified by multiple classifiers to obtain the integrated classification result of the unlabeled web page sample set, and the propagation model is used to perform forward trust propagation on each unlabeled web page sample in the unlabeled web page sample set with trust seeds to obtain the trust value of each unlabeled web page sample in the unlabeled sample set, and the cheating seed is used to perform reverse non-trust propagation on each unlabeled web page sample in the unlabeled sample set to obtain the non-trust value of each unlabeled web page sample in the unlabeled sample set.
[0112] Before selecting positive trust propagation x and the integrated classification results are credible web pages, and the reverse untrusted propagation yIntegrate web pages with the classification result of being cheated as the labeled sample set, retrain each classifier, and use it as the propagation seeds of the propagation model to perform forward trust propagation and reverse distrust propagation on the unlabeled web page sample set.
[0113] In the cheating web page detection method that integrates the propagation model and the integrated classification model of the present invention, based on the three-valued subjective logic model, a representation and propagation model of trust / distrust views reflecting the link relationship between web pages is proposed, which can more accurately reflect the trust / distrust relationship between web pages and effectively propagate it outward along the out-links / in-links. Based on the results of the common seed selection algorithms PageRank and Inverse PageRank, a seed selection algorithm that is more conducive to the representation and stronger propagation ability of trust / distrust views is proposed, and credible websites with a larger out-degree and cheating websites with a larger in-degree are selected as the seeds of the forward trust and reverse distrust propagation algorithms respectively, so as to better spread the trust view and the distrust view and obtain more credible websites and cheating websites.
[0114] Furthermore, a semi-supervised learning model that integrates the integrated classification technology and the trust / distrust propagation model is proposed. The credible web pages obtained from the forward trust propagation and the cheating web pages obtained from the reverse distrust propagation are integrated with the classification results of the integrated classification of unlabeled web pages, further improving the prediction accuracy of unlabeled samples and reducing the accumulation of errors caused by adding mispredicted samples to the labeled samples.
[0115] Specifically, the seeds selected by the high PageRank seed selection algorithm have a higher in-degree and a lower out-degree, which is not conducive to the representation of the trust view and the forward propagation of the trust view along the out-links; the seeds selected by the inverse PageRank seed selection algorithm have a higher out-degree and a lower in-degree, which is not conducive to the representation of the distrust view and the reverse propagation of the distrust view along the in-links. Therefore, the present invention comprehensively utilizes the results of the high PageRank and inverse PageRank algorithms to propose a better seed selection algorithm, so as to obtain seeds that are more conducive to the representation and stronger propagation ability of trust / distrust views.
[0116] The algorithm for detecting credible websites based on the three-valued subjective logic only uses the information of out-neighbors to represent the trust view and propagates it forward along the out-links to find more credible websites from the credible website seeds. Consider the effect of comprehensively considering the information of in-out neighbors to represent the trust / distrust view, and how to start from the cheating website seeds, along the in-links, and use distrust transfer and combination operations to perform reverse distrust propagation to obtain more cheating web pages.
[0117] The present invention effectively integrates the trust / non-trust propagation model based on three-valued subjective logic or double-end penalty with the results of integrated classification, aiming to improve the detection of cheating websites using a small number of labeled samples. It studies key issues and algorithms, and compiles a series of advanced technologies and algorithms into program modules to form a relatively systematic prototype system for improving website cheating detection by integrating the propagation model and integrated classification technology. Technically, the following three main objectives are achieved:
[0118] Integrate the results of common seed selection algorithms, the high PageRank algorithm and the inverse PageRank algorithm, and find credible websites with larger out-degrees and cheating websites with larger in-degrees as seeds for the forward trust and reverse non-trust propagation algorithms respectively, enhancing the propagation ability of the propagation model.
[0119] Based on the propagation model of three-valued subjective logic or double-end penalty, use the trust / non-trust view to describe the link relationship between websites, and apply the propagation and combination operations of the trust / non-trust view to better spread the trust view and non-trust view along the out-links and in-links respectively, obtaining more trusted and cheating websites.
[0120] Integrate the classification results of the integrated classification technology for the unlabeled sample set and the propagation results of the trust / non-trust propagation model, determine the categories of some web pages in the unlabeled sample set, and add them to the labeled data set to provide more labeled samples for the training of the integrated classification model, the view calculation of the propagation model, and the trust / non-trust propagation. It realizes obtaining a weak classifier using a small number of trusted and cheating seeds to form a labeled training set, and then combining the results of forward trust propagation and reverse non-trust propagation to obtain more trusted and cheating web pages, and adding them to the labeled sample set, further enhancing the classification performance of the classifier and the propagation ability of the propagation algorithm, thereby improving the accuracy of cheating web page detection of the integrated propagation model and integrated classification model.
[0121] When applying the cheating web page detection method of the integrated propagation model and integrated classification model provided by the present invention, it is not necessary to execute according to Figure 1 the order of the steps shown. The specific execution order of each step can be determined as needed, and the present invention does not limit this.
[0122] The above is the cheating web page detection method of the integrated propagation model and integrated classification model provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding cheating web page detection device for the integrated propagation model and integrated classification model, and the device includes:
[0123] A propagation module is configured to input the labeled sample web page set, the unlabeled sample web page set, and the web page link structure diagram into a propagation model based on ternary subjective logic or double-end penalty. Through positive trust propagation and reverse distrust propagation using the labeled sample web page set and the web page link structure diagram, the trust values and distrust values of the unlabeled sample web pages in the unlabeled sample web page set are obtained. The double-end penalty includes performing penalty operations on the trust values and distrust values at the source end and the target end of the propagation while propagating the trust values forward along the out-links and the distrust values backward along the in-links.
[0124] A classification module is configured to perform integrated classification on the unlabeled sample web pages through an integrated classification model to determine the integrated classification results of the unlabeled sample web pages. The integrated classification model is trained based on the labeled sample web page set.
[0125] A determination module is configured to determine the unlabeled sample web pages with trust values in the top first preset ranking and an integrated classification result of trustworthy web pages as trustworthy web pages, and determine the unlabeled sample web pages with distrust values in the top second preset ranking and an integrated classification result of untrustworthy web pages as cheating web pages.
[0126] An update module adds the trustworthy web pages and cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively updates the integrated classification model and the propagation model through the labeled sample web page set. The updated integrated classification model is used to detect cheating web pages in the web pages to be detected.
[0127] For the specific limitations of the cheating web page detection device integrating the propagation model and the integrated classification model, reference can be made to the limitations of the cheating web page detection method for the integration of the propagation model and the integrated classification model in the above text, which will not be elaborated here. Each module in the above cheating web page detection device integrating the propagation model and the integrated classification model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0128] The present invention also provides a computer-readable storage medium storing a computer program, which can be used to execute the above Figure 1 provided cheating web page detection method integrating the propagation model and the integrated classification model.
[0129] The present invention also provides Figure 4 a schematic structural diagram of the computer device as shown in Figure 4As shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, memory, and non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 Cheating web page detection method for the fusion propagation model and the integrated classification model provided.
[0130] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. A cheating web page detection method that combines a propagation model and an integrated classification model, characterized in that The method includes: Inputting the labeled sample web page set, the unlabeled sample web page set, and the web page link structure diagram into the propagation model based on three-valued subjective logic, and performing forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust value and the distrust value of the unlabeled sample web pages in the unlabeled sample web page set; Performing integrated classification on the unlabeled sample web pages through an integrated classification model to determine the integrated classification result of the unlabeled sample web pages; the integrated classification model is trained according to the labeled sample web page set; Determining the trustworthy web pages as the unlabeled sample web pages with the trust value in the top first preset ranking and the integrated classification result being trustworthy web pages, and determining the cheating web pages as the unlabeled sample web pages with the distrust value in the top second preset ranking and the integrated classification result being untrustworthy web pages; Adding the trustworthy web pages and the cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively updating the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web pages to be detected; Performing reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the distrust value of the unlabeled sample web pages in the unlabeled sample web page set, including: Determining the initial cheating view of each unlabeled sample web page according to the link relationship between each unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram; Determining the view matrix according to the initial cheating view of each unlabeled sample web page, the initial cheating view of the sample web page pointed to by each unlabeled sample web page, and the web page link relationship; Obtaining the cheating sample web pages in the labeled sample web page set and the web page link structure diagram, and determining the updated cheating view of each unlabeled sample web page through reverse distrust propagation and combination operations according to the view matrix and the cheating sample web pages; Quantifying the updated cheating view of each unlabeled sample web page to obtain the distrust value of each unlabeled sample web page.
2. The method according to claim 1, wherein The determining the initial cheating view of each unlabeled sample web page according to the link relationship between each unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram includes: For any unlabeled sample web page, obtaining the in-link sample web pages along the in-link direction of the unlabeled sample web page and the out-link sample web pages along the out-link direction of the unlabeled sample web page according to the link relationship between the unlabeled sample web page and its in-and-out neighbor web pages in the web page link structure diagram; If the total number of the in-link sample web pages and the out-link sample web pages is 0, determining that the unlabeled sample web page has an uncertain view; If the total number of the in-link sample web pages and the out-link sample web pages is not 0, determining the initial cheating view of the unlabeled sample web page according to the number of trustworthy sample web pages, cheating sample web pages, and uncertain category sample web pages in the in-link sample web pages and the out-link sample web pages.
3. The method according to claim 1, wherein Based on the view matrix and the cheating sample web pages, determine the updated cheating views of each unlabeled sample web page through reverse distrust propagation and combination operations, including: For any unlabeled sample web page, obtain the out-link sample web pages along the out-link direction of the unlabeled sample web page; If the number of the out-link sample web pages is 1, obtain the cheating views of the out-link sample web page and the cheating views in the view matrix, and use the propagation operation of the cheating views to determine the updated cheating views of the unlabeled sample web page; If the number of the out-link sample web pages is multiple, combine the updated cheating views of the multiple out-link sample web pages obtained by the propagation operation to obtain the updated cheating views of the unlabeled sample web page.
4. A cheating web page detection method integrating a propagation model and an integrated classification model, characterized in that, The method includes: Input the labeled sample web page set, the unlabeled sample web page set, and the web page link structure diagram into the propagation model based on double-end penalty, and perform forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust values and distrust values of the unlabeled sample web pages in the unlabeled sample web page set; the double-end penalty includes performing penalty operations on the trust values and distrust values at the source end and the target end while propagating the trust values forward along the out-links and propagating the distrust values backward along the in-links; Perform integrated classification on the unlabeled sample web pages through an integrated classification model to determine the integrated classification results of the unlabeled sample web pages; the integrated classification model is trained according to the labeled sample web page set; Determine the trustworthy web pages as the unlabeled sample web pages with the trust values in the top first preset ranking and the integrated classification results being trustworthy web pages, and determine the cheating web pages as the unlabeled sample web pages with the distrust values in the top second preset ranking and the integrated classification results being untrustworthy web pages; Add the trustworthy web pages and the cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and iteratively update the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web pages to be detected; Perform forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust values and distrust values of the unlabeled sample web pages in the unlabeled sample web page set, including: For any unlabeled sample web page, obtain the in-link sample web pages along the in-link direction of the unlabeled sample web page, and the out-link sample web pages along the out-link direction of the unlabeled sample web page from the labeled sample web page set and the web page link structure diagram; Perform source-end penalty on the unlabeled sample web page through the trust values and distrust values of the out-link sample web pages to obtain the first trust value of the unlabeled sample web page; Perform source-end penalty on the unlabeled sample web page through the trust values and distrust values of the in-link sample web pages to obtain the first distrust value of the unlabeled sample web page; Perform target-end penalty on the unlabeled sample web page according to the trust values of the out-link sample web pages, and the first trust value and the first distrust value of the unlabeled sample to obtain the trust value of the unlabeled sample. Based on the untrusted value of the incoming link sample web page, as well as the first trusted value and the first untrusted value of the unlabeled sample, perform target - end punishment on the unlabeled sample web page to obtain the untrusted value of the unlabeled sample.
5. A cheating web page detection device integrating a propagation model and an integrated classification model, characterized in that, Including: A propagation module, configured to input a labeled sample web page set, an unlabeled sample web page set, and a web page link structure diagram into a propagation model based on three - valued subjective logic. Through positive trust propagation and reverse distrust propagation using the labeled sample web page set and the web page link structure diagram, obtain the trust value and untrusted value of the unlabeled sample web pages in the unlabeled sample web page set; through reverse distrust propagation using the labeled sample web page set and the web page link structure diagram, obtain the untrusted value of the unlabeled sample web pages in the unlabeled sample web page set, including: determining the initial cheating view of each unlabeled sample web page according to the link relationship between each unlabeled sample web page and its incoming and outgoing neighbor web pages in the web page link structure diagram; determining a view matrix according to the initial cheating view of each unlabeled sample web page, the initial cheating view of the sample web page pointed to by each unlabeled sample web page, and the web page link relationship; obtaining the cheating sample web pages in the labeled sample web page set and the web page link structure diagram, and determining the updated cheating view of each unlabeled sample web page through reverse distrust propagation and combination operations according to the view matrix and the cheating sample web pages; quantifying the updated cheating view of each unlabeled sample web page to obtain the untrusted value of each unlabeled sample web page; A classification module, configured to perform integrated classification on the unlabeled sample web page through an integrated classification model to determine the integrated classification result of the unlabeled sample web page; the integrated classification model is trained according to the labeled sample web page set; A determination module, configured to determine the unlabeled sample web pages with trust values in the top first preset ranking and an integrated classification result of trusted web pages as trusted web pages, and determine the unlabeled sample web pages with untrusted values in the top second preset ranking and an integrated classification result of untrusted web pages as cheating web pages; An update module, configured to add the trusted web pages and cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and perform iterative updates on the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web page to be detected.
6. A cheating web page detection device integrating a propagation model and an integrated classification model, characterized in that, Including: A propagation module, configured to input the labeled sample web page set, the unlabeled sample web page set, and the web page link structure diagram into a propagation model based on double-end penalty, and perform forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust value and the distrust value of the unlabeled sample web pages in the unlabeled sample web page set; the double-end penalty includes performing penalty operations on the trust value and the distrust value at the source end and the target end of the propagation while propagating the trust value forward along the out-links and propagating the distrust value backward along the in-links; performing forward trust propagation and reverse distrust propagation through the labeled sample web page set and the web page link structure diagram to obtain the trust value and the distrust value of the unlabeled sample web pages in the unlabeled sample web page set, including: for any unlabeled sample web page, obtaining the in-link sample web pages along the in-link direction of the unlabeled sample web page and the out-link sample web pages along the out-link direction of the unlabeled sample web page from the labeled sample web page set and the web page link structure diagram; performing source-end penalty on the unlabeled sample web page through the trust value and the distrust value of the out-link sample web pages to obtain the first credible value of the unlabeled sample web page; performing source-end penalty on the unlabeled sample web page through the trust value and the distrust value of the in-link sample web pages to obtain the first non-credible value of the unlabeled sample web page; performing target-end penalty on the unlabeled sample web page according to the credible value of the out-link sample web pages, and the first credible value and the first non-credible value of the unlabeled sample to obtain the trust value of the unlabeled sample; performing target-end penalty on the unlabeled sample web page according to the distrust value of the in-link sample web pages, and the first credible value and the first non-credible value of the unlabeled sample to obtain the distrust value of the unlabeled sample; A classification module, configured to perform integrated classification on the unlabeled sample web pages through an integrated classification model to determine the integrated classification result of the unlabeled sample web pages; the integrated classification model is trained according to the labeled sample web page set; A determination module, configured to determine the unlabeled sample web pages with the trust value in the top first preset ranking and the integrated classification result being credible web pages as credible web pages, and determine the unlabeled sample web pages with the distrust value in the top second preset ranking and the integrated classification result being non-credible web pages as cheating web pages; An update module, configured to add the credible web pages and the cheating web pages in the unlabeled sample web page set to the labeled sample web page set, and perform iterative updates on the integrated classification model and the propagation model through the labeled sample web page set; the updated integrated classification model is used to detect cheating web pages for the web pages to be detected.
Citation Information
Patent Citations
Method for detecting search engine cheat based on small sample set
CN101350011A
Method and device for identifying cheat web-pages
CN103150369A