Method, apparatus, electronic device, and storage medium for identifying web page validity

By obtaining and analyzing the page monitoring indicators and site quality monitoring indicators of web pages, combined with deep learning models, the problem of difficulty in identifying the effectiveness of web pages in the existing technology is solved, and the accurate identification of the effectiveness of web pages is achieved, and the user's search experience is improved.

CN114036415BActive Publication Date: 2025-06-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111233683.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-06-10
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately identify the effectiveness of web pages, resulting in cheating web pages, dead-chain web pages and low-quality web pages in the web pages found in search engines, affecting the user's search experience.

Method used

By obtaining the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which they belong, and combining the deep learning model to analyze and identify these indicators to determine the effectiveness of the web page.

Benefits of technology

It realizes accurate identification of web page effectiveness, improves users' search experience in search engines, and reduces the impact of invalid web pages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036415B_ABST
    Figure CN114036415B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for identifying the validity of a web page, relating to the field of artificial intelligence technologies such as deep learning. The specific implementation solution is as follows: obtaining a target web page to be identified; obtaining page monitoring metrics of the target web page and quality monitoring metrics of the site to which the target web page belongs; and identifying the target web page based on the page monitoring metrics of the target web page and the quality monitoring metrics of the site to determine the validity of the target web page. Thus, accurate identification of the validity of the target web page is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, particularly to artificial intelligence technologies such as deep learning, and more particularly to a method, apparatus, electronic device, and storage medium for identifying the validity of web pages. Background Art

[0002] Currently, search engines are widely used. Hundreds of millions of users search for web pages that meet their needs through search engines every day. Problematic invalid web pages such as cheating web pages, dead link web pages, and low-quality web pages directly affect the user's search experience. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for identifying the validity of web pages.

[0004] According to one aspect of the present disclosure, a method for identifying the validity of a web page is provided, including: obtaining a target web page to be identified; obtaining page monitoring metrics of the target web page and quality monitoring metrics of the site to which the target web page belongs; and identifying the target web page based on the page monitoring metrics of the target web page and the quality monitoring metrics of the site to determine the validity of the target web page.

[0005] According to another aspect of the present disclosure, an apparatus for identifying the validity of a web page is provided, including: a first obtaining module for obtaining a target web page to be identified; a second obtaining module for obtaining page monitoring metrics of the target web page and quality monitoring metrics of the site to which the target web page belongs; and an identifying module for identifying the target web page based on the page monitoring metrics of the target web page and the quality monitoring metrics of the site to determine the validity of the target web page.

[0006] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for identifying the validity of a web page as described above.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to cause a computer to execute the method for identifying the validity of a web page as described above.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method for identifying the validity of a web page as described above are implemented.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic flowchart of a method for identifying the validity of a web page according to the first embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of a method for identifying the validity of a web page according to the second embodiment of the present disclosure;

[0013] Figure 3 is a schematic flowchart of a method for identifying the validity of a web page according to the third embodiment of the present disclosure;

[0014] Figure 4 is a schematic flowchart of a method for identifying the validity of a web page according to the fourth embodiment of the present disclosure;

[0015] Figure 5 is another schematic flowchart of a method for identifying the validity of a web page according to the fourth embodiment of the present disclosure;

[0016] Figure 6 is a schematic structural diagram of a device for identifying the validity of a web page according to the fifth embodiment of the present disclosure;

[0017] Figure 7 is a schematic structural diagram of a device for identifying the validity of a web page according to the sixth embodiment of the present disclosure;

[0018] Figure 8 is a block diagram of an electronic device for implementing the method for identifying the validity of a web page according to the embodiments of the present disclosure. Detailed Embodiments

[0019] The following describes exemplary embodiments of the present disclosure in conjunction with the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] The present disclosure relates to the field of computer technology, and particularly to the field of artificial intelligence technologies such as deep learning.

[0021] The following briefly describes the technical fields related to the solution of the present disclosure:

[0022] AI (Artificial Intelligence) is a discipline that studies how to make computers simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It involves technologies at both the hardware and software levels. AI hardware technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing; AI software technologies mainly include several major directions such as computer vision, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0023] DL (Deep Learning) is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed previous related technologies. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as seeing, hearing, and thinking, solves many complex pattern recognition problems, and has made great progress in related AI technologies.

[0024] Currently, the application of search engines is very extensive. Hundreds of millions of users use search engines every day to find web pages that meet their needs. Problematic invalid web pages such as cheating web pages, dead link web pages, and low-quality web pages directly affect the user's search experience. Therefore, in order to improve the user's search experience, a method that can accurately identify the validity of web pages is needed. Among them, identifying the validity of a web page can be understood as identifying whether a web page is a valid web page or an invalid web page. Among them, invalid web pages can include problematic web pages such as low-quality web pages (empty short web pages, expired web pages, invalid web pages, etc.), cheating web pages, and dead link web pages (content dead links or protocol dead links); valid web pages are web pages without the above problems. Among them, an empty short web page refers to a web page with no substantial content.

[0025] This disclosure proposes a method for identifying the validity of a web page. After obtaining the target web page to be identified, obtain the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs, and then based on the page monitoring indicators of the target web page and the quality monitoring indicators of the site, identify the target web page to determine the validity of the target web page. Thus, by combining the page monitoring indicators of the target web page itself and the quality monitoring indicators of the site to which the target web page belongs, comprehensively identifying the target web page, the accurate identification of the validity of the target web page is achieved.

[0026] The following describes a method, apparatus, electronic device, non-transitory computer-readable storage medium, and computer program product for identifying the validity of a web page according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0027] First, in combination with Figure 1 , a method for identifying the validity of a web page provided by the present disclosure will be described in detail.

[0028] Figure 1 FIG. is a flowchart of a method for identifying the validity of a web page according to a first embodiment of the present disclosure. It should be noted that, for the method for identifying the validity of a web page provided in this embodiment, the execution subject is a device for identifying the validity of a web page, hereinafter referred to as the identification device. This identification device can be an electronic device or can be configured in an electronic device to accurately identify the validity of a web page. In the embodiments of the present disclosure, the case where the identification device is configured in an electronic device is taken as an example for illustration.

[0029] Among them, the electronic device can be any stationary or mobile computing device capable of data processing, such as mobile computing devices such as laptop computers, smartphones, and wearable devices, or stationary computing devices such as desktop computers, or servers, or other types of computing devices, etc. The present disclosure does not limit this.

[0030] As Figure 1 shown, the method for identifying the validity of a web page may include the following steps:

[0031] Step 101, obtain a target web page to be identified.

[0032] Among them, the target web page can be any type of web page, such as a static web page or a dynamic web page, etc. The present disclosure does not limit this.

[0033] Step 102, obtain the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs.

[0034] Among them, the page monitoring indicators of the target web page refer to the indicators used to evaluate the quality of the target web page itself. For example, the page monitoring indicators of the target web page may include at least one of the indicators such as the relevance between the text and the title in the target web page, the normality of audio and video, the authority of the target web page, the authority sensitivity of the text in the target web page, the content richness of the text in the target web page, and the sentence fluency.

[0035] Among them, the relevance between the text and the title in the target web page indicates the level of relevance between the text and the title in the target web page. When the relevance between the text and the title in the target web page is relatively high, it indicates that the theme of the target web page is clear and legal.

[0036] The audio - video normality indicates the normal degree of the audio or video in the target web page. When the audio - video normality of the target web page is high, it means that the audio or video in the target web page plays normally, and the sound quality or picture quality is clear.

[0037] The authority of the target web page indicates the level of its authority. For example, when the target web page is a web page published by an official website, the target web page usually has authority. Correspondingly, the authority of the target web page is high.

[0038] The authority sensitivity of the main text in the target web page indicates the level of the authority sensitivity of the main text in the target web page. Among them, the authority sensitivity level, that is, the degree to which the main text in the target web page depends on authoritative official sources. When the authority sensitivity of the main text in the target web page is higher, the degree to which the main text in the target web page depends on authoritative official sources is higher.

[0039] For example, assume that an official daily newspaper publishes a web page, and the main text content in the web page is a description of today's weather. Since this web page is published by an official daily newspaper, the authority of this web page is high. However, the description of today's weather does not depend on this official daily newspaper, and any website can publish a description of the weather. Therefore, the official sensitivity of the main text in this web page is low.

[0040] The content richness and sentence fluency of the main text in the target web page indicate the content richness degree and sentence fluency degree of the main text in the target web page. Among them, when the content richness and sentence fluency of the main text in the target web page are high, it means that the main text in the target web page has the characteristics of high content quality, rich content, positive energy, strong context relevance, and smooth sentences.

[0041] Among them, the quality monitoring indicators of the site to which the target web page belongs refer to the indicators used to evaluate the quality of the site to which the target web page belongs. For example, the quality monitoring indicators of the site to which the target web page belongs can include at least one of the following indicators: the proportion of cheating web pages of the site to which the target web page belongs, the rendering failure proportion corresponding to at least one sub - link, the cheating probability, the proportion of newly added web pages, etc.

[0042] Among them, the proportion of cheating web pages is the proportion of the number of cheating web pages in all web pages of the site. Among them, a cheating web page can be understood as a web page that, according to the characteristics of the search engine, adopts targeted deception means instead of improving its own content quality, so that it obtains unfair query relevance and value importance, thereby improving its ranking in the search engine. The proportion of cheating web pages can indicate the level of the cheating possibility of the site to which the target web page belongs. When the proportion of cheating web pages is higher, the cheating possibility of this site is higher.

[0043] Sub-chains can include js (JavaScript scripting language), css (Cascading Style Sheets), images, etc. The rendering failure ratio corresponding to a sub-chain is the ratio of the number of web pages in which the sub-chain fails to render to the total number of web pages in each web page of the site. The rendering failure ratio can indicate the quality of the web pages published by the site to which the target web page belongs. When the rendering failure ratio is higher, the quality of the web pages published by the site is lower.

[0044] The cheating probability refers to the probability that the site to which the target web page belongs is a cheating site. The cheating probability can indicate the level of cheating possibility of the site to which the target web page belongs. When the cheating probability is higher, the cheating possibility of the site is higher.

[0045] The new web page ratio refers to the ratio of the increase in the number of web pages of the site to which the target web page belongs during a preset time period compared to the historical number of web pages during the time period before the preset time period, to the historical number of web pages during the time period before the preset time period. Among them, the preset time period and the time period before the preset time period can be set arbitrarily according to needs, and the present disclosure does not limit this.

[0046] For example, assume that the preset time period is within 24 hours on January 2, and the time period before the preset time period is within 24 hours on January 1. The number of web pages of a certain site within 24 hours on January 2 is 100, and the number of web pages within 24 hours on January 1 is 10. That is, the increase of 90 in the number of web pages of the site within 24 hours on January 24 compared to the historical number of web pages within 24 hours on January 1, accounts for the ratio of (100 - 10) / 10, that is, 900% of the historical number of web pages within 24 hours on January 1. Correspondingly, the new web page ratio is 900%.

[0047] The new web page ratio can indicate the level of the possibility of the productivity of the site to which the target web page belongs being stable and legal. For example, when the new web page ratio is significantly different from 1, the possibility of the productivity of the site being stable and legal is relatively low. Among them, productivity can be understood as the production capacity of the data (information) of the site, such as the ability of the site to produce web pages, that is, the number of web pages generated within a preset time period.

[0048] Step 103: Based on the page monitoring indicators of the target web page and the quality monitoring indicators of the site, identify the target web page to determine the validity of the target web page.

[0049] It can be understood that the page monitoring indicators of the target web page can be used to evaluate the quality of the target web page itself, and the quality monitoring indicators of the site to which the target web page belongs can be used to evaluate the quality of the site to which the target web page belongs. To a certain extent, the quality of the site to which the target web page belongs can reflect the effectiveness of the target web page. Then, in the embodiments of the present disclosure, the page monitoring indicators of the target web page and the quality monitoring indicators of the site can be combined to identify the target web page to determine the effectiveness of the target web page. The specific implementation process can refer to the description of the following embodiments and will not be elaborated here. Since the page monitoring indicators of the target web page itself are combined with the quality monitoring indicators of the site to which the target web page belongs to comprehensively identify the target web page, the accuracy of identifying the effectiveness of the target web page can be improved.

[0050] In summary, for the method for identifying the effectiveness of a web page provided by the embodiments of the present disclosure, after obtaining the target web page to be identified, the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs are obtained. Then, based on the page monitoring indicators of the target web page and the quality monitoring indicators of the site, the target web page is identified to determine the effectiveness of the target web page. Thus, by combining the page monitoring indicators of the target web page itself and the quality monitoring indicators of the site to which the target web page belongs to comprehensively identify the target web page, the accurate identification of the effectiveness of the target web page is realized.

[0051] Through the above analysis, it can be seen that in the embodiments of the present disclosure, the target web page can be identified based on the page monitoring indicators of the target web page and the quality monitoring indicators of the site to determine the effectiveness of the target web page. The following will be combined with Figure 2 to further illustrate the acquisition methods of the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs in the method for identifying the effectiveness of a web page provided by the present disclosure.

[0052] Figure 2 is a schematic flowchart of the method for identifying the effectiveness of a web page according to the second embodiment of the present disclosure. As Figure 2 shown, the method for identifying the effectiveness of a web page may include the following steps:

[0053] Step 201, obtain the target web page to be identified.

[0054] Among them, the target web page can be any type of web page, such as a static web page or a dynamic web page, etc. The present disclosure does not limit this.

[0055] Step 202, obtain the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs.

[0056] In an exemplary embodiment, the page monitoring metrics of the target web page include at least one of the relevance between the body text and the title, the normality of audio and video, the authority of the target web page, the authority sensitivity of the body text in the target web page, the content richness of the body text in the target web page, and the fluency of the sentences.

[0057] Next, the relevance between the body text and the title of the target web page and the method for obtaining it will be described.

[0058] The relevance between the body text and the title indicates the degree of relevance between the body text and the title in the target web page. When the relevance between the body text and the title in the target web page is high, it means that the theme of the target web page is clear and legal. The relevance between the body text and the title can specifically be a score indicating the degree of relevance between the body text and the title in the target web page.

[0059] In an exemplary embodiment, the relevance between the body text and the title can be obtained in the following way: input the target web page into the second deep semantic model to obtain the relevance between the body text and the title in the target web page.

[0060] Among them, the second deep semantic model can be a deep semantic model of any type or structure, such as a Transformer model, etc., and the present disclosure does not limit this.

[0061] Among them, the second deep semantic model can be trained by using sample web page data in a deep learning manner. Compared with other machine learning methods, deep learning performs better on large data sets. Here, the sample web page data includes multiple sample web pages, and the multiple sample web pages are labeled with the relevance between the body text and the title therein.

[0062] Specifically when training, the sample web page can be used as the input, input into the initial second deep semantic model, obtain the predicted relevance between the body text and the title of the sample web page, and combine the relevance between the body text and the title of the sample web page marked, to obtain the difference between the predicted relevance between the body text and the title of the sample web page and the marked relevance between the body text and the title of the sample web page, so as to adjust the model parameters of the initial second deep semantic model according to this difference to obtain the adjusted second deep semantic model. By continuously adjusting the model parameters of the initial second deep semantic model with multiple sample web pages to perform iterative training on the initial second deep semantic model until the accuracy of the predicted relevance between the body text and the title output by the second deep semantic model meets the preset relevance threshold, the training ends, and the trained second deep semantic model is obtained.

[0063] After training the second deep semantic model, input the target web page into the trained second deep semantic model, and the relevance between the body text and the title in the target web page can be obtained.

[0064] Through the above process, the relevance between the main body and the title in the target web page is accurately determined, laying a foundation for accurately identifying the effectiveness of the target web page subsequently.

[0065] The audio-video normality indicates the normal degree of the audio or video in the target web page. When the audio-video normality of the target web page is relatively high, it means that the audio or video in the target web page plays normally, and the sound quality or picture quality is clear. For the specific method of obtaining the audio-video normality, relevant technologies can be adopted, which will not be elaborated here.

[0066] The authority of the target web page, the authority sensitivity of the main body in the target web page, the content richness and sentence fluency of the main body in the target web page, and their obtaining methods will be described below.

[0067] The authority of the target web page indicates the level of authority of the target web page. For example, when the target web page is a web page published by an official website, the target web page usually has authority. Correspondingly, the authority of the target web page is relatively high. The authority of the target web page can specifically be a score indicating the level of authority of the target web page.

[0068] The authority sensitivity of the main body in the target web page indicates the level of authority sensitivity of the main body in the target web page. Among them, the authority sensitivity level, that is, the degree to which the main body in the target web page depends on authoritative official sources. When the authority sensitivity of the main body in the target web page is higher, the degree to which the main body in the target web page depends on authoritative official sources is higher. The authority sensitivity of the main body in the target web page can specifically be a score indicating the level of authority sensitivity of the main body in the target web page.

[0069] The content richness and sentence fluency of the main body in the target web page indicate the content richness degree and sentence fluency degree of the main body in the target web page. Among them, when the content richness and sentence fluency of the main body in the target web page are relatively high, it means that the main body in the target web page has characteristics such as high content quality, rich content, positive energy, strong context relevance, and smooth sentences. The content richness and sentence fluency of the main body in the target web page can specifically be a score indicating the content richness degree and sentence fluency degree of the main body in the target web page.

[0070] In an exemplary embodiment, the authority of the target web page, the authority sensitivity of the main body in the target web page, the content richness and sentence fluency of the main body in the target web page can be obtained through the following method: extracting the second attribute information of the main body in the target web page; inputting the second attribute information into a third deep semantic model to obtain the authority of the target web page, the authority sensitivity of the main body in the target web page, and the content richness and sentence fluency of the main body in the target web page.

[0071] Among them, the second attribute information can include any information related to the main body in the target web page, such as the main body, the article length of the main body, the theme, etc.

[0072] The third deep semantic model can be a deep semantic model of any type or structure, such as a Transformer model, etc., and the present disclosure does not limit this.

[0073] It should be noted that the authority of the target web page, the authority sensitivity of the main text in the target web page, and the content richness and sentence fluency of the main text in the target web page can each correspond to a third deep semantic model. That is, the third deep semantic models corresponding to the three dimensions of authority, authority sensitivity, and content richness and sentence fluency can be pre-trained respectively. The third deep semantic model corresponding to the authority dimension is used to obtain the authority of the target web page, the third deep semantic model corresponding to the authority sensitivity dimension is used to obtain the authority sensitivity of the main text in the target web page, and the third deep semantic model corresponding to the content richness and sentence fluency dimension is used to obtain the content richness and sentence fluency of the main text in the target web page; or, the authority of the target web page, the authority sensitivity of the main text in the target web page, and the content richness and sentence fluency of the main text in the target web page can each correspond to one of the branches of the third deep semantic model. That is, a third deep semantic model can be pre-trained. The deep semantic model includes three branches, and each branch is respectively used to obtain the authority of the target web page, the authority sensitivity of the main text in the target web page, and the content richness and sentence fluency of the main text in the target web page. The present disclosure does not limit this.

[0074] Among them, taking the authority of the target web page, the authority sensitivity of the main text in the target web page, and the content richness and sentence fluency of the main text in the target web page, each corresponding to a third deep semantic model, and taking the third deep semantic model corresponding to the authority dimension as an example, for example, the sample data of the main text in the sample web page can be used to train through deep learning to obtain the third deep semantic model corresponding to the authority dimension. Compared with other machine learning methods, deep learning performs better on large data sets. Among them, the sample data of the main text in the sample web page includes the sample attribute information of the main text in multiple sample web pages, and the sample attribute information of the main text in multiple sample web pages is labeled with the sample authority of the sample web page. Among them, the sample attribute information can include any information related to the main text in the sample web page, such as the main text in the sample web page, the article length of the main text, the theme, etc.

[0075] When specifically training, the sample attribute information of the main body in the sample web page can be used as the input, and input into the initial third-depth semantic model to obtain the predicted authority of the sample web page. By combining the labeled authority of the sample web page, the difference between the predicted authority of the sample web page and the labeled authority of the sample web page can be obtained, so as to adjust the model parameters of the initial third-depth semantic model according to this difference, and obtain the adjusted third-depth semantic model. By continuously adjusting the model parameters of the initial third-depth semantic model with the sample attribute information of the main body in multiple sample web pages, the initial third-depth semantic model is iteratively trained until the accuracy of the predicted authority output by the third-depth semantic model meets the pre-set authority threshold, and the training ends, obtaining the trained third-depth semantic model.

[0076] After training the third-depth semantic model, input the second attribute information of the main body in the target web page into the trained third-depth semantic model, and the authority of the target web page can be obtained.

[0077] The process of using the third-depth semantic model corresponding to the authority sensitivity dimension to obtain the authority sensitivity of the main body in the target web page, and using the third-depth semantic model corresponding to the content richness and sentence fluency dimension to obtain the content richness and sentence fluency of the main body in the target web page is the same as the process of using the third-depth semantic model corresponding to the authority dimension to obtain the authority of the target web page, and will not be elaborated here.

[0078] Through the above process, the authority of the target web page, the authority sensitivity of the main body in the target web page, the content richness and sentence fluency of the main body in the target web page are accurately determined, laying a foundation for accurately identifying the effectiveness of the target web page subsequently.

[0079] In an exemplary embodiment, the quality monitoring indicators of the site to which the target web page belongs include at least one of the proportion of cheating web pages, the rendering failure proportion corresponding to at least one sub-link, the cheating probability, and the proportion of newly added web pages.

[0080] The proportion of cheating web pages and its obtaining method will be described below.

[0081] The proportion of cheating web pages is the proportion of the number of cheating web pages in all web pages of the site to the total number of all web pages.

[0082] In an exemplary embodiment, the proportion of cheating web pages can be obtained in the following way: obtain all web pages in the site; use the first-depth semantic model to identify each web page to determine cheating web pages from each web page; according to the number of cheating web pages and the total number of all web pages, determine the proportion of cheating web pages.

[0083] Among them, the first deep semantic model can be a deep semantic model of any type or structure that can identify cheating web pages, and the present disclosure places no restrictions thereon.

[0084] In an exemplary embodiment, a cheating threshold can be preset in advance, and the first deep semantic model is used to score each web page in the site to which the target web page belongs. When the score of a certain web page is higher than the preset cheating threshold, it can be determined that the web page is a cheating web page. After using the first deep semantic model to identify each web page to determine the cheating web pages from each web page, the ratio of the number of cheating web pages to the total number of all web pages can be determined as the cheating web page ratio.

[0085] Through the above process, the cheating web page ratio of the site to which the target web page belongs is accurately determined, laying a foundation for accurately identifying the effectiveness of the target web page subsequently.

[0086] Next, the rendering failure ratio corresponding to at least one sub-chain and the way to obtain it will be described.

[0087] The sub-chain can include js, css, image, etc. The rendering failure ratio corresponding to the sub-chain is the ratio of the number of web pages in which the sub-chain fails to render to the total number of all web pages in the site. The rendering failure ratio can indicate the quality of the web pages published by the site to which the target web page belongs. The higher the rendering failure ratio, the lower the quality of the web pages published by the site.

[0088] In an exemplary embodiment, the rendering failure ratio corresponding to at least one sub-chain can be obtained in the following way: obtain each web page in the site; for each web page, render each sub-chain in the web page; according to the number of web pages in which each sub-chain fails to render and the total number of all web pages, determine the rendering failure ratio corresponding to each sub-chain respectively.

[0089] Specifically, each web page in the site to which the target web page belongs can be crawled by a web crawler to render each sub-chain in the web page. Since there may be dead links, uncollected pictures, etc. in each web page in the site, which may cause the situation that each sub-chain in the web page fails to render during the rendering process. In the embodiments of the present disclosure, the number of web pages in which each sub-chain fails to render can be obtained, and for each sub-chain, the ratio of the number of web pages in which the sub-chain fails to render to the total number of all web pages is determined as the rendering failure ratio corresponding to the sub-chain.

[0090] Taking the sub - chains including js, css, and image as an example, assume that the total number of web pages of the site to which the target web page belongs is 100. Among them, 30 web pages have js rendering failures, 20 web pages have css rendering failures, and 40 web pages have image rendering failures. Then, it can be determined that the rendering failure ratio corresponding to js is 30 / 100, that is, 0.3; the rendering failure ratio corresponding to css is 20 / 100, that is, 0.2; and the rendering failure ratio corresponding to image is 40 / 100, that is, 0.4.

[0091] Through the above process, the rendering failure ratios corresponding to at least one sub - chain of the site to which the target web page belongs are accurately determined, laying a foundation for accurately identifying the effectiveness of the target web page subsequently.

[0092] Next, the cheating probability and its obtaining method will be described.

[0093] The cheating probability refers to the probability that the site to which the target web page belongs is a cheating site. The cheating probability can indicate the level of cheating possibility of the site to which the target web page belongs. When the cheating probability is higher, the cheating possibility of the site is higher.

[0094] In an exemplary embodiment, the cheating probability can be obtained in the following way: obtain the first attribute information of the site and the link relationship between the site and other sites; input the first attribute information of the site and the link relationship between the site and other sites into a pre - trained first graph network model to obtain the cheating probability of the site.

[0095] Among them, the first attribute information can include any information related to the site, such as the identifier and address of the site.

[0096] Among them, the first graph network model can be trained using sample site data. The sample site data includes the first sample attribute information of multiple sample sites and the link relationship between multiple sample sites and other sites. The multiple sample sites are labeled as whether they are cheating sites or not. The first sample attribute information includes any information related to the sample site, such as the identifier and address of the sample site. Inputting the sample site data into the initial first graph network model for training can obtain the trained first graph network model, and the trained first graph network model contains the link relationship between cheating sites and other sites.

[0097] Specifically, the link relationship between the site to which the target web page belongs and other sites can be obtained based on the statistics of the entire network data, and then the first attribute information of the site to which the target web page belongs and the link relationship between the site to which the target web page belongs and the other sites can be input into a pre-trained first graph network model, so as to compare the link relationship between the site to which the target web page belongs and other sites with the link relationship between the cheating site and other sites, and obtain a score indicating the similarity of the two link relationships, which can be used as the cheating probability of the site to which the target web page belongs.

[0098] Through the above process, the cheating probability of the site to which the target web page belongs is accurately determined, laying the foundation for the subsequent accurate identification of the effectiveness of the target web page.

[0099] The proportion of newly added web pages and the method for obtaining them are described below.

[0100] The ratio of newly added web pages refers to the ratio of the number of web pages of the site to which the target web page belongs within a preset time period to the number of historical web pages in the time period before the preset time period, to the number of historical web pages in the time period before the preset time period. The preset time period and the time period before the preset time period can be set arbitrarily as needed, and the present disclosure does not impose any restrictions on this. The ratio of newly added web pages can indicate the possibility that the productivity of the site to which the target web page belongs is stable and legal. For example, when the ratio of newly added web pages is significantly different from 1, the possibility that the productivity of the site is stable and legal is low.

[0101] In an exemplary embodiment, the proportion of newly added web pages can be obtained by: obtaining the number of web pages in the site within a preset time period and the historical number of web pages in the site before the preset time period; determining the proportion of newly added web pages based on the number of web pages in the site and the historical number.

[0102] Specifically, the new web page ratio may be determined as the ratio of the increase in the number of web pages of the site to which the target web page belongs within the preset time period compared to the historical number of web pages in the time period before the preset time period to the historical number of web pages in the time period before the preset time period.

[0103] For example, assuming that the preset time period is the 24 hours of January 2, and the time period before the preset time period is the 24 hours of January 1, the number of web pages of a site in the 24 hours of January 2 is 100, and the number of web pages in the 24 hours of January 1 is 10, that is, the proportion of the increase in the number of web pages of the site in the 24 hours of January 24 compared to the historical number of web pages in the 24 hours of January 1 is (100-10) / 10, that is, 900%. Accordingly, it can be determined that the proportion of new web pages is 900%.

[0104] Through the above process, the proportion of new web pages in the site to which the target web page belongs is accurately determined, laying a foundation for the subsequent accurate identification of the effectiveness of the target web page.

[0105] Step 203, the page monitoring index of the target webpage and the quality monitoring index of the site are input into the pre-trained ensemble tree classification model to identify the target webpage.

[0106] Among them, the integrated tree classification model can be any decision tree model, such as a Random Forest model, an Adaboost (Adaptive Boosting) model, a GBDT (Gradient Boosting Decision Tree) model, etc., and the present disclosure does not impose any limitation on this.

[0107] In an exemplary embodiment, sample web page data may be used to pre-train an integrated tree classification model, wherein the sample web page data here includes page monitoring indicators of multiple sample web pages and quality monitoring indicators of the sites to which the sample web pages belong, and multiple sample web pages are labeled as valid web pages. The specific training process may refer to the relevant technology and will not be described here.

[0108] After the ensemble tree classification model is trained and the page monitoring indicators of the target web page and the quality monitoring indicators of the site are obtained, the page monitoring indicators of the target web page and the quality monitoring indicators of the site can be input into the pre-trained ensemble tree classification model to identify the target web page.

[0109] By adopting an integrated tree classification model, the validity of the target web page is identified based on the page monitoring indicators of the target web page and the quality monitoring indicators of the site, which saves computing resources when identifying the validity of the target web page and improves computing efficiency.

[0110] In summary, the method for identifying the validity of a web page provided by the embodiment of the present disclosure obtains the target web page to be identified, obtains the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs, and then inputs the page monitoring indicators of the target web page and the quality monitoring indicators of the site into a pre-trained integrated tree classification model to identify the target web page, thereby achieving accurate identification of the validity of the target web page.

[0111] Through the above analysis, it can be known that in the embodiments of the present disclosure, the target webpage can be identified based on the page monitoring indicators of the target webpage and the quality monitoring indicators of the site to determine the validity of the target webpage. In a possible implementation form, the validity of the target webpage can be further determined by combining the problem probability of the webmaster account that produced the target webpage, where the problem probability is the probability that the webmaster account is a problem account. Figure 3 , further illustrating the method for identifying the validity of a web page provided by the present invention.

[0112] Figure 3 FIG. 1 is a flow chart of a method for identifying the validity of a web page according to the third embodiment of the present disclosure. Figure 3 As shown, the method for identifying the validity of a web page may include the following steps:

[0113] Step 301, obtaining a target web page to be authenticated.

[0114] Step 302, obtaining the page monitoring index of the target webpage, the quality monitoring index of the site to which the target webpage belongs, and obtaining the problem probability of the webmaster account that produces the target webpage.

[0115] The page monitoring index includes at least one of the relevance between the text and the title, the audio and video quality, the authority of the target webpage, the authority sensitivity of the text in the target webpage, and the content richness and sentence fluency of the text in the target webpage. The process of obtaining the page monitoring index of the target webpage can refer to the description of the above embodiment, which will not be repeated here.

[0116] The quality monitoring index includes at least one of the proportion of cheating web pages, the rendering failure proportion corresponding to at least one subchain, the cheating probability, and the proportion of newly added web pages. The process of obtaining the quality monitoring index of the site to which the target web page belongs can refer to the description of the above embodiment and will not be repeated here.

[0117] The problem probability is the probability that the webmaster account is a problem account. The problem probability can indicate the possibility that the webmaster account is a problem account. The higher the problem probability, the higher the possibility that the webmaster account is a problem account.

[0118] In an exemplary embodiment, the problem probability of the webmaster who produces the target web page can be obtained in the following manner: obtaining the third attribute information of the webmaster account and the association relationship between the webmaster account and other accounts; inputting the third attribute information of the webmaster account and the association relationship between the webmaster account and other accounts into a pre-trained second graph network model to obtain the problem probability.

[0119] Among them, the third attribute information may include the webmaster account identification, registration time, account level and other any information related to the webmaster account.

[0120] The second graph network model is trained using sample account data, wherein the sample account data includes second sample attribute information of multiple sample accounts and the association relationship between multiple sample accounts and other accounts, and multiple sample accounts are marked as whether they are problem accounts. The second sample attribute information includes any information related to the sample account, such as the sample account identification, registration time, account level, etc. The sample account in the embodiment of the present disclosure refers to the sample webmaster account of the produced web page. The sample account data is input into the initial second graph network model for training to obtain the trained second graph network model, wherein the trained second graph network model contains the association relationship between the problem account and other accounts.

[0121] Specifically, the correlation relationship between the webmaster account of the target web page and other accounts can be obtained based on the statistics of the entire network data, and then the third attribute information of the webmaster account and the correlation relationship between the webmaster account and other accounts can be input into the pre-trained second graph network model, so as to compare the correlation relationship between the webmaster account of the target web page and other accounts with the correlation relationship between the problem account and other accounts, and obtain a score indicating the similarity of the two correlation relationships, which can be used as the problem probability of the webmaster account of the target web page.

[0122] Through the above process, it is possible to accurately determine the probability of problems with the webmaster account that produces the target web page, laying the foundation for subsequent accurate identification of the effectiveness of the target web page.

[0123] Step 303: Identify the target web page based on the page monitoring index of the target web page, the quality monitoring index of the site, and the problem probability.

[0124] In an exemplary embodiment, the page monitoring index of the target webpage, the quality monitoring index of the site, and the problem probability can be input into the pre-trained ensemble tree classification model to identify the target webpage. By using the ensemble tree classification model to identify the validity of the target webpage, the computing resources when identifying the validity of the target webpage can be saved and the computing efficiency can be improved.

[0125] Among them, the integrated tree classification model can be any decision tree model, such as a Random Tree model, an Adaboost model, a GBDT model, etc., and the present disclosure does not limit this.

[0126] In an exemplary embodiment, sample web page data may be used to pre-train an integrated tree classification model, wherein the sample web page data here includes page monitoring indicators of multiple sample web pages, quality monitoring indicators of the site to which the sample web pages belong, and the problem probability of the webmaster account that produces the sample web pages, and multiple sample web pages are labeled as valid web pages. The specific training process can refer to the relevant technology and will not be repeated here.

[0127] After training the integrated tree classification model and obtaining the page monitoring indicators of the target web page, the quality monitoring indicators of the site, and the problem probability of the webmaster account that produces the target web page, the page monitoring indicators of the target web page, the quality monitoring indicators of the site, and the problem probability of the webmaster account that produces the target web page can be input into the pre-trained integrated tree classification model to identify the target web page.

[0128] In summary, after obtaining the target web page to be identified, obtaining the page monitoring indicators of the target web page, the quality monitoring indicators of the site to which the target web page belongs, and the problem probability of the webmaster account that produces the target web page, combining the page monitoring indicators of the target web page itself, the quality monitoring indicators of the site to which the target web page belongs, and the problem probability of the webmaster account that produces the target web page, the target web page is comprehensively identified, which can improve the accuracy of the identification of the effectiveness of the target web page.

[0129] Through the above analysis, it can be known that in the embodiment of the present disclosure, the target webpage can be identified based on the page monitoring index of the target webpage, the quality monitoring index of the site and the problem probability of the webmaster who produced the target webpage to determine the validity of the target webpage. In a possible implementation form, the validity of the target webpage can be comprehensively determined by further combining the feedback information of the user on the target webpage. Figure 4 , further illustrating the method for identifying the validity of a web page provided by the present invention.

[0130] Figure 4 FIG. 4 is a flow chart of a method for identifying the validity of a web page according to the fourth embodiment of the present disclosure. Figure 4 As shown, the method for identifying the validity of a web page may include the following steps:

[0131] Step 401, obtaining a target web page to be authenticated.

[0132] Step 402, obtaining the page monitoring index of the target webpage, the quality monitoring index of the site to which the target webpage belongs, the problem probability of the webmaster account that produces the target webpage, and the user's feedback information on the target webpage.

[0133] The page monitoring index includes at least one of the relevance between the text and the title, the audio and video quality, the authority of the target webpage, the authority sensitivity of the text in the target webpage, and the content richness and sentence fluency of the text in the target webpage. The process of obtaining the page monitoring index of the target webpage can refer to the description of the above embodiment, which will not be repeated here.

[0134] The quality monitoring index includes at least one of the proportion of cheating web pages, the rendering failure proportion corresponding to at least one subchain, the cheating probability, and the proportion of newly added web pages. The process of obtaining the quality monitoring index of the site to which the target web page belongs can refer to the description of the above embodiment and will not be repeated here.

[0135] The problem probability is the probability that the webmaster account is a problem account. The process of obtaining the problem probability can refer to the description of the above embodiment, which will not be repeated here.

[0136] User feedback information on the target web page may include the number of clicks, likes, click-throughs, number of positive comments, number of negative comments, number of impressions of the target web page, and other data that can intuitively reflect the quality of the target web page.

[0137] Step 403: Identify the target web page based on the page monitoring index of the target web page, the quality monitoring index of the site, the problem probability and the feedback information.

[0138] In an exemplary embodiment, the page monitoring index of the target webpage, the quality monitoring index of the site, the problem probability, and the feedback information can be input into the pre-trained ensemble tree classification model to identify the target webpage. By using the ensemble tree classification model to identify the validity of the target webpage, the computing resources when identifying the validity of the target webpage can be saved and the computing efficiency can be improved.

[0139] Among them, the integrated tree classification model can be any decision tree model, such as a Random Tree model, an Adaboost model, a GBDT model, etc., and the present disclosure does not limit this.

[0140] In an exemplary embodiment, sample web page data may be used to pre-train an integrated tree classification model, wherein the sample web page data here includes page monitoring indicators of multiple sample web pages, quality monitoring indicators of the site to which the sample web pages belong, the problem probability of the webmaster account that produces the sample web pages, and user feedback information on the sample web pages, and multiple sample web pages are labeled as valid web pages. The specific training process can refer to the relevant technology and will not be repeated here.

[0141] After training the ensemble tree classification model and obtaining the page monitoring indicators of the target web page, the quality monitoring indicators of the site, the problem probability of the webmaster account that produces the target web page, and the user's feedback information on the target web page, the page monitoring indicators of the target web page, the quality monitoring indicators of the site, the problem probability of the webmaster account that produces the target web page, and the user's feedback information on the target web page can be input into the pre-trained ensemble tree classification model to identify the target web page.

[0142] By combining the page monitoring indicators of the target webpage itself, the quality monitoring indicators of the site to which the target webpage belongs, the problem probability of the webmaster account that produces the target webpage, and the user's feedback information on the target webpage, the target webpage can be comprehensively identified, which can improve the accuracy of the identification of the effectiveness of the target webpage.

[0143] In an exemplary embodiment, the target web page to be identified may be any web page in the resource library corresponding to the search engine. Accordingly, after step 403, the following may also be included:

[0144] Step 404: if the target webpage is a valid webpage, retain the target webpage in the resource library.

[0145] Step 405: if the target webpage is an invalid webpage, delete the target webpage from the resource library.

[0146] refer to Figure 5 For each target web page in the resource library corresponding to the search engine, at least one of the relevance between the text and the title of the target web page, the normality of the audio and video, the authority of the target web page, the authority sensitivity of the text in the target web page, the content richness of the text in the target web page, and the sentence flow degree can be obtained as the page monitoring index of the target web page; at least one of the proportion of cheating web pages, the rendering failure proportion corresponding to at least one sub-chain, the cheating probability, and the proportion of newly added web pages can be obtained as the quality monitoring index of the site to which the target web page belongs; the problem probability of the webmaster account that produces the target web page is obtained; or at least one of the number of clicks, display volume, likes, click-out volume, the number of positive comments, and the number of negative comments is used as the user's feedback information on the target web page. Then, the page monitoring index of the target web page, the quality monitoring index of the site to which the target web page belongs, the problem probability of the webmaster account that produces the target web page, and the user's feedback information on the target web page are input into the integrated tree classification model to identify the target web page and obtain the validity of the target web page. If the target web page is a valid web page, the target web page is retained in the resource library; if the target web page is an invalid web page, the target web page is deleted from the resource library.

[0147] By deleting invalid web pages in the resource library corresponding to the search engine, when users use the search engine to search, the web pages provided by the search engine to the users are all valid web pages, which improves the quality of the web pages provided by the search engine to the users and improves the search experience of the users. In addition, the number of web pages in the resource library is reduced, and the search engine does not need to judge the validity of the web pages in the resource library when providing web pages to users, thereby improving the efficiency of the search engine in providing web pages to users.

[0148] Combine the following Figure 6 , the device for identifying the validity of a web page provided by the present invention is described.

[0149] Figure 6 4 is a schematic diagram of the structure of a device for identifying the validity of a web page according to the fifth embodiment of the present disclosure.

[0150] like Figure 6 As shown, the webpage validity authentication device 600 provided by the present disclosure includes: a first acquisition module 601 , a second acquisition module 602 and an authentication module 603 .

[0151] The first acquisition module 601 is used to acquire the target web page to be identified;

[0152] The second acquisition module 602 is used to acquire the page monitoring index of the target webpage and the quality monitoring index of the site to which the target webpage belongs;

[0153] The identification module 603 is used to identify the target webpage based on the page monitoring index of the target webpage and the quality monitoring index of the site to determine the validity of the target webpage.

[0154] It should be noted that the webpage validity identification device provided in this embodiment can execute the webpage validity identification method of the above embodiment. The webpage validity identification device can be an electronic device or can be configured in an electronic device to accurately identify the validity of the webpage.

[0155] Among them, the electronic device can be any stationary or mobile computing device capable of data processing, such as mobile computing devices such as laptops, smart phones, wearable devices, or stationary computing devices such as desktop computers, or servers, or other types of computing devices, etc., and the present disclosure does not limit this.

[0156] It should be noted that the description of the above-mentioned embodiment of the method for identifying the validity of a web page is also applicable to the device for identifying the validity of a web page provided in the present disclosure, and will not be repeated here.

[0157] The webpage validity identification device provided by the embodiment of the present disclosure obtains the target webpage to be identified, obtains the page monitoring index of the target webpage and the quality monitoring index of the site to which the target webpage belongs, and then identifies the target webpage based on the page monitoring index of the target webpage and the quality monitoring index of the site to determine the validity of the target webpage. Thus, by combining the page monitoring index of the target webpage itself and the quality monitoring index of the site to which the target webpage belongs, the target webpage is comprehensively identified, thereby achieving accurate identification of the validity of the target webpage.

[0158] Combine the following Figure 7 , the device for identifying the validity of a web page provided by the present invention is described.

[0159] Figure 7 4 is a schematic diagram of the structure of a device for identifying the validity of a web page according to a sixth embodiment of the present disclosure.

[0160] like Figure 7 As shown, the webpage validity identification device 700 may specifically include: a first acquisition module 701, a second acquisition module 702 and an identification module 703. Figure 7 The first acquisition module 701, the second acquisition module 702 and the identification module 703 are Figure 6 The first acquisition module 601, the second acquisition module 602 and the identification module 603 have the same function and structure.

[0161] In an exemplary embodiment, the quality monitoring index includes the proportion of cheating web pages;

[0162] The second acquisition module 702 includes:

[0163] A first acquisition unit, used to acquire each web page in the site;

[0164] an identification unit, configured to identify each web page using a first deep semantic model, so as to determine a cheating web page from among the web pages;

[0165] The first determining unit is used to determine the proportion of cheating web pages according to the number of cheating web pages and the total number of web pages.

[0166] In an exemplary embodiment, the quality monitoring indicator includes a rendering failure ratio corresponding to at least one subchain;

[0167] The second acquisition module 702 includes:

[0168] A second acquisition unit, used to acquire each web page in the site;

[0169] A rendering unit, used for rendering each sub-link in each web page;

[0170] The second determining unit is used to determine the rendering failure ratio corresponding to each sub-chain according to the number of web pages corresponding to each sub-chain that failed to render and the total number of web pages.

[0171] In an exemplary embodiment, the quality monitoring indicator includes a probability of cheating;

[0172] The second acquisition module 702 includes:

[0173] A third acquisition unit, used to acquire first attribute information of the site and link relationships between the site and other sites;

[0174] A fourth acquisition unit, used for inputting the first attribute information of the site and the link relationship between the site and other sites into a pre-trained first graph network model to obtain the cheating probability of the site;

[0175] Among them, the first graph network model is trained using sample site data, the sample site data includes first sample attribute information of multiple sample sites and link relationships between multiple sample sites and other sites, and multiple sample sites are marked as cheating sites.

[0176] In an exemplary embodiment, the quality monitoring indicators include the proportion of newly added web pages;

[0177] The second acquisition module 702 includes:

[0178] A fifth acquisition unit, used to acquire the number of web pages in the site within a preset time period and the historical number of web pages in the site within a time period before the preset time period;

[0179] The third determining unit is used to determine the proportion of new web pages according to the number of web pages in the site and the historical number.

[0180] In an exemplary embodiment, the page monitoring indicators of the target web page include the relevance of the body text and the title;

[0181] The second acquisition module 702 includes:

[0182] The sixth acquisition unit is used to input the target webpage into the second deep semantic model to obtain the relevance between the text and the title in the target webpage.

[0183] In an exemplary embodiment, the page monitoring indicators of the target webpage include the authority of the target webpage, the authority sensitivity of the text in the target webpage, and the content richness and sentence fluency of the text in the target webpage;

[0184] The second acquisition module 702 includes:

[0185] An extraction unit, used for extracting second attribute information of the text in the target web page;

[0186] The seventh acquisition unit is used to input the second attribute information into the third deep semantic model to obtain the authority of the target webpage, the authority sensitivity of the text in the target webpage, and the content richness and sentence fluency of the text in the target webpage.

[0187] In an exemplary embodiment, the authentication device 700 further includes:

[0188] The third acquisition module 704 is used to obtain the problem probability of the webmaster account of the target webpage, wherein the problem probability is the probability that the webmaster account is a problem account;

[0189] The identification module 703 includes:

[0190] The first identification unit is used to identify the target web page based on the page monitoring index of the target web page, the quality monitoring index of the site and the problem probability.

[0191] In an exemplary embodiment, the third acquisition module 704 includes:

[0192] An eighth acquisition unit, used to acquire third attribute information of the webmaster account and associations between the webmaster account and other accounts;

[0193] a ninth acquisition unit, configured to input the third attribute information of the webmaster account and the association relationship between the webmaster account and other accounts into the pre-trained second graph network model to obtain a problem probability;

[0194] Among them, the second graph network model is trained using sample account data, and the sample account data includes second sample attribute information of multiple sample accounts and the association relationship between multiple sample accounts and other accounts. Multiple sample accounts are marked as to whether they are problem accounts.

[0195] In an exemplary embodiment, the authentication device 700 further includes:

[0196] The fourth acquisition module 705 is used to obtain user feedback information on the target web page;

[0197] The identification module 703 includes:

[0198] The second identification unit is used to identify the target web page based on the page monitoring index of the target web page, the quality monitoring index of the site, the problem probability and the feedback information.

[0199] In an exemplary embodiment, the identification module 703 includes:

[0200] The third identification unit is used to input the page monitoring index of the target webpage and the quality monitoring index of the site into the pre-trained integrated tree classification model to identify the target webpage.

[0201] In an exemplary embodiment, the target web page to be identified is any web page in the resource library corresponding to the search engine;

[0202] The identification device 700 further includes:

[0203] A saving module 706 is used to save the target webpage in the resource library if the target webpage is a valid webpage;

[0204] The deletion module 707 is used to delete the target webpage from the resource library when the target webpage is an invalid webpage.

[0205] It should be noted that the description of the above-mentioned embodiment of the method for identifying the validity of a web page is also applicable to the device for identifying the validity of a web page provided in the present disclosure, and will not be repeated here.

[0206] The webpage validity identification device provided by the embodiment of the present disclosure obtains the target webpage to be identified, obtains the page monitoring index of the target webpage and the quality monitoring index of the site to which the target webpage belongs, and then identifies the target webpage based on the page monitoring index of the target webpage and the quality monitoring index of the site to determine the validity of the target webpage. Thus, by combining the page monitoring index of the target webpage itself and the quality monitoring index of the site to which the target webpage belongs, the target webpage is comprehensively identified, thereby achieving accurate identification of the validity of the target webpage.

[0207] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0208] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0209] like Figure 8As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0210] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0211] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the identification method of the validity of a web page. For example, in some embodiments, the identification method of the validity of a web page may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the identification method of the validity of the web page described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the identification method of the validity of a web page by any other appropriate means (e.g., by means of firmware).

[0212] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0213] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0214] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0215] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0216] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0217] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0218] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0219] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for identifying the validity of a web page, comprising: Obtaining a target web page to be identified; Obtaining the page monitoring indicators of the target web page and the quality monitoring indicators of the site to which the target web page belongs, where the quality monitoring indicators of the site to which the target web page belongs refer to the indicators used to evaluate the quality of the site to which the target web page belongs, and the page monitoring indicators of the target web page include the relevance between the body text and the title of the target web page, the authority of the target web page, the authority sensitivity of the body text in the target web page, and the content richness and sentence fluency of the body text in the target web page. The quality monitoring indicators of the site to which the target web page belongs include the proportion of cheating web pages in the site to which the target web page belongs, the rendering failure ratio corresponding to at least one sub-link, the cheating probability, and the proportion of newly added web pages; Obtaining the third attribute information of the webmaster account and the association relationship between the webmaster account and other accounts; Inputting the association relationship into a pre-trained second graph network model to obtain a problem probability, where the problem probability is the probability that the webmaster account is a problem account; Obtaining the feedback information of the user on the target web page; Based on the page monitoring indicators of the target web page, the quality monitoring indicators of the site, the problem probability, and the feedback information, identifying the target web page to determine the validity of the target web page; The obtaining of the quality monitoring indicators of the site to which the target web page belongs includes: Obtaining the first attribute information of the site and the link relationship between the site and other sites; Inputting the link relationship into a pre-trained first graph network model to obtain the cheating probability of the site.

2. The method according to claim 1, for obtaining the quality monitoring indicators of the site to which the target web page belongs, comprising: Obtaining each web page in the site; Using a first deep semantic model to identify each web page to determine cheating web pages from each web page; Determining the proportion of cheating web pages according to the number of cheating web pages and the total number of all web pages.

3. The method according to claim 1, wherein, The obtaining of the quality monitoring indicators of the site to which the target web page belongs includes: Obtaining each web page in the site; For each web page, rendering each sub-link in the web page; Determining the rendering failure ratio corresponding to each sub-link according to the number of web pages with rendering failures corresponding to each sub-link and the total number of all web pages.

4. The method according to claim 1, wherein, The first graph network model is trained using sample site data, and the sample site data includes the first sample attribute information of multiple sample sites and the link relationship between the multiple sample sites and other sites, and the multiple sample sites are labeled as whether they are cheating sites.

5. The method according to claim 1, where obtaining the quality monitoring indicators of the site to which the target web page belongs, comprising: Obtaining the number of web pages in the site within a preset time period and the historical number of web pages in the site within the time period before the preset time period; Determining the proportion of newly added web pages according to the number of web pages in the site and the historical number.

6. The method according to any one of claims 1-5, wherein, obtaining the page monitoring metrics of the target web page includes: inputting the target web page into a second deep semantic model to obtain the relevance between the body text and the title in the target web page.

7. The method according to any one of claims 1-5, wherein, the page monitoring metrics of the target web page include the authority of the target web page, the authority sensitivity of the body text in the target web page, and the content richness and sentence fluency of the body text in the target web page; wherein, obtaining the page monitoring metrics of the target web page includes: extracting second attribute information of the body text in the target web page; inputting the second attribute information into a third deep semantic model to obtain the authority of the target web page, the authority sensitivity of the body text in the target web page, and the content richness and sentence fluency of the body text in the target web page.

8. The method according to claim 1, wherein, the second graph network model is trained using sample account data, and the sample account data includes second sample attribute information of multiple sample accounts and the association relationships between the multiple sample accounts and other accounts, and the multiple sample accounts are labeled as whether they are problem accounts.

9. The method according to any one of claims 1-5, wherein, identifying the target web page based on the page monitoring metrics of the target web page and the quality monitoring metrics of the site includes: inputting the page monitoring metrics of the target web page and the quality monitoring metrics of the site into a pre-trained ensemble tree classification model to identify the target web page.

10. The method according to any one of claims 1-5, wherein, the target web page to be identified is any web page in the resource library corresponding to the search engine; after identifying the target web page, it further includes: when the target web page is a valid web page, retaining the target web page in the resource library; when the target web page is an invalid web page, deleting the target web page from the resource library.

11. An apparatus for identifying the validity of a web page, comprising: a first acquisition module for acquiring a target web page to be identified; a second acquisition module for acquiring the page monitoring metrics of the target web page and the quality monitoring metrics of the site to which the target web page belongs, where the quality monitoring metrics of the site to which the target web page belongs refer to the metrics used to evaluate the quality of the site to which the target web page belongs, and the page monitoring metrics of the target web page include the relevance between the body text and the title of the target web page, the authority of the target web page, the authority sensitivity of the body text in the target web page, and the content richness and sentence fluency of the body text in the target web page, and the quality monitoring metrics of the site to which the target web page belongs include the proportion of cheating web pages of the site to which the target web page belongs, the rendering failure proportion corresponding to at least one sub-chain, the cheating probability, and the proportion of newly added web pages; a third acquisition module for acquiring the problem probability of the webmaster account that produced the target web page, where the problem probability is the probability that the webmaster account is a problem account; A fourth acquisition module, configured to acquire feedback information of a user on the target webpage; An identification module, configured to identify the target webpage based on the page monitoring metrics of the target webpage, the quality monitoring metrics of the site, the problem probability, and the feedback information, so as to determine the effectiveness of the target webpage; Wherein, the second acquisition module includes: A third acquisition unit, configured to acquire first attribute information of the site and the link relationship between the site and other sites; A fourth acquisition unit, configured to input the link relationship into a pre-trained first graph network model to obtain the cheating probability of the site; Wherein, the third acquisition module includes: An eighth acquisition unit, configured to acquire third attribute information of the webmaster account and the association relationship between the webmaster account and other accounts; A ninth acquisition unit, configured to input the association relationship into a pre-trained second graph network model to obtain the problem probability.

12. The apparatus according to claim 11, Wherein, The second acquisition module includes: A first acquisition unit, configured to acquire each webpage in the site; An identification unit, configured to identify each webpage by using a first deep semantic model to determine cheating webpages from each webpage; A first determination unit, configured to determine the proportion of cheating webpages according to the number of cheating webpages and the total number of all webpages.

13. The apparatus according to claim 11, Wherein, The second acquisition module includes: A second acquisition unit, configured to acquire each webpage in the site; A rendering unit, configured to render each sub-link in each webpage; A second determination unit, configured to determine the rendering failure ratio corresponding to each sub-link according to the number of webpages with rendering failure corresponding to each sub-link and the total number of all webpages.

14. The apparatus according to claim 11, Wherein, The first graph network model is trained by using sample site data, the sample site data includes first sample attribute information of multiple sample sites and the link relationship between the multiple sample sites and other sites, and the multiple sample sites are labeled according to whether they are cheating sites.

15. The apparatus according to claim 11, Wherein, The second acquisition module includes: A fifth acquisition unit, configured to acquire the number of webpages in the site within a preset time period and the historical number of webpages in the site within a time period before the preset time period; A third determination unit, configured to determine the proportion of newly added webpages according to the number of webpages in the site and the historical number.

16. The apparatus according to any one of claims 11-15, Wherein, The second acquisition module includes: A sixth acquisition unit, configured to input the target webpage into a second deep semantic model to obtain the relevance between the text and the title in the target webpage.

17. The apparatus according to any one of claims 11-15, Wherein, The second acquisition module includes: An extraction unit, configured to extract second attribute information of the text in the target webpage; A seventh acquisition unit, configured to input the second attribute information into a third deep semantic model to obtain the authority of the target web page, the authority sensitivity of the body text in the target web page, and the content richness and sentence fluency of the body text in the target web page.

18. The apparatus according to claim 11, wherein, the second graph network model is trained using sample account data, the sample account data includes second sample attribute information of multiple sample accounts and the association relationships between the multiple sample accounts and other accounts, and the multiple sample accounts are labeled as whether they are problem accounts.

19. The apparatus according to any one of claims 11-15, wherein, the authentication module includes: A third authentication unit, configured to input the page monitoring metrics of the target web page and the quality monitoring metrics of the site into a pre-trained ensemble tree classification model to authenticate the target web page.

20. The apparatus according to any one of claims 11-15, wherein, the target web page to be authenticated is any web page in the resource library corresponding to the search engine; the apparatus further includes: A storage module, configured to retain the target web page in the resource library when the target web page is a valid web page; A deletion module, configured to delete the target web page from the resource library when the target web page is an invalid web page.

21. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

23. A computer program product, including a computer program, where the steps of the method according to any one of claims 1-10 are implemented when the computer program is executed by a processor.

Citation Information

Patent Citations

  • Method and device for determining abnormal webpage, equipment, medium and program product

    CN113221035A