A web page content integrity detection system and method based on data analysis

By analyzing user browsing operation records to construct browsing transfer links and duplicate links, calculating browsing transfer rates to filter out web pages with missing content, and using fault diagnosis results for repair, the problem of the inability to detect missing web page content in existing technologies has been solved, thereby improving user experience and competitiveness.

CN115658505BActive Publication Date: 2026-05-15HAINAN INFOBAHN TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HAINAN INFOBAHN TECH SERVICE CO LTD
Filing Date
2022-10-25
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The lack of effective tools in the current technology to detect missing content on web pages leads to poor user experience and puts the user at a competitive disadvantage.

Method used

By analyzing user browsing operation records, we construct user browsing transfer links and repeat browsing links, calculate browsing transfer rate and repeat browsing rate, filter out web pages suspected of having missing content, and use the fault diagnosis results of maintenance personnel to troubleshoot the faults.

Benefits of technology

It enables timely identification and repair of missing webpage content, improving user experience and enhancing the webpage's competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658505B_ABST
    Figure CN115658505B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of webpage identification detection, in particular to a webpage content integrity detection system and method based on data analysis, which comprises the following steps: extracting user browsing operation records of each to-be-detected webpage in a time period, capturing browsing transfer operations generated between the corresponding to-be-detected webpages and repeated browsing operations generated on the corresponding to-be-detected webpages based on the user browsing operation records, and constructing corresponding first user browsing transfer links and second user browsing transfer links; locking the to-be-detected webpages suspected of having content missing, calculating the first browsing transfer rate of each target to-be-detected webpage, and completing the first calibration screening; calculating the second browsing transfer rate of each target to-be-detected webpage, and completing the second calibration screening; and assisting operation and maintenance personnel in fault screening of each target to-be-detected webpage in a target to-be-detected webpage sequence received in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of webpage recognition and detection technology, specifically to a webpage content integrity detection system and method based on data analysis. Background Technology

[0002] The inspection of a webpage includes checking for typos and missing content. While there are many existing tools for checking for typos, tools for checking for missing content are still under development. From a development perspective, missing content is often caused by errors in the initial page design, leading to some code failing to execute correctly.

[0003] Web pages with missing content not only provide a poor user experience, but also put them at a competitive disadvantage when faced with web pages containing similar content, due to the lack of content that users actually need. These web pages often require further repair and improvement. Summary of the Invention

[0004] The purpose of this invention is to provide a web page content integrity detection system and method based on data analysis to solve the problems mentioned in the background art.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a webpage content integrity detection method based on data analysis, the detection method comprising:

[0006] Step S100: Extract user browsing operation records for each webpage to be detected within a certain period of time. Based on the user browsing operation records, capture the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected, and construct the corresponding first user browsing transfer link and second user browsing transfer link respectively.

[0007] Step S200: Based on the similarity between the content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link, lock the web pages to be detected that are suspected of having missing content, set the web pages to be detected as target web pages to be detected, calculate the first browsing transfer rate for each target web page to be detected, and perform the first calibration screening of all target web pages to be detected based on the first browsing transfer rate.

[0008] Step S300: Based on the correlation distribution of the first user browsing transfer link and the second user browsing transfer link, calculate the second browsing transfer rate for each target webpage to be detected, and perform a second calibration and screening on all target webpages to be detected based on the second browsing transfer rate;

[0009] Step S400: Collect all target web pages to be detected, generate a sequence of target web pages to be detected, and input the sequence of target web pages to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target web page to be detected, extract and store the page features corresponding to various faults, and use the stored data to assist the operation and maintenance personnel in screening each target web page to be detected in the real-time received sequence of target web pages to be detected for faults.

[0010] Furthermore, step S100 includes:

[0011] Step S101: Collect all user browsing operation records for each webpage to be tested to obtain a set of user browsing operation records for each webpage to be tested; extract the user IP and the time period of each user browsing operation record for each webpage to be tested; capture the user browsing operation records with the same user IP from the set of all user browsing operation records.

[0012] Step S102: If the i-th user browsing operation record is captured in the user browsing operation record set A corresponding to a certain webpage to be detected a The j-th user browsing operation record that exists in the set B of user browsing operation records corresponding to a certain webpage b to be detected. The corresponding user IPs are the same, extract The corresponding record generation period The corresponding record generation period like lie in After that, and and If the time interval between them is less than the first interval threshold, and the current identical user IP is H, construct the first user browsing transfer link: Determine if user H has performed a browsing switch operation between a webpage to be detected and a webpage to be detected; where webpage to be detected b is the front-end webpage browsed by user H, and webpage to be detected a is the back-end webpage browsed by user H.

[0013] Step S103: If the same user IP is detected based on a certain webpage d to be detected, the xth user browsing operation record in the user browsing operation record set D corresponding to the webpage d is generated. With the yth user browsing operation record Extract separately The corresponding record generation period like lie in After that, and and If the time interval between them is less than the second interval threshold, and the current IP address of the same user is F, construct a second user browsing transfer link: It is determined that user F has performed repeated browsing operations on a certain webpage d to be detected.

[0014] Furthermore, step S200 includes:

[0015] Step S201: Extract content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link. If the similarity between the content feature information of the front-end and back-end web pages to be detected in a certain first user browsing transfer link is less than the similarity threshold, the first user browsing transfer link is removed.

[0016] In this situation, it is unlikely that a user would switch to another webpage because they cannot find the information they want on the front-end webpage due to missing content. This is because the front-end and back-end webpages are not similar in content, which means that the front-end and back-end webpages may be the corresponding webpages generated by the same user based on different search goals. Therefore, it cannot be used to determine that there is a problem of missing content on the front-end webpage being tested.

[0017] Step S202: Identify the web pages that appear as front-end web pages in each first user browsing transfer link and set them as target web pages to be detected; calculate the first browsing transfer rate for each target web page to be detected. Where u represents the total number of times the target webpage appears as a front-end webpage; m represents the total number of times the target webpage appears in all first user browsing transfer links; target webpages with a first browsing transfer rate less than the first browsing transfer rate threshold are removed.

[0018] Furthermore, step S300 includes:

[0019] Step S301: Assume that all first-user browsing transfer links appearing on a target webpage as a front-end webpage include {L1, L2, ..., L...} n}; where L1, L2, ..., L n These represent the 1st, 2nd, ..., nth first-user browsing transfer links when a target webpage appears as a front-end webpage; and the users corresponding to each first-user browsing transfer link are obtained respectively.

[0020] Step S302: If a user corresponding to a first user browsing transfer link has a corresponding second user browsing transfer link based on a target webpage to be detected, and the time period between the time period of the repeated browsing operation in the second user browsing transfer link and the time period of the time period of the record corresponding to browsing the target webpage to be detected in a first user browsing transfer link is less than the time period difference threshold, then a first user browsing transfer link is determined to be a target first user browsing transfer link for a target webpage to be detected.

[0021] Step S303: Accumulate the number k of target first-user browsing transfer links contained in all first-user browsing transfer links corresponding to a target webpage to be detected; calculate the second browsing transfer rate for each target webpage to be detected. Remove target pages whose second browsing conversion rate is less than the second browsing conversion rate threshold;

[0022] The above situation is equivalent to checking whether a user repeatedly clicked on a target webpage before browsing to that webpage. Under normal circumstances, due to network transmission speed issues, content may fail to load on a webpage. If the system detects that a user repeatedly clicked on a target webpage before browsing to that webpage, it is essentially using the user's repeated clicks to exclude page anomalies caused by the transmission network to a certain extent.

[0023] Furthermore, step S400 includes:

[0024] Step S401: Extract the first browsing transition rate e1 and the second browsing transition rate e2 corresponding to all target web pages to be detected; calculate the missing probability value P = e1 * e2 for each target web page to be detected; sort all target web pages to be detected in descending order of their corresponding missing probability values ​​to obtain the target web page sequence;

[0025] Step S402: Obtain the fault diagnosis results confirmed by the operation and maintenance personnel for each target webpage in the target webpage sequence, and extract common page features from all target webpages with the same fault diagnosis results and store them.

[0026] Step S403: Extract page features from each webpage in the real-time input target webpage sequence, and perform similarity matching between the extracted page features and the page features corresponding to various types of faults. When the similarity between the page features of a webpage to be detected and the page features corresponding to a certain type of fault is greater than the similarity threshold, feedback is given to the maintenance personnel to prioritize the investigation of a certain type of fault on the webpage to be detected.

[0027] To better implement the above method, a webpage content integrity detection system is also proposed. The system includes: a user browsing operation record collection and processing module, a target webpage screening and processing module, a target webpage calibration and screening module, a target webpage sequence generation module, and an automatic auxiliary detection module.

[0028] The user browsing operation record collection and processing module is used to extract user browsing operation records for each webpage to be detected within a certain period of time. Based on the user browsing operation records, it captures the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected, and constructs the corresponding first user browsing transfer link and second user browsing transfer link respectively.

[0029] The target webpage screening and processing module is used to lock webpages suspected of having missing content based on the similarity between the content feature information of the front-end and back-end webpages in each first user browsing transfer link. The webpage to be detected is set as the target webpage to be detected.

[0030] The target webpage calibration and screening module is used to receive data from the target webpage screening and processing module and perform the first and second calibration screenings on each target webpage.

[0031] The target webpage sequence generation module is used to receive data from the target webpage calibration and screening module, collect all the target webpages to be detected, and generate a target webpage sequence to be detected.

[0032] The automatic auxiliary detection module is used to input the sequence of target web pages to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target web page to be detected; extract and store the page features corresponding to various faults; and assist the operation and maintenance personnel in screening each target web page in the real-time received sequence of target web pages to be detected based on the stored data.

[0033] Furthermore, the user browsing operation record collection and processing module includes a first user browsing transfer link construction unit and a second user browsing transfer link construction unit;

[0034] The first user browsing transfer link construction unit is used to capture the browsing transfer operations generated by the user between the corresponding web pages to be detected, and construct the first user browsing transfer link accordingly.

[0035] The second user browsing transfer link construction unit is used to capture repeated browsing operations generated by users on the corresponding web pages to be detected, and to construct the second user browsing transfer link accordingly.

[0036] Furthermore, the target webpage calibration screening module includes a first calibration screening unit and a second calibration screening unit;

[0037] The first calibration and filtering unit is used to calculate the first browsing transition rate for each target webpage to be detected; and to perform the first calibration and filtering on all target webpages to be detected based on the first browsing transition rate.

[0038] The second calibration and screening unit is used to calculate the second browsing transition rate for each target webpage to be detected; and to perform a second calibration and screening on all target webpages to be detected based on the second browsing transition rate.

[0039] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention analyzes the user browsing data of each webpage in the test, identifies and filters out browsing shifts caused by missing content in the webpage, reflects the user competitiveness of each webpage in the actual operating context through historical browsing data, and filters and captures webpages with missing content in the context of big data based on user characteristic browsing behavior, thereby enabling timely repair of problems that occurred in the development and design stage of webpages with deficiencies. Attached Figure Description

[0040] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0041] Figure 1 This is a flowchart illustrating a webpage content integrity detection method based on data analysis according to the present invention.

[0042] Figure 2 This is a schematic diagram of the structure of a webpage content integrity detection system based on data analysis according to the present invention. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Please see Figures 1-2 The present invention provides the following technical solution:

[0045] A method for detecting webpage content integrity based on data analysis, the method comprising:

[0046] Step S100: Extract user browsing operation records for each webpage to be detected within a certain period of time. Based on the user browsing operation records, capture the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected, and construct the corresponding first user browsing transfer link and second user browsing transfer link respectively.

[0047] Step S100 includes:

[0048] Step S101: Collect all user browsing operation records for each webpage to be tested to obtain a set of user browsing operation records for each webpage to be tested; extract the user IP and the time period of each user browsing operation record for each webpage to be tested; capture the user browsing operation records with the same user IP from the set of all user browsing operation records.

[0049] Step S102: If the i-th user browsing operation record is captured in the user browsing operation record set A corresponding to a certain webpage to be detected a The j-th user browsing operation record that exists in the set B of user browsing operation records corresponding to a certain webpage b to be detected. The corresponding user IPs are the same, extract The corresponding record generation period The corresponding record generation period like lie in After that, and and If the time interval between them is less than the first interval threshold, and the current identical user IP is H, construct the first user browsing transfer link: Determine if user H has performed a browsing switch operation between a webpage to be detected and a webpage to be detected; where webpage to be detected b is the front-end webpage browsed by user H, and webpage to be detected a is the back-end webpage browsed by user H.

[0050] Step S103: If the same user IP is detected based on a certain webpage d to be detected, the xth user browsing operation record in the user browsing operation record set D corresponding to the webpage d is generated. With the yth user browsing operation record Extract separately The corresponding record generation period like lie in After that, and and If the time interval between them is less than the second interval threshold, and the current IP address of the same user is F, construct a second user browsing transfer link: Determine that user F has repeatedly browsed a certain webpage d to be detected;

[0051] Step S200: Based on the similarity between the content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link, lock the web pages to be detected that are suspected of having missing content, set the web pages to be detected as target web pages to be detected, calculate the first browsing transfer rate for each target web page to be detected, and perform the first calibration screening of all target web pages to be detected based on the first browsing transfer rate.

[0052] Step S200 includes:

[0053] Step S201: Extract content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link. If the similarity between the content feature information of the front-end and back-end web pages to be detected in a certain first user browsing transfer link is less than the similarity threshold, the first user browsing transfer link is removed.

[0054] Step S202: Identify the web pages that appear as front-end web pages in each first user browsing transfer link and set them as target web pages to be detected; calculate the first browsing transfer rate for each target web page to be detected. Where u represents the total number of times the target webpage appears as a front-end webpage; m represents the total number of times the target webpage appears in all first user browsing transition links; target webpages with a first browsing transition rate less than the first browsing transition rate threshold are removed.

[0055] Step S300: Based on the correlation distribution of the first user browsing transfer link and the second user browsing transfer link, calculate the second browsing transfer rate for each target webpage to be detected, and perform a second calibration and screening on all target webpages to be detected based on the second browsing transfer rate;

[0056] Step S300 includes:

[0057] Step S301: Assume that all first-user browsing transfer links appearing on a target webpage as a front-end webpage include {L1, L2, ..., L...} n}; where L1, L2, ..., L n These represent the 1st, 2nd, ..., nth first-user browsing transfer links when a target webpage appears as a front-end webpage; and the users corresponding to each first-user browsing transfer link are obtained respectively.

[0058] Step S302: If a user corresponding to a first user browsing transfer link has a corresponding second user browsing transfer link based on a target webpage to be detected, and the time period between the time period of the repeated browsing operation in the second user browsing transfer link and the time period of the time period of the record corresponding to browsing the target webpage to be detected in a first user browsing transfer link is less than the time period difference threshold, then a first user browsing transfer link is determined to be a target first user browsing transfer link for a target webpage to be detected.

[0059] Step S303: Accumulate the number k of target first-user browsing transfer links contained in all first-user browsing transfer links corresponding to a target webpage to be detected; calculate the second browsing transfer rate for each target webpage to be detected. Remove target pages whose second browsing conversion rate is less than the second browsing conversion rate threshold;

[0060] For example, the target webpage E1, as a front-end webpage, includes all first user browsing transfer links, such as first user browsing transfer link L1, first user browsing transfer link L2, first user browsing transfer link L3, and first user browsing transfer link L4.

[0061] Among them, the user corresponding to the first user browsing transfer link L1 is s1, the user corresponding to the first user browsing transfer link L2 is s2, the user corresponding to the first user browsing transfer link L3 is s3, and the user corresponding to the first user browsing transfer link L4 is s4.

[0062] Among them, user s1 has a corresponding second user browsing transfer link based on the target webpage E1 to be detected. That is, user s1 has repeated browsing operations on the target webpage E1 to be detected. The time difference between the time period of the record corresponding to the repeated browsing operation, that is, the second browsing operation, and the time period of the record corresponding to browsing the target webpage E1 in the first user browsing transfer link L1 is less than 1 minute. In summary, the first user browsing transfer link L1 is the target first user browsing transfer link of the target webpage E1 to be detected.

[0063] In this case, user s2 does not have a corresponding second user browsing transfer link based on the fact that the target webpage E1 does not have a corresponding second user browsing transfer link, that is, user s2 does not have a repeated browsing operation on the target webpage E1. In summary, the first user browsing transfer link L2 is not the target first user browsing transfer link of the target webpage E1.

[0064] Among them, user s3 has a corresponding second user browsing transfer link based on the target webpage E1, that is, user s3 has repeated browsing operations on the target webpage E1, and the time difference between the time period of the record corresponding to the repeated browsing operation, that is, the second browsing operation, and the time period of the record corresponding to browsing the target webpage E1 in the first user browsing transfer link L3 is greater than 1 minute. In summary, the first user browsing transfer link L3 is not the target first user browsing transfer link of the target webpage E1.

[0065] Among them, user s4 has a corresponding second user browsing transfer link based on the target webpage E1 to be detected. That is, user s4 has repeated browsing operations on the target webpage E1 to be detected. The time difference between the time period of the record corresponding to the repeated browsing operation, that is, the second browsing operation, and the time period of the record corresponding to browsing the target webpage E1 in the first user browsing transfer link L4 is less than 1 minute. In summary, the first user browsing transfer link L4 is the target first user browsing transfer link of the target webpage E1 to be detected.

[0066] The total number of target first-user browsing transfer links contained in all first-user browsing transfer links corresponding to the target webpage E1 is k = 2;

[0067] Step S400: Collect all target web pages to be detected, generate a sequence of target web pages to be detected, and input the sequence of target web pages to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target web page to be detected, extract and store the page features corresponding to various faults, and use the stored data to assist the operation and maintenance personnel in screening each target web page to be detected in the real-time received sequence of target web pages to be detected for faults;

[0068] Step S400 includes:

[0069] Step S401: Extract the first browsing transition rate e1 and the second browsing transition rate e2 corresponding to all target web pages to be detected; calculate the missing probability value P = e1 * e2 for each target web page to be detected; sort all target web pages to be detected in descending order of their corresponding missing probability values ​​to obtain the target web page sequence;

[0070] Step S402: Obtain the fault diagnosis results confirmed by the operation and maintenance personnel for each target webpage in the target webpage sequence, and extract common page features from all target webpages with the same fault diagnosis results and store them.

[0071] Step S403: Extract page features from each webpage in the real-time input target webpage sequence, and perform similarity matching between the extracted page features and the page features corresponding to various types of faults. When the similarity between the page features of a webpage to be detected and the page features corresponding to a certain type of fault is greater than the similarity threshold, feedback is given to the maintenance personnel to prioritize the investigation of a certain type of fault on the webpage to be detected.

[0072] To better implement the above method, a webpage content integrity detection system is also proposed. The system includes: a user browsing operation record collection and processing module, a target webpage screening and processing module, a target webpage calibration and screening module, a target webpage sequence generation module, and an automatic auxiliary detection module.

[0073] The user browsing operation record collection and processing module is used to extract user browsing operation records for each webpage to be detected within a certain period of time. Based on the user browsing operation records, it captures the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected, and constructs the corresponding first user browsing transfer link and second user browsing transfer link respectively.

[0074] The user browsing operation record collection and processing module includes a first user browsing transfer link construction unit and a second user browsing transfer link construction unit.

[0075] The first user browsing transfer link construction unit is used to capture the browsing transfer operations generated by the user between the corresponding web pages to be detected, and construct the first user browsing transfer link accordingly.

[0076] The second user browsing transfer link construction unit is used to capture repeated browsing operations generated by users on the corresponding web pages to be detected, and construct the second user browsing transfer link accordingly.

[0077] The target webpage screening and processing module is used to lock webpages suspected of having missing content based on the similarity between the content feature information of the front-end and back-end webpages in each first user browsing transfer link. The webpage to be detected is set as the target webpage to be detected.

[0078] The target webpage calibration and screening module is used to receive data from the target webpage screening and processing module and perform the first and second calibration screenings on each target webpage.

[0079] The target webpage calibration and screening module includes a first calibration and screening unit and a second calibration and screening unit.

[0080] The first calibration and filtering unit is used to calculate the first browsing transition rate for each target webpage to be detected; and to perform the first calibration and filtering on all target webpages to be detected based on the first browsing transition rate.

[0081] The second calibration and filtering unit is used to calculate the second browsing transition rate for each target webpage to be detected; and to perform a second calibration and filtering on all target webpages to be detected based on the second browsing transition rate.

[0082] The target webpage sequence generation module is used to receive data from the target webpage calibration and screening module, collect all the target webpages to be detected, and generate a target webpage sequence to be detected.

[0083] The automatic auxiliary detection module is used to input the sequence of target web pages to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target web page to be detected; extract and store the page features corresponding to various faults; and assist the operation and maintenance personnel in screening each target web page in the real-time received sequence of target web pages to be detected based on the stored data.

[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0085] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for detecting webpage content integrity based on data analysis, characterized in that, The detection method includes: Step S100: Extract user browsing operation records for each webpage to be detected within a certain period of time. Based on the user browsing operation records, capture the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected, and construct the corresponding first user browsing transfer link and second user browsing transfer link respectively. Step S200: Based on the similarity between the content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link, lock the web pages to be detected that are suspected of having missing content. Let the web pages to be detected be the target web pages to be detected. Calculate the first browsing transfer rate for each target web page to be detected. Perform the first calibration and screening of all target web pages to be detected based on the first browsing transfer rate. Step S300: Based on the correlation distribution of the first user browsing transfer link and the second user browsing transfer link, calculate the second browsing transfer rate for each target webpage to be detected, and perform a second calibration screening on all target webpages to be detected based on the second browsing transfer rate; Step S400: Collect all target web pages to be detected, generate a sequence of target web pages to be detected, and input the sequence of target web pages to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target web page to be detected, extract and store the page features corresponding to various faults, and use the stored data to assist the operation and maintenance personnel in screening each target web page to be detected in the real-time received sequence of target web pages to be detected for faults; Step S100 includes: Step S101: Collect all user browsing operation records for each webpage to be tested to obtain a set of user browsing operation records for each webpage to be tested; extract the user IP and the time period of each user browsing operation record for each webpage to be tested; capture the user browsing operation records with the same user IP from the set of all user browsing operation records. Step S102: If the i-th user browsing operation record is captured in the user browsing operation record set A corresponding to a certain webpage to be detected a The j-th user browsing operation record that exists in the set B of user browsing operation records corresponding to a certain webpage b to be detected. The corresponding user IPs are the same, extract The corresponding record generation period , The corresponding record generation period ,like lie in After that, and and If the time interval between them is less than the first interval threshold, and the current identical user IP is H, construct the first user browsing transfer link: It is determined that user H has performed a browsing switch operation between a certain webpage to be detected, a and a certain webpage to be detected; where the webpage to be detected, b, is the front-end webpage browsed by user H, and the webpage to be detected, a, is the back-end webpage browsed by user H. Step S103: If the same user IP is detected based on a certain webpage d to be detected, the xth user browsing operation record in the user browsing operation record set D corresponding to the webpage d is generated. With the yth user browsing operation record Extract respectively , The corresponding record generation period , ,like lie in After that, and and If the time interval between them is less than the second interval threshold, and the current IP address of the same user is F, construct a second user browsing transfer link: It is determined that user F has performed repeated browsing operations on a certain webpage d to be detected.

2. The webpage content integrity detection method based on data analysis according to claim 1, characterized in that, Step S200 includes: Step S201: Extract content feature information of the front-end and back-end web pages to be detected in each first user browsing transfer link. If the similarity between the content feature information of the front-end and back-end web pages to be detected in a certain first user browsing transfer link is less than the similarity threshold, the certain first user browsing transfer link is removed. Step S202: Identify each webpage that appears as a front-end webpage in each of the first user browsing transfer links and set it as a target webpage to be detected; calculate the first browsing transfer rate for each target webpage to be detected. Where u represents the total number of times the target webpage appears as a front-end webpage; m represents the total number of times the target webpage appears in all first user browsing transition links; target webpages with a first browsing transition rate less than the first browsing transition rate threshold are removed.

3. The webpage content integrity detection method based on data analysis according to claim 2, characterized in that, Step S300 includes: Step S301: Assume that a target webpage to be detected is the first user browsing transfer link that appears in the front-end webpage, including... ;in, These respectively represent the 1st, 2nd, and 3rd occurrences of the target webpage as a front-end webpage. There are n first-user browsing transfer links; obtain the user corresponding to each first-user browsing transfer link; Step S302: If a user corresponding to a certain first user browsing transfer link has a corresponding second user browsing transfer link based on the certain target webpage to be detected, and the time period for generating records corresponding to repeated browsing operations in the second user browsing transfer link is less than the time period difference threshold between the time period for generating records corresponding to browsing the certain target webpage in the certain first user browsing transfer link, then the certain first user browsing transfer link is determined to be the target first user browsing transfer link of the certain target webpage to be detected; Step S303: Accumulate the number k of target first-user browsing transfer links contained in all first-user browsing transfer links corresponding to a target webpage to be detected; calculate the second browsing transfer rate for each target webpage to be detected. Remove target pages whose second browsing conversion rate is less than the second browsing conversion rate threshold.

4. The webpage content integrity detection method based on data analysis according to claim 3, characterized in that, Step S400 includes: Step S401: Extract the first browsing transition rate corresponding to all target web pages to be detected. Second browsing conversion rate Calculate the missing probability value for each target webpage to be detected. Sort all target web pages to be detected in descending order of their corresponding missing probability values ​​to obtain the target web page sequence. Step S402: Obtain the fault diagnosis results confirmed by the operation and maintenance personnel for each target webpage in the target webpage sequence, and extract common page features from all target webpages with the same fault diagnosis results and store them. Step S403: Extract page features from each webpage in the real-time input target webpage sequence, and perform similarity matching between the extracted page features and the page features corresponding to various types of faults. When the similarity between the page features of a webpage to be detected and the page features corresponding to a certain type of fault is greater than the similarity threshold, feedback is given to the maintenance personnel to prioritize the investigation of the certain type of fault on the webpage to be detected.

5. A webpage content integrity detection system applying the webpage content integrity detection method based on data analysis according to any one of claims 1-4, characterized in that, The system includes: a user browsing operation record acquisition and processing module, a target webpage screening and processing module, a target webpage calibration and screening module, a target webpage sequence generation module, and an automatic auxiliary detection module; The user browsing operation record acquisition and processing module is used to extract user browsing operation records for each webpage to be detected within a certain period of time, and capture the browsing transfer operations generated by the user between the corresponding webpages to be detected and the repeated browsing operations generated on the corresponding webpages to be detected based on the user browsing operation records, and respectively construct the first user browsing transfer link and the second user browsing transfer link. The target webpage screening and processing module is used to lock the webpages that are suspected of having missing content based on the similarity between the content feature information of the front-end and back-end webpages in each first user browsing transfer link. Let the webpages be the target webpages. The target webpage calibration and screening module is used to receive data from the target webpage screening and processing module and perform the first calibration and screening and the second calibration and screening for each target webpage. The target webpage sequence generation module is used to receive data from the target webpage calibration and screening module, collect all the target webpages to be detected, and generate a target webpage sequence to be detected. The automatic auxiliary detection module is used to input the target webpage sequence to be detected into the management port; collect the fault diagnosis results confirmed by the operation and maintenance personnel for each target webpage to be detected; extract and store the page features corresponding to various faults; and assist the operation and maintenance personnel in screening each target webpage to be detected in the real-time received target webpage sequence based on the stored data.

6. The webpage content integrity detection system according to claim 5, characterized in that, The user browsing operation record collection and processing module includes a first user browsing transfer link construction unit and a second user browsing transfer link construction unit; The first user browsing transfer link construction unit is used to capture the browsing transfer operations generated by the user between the corresponding web pages to be detected, and construct the first user browsing transfer link accordingly. The second user browsing transfer link construction unit is used to capture repeated browsing operations generated by the user on the corresponding webpage to be detected, and construct the second user browsing transfer link accordingly.

7. A webpage content integrity detection system according to claim 6, characterized in that, The target webpage calibration and screening module includes a first calibration and screening unit and a second calibration and screening unit. The first calibration and screening unit is used to calculate the first browsing transition rate for each target webpage to be detected; and to perform the first calibration and screening on all target webpages to be detected based on the first browsing transition rate. The second calibration and screening unit is used to calculate the second browsing transition rate for each target webpage to be detected; A second calibration screening is performed on all target web pages to be detected based on the second browsing transfer rate.