Webpage monitoring system and method based on dynamic perception
By extracting and analyzing the multimodal features of web pages and training dynamic perception models with time series analysis technology, the problem that traditional web page monitoring is difficult to deal with dynamic content is solved, and more accurate and secure web page monitoring is achieved.
Patent Information
- Application Number
- CN202510224672.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional static web page monitoring methods are difficult to effectively handle dynamic content, and are prone to false alarms or missed reports, resulting in web page tampering, which may lead to user information leakage and poor user experience.
By obtaining historical web page monitoring images, multimodal features (visual features, text features and structural features), vectorize them, and train the web page dynamic perception model with time series analysis technology to perceive the web page tampering results in real time.
It effectively avoids the problem of web pages being tampered with, causing poor user browsing experience, and increases system security and improves the accuracy and reliability of web page monitoring.
Smart Images

Figure CN120067483A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a web page monitoring system and method based on dynamic perception. Background Art
[0002] Modern web pages usually use technologies such as JavaScript and AJAX to dynamically load content, for example: advertisement carousels, comment area updates, and real-time data displays. These dynamic contents make the state of the web page may change at any time. When users visit the same web page, the content they see may be different due to time, interaction behaviors, or changes in backend data. Although this dynamic nature improves the user experience, it also poses great challenges to web page monitoring. Traditional static web page monitoring methods (such as simple image comparison or text matching) are difficult to effectively handle dynamic content and are prone to false alarms (for example: a normal advertisement carousel may be misjudged as image tampering) or missed reports.
[0003] More seriously, web page tampering may lead to problems such as user information leakage and the inconsistency between the viewed web page content and expectations. Such misguidance will not only damage the user experience but also may have a long-term impact on the reputation and trust of the platform, and the platform risk is relatively high.
[0004] In view of this, there is an urgent need for a web page monitoring system and method based on dynamic perception to at least solve the above deficiencies. Summary of the Invention
[0005] One of the objectives of the present invention is to provide a web page monitoring system and method based on dynamic perception, which characterizes the obtained historical web page monitoring images to obtain multi-modal visual features, text features, and structural features, vectorizes the multi-modal features, and combines time series analysis technology to train a web page dynamic perception model; inputs the user's real-time accessed web page into the web page dynamic perception model to perceive the web page tampering result in real time, avoiding the problem of poor user browsing experience caused by web page tampering and increasing the system security.
[0006] The web page monitoring system based on dynamic perception provided by an embodiment of the present invention includes:
[0007] An image acquisition module, configured to acquire historical web page monitoring images;
[0008] A feature extraction module, configured to extract multi-modal features from the historical web page monitoring images and construct a multi-modal feature vector; the multi-modal features include: visual features, text features, and structural features;
[0009] A training module, configured to train a web page dynamic perception model based on the multi-modal feature vector and in combination with time series analysis technology;
[0010] An output module, configured to input the user's real-time web page access into a web page dynamic perception model to obtain web page monitoring results.
[0011] Preferably, the image acquisition module acquires historical web page monitoring images, including:
[0012] Obtain target records; the target records include: complaint records uploaded by historical users of the web page platform and active upload records of web page administrators;
[0013] Parse the target records to obtain the snapshot moment and the first screen snapshot;
[0014] Summarize the first screen snapshots belonging to the same web page tag and sort them in chronological order of the snapshot moments to obtain a first screen snapshot sequence;
[0015] Obtain the annotation personnel delivery node;
[0016] Package and deliver the first screen snapshot sequence to the annotation personnel delivery node, and assist the annotation personnel in annotating the first screen snapshots in the first screen snapshot sequence;
[0017] When all the first screen snapshot sequences under all web page tags are annotated, the historical web page monitoring images are obtained.
[0018] Preferably, the image acquisition module assists the annotation personnel in annotating the first screen snapshots in the first screen snapshot sequence, including:
[0019] Determine the first screen snapshot that the annotation personnel are currently viewing and use it as the second screen snapshot;
[0020] Determine the third screen snapshot in the first screen snapshot sequence that is similar to the second screen snapshot except for the second screen snapshot;
[0021] Based on the snapshot moment, extract the DOM structure sequence according to the second screen snapshot and the third screen snapshot;
[0022] According to the DOM structure sequence and the target web page DOM structure, obtain the DOM tree structure difference sequence;
[0023] According to the DOM tree structure difference sequence, obtain web page tampering annotation knowledge;
[0024] Visualize the web page tampering annotation knowledge to the annotation personnel.
[0025] Preferably, the image acquisition module obtains web page tampering annotation knowledge according to the DOM tree structure difference sequence, including:
[0026] Parse the DOM tree structure difference to obtain the difference tree nodes and difference branches;
[0027] Obtain the knowledge graph of web page tampering;
[0028] Based on the preset knowledge partitioning conditions and according to the DOM tree structure differences, obtain the locally partitioned graph in the knowledge graph of web page tampering;
[0029] The knowledge partitioning conditions include:
[0030] The partitioning cost is less than or equal to the preset partitioning cost threshold;
[0031] For each differential tree node in the partitioned graph nodes, a graph node with a node correlation degree greater than or equal to the preset node correlation degree threshold can be found for the corresponding differential tree node;
[0032] In the partitioned graph branches, there is a graph branch with a branch correlation degree greater than or equal to the preset branch correlation degree threshold for the differential branch corresponding to the second screen snapshot;
[0033] In the partitioned graph branches, except for the graph branches associated with the differential branches corresponding to the second screen snapshot, the source DOM tree structure differences and the target DOM tree structure differences of the remaining associated differential branches conform to the standard sequence arrangement.
[0034] Preferably, the image acquisition module also performs the following operations:
[0035] Determine the first sub-graph in the local graph according to the DOM tree structure differences corresponding to the second screen snapshot;
[0036] Determine the second sub-graph in the local graph according to the DOM tree structure differences corresponding to the third screen snapshot;
[0037] Determine the first knowledge presentation screen of the web page tampering annotation knowledge corresponding to the first sub-graph;
[0038] Extract graph features based on the preset graph feature template and according to the first sub-graph;
[0039] Determine multiple graph knowledge sorting routes according to the preset graph knowledge sorting logic and according to the graph features;
[0040] Calculate the law perception degree of the graph knowledge sorting route according to the passing order of the graph knowledge sorting route passing through the second sub-graph for the first time and the standard passing order of the second sub-graph;
[0041] Calculate the target coverage rate according to the graph coverage rate of the graph knowledge sorting route passing through the second sub-graph and the preset knowledge weight of the second sub-graph;
[0042] Sum up and calculate the law perception degree and the target coverage rate corresponding to the graph knowledge sorting route to obtain the route screening value;
[0043] Construct a sequence of knowledge presentation screens by following the second sub-graph corresponding to the second knowledge presentation screen passed by the knowledge sorting route with the largest route screening value;
[0044] Insert the first knowledge presentation screen into the head of the sequence of knowledge presentation screens, and cyclically display the knowledge presentation screens according to the preset highlighting rules based on the sequence of knowledge presentation screens.
[0045] Preferably, the image acquisition module also performs the following operations:
[0046] Obtain the puzzled screens of the annotators in the sequence of knowledge presentation screens;
[0047] Determine the third sub-graph corresponding to the puzzled screen and display it to the annotator, and obtain the first line-of-sight trajectory of the annotator in the third sub-graph;
[0048] According to the first line-of-sight trajectory, determine the second line-of-sight trajectory between the target entities of the third sub-graph that are alternately viewed, and use the second line-of-sight trajectory generated most recently from the current time as the third line-of-sight trajectory;
[0049] Determine the graph feature differences of the third sub-graph passed by the second line-of-sight trajectory other than the third line-of-sight trajectory in the third line-of-sight trajectory and the second line-of-sight trajectory;
[0050] Extract the differential description semantics of the graph feature differences and input them into the knowledge annotation model corresponding to the third sub-graph, and mark the annotation knowledge on the third sub-graph passed by the third line-of-sight trajectory.
[0051] The web page monitoring method based on dynamic perception provided by the embodiments of the present invention includes:
[0052] Step 1: Obtain historical web page monitoring images;
[0053] Step 2: Extract multi-modal features from the historical web page monitoring images and construct multi-modal feature vectors; the multi-modal features include: visual features, text features, and structural features;
[0054] Step 3: Train a web page dynamic perception model based on the multi-modal feature vectors and combined with time series analysis techniques;
[0055] Step 4: Input the user's real-time accessed web page into the web page dynamic perception model to obtain the web page monitoring result.
[0056] Preferably, Step 1: Obtain historical web page monitoring images, including:
[0057] Obtain target records; the target records include: complaint records uploaded by historical users of the web page platform and active upload records of web page administrators;
[0058] Parse the target record to obtain the snapshot moment and the first screen snapshot;
[0059] Summarize the first screen snapshots under the same web page tag and sort them in chronological order of the snapshot moments to obtain the first screen snapshot sequence;
[0060] Obtain the annotation personnel delivery node;
[0061] Package and deliver the first screen snapshot sequence to the annotation personnel delivery node, and assist the annotation personnel in annotating the first screen snapshots in the first screen snapshot sequence;
[0062] When the first screen snapshot sequences under all web page tags are all annotated, obtain the historical web page monitoring image.
[0063] Preferably, assisting the annotation personnel in annotating the first screen snapshots in the first screen snapshot sequence includes:
[0064] Determine the first screen snapshot that the annotation personnel is currently viewing and use it as the second screen snapshot;
[0065] Determine the third screen snapshot in the first screen snapshot sequence that is similar to the second screen snapshot except for the second screen snapshot;
[0066] Based on the snapshot moment, extract the DOM structure sequence according to the second screen snapshot and the third screen snapshot;
[0067] According to the DOM structure sequence and the target web page DOM structure, obtain the DOM tree structure difference sequence;
[0068] According to the DOM tree structure difference sequence, obtain the web page tampering annotation knowledge;
[0069] Visualize the web page tampering annotation knowledge to the annotation personnel.
[0070] Preferably, obtaining the web page tampering annotation knowledge according to the DOM tree structure difference sequence includes:
[0071] Parse the DOM tree structure difference to obtain the difference tree nodes and difference branches;
[0072] Obtain the web page tampering knowledge graph;
[0073] Based on the preset knowledge circle-drawing conditions, obtain the locally circled graph in the web page tampering knowledge graph according to the DOM tree structure difference;
[0074] The knowledge circle-drawing conditions include:
[0075] The circle-drawing cost is less than or equal to the preset circle-drawing cost threshold;
[0076] For each differential tree node in the circled graph nodes, a graph node with a node correlation degree greater than or equal to a preset node correlation degree threshold can be found for the corresponding differential tree node;
[0077] There is a graph branch in the circled graph branches with a branch correlation degree greater than or equal to a preset branch correlation degree threshold for the differential branch corresponding to the second screenshot;
[0078] Among the graph branches in the circled graph branches, except for the graph branches associated with the differential branch corresponding to the second screenshot, the source DOM tree structure differences and the target DOM tree structure differences of the remaining associated differential branches conform to the standard sequence arrangement;
[0079] Obtain web page tampering annotation knowledge according to the local graph.
[0080] Preferably, obtaining web page tampering annotation knowledge according to the DOM tree structure difference sequence further includes:
[0081] Determine the first sub-graph in the local graph according to the DOM tree structure difference corresponding to the second screenshot;
[0082] Determine the second sub-graph in the local graph according to the DOM tree structure difference corresponding to the third screenshot;
[0083] Determine the first knowledge presentation screen of the web page tampering annotation knowledge corresponding to the first sub-graph;
[0084] Extract graph features according to the first sub-graph based on a preset graph feature template;
[0085] Determine multiple graph knowledge sorting routes according to a preset graph knowledge sorting logic and according to the graph features;
[0086] Calculate the regularity perception degree of the graph knowledge sorting route according to the passing order of the graph knowledge sorting route passing through the second sub-graph for the first time and the standard passing order of the second sub-graph;
[0087] Calculate the target coverage rate according to the graph coverage rate of the graph knowledge sorting route passing through the second sub-graph and the preset knowledge weight of the second sub-graph;
[0088] Sum and calculate the regularity perception degree and the target coverage rate corresponding to the graph knowledge sorting route to obtain a route screening value;
[0089] Construct a knowledge presentation screen sequence based on the second knowledge presentation screens corresponding to the second sub-graphs passed through in sequence by the graph knowledge sorting route with the largest route screening value;
[0090] Insert the first knowledge presentation screen at the head of the sequence of knowledge presentation screens, and cyclically display the knowledge presentation screens according to the preset highlighting rules based on the sequence of knowledge presentation screens.
[0091] The beneficial effects of the present invention are as follows:
[0092] The present invention characterizes the obtained historical web page monitoring images to obtain multi-modal visual features, text features, and structural features, vectorizes the multi-modal features, and combines time series analysis technology to train a web page dynamic perception model; inputs the user's real-time accessed web page into the web page dynamic perception model to perceive the web page tampering result in real time, avoiding the problem of poor user browsing experience caused by web page tampering and increasing the system security.
[0093] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0094] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0095] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings:
[0096] Figure 1 It is a schematic diagram of a web page monitoring system based on dynamic perception in an embodiment of the present invention;
[0097] Figure 2 It is a schematic diagram of a web page monitoring method based on dynamic perception in an embodiment of the present invention. Detailed Embodiments
[0098] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0099] The embodiment of the present invention provides a web page monitoring system based on dynamic perception, as Figure 1 shown, including:
[0100] An image acquisition module 1 for acquiring historical web page monitoring images; the historical web page monitoring images are: web page images obtained by the web page monitoring platform in the past and marked, and the historical web page monitoring images mark the web page tampering content and tampering types;
[0101] Feature extraction module 2 is used to extract multimodal features from historical web page monitoring images and construct multimodal feature vectors; the multimodal features include: visual features, text features, and structural features; the visual features are visual information extracted from historical web page monitoring images, including: color distribution, layout structure, and the positions and shapes of key elements (such as buttons, pictures); the text features are the text content extracted from historical web page monitoring images; the structural features are the information extracted from the DOM (Document Object Model) structure of the web page corresponding to the historical web page monitoring image, such as: the hierarchical relationship of HTML tags, node attributes, and dynamically loaded content, etc.; the multimodal feature vector is a vector representation converted into a numerical form after processing visual features, text features, and structural features, and can be used as the input of a machine learning model.
[0102] Training module 3 is used to train a web page dynamic perception model based on multimodal feature vectors and combined with time series analysis technology; when training a web page dynamic perception model based on multimodal feature vectors and combined with time series analysis technology, the multimodal feature vectors are input into the YOLOv8 model, and combined with the LSTM network to learn the change relationship between multimodal feature vectors and their corresponding labeled tampering types, so as to obtain a web page dynamic perception model.
[0103] Output module 4 is used to input the user's real-time accessed web page into the web page dynamic perception model to obtain web page monitoring results.
[0104] The working principle and beneficial effects of the above technical solution are as follows:
[0105] In the present invention, the features of the obtained historical web page monitoring images are characterized to obtain multimodal visual features, text features, and structural features, the multimodal features are vectorized, and a web page dynamic perception model is trained in combination with time series analysis technology; the user's real-time accessed web page is input into the web page dynamic perception model to perceive the web page tampering result in real time, avoiding the problem of poor user browsing experience caused by web page tampering and increasing the system security.
[0106] In one embodiment, the image acquisition module acquires historical web page monitoring images, including:
[0107] Obtain target records; the target records include: complaint records uploaded by historical users of the web page platform and active upload records of web page administrators; the web page platform is: the platform that needs web page monitoring; historical users are: users who have made complaints related to web page tampering on the web page platform in history.
[0108] Parse the target records to obtain the snapshot moment and the first screen snapshot; the first screen snapshot is the web page screenshot in the target record, and the snapshot moment is the time of the web page screenshot.
[0109] Summarize the first screen snapshots belonging to the same web page tag and sort them in chronological order of the snapshot time to obtain the first screen snapshot sequence; the web page tag is a mark for classifying or identifying web pages, distinguishing different web pages or different versions of the same web page;
[0110] Obtain the delivery node of the annotator; the delivery node of the annotator is the communication node of the annotator;
[0111] Package and deliver the first screen snapshot sequence to the delivery node of the annotator, and assist the annotator in annotating the first screen snapshots in the first screen snapshot sequence; when assisting, display relevant information to the required annotator to assist them in marking the tampered content or abnormal area in the web page;
[0112] When all the first screen snapshot sequences under all web page tags are annotated, obtain the historical web page monitoring image.
[0113] The working principle and beneficial effects of the above technical solution are as follows:
[0114] The target records include historical web page user complaints and tampered web page images (the first screen snapshots) actively uploaded by web page administrators, and the tampering time (snapshot time); classifying the snapshots belonging to the same web page and displaying them according to time to obtain the first screen snapshot sequence, which is convenient for subsequent annotators to analyze the dynamic change law of web page content;
[0115] Connect to the delivery node of the annotator, package and deliver the first screen snapshot sequence and assist the annotator in annotating. When all the first screen snapshot sequences under all web page tags are annotated, obtain the historical web page monitoring image (training data), and the annotation process is more intelligent and reasonable.
[0116] In one embodiment, the image acquisition module assists the annotator in annotating the first screen snapshots in the first screen snapshot sequence, including:
[0117] Determine the first screen snapshot that the annotator is currently viewing and use it as the second screen snapshot; the second screen snapshot is a specific snapshot selected from the first screen snapshot sequence;
[0118] Determine the third screen snapshot in the first screen snapshot sequence that is similar to the second screen snapshot and except for the second screen snapshot; being similar to the second screen snapshot means that the similarity of the image features with the second screen snapshot is greater than or equal to a preset similarity threshold (for example: 70%);
[0119] Based on the snapshot time, extract the DOM structure sequence according to the second screen snapshot and the third screen snapshot; when extracting, sort the DOM document object model structure (DOM structure) of the web page corresponding to the screen snapshot in ascending order of the snapshot time;
[0120] Obtain the DOM tree structure difference sequence according to the DOM structure sequence and the DOM structure of the target web page; the DOM tree structure difference is the difference between the DOM structure and the DOM structure of the target web page, including: the added, deleted, or modified nodes and their attributes; the DOM structure of the target web page is the DOM structure preset for the web pages corresponding to the second screenshot and the third screenshot.
[0121] Obtain the web page tampering annotation knowledge according to the DOM tree structure difference sequence; the web page tampering annotation knowledge: the specific information about web page tampering obtained by analyzing the DOM tree structure difference, including the tampering type, location, and content.
[0122] Visually display the web page tampering annotation knowledge to the annotators.
[0123] The working principle and beneficial effects of the above technical solution are as follows:
[0124] Since web page tampering is dynamically changing, annotators cannot understand the changing rules of tampering based on a single second screenshot. Moreover, due to the lack of web page-related knowledge in the annotators' reserves, they cannot well discover the tampered content and mark the tampering type. Therefore, annotators need to be assisted.
[0125] During specific assistance, determine the second screenshot that the annotator is viewing. At the same time, search for the third screenshot similar to it in the first screenshot sequence, extract the DOM structures of the screenshots in the order of the snapshot time from early to late to obtain the DOM structure sequence; compare the DOM structures in the DOM structure sequence with the DOM structure of the target web page to obtain the DOM tree structure difference sequence; retrieve the web page tampering annotation knowledge according to the DOM tree structure difference sequence for the annotators to view, making the annotation process more intelligent.
[0126] In one embodiment, the image acquisition module obtains the web page tampering annotation knowledge according to the DOM tree structure difference sequence, including:
[0127] Parse the DOM tree structure difference to obtain the difference tree nodes and difference branches; the difference tree nodes are the different DOM tree nodes between the DOM structure and the DOM structure of the target web page, including: added nodes (for example, a <script> tag is inserted), deleted nodes (for example, a tag is removed), and modified nodes (for example, a The text content of the label is changed); The differential branch is: the subtree or node path that changes in the DOM tree structure difference, which is used to describe the scope and impact of the tampering behavior. For example, an entire label and its child nodes are deleted;
[0128] Obtain the web page tampering knowledge graph; The web page tampering knowledge graph is: a tampering knowledge base represented in a graph structure, where the graph nodes represent entities related to tampering (such as: DOM nodes, tampering types), and the edges represent the relationships between entities (such as: code injection causes the <script> label to change);
[0129] Based on the preset knowledge circle-drawing conditions and according to the DOM tree structure difference, obtain the locally circled graph in the web page tampering knowledge graph; The knowledge circle-drawing conditions are: the constraint conditions for selecting the local graph that adapts to explaining the DOM tree structure difference;
[0130] The knowledge circle-drawing conditions include:
[0131] The circle-drawing cost is less than or equal to the preset circle-drawing cost threshold; The circle-drawing cost is: the access cost (computing resources and time cost) of all the graph nodes and graph branches in the circle-drawing; The preset circle-drawing cost threshold is set manually in advance;
[0132] For each differential tree node in the circled graph nodes, a graph node with a node correlation degree greater than or equal to the preset node correlation degree threshold with the corresponding differential tree node can be found; The node correlation degree refers to the correlation between the graph nodes in the web page tampering knowledge graph and the differential tree nodes. For example, the correlation degree between the graph node of "code injection" and the differential tree node of the "<script> label" is 0.9. The correlation degrees of the differential tree nodes and the DOM tree nodes are defined manually in advance; The preset node correlation degree threshold is set manually, for example: 0.85;
[0133] Among the circled graph branches, there is a graph branch with a branch correlation degree greater than or equal to the preset branch correlation degree threshold that corresponds to the second screenshot; The branch correlation degree is: the semantic similarity between the semantics of the graph branch and the semantics of the differential branch. For example: 0.7; The preset branch correlation degree threshold is set manually, for example: 0.8;
[0134] Among the circled map branches, except for the map branches associated with the differential branches corresponding to the second screenshot, the source DOM tree structure differences of the differential branches associated with the remaining map branches and the target DOM tree structure differences corresponding to the second screenshot conform to the standard sequence arrangement. The source DOM tree structure difference is: the DOM tree structure difference of the parsed source of the differential branches associated with the remaining map branches; the target DOM tree structure difference is the DOM tree structure difference corresponding to the second screenshot; conforming to the standard sequence arrangement means that: after the source DOM tree structure difference and the target DOM tree structure difference are sorted according to the snapshot time of the second screenshot or the third screenshot extracted correspondingly, they are local subsequences of the DOM tree structure difference sequence;
[0135] According to the local map, obtain the knowledge of web page tampering annotation.
[0136] The working principle and beneficial effects of the above technical solution are as follows:
[0137] When obtaining the knowledge of web page tampering annotation, all relevant knowledge is displayed to the user, and the display efficiency is low. Therefore, knowledge circling conditions are introduced, and according to the DOM tree structure difference sequence, local maps are circled, and the corresponding map knowledge of the local maps is used as the knowledge of web page tampering annotation; specifically, the knowledge circling conditions are as follows:
[0138] Condition 1: The circling cost is less than the circling cost threshold; since accessing the web page tampering knowledge graph consumes computing resources and time, too high a circling cost will lead to untimely assistance, and setting this condition improves the timeliness of auxiliary presentation;
[0139] Condition 2: For each differential tree node in the circled map nodes, a map node with a node association degree greater than or equal to the preset node association degree threshold can be found corresponding to the differential tree node; the differential tree node is the DOM tree node different between each DOM structure in the DOM tree structure difference sequence and the DOM structure of the target web page, specifically all DOM tree nodes causing the web page tampering phenomenon of the second screenshot and the third screenshot, which is the basic unit for analyzing the tampering behavior and helps to locate the specific location of the tampering. Therefore, it is necessary to obtain the map nodes corresponding to each differential tree node to improve the comprehensiveness and continuity of the web page tampering annotation knowledge;
[0140] Condition 3: There is a map branch in the circled map branches whose branch association degree with the differential branch corresponding to the second screenshot is greater than or equal to the preset branch association degree threshold; since the map branches contain more detailed information, and the priorities of presenting the detailed information are different, for example: the map branch of the differential branch corresponding to the second screenshot that the annotator is currently annotating must be presented with the highest priority;
[0141] Condition 4: Among the circled graph branches, except for the graph branches associated with the differential branches corresponding to the second screenshot, the source DOM tree structure differences and the target DOM tree structure differences of the remaining associated differential branches conform to the standard sequence arrangement; since it is dynamically perceived tampering information, the graph branches associated with the differential branches corresponding to the third screenshot closer to the second screenshot in time are circled with higher priority. Therefore, it is constrained that the source DOM tree structure differences and the target DOM tree structure differences are local subsequences of the DOM tree structure difference sequence after being sorted according to the snapshot moments of the second screenshot or the third screenshot corresponding to the extraction. In the case of limited circled resources, more web page tampering annotation knowledge with higher relevance is presented, improving the utilization rate of circled resources.
[0142] In one embodiment, the image acquisition module further performs the following operations:
[0143] Determine the first sub-graph in the local graph according to the DOM tree structure difference corresponding to the second screenshot; the first sub-graph is: the part of the local graph corresponding to the knowledge causing the DOM tree structure difference corresponding to the second screenshot;
[0144] Determine the second sub-graph in the local graph according to the DOM tree structure difference corresponding to the third screenshot; the second sub-graph is: the part of the local graph corresponding to the knowledge causing the DOM tree structure difference corresponding to the third screenshot;
[0145] Determine the first knowledge presentation screen of the web page tampering annotation knowledge corresponding to the first sub-graph; the first knowledge presentation screen is: the presentation screen area when the knowledge corresponding to the first sub-graph is visually presented;
[0146] Extract graph features based on a preset graph feature template and according to the first sub-graph; the graph features are: the edge graph entities and edge graph relationships of the first sub-graph;
[0147] Determine multiple graph knowledge sorting routes according to a preset graph knowledge sorting logic and according to the graph features; the preset graph knowledge sorting logic is set by personnel with web page tampering knowledge teaching experience and includes: which part of the knowledge to learn first and which part of the knowledge to learn next;
[0148] Calculate the law perception degree of the graph knowledge sorting route according to the passing order when the graph knowledge sorting route first passes through the second sub-graph and the standard passing order of the second sub-graph; the standard passing order of the second sub-graph is determined according to the sequence position interval between the DOM tree structure difference corresponding to the second sub-graph in the DOM tree structure difference sequence and the DOM tree structure difference corresponding to the second sub-graph. For example, if the interval is 1, the standard passing order is 1; the law perception degree is the sum of the absolute values of the passing order of each second sub-graph minus its corresponding standard passing order. For example, 16;
[0149] Calculate the target coverage rate according to the coverage rate of the second sub-graph and the preset knowledge weight of the second sub-graph along the graph knowledge sorting route; the preset knowledge weight of the second sub-graph is the reciprocal of the standard passing order value of the second sub-graph.
[0150] Sum up the law perception degree and the target coverage rate corresponding to the graph knowledge sorting route to obtain the route screening value.
[0151] Construct a knowledge presentation screen sequence based on the second knowledge presentation screens corresponding to the second sub-graphs passed in sequence by the graph knowledge sorting route with the largest route screening value.
[0152] Insert the first knowledge presentation screen at the head of the sequence of the knowledge presentation screen sequence, and based on the preset highlighting rules, display the knowledge presentation screens in a loop according to the knowledge presentation screen sequence. Displaying the knowledge presentation screens in a loop according to the preset highlighting rules is as follows: when looping to a certain knowledge presentation screen in the knowledge presentation screen sequence, highlight it (for example: highlight the knowledge background color).
[0153] The working principle and beneficial effects of the above technical solution are as follows:
[0154] The web page tampering annotation knowledge presented to annotators has a viewing logic. If all the web page tampering annotation knowledge is directly shown to the annotators, the learning effect is not good. Therefore, first determine the first sub-graph corresponding to the second screen snapshot to be viewed first. The knowledge of the first sub-graph has the highest relevance to the second screen snapshot and needs to be viewed first to determine the first knowledge presentation screen. Then, extract the graph features (edge graph entities and edge graph relationships) of the first sub-graph in the local graph. Based on the graph knowledge sorting logic set by personnel with web page tampering knowledge teaching experience, determine multiple graph knowledge sorting routes. The standard passing order of the second sub-graph is the ideal access order that can realize the comparison of dynamic web page tampering knowledge. However, the coherence of knowledge understanding also needs to be considered. Therefore, according to the passing order of the graph knowledge sorting route passing through the second sub-graph for the first time and the standard passing order of the second sub-graph, calculate the regularity perception degree of the graph knowledge sorting route. The greater the regularity perception degree, the more preferentially the corresponding graph knowledge sorting route is selected. In addition, the graph coverage rate of the graph knowledge sorting route passing through the second sub-graph also needs to be considered. The knowledge weights of different second sub-graphs are also different. Assign the knowledge weight to the graph coverage rate to obtain the target coverage rate. Sum up the regularity perception degree and the target coverage rate corresponding to the graph knowledge sorting route to obtain the route screening value. Sort the second knowledge presentation screens corresponding to the second sub-graphs presented in the order of the second sub-graphs passed by the graph knowledge sorting route with the largest route screening value to form a knowledge presentation screen sequence. Insert the first knowledge presentation screen at the head of the sequence and display it in a loop. The knowledge screen displayed in the loop is highlighted, so that the annotators can view the web page tampering knowledge in a reasonable logical order, improve the understanding effect, and further improve the annotation quality of the auxiliary annotation.
[0155] In one embodiment, when displaying the knowledge presentation screens in a loop according to the knowledge presentation screen sequence, it further includes:
[0156] Obtain the puzzled screens of the annotators in the knowledge presentation screen sequence; the puzzled screens are the knowledge presentation screens that the annotators show puzzled expressions when viewing;
[0157] Determine the third sub-graph corresponding to the puzzled screen and show it to the annotators, and obtain the first line-of-sight trajectory of the annotators in the third sub-graph;
[0158] According to the first line-of-sight trajectory, determine the second line-of-sight trajectory between the target entities of the third sub-graph to be alternately viewed, and use the second line-of-sight trajectory generated closest to the current moment as the third line-of-sight trajectory; alternately viewing the target entities of the third sub-graph means that there is more than one line-of-sight trajectory (the second line-of-sight trajectory) connecting the target entities of the third sub-graph being viewed;
[0159] Determine the atlas feature differences of the third sub-atlas passed by the third line-of-sight trajectory and the second line-of-sight trajectory other than the third line-of-sight trajectory; the atlas feature differences are: different third sub-atlas entities and different third sub-branches;
[0160] Extract the differential description semantics of the atlas feature differences and input them into the knowledge annotation model corresponding to the third sub-atlas, and mark the corresponding annotation knowledge on the third sub-atlas passed by the third line-of-sight trajectory. The differential description semantics describe the atlas differences between the third sub-atlas entity where the difference starts and the third sub-atlas entity where the difference ends; the knowledge annotation model corresponding to the third sub-atlas is an AI model that automatically annotates the third sub-atlas trained based on the construction principle explanation of the third sub-atlas by the constructor of the third sub-atlas.
[0161] The working principle and beneficial effects of the above technical solution are:
[0162] When the annotator views the knowledge presentation screen in the order of the understanding logic, they may be puzzled by the content items therein. Therefore, based on the facial expression analysis technology, determine the puzzled screen of the annotator and extract the third sub-atlas corresponding to the puzzled screen; the third sub-atlas well sorts out the knowledge outline corresponding to the puzzled screen of the annotator, which is convenient for quickly locating the knowledge that the annotator does not understand later;
[0163] Next, the annotator will view the third sub-atlas. Based on the line-of-sight extraction technology, obtain the first line-of-sight trajectory of the annotator in the third sub-atlas and determine that there is more than one line-of-sight trajectory connecting the target entities of the third sub-atlas being viewed, that is, the second line-of-sight trajectory; Alternate viewing means that the annotator views the area of the third sub-atlas corresponding to the second line-of-sight trajectory more than once, indicating that the annotator has repeated viewing behavior and needs to be annotated;
[0164] Determine the atlas feature differences between the third line-of-sight trajectory and the remaining second line-of-sight trajectories. The atlas feature differences represent the atlas entity differences and atlas branch differences of the repeatedly viewed third sub-atlas, including the differential knowledge outline information of repeated comparison. The more likely it is to need annotation. Therefore, introduce the knowledge annotation model corresponding to the third sub-atlas, input the differential description semantics into the model, and mark the output annotation in real time on the third sub-atlas passed by the latest viewed third line-of-sight trajectory, realizing the automatic annotation of the knowledge of the puzzled atlas of the annotator and further assisting the annotator in understanding.
[0165] The embodiment of the present invention provides a web page monitoring method based on dynamic perception, as Figure 2 shown, including:
[0166] Step 1: Obtain historical web page monitoring images;
[0167] Step 2: Extract multi-modal features from the historical web page monitoring images and construct a multi-modal feature vector; the multi-modal features include: visual features, text features, and structural features;
[0168] Step 3: Based on the multi-modal feature vector and combined with time series analysis techniques, train a web page dynamic perception model;
[0169] Step 4: Input the user's real-time accessed web page into the web page dynamic perception model to obtain the web page monitoring result.
[0170] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A webpage monitoring system based on dynamic perception, characterized in that: include: An image acquisition module is used to acquire historical web page monitoring images; A feature extraction module is used to extract multimodal features from historical web page monitoring images and construct a multimodal feature vector; Multimodal features include: visual features, text features, and structural features; The training module is used to train the web page dynamic perception model based on multimodal feature vectors and time series analysis technology; The output module is used to input the user's real-time visited web pages into the web page dynamic perception model to obtain the web page monitoring results.
2. The webpage monitoring system based on dynamic perception as claimed in claim 1, characterized in that: The image acquisition module acquires historical web monitoring images, including: Obtain target records; target records include: complaint records uploaded by historical users of the web platform and records actively uploaded by web administrators; Parse the target record to obtain the snapshot time and the first screen snapshot; The first screen snapshots belonging to the same web page tag are aggregated and sorted in chronological order of the snapshot moments to obtain a first screen snapshot sequence; Get the labeler delivery node; Packing the first screenshot sequence and delivering it to the labeler delivery node, and assisting the labeler to label the first screenshot in the first screenshot sequence; When the first screen snapshot sequences under all web page tags are marked, a historical web page monitoring image is obtained.
3. The webpage monitoring system based on dynamic perception as claimed in claim 2, characterized in that: The image acquisition module assists the annotator in annotating the first screenshot in the first screenshot sequence, including: Determine the first screen snapshot currently being viewed by the annotator and use it as the second screen snapshot; determining a third screenshot in the first screenshot sequence that is different from the second screenshot and is similar to the second screenshot; Extracting a DOM structure sequence based on the snapshot time according to the second screenshot and the third screenshot; According to the DOM structure sequence and the DOM structure of the target web page, a DOM tree structure difference sequence is obtained; According to the DOM tree structure difference sequence, the web page tampering annotation knowledge is obtained; Visually display web page tampering annotation knowledge to annotators.
4. The webpage monitoring system based on dynamic perception as claimed in claim 3, characterized in that: The image acquisition module obtains web page tampering annotation knowledge based on the DOM tree structure difference sequence, including: Analyze the DOM tree structure differences and obtain the difference tree nodes and difference branches; Obtain the knowledge graph of web page tampering; Based on the preset knowledge demarcation conditions and the differences in DOM tree structures, obtain the demarcated partial graph in the webpage tampering knowledge graph; Knowledge zoning conditions include: The demarcation cost is less than or equal to the preset demarcation cost threshold; Each difference tree node in the circled graph node can find a graph node whose node correlation with the corresponding difference tree node is greater than or equal to a preset node correlation threshold; Among the circled atlas branches, there is an atlas branch whose branch correlation degree with the difference branch corresponding to the second screenshot is greater than or equal to a preset branch correlation degree threshold; Among the circled graph branches, except for the graph branch associated with the difference branch corresponding to the second screenshot, the source DOM tree structure differences of the difference branches associated with the remaining graph branches and the target DOM tree structure differences corresponding to the second screenshot conform to the standard sequence arrangement; Based on the local graph, the web page tampering annotation knowledge is obtained.
5. The webpage monitoring system based on dynamic perception as claimed in claim 4, characterized in that: The image acquisition module also performs the following operations: Determine a first sub-graph in the local graph according to a DOM tree structure difference corresponding to the second screenshot; Determine a second sub-graph in the local graph according to a DOM tree structure difference corresponding to the third screenshot; Determine a first knowledge presentation screen of webpage tampering annotation knowledge corresponding to the first sub-graph; Extracting graph features based on a preset graph characterization template and the first sub-graph; According to the preset graph knowledge combing logic and graph features, multiple graph knowledge combing routes are determined; According to the order in which the graph knowledge combing route passes through the second sub-graph for the first time and the standard order in which the graph knowledge combing route passes through the second sub-graph, the regularity perception of the graph knowledge combing route is calculated; According to the graph knowledge, the route is combed through the graph coverage of the second sub-graph and the knowledge weight preset in the second sub-graph, and the target coverage is calculated; Sum and calculate the regularity perception and target coverage of the route sorting graph knowledge to obtain the route screening value; Based on the graph knowledge with the largest route screening value, the second knowledge presentation screen corresponding to the second sub-graph that the route passes through in sequence is sorted out to construct a knowledge presentation screen sequence; The first knowledge presentation screen is inserted into the sequence head of the knowledge presentation screen sequence, and based on a preset highlighting rule, the knowledge presentation screens are cyclically displayed according to the knowledge presentation screen sequence.
6. A webpage monitoring method based on dynamic perception, characterized in that: include: Step 1: Obtain historical web monitoring images; Step 2: Extract multimodal features from historical web page monitoring images and construct multimodal feature vectors; Multimodal features include: visual features, text features, and structural features; Step 3: Based on multimodal feature vectors and combined with time series analysis technology, train a web page dynamic perception model; Step 4: Input the user's real-time visited web pages into the web page dynamic perception model to obtain the web page monitoring results.
7. The webpage monitoring method based on dynamic perception according to claim 6, characterized in that: Step 1: Obtain historical web monitoring images, including: Obtain target records; target records include: complaint records uploaded by historical users of the web platform and records actively uploaded by web administrators; Parse the target record to obtain the snapshot time and the first screen snapshot; The first screen snapshots belonging to the same web page tag are aggregated and sorted in chronological order of the snapshot moments to obtain a first screen snapshot sequence; Get the labeler delivery node; Packing the first screenshot sequence and delivering it to the labeler delivery node, and assisting the labeler to label the first screenshot in the first screenshot sequence; When the first screen snapshot sequences under all web page tags are marked, a historical web page monitoring image is obtained.
8. The webpage monitoring method based on dynamic perception according to claim 7, characterized in that: The auxiliary annotator annotates the first screenshot in the first screenshot sequence, including: Determine the first screen snapshot currently being viewed by the annotator and use it as the second screen snapshot; determining a third screenshot in the first screenshot sequence that is different from the second screenshot and is similar to the second screenshot; Extracting a DOM structure sequence based on the snapshot time according to the second screenshot and the third screenshot; According to the DOM structure sequence and the DOM structure of the target web page, a DOM tree structure difference sequence is obtained; According to the DOM tree structure difference sequence, the web page tampering annotation knowledge is obtained; Visually display web page tampering annotation knowledge to annotators.
9. The webpage monitoring method based on dynamic perception according to claim 8, characterized in that: According to the DOM tree structure difference sequence, the webpage tampering annotation knowledge is obtained, including: Analyze the DOM tree structure differences and obtain the difference tree nodes and difference branches; Obtain the knowledge graph of web page tampering; Based on the preset knowledge demarcation conditions and the differences in DOM tree structures, obtain the demarcated partial graph in the webpage tampering knowledge graph; Knowledge zoning conditions include: The demarcation cost is less than or equal to the preset demarcation cost threshold; Each difference tree node in the circled graph node can find a graph node whose node correlation with the corresponding difference tree node is greater than or equal to a preset node correlation threshold; Among the circled atlas branches, there is an atlas branch whose branch correlation degree with the difference branch corresponding to the second screenshot is greater than or equal to a preset branch correlation degree threshold; Among the circled graph branches, except for the graph branch associated with the difference branch corresponding to the second screenshot, the source DOM tree structure differences of the difference branches associated with the remaining graph branches and the target DOM tree structure differences corresponding to the second screenshot conform to the standard sequence arrangement; Based on the local graph, the web page tampering annotation knowledge is obtained.
10. The webpage monitoring method based on dynamic perception according to claim 9, characterized in that: According to the DOM tree structure difference sequence, the webpage tampering annotation knowledge is obtained, which also includes: Determine a first sub-graph in the local graph according to a DOM tree structure difference corresponding to the second screenshot; Determine a second sub-graph in the local graph according to a DOM tree structure difference corresponding to the third screenshot; Determine a first knowledge presentation screen of webpage tampering annotation knowledge corresponding to the first sub-graph; Extracting graph features based on a preset graph characterization template and the first sub-graph; According to the preset graph knowledge combing logic and graph features, multiple graph knowledge combing routes are determined; According to the order in which the graph knowledge combing route passes through the second sub-graph for the first time and the standard order in which the graph knowledge combing route passes through the second sub-graph, the regularity perception of the graph knowledge combing route is calculated; According to the graph knowledge, the route is combed through the graph coverage of the second sub-graph and the knowledge weight preset in the second sub-graph, and the target coverage is calculated; Sum and calculate the regularity perception and target coverage of the route sorting graph knowledge to obtain the route screening value; Based on the graph knowledge with the largest route screening value, the second knowledge presentation screen corresponding to the second sub-graph that the route passes through in sequence is sorted out to construct a knowledge presentation screen sequence; The first knowledge presentation screen is inserted into the sequence head of the knowledge presentation screen sequence, and based on a preset highlighting rule, the knowledge presentation screens are cyclically displayed according to the knowledge presentation screen sequence.
Citation Information
Patent Citations
Hijacked website detection method and system
CN107911360A
Text and image-based multi-modal harmful link identification
CN114662033A
Pre-training method and device for multi-task model of webpage and electronic equipment
CN116049597A
Video similarity detection method, apparatus, and device
US20220172476A1