Web page intelligent sensing method
Through the intelligent perception method of web pages, AI and multimodal vision technology are used to realize intelligent identification and dynamic routing adjustment of web page elements, solving the problems of inefficiency and high cost caused by page changes in traditional web development, and improving response speed and user experience.
Patent Information
- Application Number
- CN202510048249.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-13
AI Technical Summary
In traditional web development, testing and operations, in the face of page changes or function iterations, it is often necessary to manually adjust the operation process or redevelop adaptation scripts, which are inefficient and expensive.
A web page intelligent perception method is adopted to realize intelligent identification and dynamic routing adjustment of web page elements through AI adaptive intelligent analysis, multimodal visual perception and large-modal dynamic routing strategies.
It improves the response speed to business changes, enhances the ability to identify changes in web page elements, optimizes user experience, reduces operating costs, reduces the need for manual intervention, and improves the overall intelligence level of the system.
Smart Images

Figure CN119996284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent arrangement and preparation of Web pages, and in particular to a method for intelligent perception of Web pages. Background Art
[0002] With the rapid development of Internet technology, Web technology has become the main platform for people to obtain information and interact. Web intelligent perception technology, as an emerging technology, aims to locate and automatically execute website elements by analyzing Web page structure data. At present, the development of big data and machine learning technology has provided strong technical support for Web intelligent perception.
[0003] With the rapid development of Internet technology, the complexity and dynamism of Web pages continue to increase. In traditional Web development, testing and operation, when faced with page changes or function iterations, it is often necessary to manually adjust the operating procedures or redevelop adaptation scripts, which is inefficient and costly.
[0004] Therefore, in order to address the above-mentioned problems in traditional Web development, testing and operation, when faced with page changes or function iterations, it is often necessary to manually adjust the operating procedures or redevelop adaptation scripts, which is inefficient and costly. A Web page intelligent perception method can be designed. Summary of the invention
[0005] In order to overcome the problems of inefficiency and high cost in traditional Web development, testing and operation, when faced with page changes or function iterations, it is often necessary to manually adjust the operating procedures or redevelop adaptation scripts.
[0006] The technical solution of the present invention is: a method for intelligent perception of a Web page, comprising the following steps:
[0007] Step 1: AI Adaptive Intelligent Analysis
[0008] 1.1 Get the type, content and attributes of the Dom element of the current Web page;
[0009] 1.2 Get the top node of the DOM tree of the page and perform a deep traversal;
[0010] 1.3 Multi-dimensional intelligent extraction of page elements;
[0011] 1.4 Preprocess the collected page data;
[0012] 1.5 Use AI adaptive algorithms to intelligently analyze the collected data;
[0013] Step 2: Large Model Dynamic Routing Strategy
[0014] 2.1 Use the collected data for cleaning;
[0015] 2.2 Generate target task pairs and perform data enhancement;
[0016] 2.3 Use the Embedding vector model to generate text embedding for the target text content;
[0017] 2.4 Fine-tune the large language model for Lora tasks to improve the model's routing capabilities for target tasks;
[0018] 2.5 Embed fine-tuned large language models in AI workflows;
[0019] 2.6 Accept user input and perform vector conversion and semantic analysis;
[0020] 2.7 Realize dynamic routing identification and logic control of multiple Web pages through automated workflow;
[0021] Step 3: Multimodal visual perception
[0022] 3.1 Input the visual features, text embeddings, and structured features extracted in step 1 into the multimodal visual model;
[0023] 3.2 Combine visual and text features and perform deep fusion through multimodal models;
[0024] 3.3 Collecting target task visual data in the automated perception process;
[0025] 3.4 Using multimodal visual models to calculate the similarity of Web elements;
[0026] 3.5 Output the calculation results and the coordinate location of the elements.
[0027] Preferably, the collected data is intelligently analyzed using an AI adaptive algorithm, including full strategy matching, Dom content matching, Dom selector matching, Dom tree-based Path matching, and analyzing the path of the target element in the DOM tree.
[0028] As a preferred embodiment, the full strategy matching is to intelligently select the matching strategy that best suits the current page structure by deeply analyzing the overall layout of the web page and the relationship between elements.
[0029] Preferably, Dom content matching is based on the features of element content, and the algorithm can perform intelligent identification and matching.
[0030] As a preference, Dom selector matching leverages the power of CSS selectors, and the algorithm can accurately locate specific elements on the page.
[0031] Preferably, the Path matching based on the Dom tree is performed by analyzing the path of the element in the DOM tree.
[0032] Preferably, the layout and display mode of the elements can be intelligently determined by analyzing the path algorithm of the elements in the DOM tree.
[0033] Preferably, multimodal visual perception includes page screenshot collection and multimodal image algorithm, the page screenshot collection is to collect the image of the page in real time, and the multimodal image algorithm is to combine image recognition technology and machine learning algorithm.
[0034] As a preferred option, the large model dynamic routing strategy is large model data fine-tuning technology, AI large model dynamic routing algorithm and AI model monitoring and behavior prediction. The large model data fine-tuning technology trains the AI model to identify and process Web dynamic page elements. When the AI large model dynamic routing algorithm is running, the algorithm dynamically adjusts the element matching strategy according to the output of the AI model. The AI model monitoring and behavior prediction algorithm monitors the usage of page elements and user behavior, and can dynamically adjust the routing to optimize the user experience.
[0035] The beneficial effects of the present invention are as follows: through the AI adaptive intelligent analysis algorithm, the system can quickly adapt to system upgrades and functional changes without manually rebuilding the process, thereby significantly improving the response speed to business changes; the combination of multimodal visual perception technology and AI algorithm improves the ability to recognize changes in Web page elements, allowing the system to more accurately capture and respond to changes in interface elements; the AI large model dynamic routing strategy can dynamically adjust the routing according to user behavior and the usage of page elements, optimize user experience, and improve business efficiency; reduce the repeated configuration required due to system changes, thereby reducing the waste of time and resources and reducing the operating costs of the enterprise; the system can automatically Learning and adaptive capabilities can automatically adjust to suit different web page designs and structures, reducing the need for human intervention; technical solutions support fine-tuning of business processes, making the system more flexible and able to quickly adapt to changes in market and business needs; through precise element matching and path recognition, the system can more accurately locate and operate web page elements, reducing the possibility of erroneous operations; integrated AI technology improves the overall intelligence level of the system, allowing the system to handle complex web page elements and user interactions more intelligently; by predicting and adapting to user behavior, the system can operate more stably and reliably, reducing failures caused by human errors or system inadaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 What is shown is a schematic diagram of a Web page intelligent perception process of a Web page intelligent perception method of the present invention;
[0037] Figure 2What is shown is a schematic diagram of the intelligent perception principle of a Web page in a Web page intelligent perception method of the present invention. DETAILED DESCRIPTION
[0038] The present invention will be further described below in conjunction with the embodiments.
[0039] See also Figure 1-2 , obtain the type, content and attributes of the Dom element of the current Web page; obtain the top-level node of the DOM tree of the page and perform a deep traversal; intelligently extract page elements in multiple dimensions; pre-process the collected page data; use the AI adaptation algorithm to intelligently analyze the collected data; the method of dynamically selecting the best matching strategy by globally analyzing the page layout and the relationship between elements significantly improves the matching efficiency and accuracy; based on the characteristics of the element content, the algorithm can perform intelligent identification and matching. Even in the absence of a clear selector, the algorithm can identify and match elements through the semantic information of the content; using the powerful function of the CSS selector, the algorithm can accurately locate specific elements on the page. This method is particularly effective when dealing with elements with clear CSS classes or IDs; by combining DOM path parsing with neighboring element relationship analysis, the precise positioning and layout optimization of deeply nested elements are achieved. This multi-dimensional analysis strategy is unique; comprehensively using content, selectors, paths and neighboring element matching strategies, a multi-level, multi-dimensional intelligent matching framework is constructed to adapt to complex page structures; using the collected data for cleaning; generating target task pairs and performing data enhancement; Use the Embedding vector model to generate text embeddings for the target text content; fine-tune the large language model for Lora tasks to improve the model's routing capabilities for the target tasks; embed the fine-tuned large language model in the AI workflow, use the AI large model to monitor the page status in real time, dynamically adjust the page element matching strategy, and ensure that the system can adapt to changes in complex scenarios; access user input and perform vector conversion and semantic analysis; model user behavior through machine learning, dynamically optimize the layout and routing strategy of page elements, and the dynamic adjustment mechanism with user experience as the core can significantly improve the interaction efficiency; input the extracted visual features, text embeddings, and structured features into the multimodal visual model; deeply integrate traditional text analysis with image recognition technology through real-time screenshot collection and multimodal analysis to provide more comprehensive information support for page element matching; collect target task visual data in the automated perception process; based on the combination of image recognition and machine learning, it is possible to extract and analyze elements from page screenshots. This method of directly extracting page structure from visual data is an important technical breakthrough, especially for pages with dynamic or no fixed structure; output calculation results and element coordinate positioning.
[0040] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of those skilled in the art without departing from the purpose of the present invention.
Claims
1. A method for intelligent perception of a web page, comprising the following steps: Step 1: AI Adaptive Intelligent Analysis 1.1 Get the type, content and attributes of the Dom element of the current Web page; 1.2 Get the top node of the DOM tree of the page and perform a deep traversal; 1.3 Multi-dimensional intelligent extraction of page elements; 1.4 Preprocess the collected page data; 1.5 Use AI adaptive algorithms to intelligently analyze the collected data; Step 2: Large Model Dynamic Routing Strategy 2.1 Use the collected data for cleaning; 2.2 Generate target task pairs and perform data enhancement; 2.3 Use the Embedding vector model to generate text embedding for the target text content; 2.4 Fine-tune the large language model for Lora tasks to improve the model's routing capabilities for target tasks; 2.5 Embed fine-tuned large language models in AI workflows; 2.6 Accept user input and perform vector conversion and semantic analysis; 2.7 Realize dynamic routing identification and logic control of multiple Web pages through automated workflow; Step 3: Multimodal visual perception 3.1 Input the visual features, text embeddings, and structured features extracted in step 1 into the multimodal visual model; 3.2 Combine visual and text features and perform deep fusion through multimodal models; 3.3 Collecting target task visual data in the automated perception process; 3.4 Using multimodal visual models to calculate the similarity of Web elements; 3.5 Output the calculation results and the coordinate location of the elements.
2. The method for intelligent perception of web pages according to claim 1, characterized in that: The collected data is intelligently analyzed using AI adaptive algorithms, including full strategy matching, Dom content matching, Dom selector matching, Dom tree-based Path matching, and analyzing the path of the target element in the DOM tree.
3. The method for intelligent perception of web pages according to claim 1, characterized in that: Full strategy matching can intelligently select the matching strategy that best suits the current page structure by deeply analyzing the overall layout of the web page and the relationships between elements.
4. The method for intelligent perception of web pages according to claim 1, characterized in that: Dom content matching is based on the characteristics of element content, and the algorithm can perform intelligent identification and matching.
5. The method for intelligent perception of web pages according to claim 1, characterized in that: Dom selector matching uses the power of CSS selectors to accurately locate specific elements on the page.
6. The method for intelligent perception of web pages according to claim 1, characterized in that: Path matching based on the Dom tree is done by analyzing the path of the element in the DOM tree.
7. The method for intelligent perception of Web pages according to claim 1, characterized in that: By analyzing the path algorithm of elements in the DOM tree, the layout and display mode of elements can be intelligently determined.
8. The method for intelligent perception of Web pages according to claim 1, characterized in that: Multimodal visual perception includes page screenshot collection and multimodal image algorithm. Page screenshot collection is to collect images of the page in real time, and multimodal image algorithm is to combine image recognition technology and machine learning algorithm.
9. The method for intelligent perception of Web pages according to claim 1, characterized in that: The big model dynamic routing strategy includes big model data fine-tuning technology, AI big model dynamic routing algorithm and AI model monitoring and behavior prediction. Big model data fine-tuning technology trains AI models to identify and process Web dynamic page elements. When the AI big model dynamic routing algorithm is running, the algorithm dynamically adjusts the element matching strategy according to the output of the AI model. The AI model monitoring and behavior prediction algorithm monitors the usage of page elements and user behavior, and can dynamically adjust the routing to optimize the user experience.
Citation Information
Patent Citations
Supply chain monitoring system based on intelligent perception and optimal transmission model
CN114500604A
Multimedia page generation method and device, equipment, medium and program product
CN117011875A
Intelligent digital human control method and device, electronic equipment and storage medium
CN118797101A
Large model driven Web task automatic execution method and system
CN119248379A
Cited By
Web business process intelligent automatic execution system based on deep learning
CN120873313A
Web interface element identification method, system, equipment and medium
CN120894792A
Web-RPA script migration method based on large language model
CN120973416A