Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "XPath" patented technology

XPath (XML Path Language) is a query language for selecting nodes from an XML document. In addition, XPath may be used to compute values (e.g., strings, numbers, or Boolean values) from the content of an XML document. XPath was defined by the World Wide Web Consortium (W3C).

Intelligent element positioning method and system based on AI and dynamic feature library

The invention discloses an intelligent element positioning method and system based on AI and a dynamic feature library, and the method comprises the steps: carrying out the anomaly detection and abnormal data capture, executing an element positioning operation through an automatic script, and recording a page DOM tree structured snapshot when the positioning is judged to be failed; multi-modal feature extraction: analyzing the DOM tree, and extracting an XPath / CSS path and semantic attributes of a target element; performing dynamic feature library retrieval and matching, calculating DOM structure similarity, and taking a high-similarity rule as a correction basis if the high-similarity rule exists; an AI intelligent rule is generated, and candidate positioning rules are generated through the classification model; performing script correction and real-time feedback, replacing the failure positioning statement and updating the dynamic feature library; and generating a structured report, and recording abnormal data. According to the method, the AI is combined with the dynamic feature library, so that the intelligent correction of element positioning in the automatic test process is realized, and the stability and efficiency of the automatic test are improved.
Owner:SHANGHAI TIANHAO INFORMATION TECH CO LTD

RPA code generation method based on HTML and image

The invention discloses an RPA code generation method based on an HTML and an image, and the method comprises the following steps: S1, receiving multi-modal input data provided by a user, the multi-modal input data comprising an HTML structure file of a target business system, an interface screenshot image and a natural language operation demand description; s2, performing standardized analysis on the HTML structure file, extracting DOM tree structure information, and performing visual feature enhancement processing on the interface screenshot image; and S3, inputting the processed HTML structure information, the interface screenshot image and the natural language operation requirement into a pre-trained visual language model, and outputting a target XPath path and a Python code sequence through the visual language model. According to the method, comprehensive understanding of a complex service system is realized by combining HTML structure information and an interface image, so that efficient and accurate RPA script codes capable of being executed on multiple platforms are automatically generated.
Owner:HANGZHOU BRANCH INTELLIGENT TECH CO LTD

Intelligent news crawling system and method based on SpiderFlow

The invention relates to the technical field of computer network data acquisition and processing, in particular to an intelligent news crawling system and method based on SpiderFlow, and the system comprises a dynamic time control module, an intelligent duplicate removal module, a paging processing module, a fault-tolerant data extraction module and a dynamic parameter configuration module. The dynamic time control module generates a time variable through a date calculation engine to realize accurate limitation of a crawling time range; the intelligent deduplication module adopts a three-level deduplication mechanism to ensure data uniqueness; the paging processing module supports static parameter and dynamic AJAX loading dual-mode paging, and realizes full-amount acquisition in combination with a self-increasing page number iterator and an intelligent termination strategy; the fault-tolerant data extraction module is integrated with CSS / XPath double parsers, and data integrity is guaranteed through field missing detection and messy code transcoding. Through multi-module cooperation and an intelligent algorithm, crawling efficiency, data quality and system stability are remarkably improved, and challenges such as webpage structure change and anti-crawling limitation are effectively handled.
Owner:XIAMEN BEST DIGITAL TECH CO LTD

Automatic operation backtracking and digital person continuous talk control method and system oriented to multi-modal interaction

The invention relates to a multi-modal interaction-oriented automatic operation backtracking and digital person continuous talking control method and system, and the system comprises an operation snapshot module which records operation nodes through a semantic anchor point marking technology, such as an interface XPath page number and a PPT page number; the interruption detection module is used for capturing a user interruption instruction in real time; the progress prediction engine dynamically adjusts a follow-up demonstration path based on a program; and the continuous talk controller is used for coordinating RPA operation execution, digital human rendering and voice synthesis through a three-level priority thread pool. The problems of inaccurate operation flow backtracking and multi-modal thread conflict after demonstration interruption are solved, the breakpoint continuous talk error is obviously lower than that of a traditional scheme, and the method is suitable for PPT demonstration, cockpit explanation, business system training, course explanation and 3D modeling guidance scenes.
Owner:WUXI GANGWAN NETWORK TECH

Large language model XPath generation method based on hierarchical composite reward reinforcement learning

The invention discloses a large language model XPath generation method based on hierarchical composite reward reinforcement learning, and the method specifically comprises the following steps: 1, obtaining HTML (Hypertext Markup Language) source codes and page element information of a target webpage, carrying out the data cleaning, and obtaining structured data containing a DOM (Document Object Model) hierarchical sequence structure and an element attribute value; performing data annotation on the structured data after data cleaning to obtain an annotated data set; 2, selecting a basic model, performing supervision and fine tuning on the basic model by using the annotation data set, and taking the basic model subjected to supervision and fine tuning as a strategy model; constructing a layered composite reward function to perform reinforcement learning fine tuning, so that the output hierarchy of the strategy model is aligned with the input DOM hierarchy, and obtaining a final model subjected to two-stage fine tuning; and 3, generating a standard XPath character string, and outputting structured data matched with the input DOM hierarchy to display layer-by-layer construction logic of the standard XPath character string. According to the method, the stable XPath can be generated, and the generation process is completely transparent and traceable.
Owner:HANGZHOU BRANCH INTELLIGENT TECH CO LTD

AI-driven self-healing test automation system for WDIO with adaptive detection and predictive recovery

An AI-driven, self-healing test automation system for WebDriverIO (WDIO), comprising: a. an error catching layer configured to detect when a test step fails due to a missing, moved, or modified UI element; b. an AI-based detection engine configured to analyze the current Document Object Model (DOM), historical test execution data, and element attributes including XPath, CSS selectors, labels, and inner text to identify and evaluate potential alternative UI elements; c. a predictive recovery module that uses pre-trained models such as Long Short-Term Memory (LSTM) networks or transformer-based architectures to predict and select an appropriate recovery action based on the current test context, including repeating the operation, jumping to the next logical step, or redirecting the test flow; d. a dynamic healing process that automatically applies the selected recovery action without requiring manual updates to the test code; e. a feedback learning loop configured to update and retrain the detection and prediction models based on user validation, success metrics, and execution results, thereby improving the accuracy and adaptability of the system over time.
Owner:SRINIVAS SRIKANTH MCKINNEY

Method and system for automatically extracting news detail page XPath based on large model

The invention belongs to the technical field of computers, and discloses a news detail page XPath automatic extraction method and system based on a large model, and the method comprises the steps: determining a large model base through multi-aspect evaluation; the method comprises the following steps: preprocessing an HTML page, designing a structured prompt template and constructing a thinking chain CoT data set; utilizing a Lora technology fine tuning model, and adopting a smooth loss function and a dual reward mechanism of fusion format and content evaluation to optimize; and performing format calibration on the XPath and the JSON output by the model, and extracting and verifying the validity of the URL of the detail page by using the calibrated XPath. Compared with a traditional template-based information extraction method, the method does not depend on the stability of a webpage structure, has higher generalization ability and fault tolerance, remarkably improves the information extraction precision, enhances the complex webpage adaptability, improves the overall robustness and universality of a system, and is suitable for large-scale popularization and application. And a stable extraction effect can still be kept in news websites with frequent structure change.
Owner:SHENZHEN WANGLIAN ANRUI NETWORK TECH CO LTD

Webpage data extraction method based on large language model

The invention discloses a webpage data extraction method based on a large language model, which comprises the following steps of: generating an Xpath sequence webpage grabber by utilizing the large language model, and processing diversified and variable network environments through a two-stage framework: in the first stage, inversely checking and removing HTML (Hypertext Markup Language) noise by utilizing LLM (Language Language Model) information extraction capability, and in the second stage, inversely checking and removing HTML (Hypertext Markup Language Model) noise; self-adaptive generation of an Xpath action sequence is carried out according to the hierarchical structure of the HTML; according to the combination of an external evaluation mechanism and a local evaluation mechanism of LLMs, a plurality of Xpath action sequences generated on different webpages in one stage are integrated, and a general grabber specific to a website is generated. According to the method, a baseline method is always exceeded under zero sample setting, higher efficiency is shown in large-scale webpage information extraction tasks, the method can quickly adapt to different website and task requirements, dependence on LLMs is reduced when similar tasks are processed, and therefore the efficiency of processing a large number of webpage tasks is improved.
Owner:HANGZHOU DIANZI UNIV

Website data analysis method and device based on large model technology

The invention discloses a website data analysis method and device based on a large model technology, and belongs to the field of fusion of web crawlers and artificial intelligence. According to the method, a local knowledge base drives a large model to generate a precise XPath rule, and the precise XPath rule is packaged into a decoupling plug-in and integrated into an existing crawler framework; a dynamic feedback optimization mechanism is innovatively designed, rule validity is monitored in real time by using a composite verifier, and XPath reconstruction and knowledge base updating are automatically triggered. According to the method, the high concurrency performance of a traditional crawler is reserved, the rule generation efficiency and the system robustness are remarkably improved, the manual maintenance cost is reduced, and the method is suitable for a website data collection scene with a complex structure or frequent updating.
Owner:XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD

Page data extraction method, page automated testing method

The present application provides a method for extracting page data and a method for automated page testing. The method for extracting page data is used to extract the tree-structured data corresponding to the tree-like content displayed on a web page, and includes: obtaining the relative xpath of the root node in the tree-structured data in the page to be tested; using the relative xpath of the root node as the current relative xpath, and using a preset recursive call method to obtain a list of relative xpaths, where the list of relative xpaths includes the relative xpath of each node arranged in depth-first order in the tree-structured data; converting the list of relative xpaths into a dictionary. The page extraction method can automatically extract the tree-structured data corresponding to the tree-like content and return it in the form of a dictionary, which is convenient for subsequent automated page testing.
Owner:WUHAN SIPU TECH CO LTD

Data burying point testing method and device based on natural language intelligent identification control

The application discloses a data burying point test method and device based on natural language intelligent identification control, and comprises the following steps: identifying a target control based on a natural language script, the target control comprising XPath and a first target parameter; determining whether a target element exists in a current page element structure based on the XPath, the target element comprising a second target parameter, and the target element being at least one; if the target element exists, determining whether the target control matches the target element based on the first target parameter and the second target parameter; in the case that the target control matches the target element, obtaining data burying points corresponding to the target element, and verifying the data burying points. In the above process, the target control can be identified through natural language, and it is determined whether the target element exists based on the XPath in the target control, front-end page operation is performed, and automatic verification of burying point data is performed, so that the element attribute and the decomposed page module do not need to be known in advance, and the matching efficiency is improved.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Webpage information extraction and classification method and device

The application provides a webpage information extraction and classification method and device, and belongs to the technical field of artificial intelligence. The webpage information extraction and classification method comprises the following steps: converting the source code of a target webpage into a dom tree; processing each node of the dom tree to obtain four feature matrices, i.e., a text feature matrix, an Xpath feature matrix, a layout feature matrix and a visual feature matrix; inputting the four feature matrices into an encoding network respectively to obtain four representation vectors; performing feature fusion on the four representation vectors to obtain fused features; inputting the fused features into a classification network to obtain and store the classification result of an information unit of the target webpage, wherein the classification result comprises at least one of the following: a table, a form to be filled, a text unit that needs to be linked to a next webpage for display, a navigation bar or a display bar, pure text, an advertisement and useless information. The technical scheme of the application can improve the universality of the webpage information extraction and classification scheme.
Owner:CHINA MOBILE COMM LTD RES INST +1

Vertical website information extraction method, device, equipment and medium based on large model

The present invention discloses a method, device, equipment and medium for extracting vertical website information based on a large model. According to the technical solution provided by the present invention, a large language model is used to extract the first attribute text information corresponding to the target attribute from the seed web page selected from the vertical field website; the correct node is screened from the node corresponding to the information, and the absolute path expression of the XPath of the correct node is determined; the anchor node is determined from the DOM tree based on the absolute path expression, and the XPath final expression is constructed based on the relative position of the correct node and the anchor node; the second attribute text information corresponding to the target attribute is extracted from the vertical field website using the XPath final expression. Through the present invention, the correct node and the anchor node are determined from the seed web page in the vertical field website, and the XPath final expression derived from the relative position of the two is used to extract the target information from the website, thereby achieving a lower cost and more accurate extraction of the target information without the need for model training.
Owner:PEKING UNIV

Method for extracting and merging EXCEL table data in XFA based on JAVA language

PendingCN121882001ANatural language data processingJavaProgram Efficiency
The invention belongs to the technical field of table data processing, and discloses a method for extracting and merging EXCEL table data in XFA based on a JAVA language, and the method for extracting and merging table data comprises the following steps: step 1, analyzing an XFA form; step 2, analyzing the XFAXML data; step 3, constructing an EXCEL table template; and step 4, filling the data into Exce l. According to the method, the compatibility of various PDF forms is realized by combining libraries such as Apache PDFBox and the like, XFA data are extracted from the libraries, the XFAXML data are accurately and efficiently processed by utilizing a JavaDOM analyzer and XPath, complex data models and relation requirements are met, styles, formats and placeholders can be flexibly adjusted by customizing an Exce l template, a modular processing mode is adopted for a plurality of XFA data sources, and the data processing efficiency is greatly improved. According to the method and the system, the data are integrated into different areas of the Exce l workbook, resource management and exception handling are emphasized, the program efficiency and the system stability are improved, and the method and the system have excellent expansibility, can quickly adapt to new scenes and requirements, and realize quick response to new challenges.
Owner:BEI JING ZHONG YAN CHUANG XIN KE JI YOU XIAN GONG SI

A text chunking method and device based on HTML node path resolution

Embodiments of the present specification relate to the technical field of text processing, and provide a text blocking method and device based on HTML node path analysis, comprising: performing initial blocking on the HTML document to obtain an initial HTML document blocking result, and recording an XPath path expression corresponding to each piece of text in the target HTML document and an initial HTML document block to which the text belongs; performing preprocessing on each initial HTML document block to obtain a plurality of initial pure text document blocks; performing merging or segmentation operations on the initial pure text document blocks to obtain a plurality of final pure text document blocks; and blocking the target HTML document according to the XPath path expression corresponding to each piece of text in the final pure text document blocks and the initial HTML document block to which the text belongs, to obtain a final HTML document blocking result. Through the embodiments of the present specification, the accuracy of HTML text blocking can be improved.
Owner:CHINA EVERBRIGHT BANK

Browser extension with automation testing support

Described are methods and corresponding systems for generating and using selectors during web development. In some implementations, one or more natural language statements are obtained as input to a software application, for example, a web browser extension. The one or more input statements are analyzed, using natural language processing, to identify a first web element of a webpage and an action to be performed with respect to the first web element. A selector is then generated based on one or more attributes of the first web element. The selector operates as an address of the first web element and can, for example, be an XPath or CSS selector. To provide a user with access to the selector, the selector can be displayed on a user interface and / or saved to an output file. In some instances, the selector is generated as part of program code executable to perform the action.
Owner:SAUCE LABS

Process aided design method based on matching degree calculation and parameter pushing

The invention provides a process aided design method based on matching degree calculation and parameter pushing, which comprises the following steps of: firstly, constructing different process file label systems according to different product category libraries and process method libraries, and labeling each process file by using an Xpath language and a regular matching method; secondly, acquiring a requirement input by a user, comparing a process file matching degree calculation code with a process file in a corresponding process file label system, and outputting a matching degree; then, sorting the process files from high to low according to the matching degrees, and displaying a plurality of process files with the highest matching degree to a user; and finally, based on the technological parameter recommendation codes of the technological parameter library, matching technological parameter items and parameter values in the existing technological parameter library, and displaying the technological parameter items and parameter values to a user for reference. According to the method, the user can be helped to quickly find the process file which is most matched with the demand, the user is helped to carry out product process design through historical knowledge pushing and parameter matching, and the design efficiency and quality are improved.
Owner:SHANGHAI AEROSPACE EQUIPMENTS MANUFACTURER CO LTD +1

A method and device for generating an automated test script, and a storage medium

The present disclosure provides a method and device for generating an automatic test script and a storage medium, including: obtaining an initial automatic test script and parsing the initial automatic test script to extract an initial XPath path; obtaining content of a webpage to be tested, parsing the content of the webpage to generate a DOM tree of the webpage to be tested; determining a corresponding test script data table based on the initial XPath path; establishing a webpage element matrix model based on the DOM tree; optimizing the initial XPath path based on the test script data table and the webpage element matrix model to obtain a target XPath path, and generating a target UI automatic test script based on the target XPath path. Thus, the present disclosure can automatically adjust the initial XPath path to obtain the target XPath path through the test script data table and the webpage element matrix model, without manual adjustment, thereby improving the generation efficiency and accuracy of the test script, and further improving the efficiency and accuracy of UI testing.
Owner:CHINA MOBILE GRP GUANGDONG CO LTD +1

A method and system for inferring a target UI element position based on a UI operation trajectory

The present application relates to the technical field of software test automation, and particularly relates to a method and system for inferring the position of a target UI element based on UI operation trajectories, comprising: when an XPath positioning operation performed by a test case on a target UI element in a page under test fails, constructing a first operation trajectory according to a preceding UI operation that has been successfully performed, matching and determining a candidate set of position information of the target UI element in combination with all historical successful operation trajectories, extracting interactive UI elements and sorting them according to preset conditions, and determining the target UI element and its final position information. When the XPath positioning fails, the present application can automatically infer the position of the target element according to the preceding operation trajectory and the historical operation trajectory to restore the test execution, significantly improving the stability and continuity of the automated test, reducing the dependence on manual maintenance, and simultaneously having good self-adaptive ability to UI interface changes.
Owner:ICLOUDSHIELD SECURITY TECHNOLOGY CO LTD

XPath path positioning-based HTML (Hypertext Markup Language) document sentence-level intelligent labeling method and system

The invention discloses an XPath path positioning-based HTML (Hypertext Markup Language) document sentence-level intelligent labeling method and system, relates to the technical field of computer document processing and data labeling, constructs a'purification-sandbox-event interception 'three-layer security architecture, and provides a reliable guarantee for processing an untrusted HTML document through DOMPurity purification, iframe sandbox isolation and event interception mechanisms; sentence boundaries in the HTML document are recognized, and the labeling granularity is improved from a traditional paragraph level to a sentence level; the XPath path technology is used for HTML document sentence-level labeling and positioning, and the positioning precision and stability of labeling are improved by generating an accurate and flexible positioning expression; and a joint index mechanism of the XPath path and the text abstract is designed, so that the technical problem that the annotation is easy to lose after the document structure is changed is solved, and the annotation recovery success rate is improved. According to the method and the system, sentence-level refined and secure intelligent labeling on the HTML document is realized.
Owner:HEFEI DAZHIHUI CAIHUI DATA TECH CO LTD

An intelligent element positioning method and system based on AI and dynamic feature library

The application discloses an intelligent element positioning method and system based on AI and a dynamic feature library, which comprises: abnormality detection and abnormal data capture, element positioning operation is performed through an automatic script, and a page DOM tree structure snapshot is recorded when positioning fails; multi-modal feature extraction, DOM tree analysis, extraction of XPath / CSS path and semantic attributes of a target element; dynamic feature library retrieval and matching, calculation of DOM structure similarity, and use of high similarity rules as a correction basis; AI intelligent rule generation, generation of candidate positioning rules through a classification model; script correction and real-time feedback, replacement of invalid positioning statements and updating of the dynamic feature library; and generation of a structured report, recording of abnormal data. The application realizes intelligent correction of element positioning in an automatic test process through the combination of AI and a dynamic feature library, and improves the stability and efficiency of automatic testing.
Owner:SHANGHAI TIANHAO INFORMATION TECH CO LTD

A fine-grained internet of things device automatic identification method based on firmware emulation

This invention provides a fine-grained automatic identification method for IoT devices based on firmware simulation, relating to the field of IoT technology. The method includes: automatically collecting IoT device tags from aggregation websites using web crawling technology and constructing an IoT device tag library; automatically collecting IoT device firmware using XPath expressions combined with an ASCII-based firmware search algorithm; simulating the obtained device firmware, accessing the simulated device IP address and extracting the device fingerprint; automatically generating and matching regular expressions for positive samples based on the device fingerprint; using a binary tree-based model regular expression matching strategy for efficient and accurate model identification; extracting features from web files in the IoT device firmware using BinWalk; training the extracted features using a decision tree model; designing a tree-based rule-based scanning and identification strategy; converting the decision tree model into a rule tree; and identifying the version number of the IoT device, thereby improving the accuracy of identification.
Owner:NORTHEASTERN UNIV AT QINHUANGDAO

Method to decode and discover using XML xpath queries

A method for creating a standardized XML format including the steps of (1) defining an event element, (2) creating a template file having the event element; (3) identifying a message on a shared bus; (4) determining a bus message type of the message; and (5) creating an output XML file based on the template file when the value of the bus message type is equal to the value of the event user identification attribute.
Owner:AERONIX

Template-based web page element positioning method, device and equipment and storage medium

The application relates to the development of auxiliary technology, and discloses a template-based WEB page element positioning method, device, equipment and medium. The method comprises the following steps: using a marking tool to select a region of interest in a screenshot of a WEB page currently displayed by a browser, and generating a template definition file of the region of interest; obtaining a target screenshot of a target WEB page in the browser, and converting corresponding elements in the target screenshot into a conversion matrix list by using elements with an anchor type identified by a type field in the template definition file; calculating the positions of elements with the same element name and a value type identified by the type field in the target screenshot according to the conversion matrix list, obtaining a text box of the value type element corresponding to the positions, and obtaining an xpath of the value type element; and outputting corresponding elements in the target page screenshot according to the xpath of the value type element by using an RPA system. The application can improve the accuracy and efficiency of information acquisition in a WEB page of a hospital system.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method and device for generating data verification rule of shared document, equipment and medium

The invention provides a method and a device for generating a data verification rule of a shared document, equipment and a medium. The method comprises the following steps: S1, receiving related configuration of each data node based on an industry specification through a configuration interface; s2, obtaining related configuration of the shared document according to the type of the selected shared document of which the verification rule is to be generated; s3, sorting the node set, arranging non-necessary nodes in the front, and arranging short nodes xpath in the front; s4, creating a related set; and S5, traversing the node set, generating a verification rule of each node, and further generating a data verification rule of the shared document. After configuration based on industry shared document specification requirements, the data verification rules of all the shared documents are automatically generated by analyzing the configuration, and the technical blank of automatically generating the data verification rules of the shared documents is filled.
Owner:FUJIAN ECAN INFORMATION TECH CO LTD

Webpage login element identification and test method based on artificial intelligence and related equipment

The invention provides a webpage login element identification and test method based on artificial intelligence and related equipment, and relates to the technical field of network security attack and defense. The method comprises the following steps: acquiring a complete webpage source code file of a target login webpage, traversing the content of the webpage source code file, and storing each form element into an XPath file; the webpage source code file and the Xpath file are input into a preset large language model for recognition processing, and a recognition result is obtained; when the verification code image is a dynamic verification code image, identifying the dynamic verification code image to obtain an operation target and an operation area; performing automatic combination matching on the operation target and the operation area to obtain a verification data combination, performing automatic filling processing on the user name field and the password field to obtain a login data combination, and performing browser instance generation on the login data combination and the verification data combination to obtain a plurality of browser instances; in each thread, the browser instance is utilized to test and crack the target login webpage to obtain a test result, so that the red team attack is quickly and automatically tested.
Owner:BEIJING TIMES XINWEI INFORMATION TECH CO LTD

A method, device and equipment for generating PDF based on webpage and storage medium

The application provides a method, device and equipment for generating a PDF based on a webpage, and a storage medium. The method comprises the following steps: obtaining a webpage access path and an Xpath path of core content of a webpage to be accessed; sending an access request to the webpage based on the webpage access path, and obtaining feedback webpage data; determining a webpage element node tree of the webpage data; determining non-core data that needs to be deleted in the webpage data based on content in the Xpath path and the webpage element node tree, wherein the non-core data is data other than core data and display-related data in the webpage data; deleting the non-core data in the webpage; and generating a corresponding PDF file based on the core data displayed in the webpage data. The method for generating a PDF based on a webpage can directly convert webpage content into a PDF file, and the layout is normal, and the file content can be enlarged without distortion.
Owner:BEIJING TOPSEC NETWORK SECURITY TECH +2

Web page element search method, device and computing equipment

The present invention discloses a web page element search method, device, and computing equipment. The web page element search method includes the following steps: in response to a user's request to search for a web page element on a Chrome browser, sending the domain name and URL of the current site to a server, then obtaining a template returned by the server, wherein the template records a target field and an extraction rule for the target field; when the extraction rule for the target field is absolute positioning, extracting the target field in the current page using one of an Xpath selector, a CSS selector, and an ID selector; when the extraction rule for the target field is relative positioning, extracting a target container in the current page, and extracting the target field in the target container; and using the extracted target field as a search result for the web page element. The present invention also discloses a corresponding computing device and apparatus.
Owner:HAINAN CHEZHIYITONG INFORMATION TECH CO LTD

A method, system, computer, and storage medium for structured extraction of multimodal web page data based on visual features and large models.

This invention discloses a method, system, computer, and storage medium for multimodal webpage data structured extraction based on visual features and a large model. The method involves acquiring webpage screenshots and HTML source code via a browser; normalizing, enhancing, and segmenting the screenshots to output preprocessed screenshots and coordinates of the core content area; cleaning the HTML, adding XPath and hierarchical structured text; inputting both into a multimodal large model to generate an initial XPath; performing bimodal validation of the initial XPath using both code and visual rules; if either fails, retrying based on the same preprocessed data until the limit is reached or success is achieved; periodic monitoring is set up, re-collecting and preprocessing data, comparing visual and code features with historical preprocessed data corresponding to the previous valid XPath; if any feature changes, repeating the aforementioned steps to generate a new XPath. This invention enables automatic XPath generation and adaptive updates, reducing maintenance costs.
Owner:SUZHOU AEROSPACE INFORMATION RES INST

Attribute information disambiguation method and system for html document

The application belongs to the technical field of network information processing, and specifically provides a kind of html document attribute information disambiguation method and system, wherein the method comprises: html document is parsed into text data and table data and is stored by XPath;Attribute key and information value are obtained by using rule extraction or model extraction for line-by-line processing;Attribute key in text data and table data is disambiguated using context information respectively.The scheme converts html document into chapter and table, and then disambiguates text and table using some context information, improving the accuracy of information extraction.
Owner:ZHONGKE FANYU TECH