Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12 results about "Web extraction" patented technology

Computer-implemented system and method for providing website navigation recommendations

A system and method for providing Website navigation recommendations is provided. A Web page of interest is identified as a destination Web page. A domain of Web pages related to the destination Web page is determined. Information is extracted from each Web page in the domain and a recommendation comprising instructions for navigating to the destination Web page is generated based on the extracted information.
Owner:PALO ALTO RESEARCH CENTER INC

Traffic facility attribute mining method and system based on multi-modal network open source data

The invention relates to a traffic facility attribute mining method and system based on multi-modal network open source data. The method comprises the following steps: extracting traffic facility information from a webpage to construct a knowledge graph; extracting positions and appearances of traffic facilities from the images, and analyzing streetscape images to obtain attributes such as road traffic; collecting map images, tiles and vector data to construct a road network topology and attribute database; associating webpage texts, pictures, streetscape images and network map multi-source data according to the spatial position of the traffic facility; through comprehensive and deep attribute mining, different modal data are integrated to improve the accuracy and reliability of attribute mining, the real-time and dynamic updating capability, the convenient visualization and interaction operation, the data sharing and integration convenience, and the traffic facility management efficiency and collaboration are improved. According to the method, comprehensive, accurate and real-time traffic facility attribute data support can be provided for urban traffic planning, traffic management and intelligent traffic system construction, so that the efficiency and the intelligent level of traffic facility management are remarkably improved.
Owner:NAT UNIV OF DEFENSE TECH

Service monitoring data enhancement method, system and device based on LLM and retrieval enhancement generation and storage medium

The invention discloses a service monitoring data enhancement method, system and device based on LLM and retrieval enhancement generation and a storage medium, and relates to the technical field of service monitoring and exception handling. The method comprises the steps that user exception description is received, a webpage is retrieved through semantic expansion, and triple output structured data is extracted; and comparing with a local knowledge base, processing and storing the new content classification vector, and updating the index. Performing preliminary screening and reordering by combining exception description and an update library, taking a front list as a context to enable a large language model to generate analysis, evaluation and a strategy, and performing rule check and output; a result is verified, scene variants are generated according to exception types, after duplicate removal, the scene variants, strategies and metadata are structurally stored in a knowledge base, and continuous learning of a closed loop is completed; according to the method, high-quality and high-timeliness service exception training data and processing strategies can be automatically generated, the illusion problem of a large language model in actual operation and maintenance is effectively solved, and the fault response speed and the decision accuracy are improved.
Owner:GUANGXI POWER GRID CORP

News pushing method and device, computer program product and news pushing system

The invention provides a news pushing method and device, a computer program product and a news pushing system. The method comprises the following steps: extracting a plurality of initial news from a webpage; news in which the target object is interested is screened out from the multiple pieces of initial news, multiple pieces of target news are obtained, and the number of the target news is smaller than or equal to the number of the initial news; the target news is sorted according to the importance degree of the target news, the sorted target news is obtained, and the importance degree is the degree of the influence of the target news on the target object; the sorted first N pieces of target news are pushed to the target object, and N is larger than or equal to 1. According to the method, the problems that in the prior art, news screening and sorting depend on manual screening, expert evaluation or simple statistical analysis, obvious limitation and defects exist, and news cannot be accurately recommended to the user are solved.
Owner:PEKING UNIV CHONGQING RES INST OF BIG DATA +1

Product weight prediction device and product weight prediction method

A product weight prediction device disclosed in the present document comprises: a communication interface for receiving user input data related to an order of a product; and at least one processor for extracting product information of the product from a web page on the basis of the user input data, generating preprocessed data by embedding the product information, and predicting a weight of the product by inputting the preprocessed data to a learning model, an input value of which is the preprocessed data and an output value of which is a predicted weight of the product.
Owner:SHAASHOP INC

A webpage data collection and dynamic monitoring method, device and equipment

This invention discloses a method, apparatus, and device for webpage data acquisition and dynamic monitoring. The method includes responding to a data monitoring task, initializing a simulated browser according to the task cycle corresponding to the data monitoring task, and accessing the webpages on the list to be monitored specified by the data monitoring task; traversing the webpages on the list to be monitored and extracting webpage element data; if the webpage element data is plain text, extracting the text corresponding to the plain text element; if the webpage element data contains image elements, performing optical character recognition on the images corresponding to the image elements to generate image recognition results; using the text and image recognition results, constructing a blacklist and performing fuzzy matching with the list to be compared, and pushing early warning information according to the fuzzy matching results. This method, combined with the simulated browser, performs categorized identification and simulated matching of webpage elements, improving identification accuracy while ensuring monitoring efficiency.
Owner:创优数字科技(广东)有限公司

Model ensemble for matching nearest script to a new script page

A method including extracting a number of page features from a web page. The number of page features represent an executable logic of the web page. The method also includes embedding, by a page feature embedding model, the number of page features to generate a page vector data structure. The method also includes comparing, by a comparison model, the page vector data structure and a number of script vector data structures to identify a selected script. Each of the number of script vector data structures is generated by a script feature embedding model processing computer executable program code of a corresponding script for performing a computer function on a web page. The method also includes presenting the selected script.
Owner:INTUIT INC

A web page content integrity detection system and method based on data analysis

The application relates to the technical field of webpage identification detection, in particular to a webpage content integrity detection system and method based on data analysis, which comprises the following steps: extracting user browsing operation records of each to-be-detected webpage in a time period, capturing browsing transfer operations generated between the corresponding to-be-detected webpages and repeated browsing operations generated on the corresponding to-be-detected webpages based on the user browsing operation records, and constructing corresponding first user browsing transfer links and second user browsing transfer links; locking the to-be-detected webpages suspected of having content missing, calculating the first browsing transfer rate of each target to-be-detected webpage, and completing the first calibration screening; calculating the second browsing transfer rate of each target to-be-detected webpage, and completing the second calibration screening; and assisting operation and maintenance personnel in fault screening of each target to-be-detected webpage in a target to-be-detected webpage sequence received in real time.
Owner:HAINAN INFOBAHN TECH SERVICE CO LTD

Multi-modal automated evaluation for improved accessibility

One example method includes a machine-learning (ML) model receiving a first input that includes images that have been extracted from a web page and a second input that includes alt-texts that have been extracted from the web page. The alt-texts describe the images. The ML model converts the images into a first embedding representation and converts the alt-texts into a second embedding representation. Based on the first and second embedding representations, a similarity score between the images and the alt-texts is calculated. The similarity score specifies how accurately each of the alt-texts describe the images. The one of the alt-texts having the highest similarity score is then selected.
Owner:DELL PROD LP

Data extraction using llms

PendingUS20260187174A1Data ingestionEngineering
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving information identifying a domain to be analyzed and identifying an entity referenced by the domain. The domain is queried, and a plurality of web pages located within the domain are received. The plurality of web pages is inputted into an artificial intelligence system that includes a large language model which extracts first content from a first web page among the plurality of web pages. The artificial intelligence system extracts second content from a second web page, the second content in a second format that differs from the first content. The artificial intelligence system generates third content representing a characterization of the entity based on the extracted first and second content. The generated characterization is an interpretation of the extracted first and second content.
Owner:GOOGLE LLC

Framework for exposing context-driven services within a web browser

Systems and methods for securely exposing context-driven services within a web browser. An example method includes receiving manifests from hubs apps (e.g., remote services). The manifests define requested context types for the hub apps. When the web browser loads a web page, the web browser may execute context extractors to extract context from the web page. The context extractors that are executed are based on the context types requested by the hub apps. The extracted context is then sent to the corresponding hub apps without providing the hub apps direct access to the web page. For instance, the hub apps do not have access to the document object model (DOM) of the web page and the hub apps cannot inject data into the web page.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Verified entity attributes

Systems and methods enable an entity to certify a web page address as being linked to the entity. The web page address includes semantic web mark-up identified attributes for the entity. A system may extract the attributes from the web page for the entity and use the attributes to generate an information card for the entity. The certification process ensures that the attributes are accurate, so that information cards generated for the entity are of high quality and reliable. Implementations may also simplify maintenance and quality assurances processes for an entity repository.
Owner:GOOGLE LLC