Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

10 results about "Headless browser" patented technology

A headless browser is a web browser without a graphical user interface. Headless browsers provide automated control of a web page in an environment similar to popular web browsers, but are executed via a command-line interface or using network communication. They are particularly useful for testing web pages as they are able to render and understand HTML the same way a browser would, including styling elements such as page layout, colour, font selection and execution of JavaScript and AJAX which are usually not available when using other testing methods.

Dynamic webpage content complete acquisition system and acquisition method based on browser rendering engine

The invention discloses a dynamic webpage content complete collection system and method based on a browser rendering engine. A CDP controller main module calls a command to start a headless browser process, and a network monitoring module registers a Network. Response Received event monitor; a DOM (Document Object Model) monitoring module starts a Mutation Observer to monitor the change of a DOM (Document Object Model) tree; the interactive simulation module preloads a mouse click and scroll event instruction set; capturing JSON / XML data returned by an interface in real time by using a network monitoring module, intercepting API response data, generating a DOM snapshot by using a DOM monitoring module, and extracting key structure changes; the DOM monitoring module provides a complete rendering HTML, the network monitoring module provides request metadata, and the interactive simulation module provides an operation time sequence log and packages generated complete webpage data into a WARC file. According to the method, document collection is upgraded to application-level interaction, and a technical implementation path is provided for collection scenes of dynamic data-intensive websites (such as financial public opinion monitoring and e-commerce price tracking) in which Internet resources are stored for a long time.
Owner:SUZHOU JIATU SOFTWARE CO LTD

Small language webpage adaptive acquisition method and device based on large language model

The invention relates to the technical field of artificial intelligence, and discloses a small language webpage adaptive collection method and device based on a large language model.The method comprises the steps that a target small language webpage is loaded and rendered through a headless browser, and a document object model tree structure is obtained; carrying out analysis and semantic annotation on the document object model tree structure based on a large language model, identifying dynamic content nodes and generating a self-adaptive content positioning rule; based on a self-adaptive content positioning rule, performing semantic vectorization on the at least two continuously collected webpage contents, and calculating semantic similarity; querying a preset frequency mapping rule according to the semantic similarity, dynamically adjusting the initiation frequency of a subsequent acquisition request, and generating an acquisition strategy; and executing the acquisition strategy to access the target webpage, and injecting an operation sequence for simulating human interaction behaviors in the acquisition process. Complete collection of the dynamic content of the small language webpage can be achieved, and continuity and stability of the collection process are guaranteed.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD

Method and apparatus for implementing a headless browser cluster

The present disclosure relates to a method and device for implementing a headless browser cluster. The method comprises: assigning a corresponding accessible IP address to each server in a plurality of servers; wherein the plurality of servers form a cluster; installing a headless browser in each server; after receiving a headless browser starting instruction, starting a preset number of headless browser instances, the started headless browser instances forming a headless browser instance pool; after detecting that a headless browser instance acquisition instruction is received, acquiring a target headless browser instance available from the headless browser instance pool; acquiring a network socket endpoint address of the target headless browser instance; and sending the network socket endpoint address of the target headless browser instance to a client.
Owner:WIRELESS LIFE (HANGZHOU) INFORMATION TECH CO LTD

Distributed document generation method based on message queue and headless browser

The invention relates to the technical field of computers, and discloses a distributed document generation method based on a message queue and a headless browser, and the method comprises the steps: receiving a document generation request, packaging the document generation request into a standard task object, and verifying the validity of a rendering target; splitting an original task into sub-tasks, and calculating task fingerprints and time-carrying task fingerprints; the core parameters and the service data loads are separately stored; converting the effective sub-tasks into rendering instructions, and performing queue arrangement according to the calling mode of the sub-tasks; and taking a headless browser as a rendering processing unit, continuously consuming a rendering instruction from the message queue, carrying out webpage rendering, converting the rendered webpage into a document, and responding to a document generation request. The headless browser is used as a core rendering processing unit, high-fidelity restoration of webpage content is achieved, and the visual quality and accuracy of the generated document are improved.
Owner:HUIDIAN TECHNOLOGY (SUZHOU) CO LTD

A large PDF file intelligent generation method with front-end and back-end cooperation

The application discloses a front-rear end cooperative large PDF file intelligent generation method, relates to the technical field of computer software, and comprises the following steps: front-end rendering: the front end renders the work order data page by combining a headless browser with a Vue framework and an ElementUI component library, and dynamically displays structured work order details through a responsive data binding mechanism; data fragmentation: large work order data sets are split into N subtasks based on a dynamic fragmentation strategy; and small PDF files are generated in parallel: a containerized rendering cluster based on Kubernetes deployment management is used to allocate rendering tasks to each container; the small PDF files generated by the application through careful design with the help of the Vue and ElementUI frameworks are more attractive in terms of text layout, color matching and graphic display, and the large file after merging is more beautiful than the file generated by a traditional back end, thereby improving user experience.
Owner:BEIJING ZHIXING TONGDE INFORMATION TECH CO LTD

Phishing webpage intelligent detection system and method based on proxy AI

PendingCN121750321ASecuring communicationEngineeringPhishing
The invention provides a proxy AI-based phishing webpage intelligent detection system and method, and the system is characterized in that the system comprises a link input module which is used for receiving a to-be-detected URL, and carrying out the preliminary filtering; the headless browser control module is used for controlling a headless browser to load a page and managing the life cycle of the page; the page content extraction module is used for extracting page information from the headless browser; manual analysis requirements are reduced through automatic interaction, the detection efficiency is improved, and the method is suitable for large-scale mail security scanning scenes; an interaction path and a decision process can be recorded, a transparent log is provided for security analysis, and subsequent optimization and auditing are facilitated.
Owner:BEIJING DIRECTION BIAO INFORMATION TECHNOLOGY CO LTD

A method and system for determining the homology of illegal websites

PendingCN122293403AWeb siteEngineering
This invention provides a method and system for determining the homology of illegal websites, relating to the field of network monitoring technology. The determination method includes: using a headless browser to realistically render the entry addresses of two websites to be determined, and recursively crawling all pages and link information within the corresponding two-level backlink range, starting from the entry address, to obtain the corresponding raw collected data; performing data normalization and cleaning on the raw collected data to obtain the set of second-level backlink domains and statistical data of the two websites to be determined; obtaining multi-feature data based on the set of second-level backlink domains and statistical data; and fusing the various multi-feature data to obtain homology probability data. This invention elevates homology determination from a single, superficial similarity judgment to a comprehensive quantitative assessment based on multiple evidence such as structure, attribution, and content, thereby significantly improving the accuracy of homology association for illegal websites disguised by methods such as domain switching, mirroring, and template reuse.
Owner:CHINA RAILWAY ERYUAN ENGINEERING GROUP CO LTD

Webpage main body information extraction method and device and medium

The invention discloses a webpage main body information extraction method and device and a medium, and relates to the technical field of Internet information processing. The method comprises the following steps: acquiring a document object model (DOM) tree of a webpage and visual rendering information of at least one HTML element through a headless browser; based on the visual rendering information, dividing the webpage into at least one visual block by using a visual separation algorithm, and calculating a visual feature score of the visual block; performing weighted fusion on the visual feature scores of the DOM nodes in the DOM tree and the structural feature scores of the DOM nodes to obtain comprehensive weight scores of the DOM nodes; and determining the DOM node with the comprehensive weight score greater than a preset weight threshold as a container node of the main body content, so as to perform information extraction on the container node and obtain main body information of the webpage. Therefore, webpage main body information extraction with high accuracy, good robustness and excellent efficiency can be realized.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

A markdown document to word method and system based on middleware preprocessing

The application discloses a Markdown document to Word conversion method and system based on middleware preprocessing, and relates to the technical field of document format conversion and data processing. The application aims to solve the technical problem that Mermaid code and LaTeX formula in AI generated Markdown content cannot be directly converted into a Word document. The method comprises the following steps: a conversion service receives a Markdown source text; a chart processor is used to parse and extract Mermaid code blocks, a headless browser rendering engine is called to directly render the Mermaid code blocks to generate a picture file, and a Markdown picture reference string pointing to the picture file is constructed to replace the code blocks in the source text; a formula processor is used to extract LaTeX formula and perform delimiter standardization processing; finally, a document conversion engine is called to convert the processed intermediate state text into a Word document stream, during which the generated picture is automatically embedded and the formula is parsed into a Word editable OMML formula object. Through the automatic preprocessing pipeline, the application realizes lossless and high-fidelity conversion from a complex technical document to an Office document.
Owner:HUAHAI TECHNICAL SERVICES (YUNNAN) CO LTD

Small language webpage self-adaptive collection method and device based on large language model

This application relates to the field of artificial intelligence technology and discloses an adaptive acquisition method and apparatus for minority language web pages based on a large language model. The method includes: loading and rendering a target minority language web page through a headless browser to obtain a document object model tree structure; parsing and semantically annotating the document object model tree structure based on the large language model, identifying dynamic content nodes and generating adaptive content location rules; semantically vectorizing the content of at least two consecutively acquired web pages based on the adaptive content location rules and calculating semantic similarity; dynamically adjusting the initiation frequency of subsequent acquisition requests according to a preset frequency mapping rule based on semantic similarity to generate an acquisition strategy; executing the acquisition strategy to access the target web page, and injecting an operation sequence to simulate human interaction behavior during the acquisition process. This application can achieve complete acquisition of dynamic content of minority language web pages and ensure the continuous and stable acquisition process.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD