Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

31 results about "Headless browser" patented technology

A headless browser is a web browser without a graphical user interface. Headless browsers provide automated control of a web page in an environment similar to popular web browsers, but are executed via a command-line interface or using network communication. They are particularly useful for testing web pages as they are able to render and understand HTML the same way a browser would, including styling elements such as page layout, colour, font selection and execution of JavaScript and AJAX which are usually not available when using other testing methods.

Dynamic webpage content complete acquisition system and acquisition method based on browser rendering engine

The invention discloses a dynamic webpage content complete collection system and method based on a browser rendering engine. A CDP controller main module calls a command to start a headless browser process, and a network monitoring module registers a Network. Response Received event monitor; a DOM (Document Object Model) monitoring module starts a Mutation Observer to monitor the change of a DOM (Document Object Model) tree; the interactive simulation module preloads a mouse click and scroll event instruction set; capturing JSON / XML data returned by an interface in real time by using a network monitoring module, intercepting API response data, generating a DOM snapshot by using a DOM monitoring module, and extracting key structure changes; the DOM monitoring module provides a complete rendering HTML, the network monitoring module provides request metadata, and the interactive simulation module provides an operation time sequence log and packages generated complete webpage data into a WARC file. According to the method, document collection is upgraded to application-level interaction, and a technical implementation path is provided for collection scenes of dynamic data-intensive websites (such as financial public opinion monitoring and e-commerce price tracking) in which Internet resources are stored for a long time.
Owner:SUZHOU JIATU SOFTWARE CO LTD

Method and device for acquiring OG data of multiple social media platforms through links

The invention provides a method and device for acquiring OG data of multiple social media platforms through links, and belongs to the technical field of data acquisition and processing. The method comprises the following steps: acquiring URL links from a plurality of social media platforms and common webpages as source data, wherein the source data correspond to webpage OG data lt; carrying out metagt; the label form is embedded into the lt of the HTML; carrying out headt; a label; performing validity check and standardization on the URL link, capturing webpage content by using an asynchronous HTTP (Hyper Text Transport Protocol) request library, and loading a complete HTML (Hyper Text Markup Language) on a dynamic webpage through a headless browser; the HTML positioning result is analyzed; carrying out metagt; tag: extracting an OG standard field and a platform extension field, and deducing a missing field from other HTML tags; mapping each platform OG field into a unified standard field; and outputting the OG data containing fields such as title and the like through the standardized API interface. According to the method, the multi-platform OG data acquisition efficiency, accuracy and safety are improved, good expansibility is achieved, and the cross-platform data integration requirement is met.
Owner:ONE NETWORK INTEROPERABILITY (BEIJING) TECH CO LTD

Small language webpage adaptive acquisition method and device based on large language model

The invention relates to the technical field of artificial intelligence, and discloses a small language webpage adaptive collection method and device based on a large language model.The method comprises the steps that a target small language webpage is loaded and rendered through a headless browser, and a document object model tree structure is obtained; carrying out analysis and semantic annotation on the document object model tree structure based on a large language model, identifying dynamic content nodes and generating a self-adaptive content positioning rule; based on a self-adaptive content positioning rule, performing semantic vectorization on the at least two continuously collected webpage contents, and calculating semantic similarity; querying a preset frequency mapping rule according to the semantic similarity, dynamically adjusting the initiation frequency of a subsequent acquisition request, and generating an acquisition strategy; and executing the acquisition strategy to access the target webpage, and injecting an operation sequence for simulating human interaction behaviors in the acquisition process. Complete collection of the dynamic content of the small language webpage can be achieved, and continuity and stability of the collection process are guaranteed.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD

Making remote procedure calls over QUIC protocol, and applications thereof

PendingUS20260037350A1Interprogram communicationTransmissionRemote controlProcedure calls
A computer-implemented method is provided that enables remote control of a headless browser. A message from a client to open a connection to a headless browser running on a device remote from the client is received. A command to provision the browser is sent to the headless browser. The command may configure the browser to appear as if the browser is controlled by a human. A command for controlling the headless browser is received from the client. The command is sent to the headless browser for execution. A response to the command is received from the headless browser. Finally, the response is forwarded to the client.
Owner:OXYLABS UAB

Web honeypot automatic script identification and countering method based on large language model and related equipment

The invention discloses a Web honeypot automatic script identification and countering method based on a large language model and related equipment, and the method comprises the steps: obtaining a real IP address of an attack request through a WebRTC protocol; analyzing the attack request through a headless browser and a large language model, judging whether the attack request is an automatic script behavior or not, countering the attack request if the attack request is the automatic script behavior, and guiding the attack request to a honeypot if the attack request is not the automatic script behavior. The method can improve the accuracy of automatic script recognition, improves the honeypot stability, and can be widely applied to the technical field of network security.
Owner:GUANGZHOU UNIVERSITY

Method and apparatus for implementing a headless browser cluster

The present disclosure relates to a method and device for implementing a headless browser cluster. The method comprises: assigning a corresponding accessible IP address to each server in a plurality of servers; wherein the plurality of servers form a cluster; installing a headless browser in each server; after receiving a headless browser starting instruction, starting a preset number of headless browser instances, the started headless browser instances forming a headless browser instance pool; after detecting that a headless browser instance acquisition instruction is received, acquiring a target headless browser instance available from the headless browser instance pool; acquiring a network socket endpoint address of the target headless browser instance; and sending the network socket endpoint address of the target headless browser instance to a client.
Owner:WIRELESS LIFE (HANGZHOU) INFORMATION TECH CO LTD

Method and device for evaluating website function accessibility in IPv6 single-stack environment

The invention discloses a website function accessibility evaluation method and device in an IPv6 single-stack environment, and belongs to the technical field of network function detection. The invention aims to solve the problem that in the prior art, website resource dependence and function accessibility cannot be comprehensively detected in an IPv6 single-stack environment. The method comprises the following steps: dynamically creating an isolated IPv6 single-stack network environment, starting a headless browser to capture all resource requests of a webpage, constructing a resource dependency graph, carrying out IPv6 reachability test on each resource, calculating a weighted accessibility score and generating a visual report. According to the method, the IPv6 single-stack user access behavior can be accurately simulated, the target website and all dependent resources of the target website can be comprehensively evaluated, a quantitative accessibility score and a visual dependency relationship visualization graph are provided, and accurate and operable diagnosis information is provided for website developers and operation and maintenance personnel.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Distributed document generation method based on message queue and headless browser

The invention relates to the technical field of computers, and discloses a distributed document generation method based on a message queue and a headless browser, and the method comprises the steps: receiving a document generation request, packaging the document generation request into a standard task object, and verifying the validity of a rendering target; splitting an original task into sub-tasks, and calculating task fingerprints and time-carrying task fingerprints; the core parameters and the service data loads are separately stored; converting the effective sub-tasks into rendering instructions, and performing queue arrangement according to the calling mode of the sub-tasks; and taking a headless browser as a rendering processing unit, continuously consuming a rendering instruction from the message queue, carrying out webpage rendering, converting the rendered webpage into a document, and responding to a document generation request. The headless browser is used as a core rendering processing unit, high-fidelity restoration of webpage content is achieved, and the visual quality and accuracy of the generated document are improved.
Owner:HUIDIAN TECHNOLOGY (SUZHOU) CO LTD

A method and system for generating a dynamic deception domain name honeypot for active defense

The application provides a kind of active defense-oriented dynamic deception domain name honeypot generation method and system, it is related to network security active deception technical field, specific scheme: based on headless browser access and save the web page of real business, the type of the web page is analyzed and different strategies are recorded to obtain web variables, based on re-renderer component rendering the web variables are cloned to obtain simulation web page, the simulation web page is packaged and combined with known container template to build container image to obtain honeypot device container;Based on the honeypot device container, attack logs are collected and marked, and the optimal cycle is obtained based on the multi-objective optimization algorithm, and the cycle of the false domain name bound to the honeypot device container is dynamically adjusted based on the optimal cycle to obtain the domain name honeypot.The application improves the problem that domain name is difficult to adapt to the change of attacker strategy and is difficult to dynamically optimize the defense strategy according to the behavior of the attacker.
Owner:GUANGZHOU UNIVERSITY

A method for assisting large model networking query

PendingCN122633930AData scrapingOriginal data
The application discloses a kind of methods for assisting large model networking query, the method is by receiving and encapsulating as standardization demand to parse user query request, matching dynamic residential proxy IP pool, IP dynamic rotation and load balancing mechanism are started, generate standardization collection task, realize multi-platform data scraping through multi-tool cooperation, anti-crawling avoidance and high concurrency scheduling, clean, double check and structure transformation are carried out to original data, add standardization field annotation, push structured data to large model and monitor data satisfaction degree;The application is through dynamic residential proxy IP pool and IP rotation mechanism, avoid regional restriction and anti-crawling interception, with headless browser, API identification technology, adapt to dynamic and waterfall page, improve collection integrity and efficiency, after cleaning, check and structure transformation, combined with automatic completion query condition and daily optimization mechanism, the accuracy, timeliness and stability of large model real-time query response are greatly improved.
Owner:JIANGSU LINGJIANG INFORMATION TECH CO LTD

A large PDF file intelligent generation method with front-end and back-end cooperation

The application discloses a front-rear end cooperative large PDF file intelligent generation method, relates to the technical field of computer software, and comprises the following steps: front-end rendering: the front end renders the work order data page by combining a headless browser with a Vue framework and an ElementUI component library, and dynamically displays structured work order details through a responsive data binding mechanism; data fragmentation: large work order data sets are split into N subtasks based on a dynamic fragmentation strategy; and small PDF files are generated in parallel: a containerized rendering cluster based on Kubernetes deployment management is used to allocate rendering tasks to each container; the small PDF files generated by the application through careful design with the help of the Vue and ElementUI frameworks are more attractive in terms of text layout, color matching and graphic display, and the large file after merging is more beautiful than the file generated by a traditional back end, thereby improving user experience.
Owner:BEIJING ZHIXING TONGDE INFORMATION TECH CO LTD

Phishing webpage intelligent detection system and method based on proxy AI

The invention provides a proxy AI-based phishing webpage intelligent detection system and method, and the system is characterized in that the system comprises a link input module which is used for receiving a to-be-detected URL, and carrying out the preliminary filtering; the headless browser control module is used for controlling a headless browser to load a page and managing the life cycle of the page; the page content extraction module is used for extracting page information from the headless browser; manual analysis requirements are reduced through automatic interaction, the detection efficiency is improved, and the method is suitable for large-scale mail security scanning scenes; an interaction path and a decision process can be recorded, a transparent log is provided for security analysis, and subsequent optimization and auditing are facilitated.
Owner:BEIJING DIRECTION BIAO INFORMATION TECHNOLOGY CO LTD

Data processing method and device, electronic equipment, storage medium and program product

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: in response to a received first operation request for a front-end page, generating a screenshot request and sending the screenshot request to middleware service; the screenshot request is generated based on screenshot operation on at least part of page information including third-party resource information in the front-end page; the middleware service starts a headless browser based on the screenshot request, and loads the front-end page through the headless browser; the middleware service performs screenshot on at least part of page information of the headless browser to obtain a target image, and sends the target image to the front-end page; and the front-end page generates target content corresponding to the first operation request based on the target image. According to the data processing method and device, the electronic equipment, the storage medium and the program product, screenshot of cross-domain resources can be achieved.
Owner:BOE TECHNOLOGY GROUP CO LTD

Webpage snapshot signature method and device

The embodiment of the invention discloses a signature method and device for a webpage snapshot, and the method comprises the following steps: loading a target webpage in a uniformly configured headless browser environment, automatically detecting and removing dynamic contents in the target webpage according to a preset dynamic element recognition rule, and generating purified webpage document object model (DOM) data; in the webpage state after the de-dynamic processing, respectively collecting standardized webpage DOM data and a webpage rendering image snapshot; hash abstracts of the webpage DOM data and the webpage rendering image snapshot are calculated respectively, and two abstract values obtained through calculation are combined to generate a webpage snapshot signature; and storing the webpage DOM data, the webpage rendering image snapshot, the webpage snapshot signature and the corresponding metadata information in a secure storage medium together. According to the embodiment of the invention, the consistency and reproducibility of webpage snapshot contents can be improved, the tamper-proof capability of webpage evidence storage is enhanced, and the accuracy and robustness of tamper detection are improved.
Owner:WUXI BAISHANG ZHONGWANG DATA TECHNOLOGY CO LTD

Business mail automatic processing method and system based on robot process automation

The invention discloses a business mail automatic processing method and system based on robot process automation. The method comprises the following steps: automatically analyzing a theme, a text and an attachment of a received mail through a natural language processing technology, extracting key information such as an enterprise ID (Identity), a chemical identifier and a mail type, and converting the key information into structured data; identifying task types according to the data and selecting corresponding automatic processes; a headless browser is used for simulating and logging in a third-party system, and automatic processing including attachment downloading, analysis, screenshot and multi-language translation is executed; the execution state is monitored in real time, the success rate and the average time are counted to optimize the performance, and meanwhile, various abnormities are captured and reports are generated to trigger the processing flow; and cleaning the resources after execution is finished, and generating a notification mail based on a result to send and archive. According to the scheme, full-process automation of mail processing is realized, and efficiency and accuracy are remarkably improved.
Owner:HANGZHOU REACH PROD TECH CO LTD

Front-end and rear-end collaborative intelligent generation method for large PDF (Portable Document Format) file

The invention discloses a front-end and rear-end collaborative intelligent generation method for a large PDF (Portable Document Format) file, which relates to the technical field of computer software and comprises the following steps: front-end rendering: a front end renders a work order data page through a headless browser in combination with a Vue framework and an Element UI (User Interface) component library, and dynamically displays structured work order details through a responsive data binding mechanism; data fragmentation: segmenting the large-scale work order data set into N subtasks based on a dynamic fragmentation strategy; generating small PDF files in parallel; deploying a managed containerized rendering cluster based on Kubernetes, and distributing a rendering task for each container; according to the method, the front end is elaborately designed by means of Vue and Element UI frames, the generated small PDF file is more attractive in character typesetting, color matching and graphic display, the attractiveness of the combined large file is far better than that of a traditional file generated at the rear end, and the user experience is improved.
Owner:BEIJING ZHIXING TONGDE INFORMATION TECH CO LTD

Active defense-oriented dynamic deception domain name honey point generation method and system

The invention provides an active defense-oriented dynamic deception domain name honey point generation method and system, and relates to the technical field of network security active deception, and the specific scheme is as follows: accessing and storing a webpage of a real service based on a headless browser, analyzing the type of the webpage, and recording by adopting different strategies to obtain webpage variables; rendering the webpage variable based on a re-renderer component, then performing page cloning to obtain a simulation webpage, packaging the simulation webpage, and constructing a container mirror image in combination with a known container template to obtain a honey point equipment container; and collecting and marking an attacker request based on the honey point device container to obtain an attack log, obtaining an optimal period based on a multi-objective optimization algorithm, and dynamically adjusting the period of binding the false domain name to the honey point device container based on the optimal period to obtain the domain name honey point. The problems that the domain name is difficult to adapt to the strategy change of an attacker and the defense strategy is difficult to dynamically optimize according to the behavior of the attacker are solved.
Owner:GUANGZHOU UNIVERSITY

Execution method and system of webpage operation task, storage medium and computer equipment

The invention discloses a webpage operation task execution method and system, a storage medium and computer equipment, and the method comprises the steps that a user interaction interface responds to a task creation instruction, creates a webpage operation task, and adds the webpage operation task into a task queue; monitoring the task queue by each service instance in a back-end service instance cluster, and when monitoring that a newly added webpage operation task exists in the task queue, judging whether to obtain an execution permission of the webpage operation task or not based on a self state and a preemption state of the webpage operation task; and for the target service instance which obtains the execution permission, the target service instance updates the state of the target service instance to a working state, starts a target headless browser in a task execution unit based on the webpage operation task, and executes a target webpage operation through the target headless browser until the webpage operation task is completed. And closing the target headless browser, and updating the state of the target headless browser to be an idle state.
Owner:BEIJING WATERDROP TECH GRP CO LTD

Data quality inspection method and device, electronic equipment and storage medium

The embodiment of the invention discloses a data quality inspection method and device, electronic equipment and a storage medium, and relates to the technical field of data quality intelligent inspection. The method comprises the steps that a headless browser technology, a User-Agent camouflage technology and window simulation are adopted to collect a first screenshot and a second screenshot corresponding to key indexes in a target billboard area of a first terminal and a target billboard area of a second terminal respectively; performing standardization processing and key area enhancement on the first screenshot and the second screenshot; the standardization processing comprises size normalization, color unification and background purification; calling a deep learning model to intelligently identify the processed first screenshot and second screenshot to obtain first terminal data and second terminal data, and comparing to obtain a difference value; and setting a difference threshold value, and automatically triggering an alarm mechanism when the difference value exceeds the difference threshold value. According to the invention, the problems of data inconsistency, difficulty in troubleshooting and lack of real-time monitoring in the prior art are solved.
Owner:SHENZHEN SHUZHI XINCHENG TECHNOLOGY CO LTD

A method and system for determining the homology of illegal websites

PendingCN122293403AWeb siteEngineering
This invention provides a method and system for determining the homology of illegal websites, relating to the field of network monitoring technology. The determination method includes: using a headless browser to realistically render the entry addresses of two websites to be determined, and recursively crawling all pages and link information within the corresponding two-level backlink range, starting from the entry address, to obtain the corresponding raw collected data; performing data normalization and cleaning on the raw collected data to obtain the set of second-level backlink domains and statistical data of the two websites to be determined; obtaining multi-feature data based on the set of second-level backlink domains and statistical data; and fusing the various multi-feature data to obtain homology probability data. This invention elevates homology determination from a single, superficial similarity judgment to a comprehensive quantitative assessment based on multiple evidence such as structure, attribution, and content, thereby significantly improving the accuracy of homology association for illegal websites disguised by methods such as domain switching, mirroring, and template reuse.
Owner:CHINA RAILWAY ERYUAN ENGINEERING GROUP CO LTD

A remote headless browser and related protocols

PCT designated stageWO2026027533A1TransmissionRemote controlEngineering
A computer-implemented method is provided that enables remote control of a headless browser. A message from a client to open a connection to a headless browser running on a device remote from the client is received. A command to provision the browser is sent to the headless browser. The command may configure the browser to appear as if the browser is controlled by a human. A command for controlling the headless browser is received from the client. The command is sent to the headless browser for execution. A response to the command is received from the headless browser. Finally, the response is forwarded to the client.
Owner:OXYLABS UAB

Webpage main body information extraction method and device and medium

The invention discloses a webpage main body information extraction method and device and a medium, and relates to the technical field of Internet information processing. The method comprises the following steps: acquiring a document object model (DOM) tree of a webpage and visual rendering information of at least one HTML element through a headless browser; based on the visual rendering information, dividing the webpage into at least one visual block by using a visual separation algorithm, and calculating a visual feature score of the visual block; performing weighted fusion on the visual feature scores of the DOM nodes in the DOM tree and the structural feature scores of the DOM nodes to obtain comprehensive weight scores of the DOM nodes; and determining the DOM node with the comprehensive weight score greater than a preset weight threshold as a container node of the main body content, so as to perform information extraction on the container node and obtain main body information of the webpage. Therefore, webpage main body information extraction with high accuracy, good robustness and excellent efficiency can be realized.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Map screen capture method and system, electronic equipment and storage medium

The invention provides a map screen capture method and system, electronic equipment and a storage medium. The method comprises the steps that screen capture configuration parameters are acquired; in response to the screen capture instruction, performing multiple times of screen capture on the map by using the screen capture configuration parameters and the headless browser to obtain a plurality of map screen capture pictures; and splicing the plurality of map screenshot pictures to obtain a final map screenshot picture. According to the scheme, the map is subjected to multiple times of screen capture through the screen capture configuration parameters and the headless browser, and the multiple map screen capture pictures obtained through multiple times of screen capture are spliced to obtain the large-range final map screen capture picture.
Owner:BEIJING CHINA INDEX SHIZHENG INFORMATION

Page translation content detection method, related device and detection tool

The invention discloses a page translation content detection method, a related device and a detection tool, and relates to the field of artificial intelligence, the detection tool comprises a browser plug-in layer, a headless browser, a content extraction engine, a preset language model and a report generation module, and the method comprises the steps that the browser plug-in layer triggers page initialization loading and traverses a page; in the traversing process, page DOM change is triggered through the headless browser, and after change is completed, the current page is recorded; the content extraction engine extracts and cleans text content of a current page to obtain translated text content, the preset language model analyzes the translated text content and marks translated abnormal content to obtain a detection result, and the report generation module generates a translation detection report according to the detection result. According to the method, page DOM change is automatically triggered, detection is carried out after change is completed, automatic detection is carried out through the preset language model, omission is reduced, the translation detection report is output, translation abnormity can be conveniently determined, and the test efficiency is effectively improved.
Owner:AGRICULTURAL BANK OF CHINA

Video generation using a headless browser

Techniques for capturing videos of content displayed and / or generated by a web-based application are described herein. The videos may be captured by automating screenshots of the content within a headless browser at a fixed frame interval. The headless browser may be used to automate control of the web-based application to generate the content and capture screenshots of the content as the content would appear if the web-based application was being accessed through a traditional web browser graphical user interface. The screenshots may be captured at specific frame intervals while the content is being generated. Additionally, the headless browser or a server executing headless browser may wait for the web-based application to load individual frames of the content before capturing screenshots. The captured screenshots may then be combined to generate a video of the content.
Owner:ZOOX INC

Operation data report generation method and device, electronic equipment and storage medium

The embodiment of the invention discloses an operation data report generation method and device, electronic equipment and a storage medium, historical operation data of a cloud application firewall is obtained by responding to a request for generating an operation data report for the cloud application firewall, the historical operation data is written into a pre-created webpage file, and the operation data report is generated. Executing the webpage file through the headless browser, calling the data analysis module to analyze based on the historical operation data according to the calling path, drawing a result obtained based on the analysis of the historical operation data in the headless browser to obtain a target webpage, converting the target webpage through the headless browser, generating an operation data report, and displaying the operation data report. The running data report is generated in the mode that the headless browser is introduced to process the target webpage, automatic generation of the running data report can be achieved, the generation efficiency of the running data report is effectively improved, and the method and device are widely applied to scenes such as the cloud technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Subprogram testing method and device, equipment and storage medium

The invention discloses a subprogram testing method and device, equipment and a storage medium, and the method comprises the steps: obtaining a program code of a to-be-tested subprogram, and converting the program code into a target webpage code; calling a headless browser in a continuous integration environment, running the subprogram based on the target webpage code, and rendering a graphic user page of the subprogram; and performing an end-to-end test on the subprogram in a running state based on the graphic user page of the subprogram by using an automatic test script and the headless browser to obtain a test result of the subprogram. By means of the method and device, automatic testing of the subprograms can be achieved, and the testing efficiency and testing flexibility of the subprograms are improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A markdown document to word method and system based on middleware preprocessing

The application discloses a Markdown document to Word conversion method and system based on middleware preprocessing, and relates to the technical field of document format conversion and data processing. The application aims to solve the technical problem that Mermaid code and LaTeX formula in AI generated Markdown content cannot be directly converted into a Word document. The method comprises the following steps: a conversion service receives a Markdown source text; a chart processor is used to parse and extract Mermaid code blocks, a headless browser rendering engine is called to directly render the Mermaid code blocks to generate a picture file, and a Markdown picture reference string pointing to the picture file is constructed to replace the code blocks in the source text; a formula processor is used to extract LaTeX formula and perform delimiter standardization processing; finally, a document conversion engine is called to convert the processed intermediate state text into a Word document stream, during which the generated picture is automatically embedded and the formula is parsed into a Word editable OMML formula object. Through the automatic preprocessing pipeline, the application realizes lossless and high-fidelity conversion from a complex technical document to an Office document.
Owner:HUAHAI TECHNICAL SERVICES (YUNNAN) CO LTD

Video synthesis method, device, equipment, medium and product

The invention provides a video synthesis method and device, equipment, a medium and a product, the method is applied to a server, and the method comprises the following steps: receiving a video synthesis request, the video synthesis request comprising a replacement visual content and an identifier of a target video template, the target video template is obtained by converting a closed-source original video template into an open-source format file; based on the identifier of the target video template, loading the target video template to a headless browser of a server through a preset browser interaction protocol; and performing video synthesis based on the target video template and the replacement visual content through the headless browser to obtain a target video. According to the invention, the dependence of the video synthesis process on a specific platform can be reduced, and the autonomous controllability is improved.
Owner:IFLYTEK CO LTD

Small language webpage self-adaptive collection method and device based on large language model

This application relates to the field of artificial intelligence technology and discloses an adaptive acquisition method and apparatus for minority language web pages based on a large language model. The method includes: loading and rendering a target minority language web page through a headless browser to obtain a document object model tree structure; parsing and semantically annotating the document object model tree structure based on the large language model, identifying dynamic content nodes and generating adaptive content location rules; semantically vectorizing the content of at least two consecutively acquired web pages based on the adaptive content location rules and calculating semantic similarity; dynamically adjusting the initiation frequency of subsequent acquisition requests according to a preset frequency mapping rule based on semantic similarity to generate an acquisition strategy; executing the acquisition strategy to access the target web page, and injecting an operation sequence to simulate human interaction behavior during the acquisition process. This application can achieve complete acquisition of dynamic content of minority language web pages and ensure the continuous and stable acquisition process.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD