Script code generation method and device, electronic equipment and storage medium

CN122837804APending Publication Date: 2026-09-29BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611123023.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但这种方式需要耗费很大的人力和时间成本,而且效率低下

Benefits of technology

[0018]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837804A_ABST
    Figure CN122837804A_ABST
Patent Text Reader

Abstract

This disclosure provides a script code generation method, apparatus, electronic device, and storage medium, relating to artificial intelligence fields such as natural language processing, large language models, and front-end web development. The method may include: obtaining user-generated request description information for a target page, the request description information including: a description of a request to add a predetermined function to the target page; obtaining the page document object model (DOM) information of the target page, and determining target summary information of the target page based on the DOM information; generating first input information based on the request description information and the target summary information; generating corresponding script code using a large language model based on the first input information; determining target code based on the script code; and using the target code to implement the predetermined function. Applying the solution described in this disclosure can save manpower and time costs, and improve processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing, large language models and front-end web development, and especially to script code generation methods, devices, electronic devices and storage media. Background Technology

[0002] When business users access various business systems using browsers, they often encounter situations where page functionality is incomplete, such as missing batch operation buttons, forms requiring repeated filling, and data needing to be manually copied and exported. Traditionally, professional developers need to manually write the relevant script code to overcome these functionalities. However, this method is costly in terms of manpower and time, and is inefficient. Summary of the Invention

[0003] This disclosure provides methods, apparatus, electronic devices, and storage media for generating script code.

[0004] A script code generation method, comprising:

[0005] Obtain user request description information for the target page, the request description information including: a description of the request to add a pre-defined function to the target page;

[0006] Obtain the page document object model information of the target page, and determine the target summary information of the target page based on the page document object model information;

[0007] First input information is generated based on the requirement description information and the target summary information. Based on the first input information, corresponding script code is generated using a large language model. Target code is determined based on the script code. The target code is used to implement the predetermined function.

[0008] A script code generation device includes: a requirement acquisition module, an information acquisition module, and a code generation module;

[0009] The requirement acquisition module is used to acquire the requirement description information sent by the user for the target page. The requirement description information includes: description information of requesting the addition of a pre-defined function on the target page.

[0010] The information acquisition module is used to acquire the page document object model information of the target page, and determine the target summary information of the target page based on the page document object model information;

[0011] The code generation module is used to generate first input information based on the requirement description information and the target summary information, generate corresponding script code using a large language model based on the first input information, determine target code based on the script code, and the target code is used to implement the predetermined function.

[0012] An electronic device, comprising:

[0013] At least one processor; and

[0014] A memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described above.

[0016] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.

[0017] A computer program product includes a computer program / instructions that, when executed by a processor, implement the method described above.

[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0019] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0020] Figure 1 This is a flowchart of an embodiment of the script code generation method described in this disclosure;

[0021] Figure 2 This is a flowchart illustrating an embodiment of the method for filtering DOM elements as described in this disclosure;

[0022] Figure 3 This is a flowchart of an embodiment of the method for generating target summary information as described in this disclosure;

[0023] Figure 4 This is a flowchart of an embodiment of the method for determining target code based on script code as described in this disclosure;

[0024] Figure 5 This is a schematic diagram of the composition structure of the first embodiment 500 of the script code generation device described in this disclosure;

[0025] Figure 6This is a schematic diagram of the composition structure of the second embodiment 600 of the script code generation device described in this disclosure;

[0026] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0028] Furthermore, it should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0029] Figure 1 This is a flowchart illustrating an embodiment of the script code generation method described in this disclosure. Figure 1 As shown, the specific implementation methods are as follows.

[0030] In step 101, the user's request description information for the target page is obtained. The request description information includes: a description of the request to add a pre-defined function to the target page.

[0031] In step 102, the document object model (DOM) information of the target page is obtained, and the target summary information of the target page is determined based on the page DOM information.

[0032] In step 103, first input information is generated based on the requirement description information and target summary information. Based on the first input information, corresponding script code is generated using a large language model. Target code is determined based on the script code, and the target code is used to implement the predetermined function.

[0033] The solution described in the above method embodiment can automatically generate the required target code based on the user's request description information and the target page's target summary information, thereby saving manpower and time costs and improving processing efficiency. Moreover, by leveraging the powerful reasoning ability of the large language model, the accuracy of the obtained target code can be improved.

[0034] The target page, requirement description information, etc., in the embodiments described in this disclosure are not targeted at any specific user and are not intended to reflect the personal information of any specific user. The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, involved in the technical solutions of this disclosure comply with relevant laws and regulations and do not violate public order and good morals.

[0035] For a target page, users can submit a request description, such as "Add a button at the top of the page to export financial news." Correspondingly, the page's DOM information can be obtained, and the target summary information can be determined based on this DOM information.

[0036] In some embodiments of this disclosure, DOM elements that are not visible to the user can be filtered out from the DOM elements in the page DOM information first, and the remaining DOM elements can be determined as candidate elements. Then, M types of target elements of a predetermined type can be extracted from the candidate elements, and the element summary of each target element can be obtained respectively, where M is a positive integer greater than 1. In addition, the plain text context of the target page can be obtained. Furthermore, target summary information can be generated based on the element summary and the plain text context.

[0037] By filtering out DOM elements invisible to the user, all subsequent extracted target elements are valid elements that the user can perceive, avoiding interference from invalid elements in subsequent processing. By extracting multiple predefined target elements from candidate elements and obtaining element summaries for each target element, multi-dimensional and categorized extraction of page structured information can be achieved. This allows the semantic information of different types of elements to be fully preserved and expressed differently, providing rich and accurate page structure information for subsequent code generation. In addition, by obtaining the plain text context of the target page, unstructured business semantic information in the page can be effectively supplemented, enabling the large language model to understand the scenarios and contexts involved in user needs in conjunction with the overall semantics of the page, further improving the accuracy of the obtained script code.

[0038] In some embodiments of this disclosure, the method of filtering out DOM elements that are not visible to the user from each DOM element in the page DOM information may include: for any DOM element, performing the following first processing: determining whether the DOM element is located within a predetermined container; if it is located within the predetermined container, filtering out the DOM element; if it is not located within the predetermined container, determining whether the width of the DOM element is 0; if the width is 0, filtering out the DOM element; if the width is not 0, determining whether the DOM element matches a predetermined hidden style; if the hidden style matches, filtering out the DOM element.

[0039] Accordingly, Figure 2This is a flowchart illustrating an embodiment of the method for filtering DOM elements as described in this disclosure. Figure 2 As shown, the specific implementation methods are as follows.

[0040] In step 201, each DOM element is processed according to the methods shown in steps 202-206.

[0041] In step 202, it is determined whether the DOM element is located within a predetermined container. If so, step 203 is executed; otherwise, step 204 is executed.

[0042] For example, it can determine whether the DOM element is located within a container of tags such as script, style, noscript, or template.

[0043] In step 203, the DOM element is filtered out, and then the process ends.

[0044] In step 204, determine whether the width of the DOM element is 0. If it is, proceed to step 203; otherwise, proceed to step 205.

[0045] For example, you can determine whether the width of a DOM element is 0 by using the getBoundingClientRect() function to get the element's bounding rectangle.

[0046] In step 205, it is determined whether the DOM element matches the predefined hidden style. If yes, step 203 is executed; otherwise, step 206 is executed.

[0047] For example, you can determine whether a DOM element has one of the following hidden styles by getting the computed style function (getComputedStyle()): display === none, visibility === hidden, or opacity === 0.

[0048] In step 206, the DOM element is identified as a candidate element, and then the process ends.

[0049] By adopting the above three-level filtering method, DOM elements that are not visible to the user can be identified and filtered out in a comprehensive and accurate manner, so that the candidate elements that are ultimately retained are all valid elements that the user can perceive, thereby improving the efficiency of subsequent processing and the accuracy of processing results.

[0050] For each candidate element, M types of target elements of a predetermined type can be extracted, and the element summary of each target element can be obtained. In addition, the plain text context of the target page can also be obtained.

[0051] In some embodiments of this disclosure, the target element may include: interactive element, form structure, table structure, and image, and the element summary may include: interactive element summary, form structure summary, table structure summary, and image summary.

[0052] Accordingly, five types of structured information can be obtained, as shown below:

[0053] 1) Summary of interactive elements

[0054] It can iterate through each candidate element and identify interactive elements, such as buttons, input boxes, and select boxes, using a multi-dimensional interactivity judgment algorithm based on artificial intelligence (AI). For each interactive element, it can obtain a summary of the interactive element, including the tag name, visible text content, adjacent context, parent path, and key attributes.

[0055] 2) Form Structure Summary

[0056] It can extract all form structures and, for each form structure, retrieve its internal visible input fields, including field labels, field selectors, field attributes, and current values, as a form structure summary.

[0057] 3) Table Structure Summary

[0058] It can extract all table structures and, for each table structure, obtain table structure summaries such as header text, header visibility, and row count estimation.

[0059] 4) Image Summary

[0060] It can extract all visible images and, for each image, extract image summaries such as alternative text (alt) and source address (src).

[0061] 5) Plain text context

[0062] For example, it can clone the page body (document.body), remove noisy nodes such as script, style, noscript, scalable vector graphics (SVG), inline frames (iframe), canvas (canvas), template, external resource links (link), and metadata (meta), and extract the cleaned plain text context.

[0063] Through the above processing, multi-dimensional and categorized extraction of structured information of the page can be achieved, so that the semantic information of different types of elements can be fully preserved and expressed in a differentiated manner. At the same time, the element summaries of different types of elements have different focuses, which enables the large language model to accurately understand the page layout and function according to the characteristics of different types of elements, thereby improving the adaptability and accuracy of the generated script code to the target page.

[0064] In some embodiments of this disclosure, a corresponding Cascading Style Sheets (CSS) selector can be generated for a predetermined target element, and the CSS selector can be added to the element summary of the corresponding target element. In addition, the obtained element summary can be compressed, and the required target summary information can be generated based on the compression result and the plain text context.

[0065] The specific target elements included in the predefined list can be determined according to actual needs, such as interactive elements. In the scheme described in this disclosure, a hierarchical priority strategy can be used to generate the corresponding CSS selector for each target element.

[0066] Specifically, the most stable CSS selector can be generated for the target element according to the priority of ID (id) > name (name) > tag name.class name (tag.class) > tag name (tag).

[0067] For example, if the target element has a globally unique id attribute, an "id selector (#id)" can be generated directly. If the target element has a name attribute, an attribute selector (tag[name="xxx"]) can be generated. If the target element has a class attribute, a "multi-class name tag selector (tag.class1.class2)" can be generated by taking the first two class names. If none of the above are satisfied, the tag name of the target element can be used as a fallback selector.

[0068] By adopting the above processing method, the generated CSS selectors can remain robust when the page undergoes minor changes, and the probability of script code becoming invalid due to target page upgrades can be reduced.

[0069] In addition, the element summaries can be compressed. In some embodiments of this disclosure, for each type of target element, at most P element summaries corresponding to the first P target elements can be retained in order of their appearance position on the target page, where P is a positive integer greater than 1. For the text fragments in each element summary, at most Q characters can be retained, where Q is a positive integer greater than 1. For the input fields in each form structure summary, at most R input fields can be retained, where R is a positive integer greater than 1.

[0070] By employing the above processing method, a three-layer truncation mechanism can prioritize the retention of key target elements appearing at the beginning of the page, key text fragments in the summaries of each element, and key input fields in each form within a limited large language model context window. This effectively controls the length of the first input information and avoids the problem of exceeding the context window limit due to overly complex page structures, too many elements, or excessively long text. At the same time, it preserves the most critical and relevant structured information on the page that meets the user's needs to the greatest extent possible, balancing information integrity and length constraints, and improving the stability and versatility of the solution in large and complex page scenarios.

[0071] Based on the above introduction, Figure 3 This is a flowchart illustrating an embodiment of the method for generating target summary information as described in this disclosure. Figure 3 As shown, the specific implementation methods are as follows.

[0072] In step 301, DOM elements that are not visible to the user are filtered out from the DOM elements in the page DOM information, and the remaining DOM elements are determined as candidate elements.

[0073] In step 302, M types of target elements are extracted from the candidate elements, the element summary of each target element is obtained, and the plain text context of the target page is obtained.

[0074] In step 303, a corresponding CSS selector is generated for the predetermined target element, and the CSS selector is added to the element summary of the corresponding target element.

[0075] In step 304, the element digest is compressed.

[0076] For example, three levels of compression thresholds can be set, and the three levels can be executed sequentially with decreasing priority.

[0077] First, for each type of target element, the element summary of the first P target elements can be retained according to their appearance position on the target page, from first to last. P is a positive integer greater than 1, and the specific value can be determined according to actual needs, such as 40. That is to say, in the first level, "element number compression" can be performed, setting "maximum number per category (MAX_EACH) = 40", that is, a maximum of 40 target elements per category (interactive elements, form structures, table structures, images) can be extracted, and those exceeding this limit will be truncated.

[0078] Then, for each element's summary text segment, up to the first Q characters can be retained, where Q is a positive integer greater than 1, and the specific value can be determined according to actual needs, such as 80. That is to say, in the second level, "text length compression" can be performed, setting "maximum text length (MAX_TEXT) = 80", and each text field (such as visible text content, table header text, etc.) can retain a maximum of 80 characters, and will be truncated if it exceeds this limit.

[0079] Next, for each form structure summary, up to the first R input fields can be retained, where R is a positive integer greater than 1. The specific value of R can be determined according to actual needs, such as 10. In other words, "form field compression" can be performed in the third level, setting "maximum number of fields (MAX_FIELDS) = 10". Each form structure can retain a maximum of 10 input fields, and fields exceeding this limit will be truncated.

[0080] In step 305, target summary information is generated based on the compression result and the plain text context.

[0081] For example, the compressed element summaries can be integrated and spliced ​​with the plain text context to form structured target summary information.

[0082] After obtaining the target summary information, the first input information corresponding to the large language model can be generated based on the requirement description information and the target summary information. Based on the first input information, the corresponding script code can be generated using the large language model.

[0083] For example, if the user has not sent any other request descriptions within the recent scheduled time period besides the current one, then the first input information can be generated based on the current request description, target summary information, and scheduled prompts. If the user has sent other request descriptions within the recent scheduled time period, then the first input information can be generated based on the current request description, the most recent T historical request descriptions, target summary information, and scheduled prompts, where T is a positive integer and its specific value can be determined according to actual needs. Furthermore, the first input information can be input into a large language model to obtain the script code generated by the large language model.

[0084] In some embodiments of this disclosure, during the process of generating script code using a large language model, a block-appending streaming display method can also be adopted to display the generated script code to the user in real time.

[0085] For example, a full-screen Monaco code editing window can be injected into a tab on the target page using the context function `tabPresenter.executeJavaScript(tabId,buildCodeModalScript(...))`, overwriting the original page content for focused editing. The window includes an editor that supports syntax highlighting and code folding for lightweight interpreted programming languages ​​(JavaScript), displays an "AI-generated code" label at the top, and provides save and close buttons at the bottom. During the script code generation process using a large language model, a block-appending streaming display method can be used to show the generated script code to the user in real time. This can be achieved by setting a typing buffer that refreshes every 12 milliseconds, sending up to 5 characters per refresh, creating a typewriter-like visual effect that allows the user to observe the script code generation process simultaneously. After generation is complete, an update message can be sent for final full synchronization, ensuring that the script code in the editor matches the actually generated script code.

[0086] Based on the script code, the target code can be further determined. In some embodiments of this disclosure, in response to the completion of script code generation, a retry parameter can be set, initially set to 0, and the following second process can be performed: static syntax parsing verification is performed on the newly generated script code; in response to the verification being passed, the newly generated script code is determined as the target code; in response to the verification failing and the retry parameter being less than L, correction prompt information can be generated for the newly generated script code; and second input information can be generated based on the requirement description information, target summary information, and correction prompt information; then, based on the second input information, the script code can be regenerated using a large language model; and the value of the retry parameter can be incremented by 1, and then the second process can be repeated, where L is a positive integer; in response to the verification failing and the retry parameter being equal to L, an error prompt information can be displayed to the user.

[0087] Accordingly, Figure 4 This is a flowchart illustrating an embodiment of the method for determining target code based on script code as described in this disclosure. Figure 4 As shown, the specific implementation methods are as follows.

[0088] In step 401, in response to the completion of script code generation, the retry parameter is set, with an initial value of 0.

[0089] In step 402, the newly generated script code is statically parsed and verified, and it is determined whether the verification passes. If it does, step 403 is executed; otherwise, step 404 is executed.

[0090] For example, the new function constructor can be used to perform static syntax parsing on the newly generated script code, capture syntax errors (SyntaxError), and record error types and descriptions. If syntax errors are found, it can be determined that validation failed.

[0091] In step 403, the newly generated script code is identified as the target code, and then the process ends.

[0092] In step 404, it is determined whether the retry parameter is less than L. If so, step 405 is executed; otherwise, step 406 is executed.

[0093] The specific value of L can be determined according to actual needs, such as 3.

[0094] In step 405, a correction prompt message is generated for the newly generated script code, and a second input message is generated based on the requirement description information, target summary information and correction prompt message. Based on the second input message, the script code is regenerated using the large language model, and the value of the retry parameter is incremented by 1. Then, step 402 is executed.

[0095] Error feedback can be constructed, and a correction prompt message can be generated for the previous error response of the large language model. This message can include the error correction information generated based on the error type and error description information of the existing syntax error.

[0096] The system can generate second input information based on the requirement description, target summary information, and correction prompts. This second input information can then be input into a large language model to obtain regenerated script code.

[0097] In step 406, an error message is displayed to the user, and then the process ends.

[0098] In addition, the newly generated script code can be saved for manual review, etc.

[0099] By adopting the above processing method, a self-correcting closed loop of code can be formed through static syntax parsing and verification, error feedback mechanism and limited number of automatic retry error correction, so as to automatically identify and fix syntax errors in script code, effectively improve the automation level of code generation process and the accuracy of output code.

[0100] Furthermore, it can be seen that the scheme described in this disclosure adopts the pipeline pattern of the "extraction-generation-verification" reasoning and action collaborative framework (ReAct, Reasoning + Acting), and accordingly, a finite state machine can be maintained to drive the entire implementation process.

[0101] The complete states of the finite state machine may include: idle, extracting structured summaries, generating code, validating in the sandbox, successful validation, and failed validation with the maximum number of retries reached.

[0102] The triggering conditions and guarding conditions for state transitions can be:

[0103] 1) idle→extracting: The trigger condition is that the user sends a description of the request, and the guard condition is that the target page has been fully loaded;

[0104] 2) Extracting → Generating: The trigger condition is that the structured summary extraction is completed, and the guard condition is that the extraction result (i.e., the target summary information) is not empty;

[0105] 3) generating→validating: The trigger condition is the end of the streaming output of the large language model (receiving the stream completion marker), and the guard condition is that the generated code is not empty;

[0106] 4) validating→success: The trigger condition is that the static syntax parsing validation passes;

[0107] 5) validating→generating: Triggered by static syntax parsing validation failure, guarded by the current retry parameter. <L;

[0108] 6) validating→error: The trigger condition is that static syntax parsing validation fails, and the guard condition is that the current retry parameter is ≥L.

[0109] By binding the ReAct pipeline to a finite state machine, state transitions can be controlled by clear triggering and guarding conditions, enabling the orderly flow of target summary information extraction, script code generation, static syntax parsing and verification, and automatic retry processes. This avoids multi-stage task concurrency conflicts, ensures clear state boundaries and complete branch logic, and effectively improves the operational stability and process controllability of the entire implementation process.

[0110] In some embodiments of this disclosure, one or all of the following may also be performed: updating the target code according to the user's editing operation on the target code; and, in response to obtaining the user's debugging instructions, displaying the interface effect of the target page after adding the predetermined function according to the target code.

[0111] In other words, users can manually edit the target code and debug it directly on the target page, achieving a "what you see is what you get" development experience.

[0112] In practical applications, an independent code version can be maintained for each target code, and a three-level priority strategy can be used to determine the valid code. The priority levels are as follows, from highest to lowest: First priority - user-edited version, i.e., the target code has been manually edited and saved by the user; Second priority - page cached version, i.e., the target code has been cached by the target page through message passing (postMessage) even if the user has not manually edited the target code; Third priority - target code, i.e., the target code obtained directly from the output of the large language model.

[0113] A user can issue a debugging command by clicking the "Debug" button on the predefined interface. Accordingly, the target code can be obtained according to the three-level priority strategy mentioned above. Then, the browser's developer tools can be opened in the target page's tab. The target code is then packaged into an Immediately Invoked Function Expression (IIFE) with complete exception handling and execution status reporting. The IIFE can then be executed in the real DOM context of the target page. The execution result can be output through the browser's developer tools console, allowing the user to directly observe the actual interface effect after the predefined function is added to the target page. At the same time, the target code runs directly in the real page environment, and its behavior is completely consistent with the actual effect after it is finally released as a plugin, ensuring the effectiveness and reliability of the debugging results.

[0114] In some embodiments of this disclosure, in response to receiving a user's release instruction, a configuration dialog box can be displayed, and the plugin configuration information entered by the user in the configuration dialog box can be obtained. Then, the target code can be released as a persistent script plugin according to the plugin configuration information.

[0115] For example, after the user confirms the target code and debugs it successfully, they can click the "Publish as Script" button on the scheduled interface. Accordingly, a configuration dialog box will be displayed, requiring the user to fill in the following information: 1) Plugin name: a unique identifier that only allows letters, numbers, underscores, and hyphens; 2) Uniform Resource Locator (URL) matching rules: wildcard mode is supported, such as "*: / / *.example.com / *", one rule per line; 3) Function description: optional, for unified management and identification in the plugin asset center.

[0116] Based on the plugin configuration information mentioned above, the target code can be published as a persistent script plugin. For example, the uniqueness of the plugin name can be verified. After successful verification, a subdirectory named after the plugin name can be created in the user data directory. Then, the target code can be written to a predefined file, and the 256-bit SecureHash Algorithm (SHA-256) hash value of the target code can be obtained. The first 16 hexadecimal characters are taken as the version identifier. Furthermore, a script item configuration object can be generated based on the plugin configuration information, entry file path, version identifier, and other fields. This configuration object can be written to a centralized configuration file to complete the registration, ultimately generating a persistent script plugin that can be stored for a long time.

[0117] Subsequently, automatic plugin injection can be implemented. That is, each time the target page is loaded, the following matching injection process can be automatically executed: obtain the URL of the target page, traverse the script plugins that are enabled, match the URL matching rules of each script plugin with the URL of the target page, and inject and execute the target code of the successfully matched script plugin into the target page, thereby enabling the script plugin to enhance the functionality of the target page.

[0118] When the script plugin is enabled or its code changes, the injection logic of the matching page (such as the target page mentioned above) can be automatically refreshed to ensure that the changes take effect in real time.

[0119] As can be seen, the solution described in this disclosure adopts an asset-based approach, solidifying AI capabilities into persistent script plugins that can be automatically injected via URL matching. These plugins take effect as soon as the page loads, without requiring users to manually trigger or re-describe their needs. Furthermore, the script plugins can execute accurately according to preset logic, avoiding the operational illusion and latency costs associated with real-time inference.

[0120] In some embodiments of this disclosure, the script plugin may also be stored in the plugin asset center, and users may be able to perform one or any combination of the following operations on the script plugin: enable operation, disable operation, online code editing operation, and delete operation. In addition, the script plugin may include metadata and code, and the metadata and code may be stored in the plugin asset center using a separate storage architecture.

[0121] All script plugins can be centrally managed in the plugin asset center for full lifecycle management. Users can perform one or any combination of the following operations on script plugins: 1) Enable or disable: Disabling the script plugin will immediately stop injecting into the matching page, achieving a second-level on / off switch for functionality; 2) Online code editing: Edit the script source code online; 3) Deletion: Simultaneously clean up locally stored script files and configuration items to avoid remnants.

[0122] Furthermore, the plugin asset center enables standardized distribution and team sharing of script plugins. For example, multiple script plugins can be packaged into a compressed file (zip) and deployed to users' local machines with a single click via batch import. This allows best practices within an enterprise to be quickly replicated and promoted across departments, forming reusable organizational-level knowledge assets.

[0123] In addition, the plugin asset center can adopt a metadata-code separation storage architecture. For example, the specific design method can be as follows: (1) Metadata layer: Metadata can be stored in the scripts table of the local lightweight embedded relational database (SQLite). Each record can include: id (auto-incrementing primary key), plugin name, function description, URL matching rules, enabled boolean value, version identifier, creation timestamp (createdAt), update timestamp (updatedAt), etc.; (2) Code layer: The code of the script plugin can be stored in the scripts / directory of the local file system as an independent code file; (3) Association mechanism: The version identifier in the metadata can be associated with the file name of the code file. Metadata changes (such as enabling status switching) do not affect the code file, while code changes trigger the update of the version identifier and update timestamp in the metadata; (4) Batch export: The metadata and corresponding code files of one or more selected script plugins can be packaged into zip format for easy distribution and deployment across teams.

[0124] By storing metadata and code separately, system coupling can be reduced, making it easier to manage and maintain metadata and code independently, and reducing storage redundancy.

[0125] In some embodiments of this disclosure, each script plugin in the plugin asset center may have a corresponding version identifier. The version identifier may be the first W hexadecimal characters extracted from the SHA-256 hash value of the script plugin, where W is a positive integer greater than 1, and the specific value may be determined according to actual needs, such as 16. In addition, the version identifier may be written to the version history. In response to the update of the script plugin (such as when the user performs an online code editing operation), the version identifier may be regenerated and written to the version history. The version identifiers in the version history are saved in the form of a version chain. Furthermore, the content snapshots of the script plugins corresponding to different version identifiers may also be saved separately.

[0126] For example, the SHA-256 hash value can be calculated based on the 8-bit Unicode Transformation Format (UTF-8) text content of the script plugin, and the first 16 hexadecimal characters can be used as the version identifier. Furthermore, each time a newly generated version identifier is written to the version history, a corresponding record can be appended. This record can include fields such as the version identifier, timestamp, and previous version identifier (previousHash), where previousHash points to the previous version identifier, forming an immutable version chain. In addition, snapshots of the script plugin content corresponding to different version identifiers can be saved separately.

[0127] By saving version identifiers in a version chain and saving content snapshots corresponding to different version identifiers, every change to the script plugin can be traced, compared (e.g., by comparing the content snapshots of two versions, the changed lines of code can be highlighted), and rolled back. This makes it easier to quickly locate the change points when problems occur during updates and improves version consistency during team collaboration.

[0128] The following specific examples further illustrate the solution described in this disclosure.

[0129] When users need to add new features to the target page, they can click the AgentBridge development function button to enter the initial development page, and then select "Plugin Development" to enter the plugin development page.

[0130] The plugin development page offers three development options. Users can choose "Intelligent Script Plugin Generation," enter the target webpage address in the pop-up window, and click the OK button.

[0131] The system will automatically access the target page and pop up a sidebar dialog panel. Users can enter their requirements in the dialog panel, such as "Add a button at the top of the page to export financial news." The system will display the generation progress in real time, including "extraction-generation-verification," and can present the generated script code in a real-time, segmented appending streaming manner.

[0132] After verifying the script code, users can directly click the debug button to run the script code on the current page and view the actual effect. Once debugging is successful, users can issue a release command and enter plugin configuration information in the configuration dialog box. Accordingly, the system can generate persistent script plugins based on the plugin configuration information. Debugging and release are seamlessly integrated, making the transformation from requirements to usable plugins a one-step process.

[0133] The entire process described above typically takes only minutes, and users only need to perform some simple operations, which greatly reduces the technical threshold for feature customization and allows ordinary users to quickly create their own plugins.

[0134] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure. Furthermore, for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0135] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.

[0136] Figure 5 This is a schematic diagram of the structural composition of the first embodiment 500 of the script code generation device described in this disclosure. Figure 5 As shown, it includes: a requirement acquisition module 501, an information acquisition module 502, and a code generation module 503.

[0137] The requirement acquisition module 501 is used to acquire the requirement description information sent by the user for the target page. The requirement description information includes: description information of requesting the addition of a pre-defined function on the target page.

[0138] The information acquisition module 502 is used to acquire the page DOM information of the target page and determine the target summary information of the target page based on the page DOM information.

[0139] The code generation module 503 is used to generate first input information based on the requirement description information and target summary information, generate corresponding script code based on the first input information using a large language model, determine the target code based on the script code, and use the target code to implement the predetermined function.

[0140] In some embodiments of this disclosure, the method by which the information acquisition module 502 determines the target summary information of the target page based on the page DOM information may include: filtering out DOM elements that are not visible to the user from each DOM element in the page DOM information, and determining the remaining DOM elements as candidate elements; extracting M types of target elements from the candidate elements, obtaining the element summary of each target element respectively, where M is a positive integer greater than 1, and obtaining the plain text context of the target page; and generating target summary information based on the element summary and the plain text context.

[0141] In some embodiments of this disclosure, the information acquisition module 502 may filter out user-invisible DOM elements from each DOM element in the page DOM information in the following manner: for any DOM element, perform the following first processing: determine whether the DOM element is located within a predetermined container; if it is located within the predetermined container, filter out the DOM element; if it is not located within the predetermined container, determine whether the width of the DOM element is 0; if the width is 0, filter out the DOM element; if the width is not 0, determine whether the DOM element matches a predetermined hidden style; if the hidden style matches, filter out the DOM element.

[0142] In some embodiments of this disclosure, the target element may include: interactive elements, form structures, table structures, and images; correspondingly, the element summary may include: interactive element summary, form structure summary, table structure summary, and image summary.

[0143] In some embodiments of this disclosure, the method by which the information acquisition module 502 generates target summary information based on element summary and plain text context may include: generating a corresponding CSS selector for a predetermined target element, adding the CSS selector to the element summary of the corresponding target element, compressing the element summary, and generating target summary information based on the compression result and plain text context.

[0144] In some embodiments of this disclosure, the information acquisition module 502 may compress the element digests in the following ways: for each type of target element, retain at most P element digests corresponding to the first P target elements in order of their appearance position on the target page, where P is a positive integer greater than 1; for the text fragments in each element digest, retain at most Q characters, where Q is a positive integer greater than 1; and for the input fields in each form structure digest, retain at most R input fields, where R is a positive integer greater than 1.

[0145] In some embodiments of this disclosure, during the process of generating script code using a large language model, the code generation module 503 may also adopt a block-appending streaming display method to display the generated script code to the user in real time.

[0146] In some embodiments of this disclosure, the code generation module 503 may determine the target code based on the script code in the following ways: in response to the completion of script code generation, a retry parameter is set, initially set to 0, and the following second processing is performed: static syntax parsing verification is performed on the newly generated script code; in response to the verification being passed, the newly generated script code is determined as the target code; in response to the verification failing and the retry parameter being less than L, correction prompt information is generated for the newly generated script code; second input information is generated based on the requirement description information, target summary information, and correction prompt information; script code is regenerated using a large language model based on the second input information; the value of the retry parameter is incremented by 1; and then the second processing is repeated, where L is a positive integer; in response to the verification failing and the retry parameter being equal to L, an error prompt information is displayed to the user.

[0147] Figure 6 This is a schematic diagram of the structural composition of the second embodiment 600 of the script code generation device described in this disclosure. Figure 6 As shown, it includes: a requirement acquisition module 501, an information acquisition module 502, a code generation module 503, and a post-processing module 504.

[0148] The post-processing module 504 can update the target code according to the editing operation performed by the user on the target code, and / or, in response to obtaining the user's debugging instructions, can display the interface effect of the target page after adding the predetermined function according to the target code.

[0149] In some embodiments of this disclosure, the post-processing module 504, in response to receiving a user's release instruction, can display a configuration dialog box and obtain the plugin configuration information entered by the user in the configuration dialog box. Then, based on the plugin configuration information, it can release the target code as a persistent script plugin.

[0150] In some embodiments of this disclosure, the post-processing module 504 may also store the script plugin in the plugin asset center and support users to perform one or any combination of the following operations on the script plugin: enable operation, disable operation, online code editing operation, and delete operation; wherein, the script plugin includes metadata and code, and the metadata and code are stored in the plugin asset center using a separate storage architecture.

[0151] In some embodiments of this disclosure, the script plugin has a corresponding version identifier. The version identifier is the first W hexadecimal characters extracted from the SHA-256 hash value of the script plugin, where W is a positive integer greater than 1. The post-processing module 504 can write the version identifier into the version history. In addition, in response to an update of the script plugin, the post-processing module 504 can regenerate the version identifier and write the regenerated version identifier into the version history. Each version identifier in the version history is saved in a version chain, and snapshots of the content of the script plugin corresponding to different version identifiers can be saved separately.

[0152] The specific workflow of each of the above device embodiments can be found in the relevant descriptions in the foregoing method embodiments, and will not be repeated here.

[0153] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly in areas such as natural language processing, large language models, and front-end web development. Artificial intelligence is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. Artificial intelligence hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0154] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0155] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0156] like Figure 7As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0157] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0158] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the methods described herein by any other suitable means (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0163] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0164] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0165] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A script code generation method, comprising: Obtain user request description information for the target page, the request description information including: a description of the request to add a pre-defined function to the target page; Obtain the page document object model information of the target page, and determine the target summary information of the target page based on the page document object model information; First input information is generated based on the requirement description information and the target summary information. Based on the first input information, corresponding script code is generated using a large language model. Target code is determined based on the script code. The target code is used to implement the predetermined function.

2. The method according to claim 1, wherein, The step of determining the target summary information of the target page based on the page document object model information includes: Filter out the document object model elements that are not visible to the user from each document object model element in the page document object model information, and determine the remaining document object model elements as candidate elements; M types of target elements are extracted from the candidate elements, and the element summary of each target element is obtained, where M is a positive integer greater than 1, and the plain text context of the target page is obtained. The target summary information is generated based on the element summary and the plain text context.

3. The method according to claim 2, wherein, The step of filtering out document object model elements that are not visible to the user from each document object model element in the page document object model information includes: For any document object model element, perform the following first processing: Determine whether the document object model element is located within a predetermined container, and filter out the document object model element if it is located within the predetermined container; In response to the fact that the document object model element is not located within the predetermined container, it is determined whether the width of the document object model element is 0. In response to the fact that the width is 0, the document object model element is filtered out. In response to the width not being 0, determine whether the document object model element matches the predetermined hidden style; in response to the hidden style being matched, filter out the document object model element.

4. The method according to claim 2, wherein, The target elements include: interactive elements, form structures, table structures, and images; The element summary includes: interactive element summary, form structure summary, table structure summary, and image summary.

5. The method according to claim 4, wherein, The step of generating the target summary information based on the element summary and the plain text context includes: Generate a corresponding Cascading Style Sheet (CSS) selector for the predetermined target element, and add the CSS selector to the element summary of the corresponding target element; The element digest is compressed; The target summary information is generated based on the compression result and the plain text context.

6. The method according to claim 5, wherein, The compression process for the element digest includes: For each type of target element, at most P element summaries are retained according to their appearance positions on the target page, where P is a positive integer greater than 1. For each text segment in the element summary, retain at most the first Q characters, where Q is a positive integer greater than 1; For each form structure summary, retain at most the first R input fields, where R is a positive integer greater than 1.

7. The method according to claim 1, further comprising: During the process of generating the script code using the large language model, a block-appending streaming display method is adopted to display the generated script code to the user in real time.

8. The method according to claim 1, wherein, The step of determining the target code based on the script code includes: Upon completion of the script code generation, a retry parameter is set, initially set to 0, and the following second process is executed: Perform static syntax parsing and verification on the newly generated script code; In response to the confirmation that the verification is successful, the newly generated script code is identified as the target code; In response to the determination that the verification failed and that the retry parameter is less than L, a correction prompt message is generated for the newly generated script code, and a second input message is generated based on the requirement description information, the target summary information and the correction prompt message. Based on the second input message, the script code is regenerated using the large language model, and the value of the retry parameter is incremented by 1. Then the second process is repeated. L is a positive integer. In response to determining that the verification failed and that the retry parameter is equal to L, an error message is displayed to the user.

9. The method of claim 1, further comprising one or all of the following: Update the target code based on the editing operation performed by the user on the target code; In response to receiving the user's debugging instructions, the interface effect of adding the predetermined function to the target page is displayed according to the target code.

10. The method according to claim 1, further comprising: In response to receiving the user's release command, a configuration dialog box is displayed, and the plugin configuration information entered by the user in the configuration dialog box is obtained; Based on the plugin configuration information, the target code is published as a persistent script plugin.

11. The method of claim 10, further comprising: The script plugin is stored in the plugin asset center, and the user is allowed to perform one or any combination of the following operations on the script plugin: enable operation, disable operation, online code editing operation, and delete operation; The script plugin includes metadata and code, and the metadata and code are stored in the plugin asset center using a separate storage architecture.

12. The method according to claim 11, wherein, The script plugin has a corresponding version identifier. The version identifier is the first W hexadecimal characters extracted from the 256-bit hash value of the script plugin's secure hash algorithm, where W is a positive integer greater than 1. The version identifier is written into the version history. In response to an update to the script plugin, a version identifier is regenerated and written to the version history. Each version identifier in the version history is saved in a version chain, and content snapshots of the script plugin corresponding to different version identifiers are saved respectively.

13. A script code generation apparatus, comprising: The module includes a requirements gathering module, an information gathering module, and a code generation module. The requirement acquisition module is used to acquire the requirement description information sent by the user for the target page. The requirement description information includes: description information of requesting the addition of a pre-defined function on the target page. The information acquisition module is used to acquire the page document object model information of the target page, and determine the target summary information of the target page based on the page document object model information; The code generation module is used to generate first input information based on the requirement description information and the target summary information, generate corresponding script code using a large language model based on the first input information, determine target code based on the script code, and the target code is used to implement the predetermined function.

14. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.

16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-12.