Multi-source heterogeneous hybrid code unified standardization method oriented to intelligent model service

By combining large language models and reinforcement learning algorithms, multi-source heterogeneous mixed code for intelligent model services is solved, and the problem of low code integration efficiency and difficult to guarantee consistency is achieved, efficient code integration and explanation document generation is achieved, and development efficiency and code reliability are improved.

CN120085845AInactive Publication Date: 2025-06-03ZHEJIANG UNIV

Patent Information

Application Number
CN202510579717.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate multi-source heterogeneous mixed code, especially in multi-person collaborative development scenarios for intelligent model services, resulting in low efficiency and consistency of code integration.

Method used

The method of combining large language model and reinforcement learning algorithm is adopted to integrate the code blocks and explanation documents of the intelligent model serving each subtask and the explanation documents for the large language model, and the traceability optimization strategy of reinforcement learning is used to optimize and update the prompt words until an integrated code and explanation document without error is generated.

Benefits of technology

It realizes unified integration and document generation of multi-source heterogeneous mixed codes under multi-person collaboration, improves the algorithm development efficiency of intelligent model services, ensures the consistency and reliability of codes, and reduces the cost of bug fixes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085845A_ABST
    Figure CN120085845A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent model service-oriented multi-source heterogeneous hybrid code unified standardization method, which aims at processing unified standardization and management of multi-source heterogeneous codes of the same intelligent model service development task, and comprises the following steps of: extracting names and meanings of key variables and methods from each subtask code block and description file; performing unified naming on related variables and methods in combination with a large language model; naming replacement of the code blocks is achieved in combination with a large language model, and the code blocks are integrated according to the description file; and combining the big language model with the integrated code and description file, and outputting the code and description file after unified specification. In order to optimize a code integration result, a traceability optimization strategy based on reinforcement learning is constructed, error report information is integrated to realize code reorganization, and the flexibility of a code development process and the multi-person cooperation efficiency oriented to intelligent model service can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language model code integration, and particularly relates to a unified specification method for multi-source heterogeneous hybrid codes for intelligent model services. Background Art

[0002] The development of intelligent model services aims to develop intelligent models for specific tasks in application scenarios that require model support, such as the development of data-driven fault monitoring models for industrial processes, the development of traffic flow prediction models for traffic data, etc. The implementation process requires multi-person collaboration to complete modules such as data preprocessing, model structure construction, inference optimization, and result visualization. Therefore, the code and description files have the characteristics of multi-source heterogeneity, which poses challenges to the final implementation of code integration.

[0003] Large language models have received extensive attention nowadays and have been widely applied in fields such as code completion and knowledge Q&A. However, for the task of integrating multi-source heterogeneous hybrid codes for intelligent model services, there are still few works that can effectively achieve it.

[0004] Existing related research works, such as the Chinese patent application with the publication number CN119248250A, which provides a code development assistance system and method based on an AI code assistance large model, and the Chinese patent application with the publication number CN119537254A, which provides a fuzz testing method and system for code generation and automatic program repair based on a large model, both focus on using the generation ability of large language models to perform overall generation and modification of project codes. However, although the above-mentioned existing technical solutions are relatively efficient to execute, they rely on the capabilities of large language models. However, in the process of handling actual tasks, especially in the application background of some specific domain-specific tasks, human participation is the key to improving the effectiveness of the model. Summary of the Invention

[0005] In view of the above, the present invention provides a unified specification method for multi-source heterogeneous hybrid codes for intelligent model services, which can be applied to the intelligent model development application scenario of multi-person collaboration, and realizes the unified integration of sub-code blocks with multi-source heterogeneous characteristics and the generation of description documents.

[0006] A unified specification method for multi-source heterogeneous hybrid codes for intelligent model services includes the following steps: (1) Input the code blocks and description documents of each subtask of the intelligent model service into a large language model together with prompt words for code integration, and generate integrated code and an overall description document; (2) Input the integrated code into a test environment to run and obtain a running result; if an error occurs and the test fails, further input the error information into the large language model for status label mapping to generate the status label of this error; (3)Optimize and update the prompt according to the status label using the reinforcement learning-based traceability optimization strategy, and input the new prompt into the large language model to guide it to correct the integrated code. Return to step (2) until it passes the test without errors, and output the final integrated code and the overall documentation.

[0007] Furthermore, the prompt specifies the overall task requirements and the sub-requirements corresponding to each subtask. The optimized and updated prompt adds error information, status labels, and recommended optimization actions on this basis.

[0008] Furthermore, the code block of the subtask contains variables and methods related to the subtask, and the documentation of the subtask contains the meanings of these related variables and methods.

[0009] Furthermore, the specific implementation process of the large language model for code integration in step (1) is as follows: 1.1 Identify the similar variables and methods among the code blocks of each subtask according to the descriptions of the meanings of variables and methods in the subtask documentation, and perform unified naming and naming replacement for these variables and methods; 1.2 Infer the logical relationships between each variable, method, and code block according to the requirement content in the task input prompt and the documentation, and integrate the code blocks of each subtask according to the inferred logical relationships to generate the integrated code; 1.3 Further integrate the documentation of each subtask in combination with the inferred logical relationships to generate the overall documentation.

[0010] Furthermore, the overall documentation combines the documentation of each sub-code, and contains the inferred logical relationships, the names and meanings of the variables and methods after unified naming.

[0011] Furthermore, in step (2), the large language model maps the input error information according to the corresponding prompt to determine that the status label of this error belongs to one of the six error statuses of syntax error, runtime error, logical error, type error, name error, and import error.

[0012] Further, the specific implementation of step (3) is as follows: Obtain the Q-table (a core data structure in reinforcement learning for storing state-action values (Q-values)), which records the values of various optimization actions in different error states and the test-pass state. The optimization actions include four categories: unifying variable method naming, logical connection, code combination, and no operation; Search the Q-table according to the status label of the current error to find the optimization action with the corresponding maximum value, and add the name of the optimization action, the error message, and the status label to the prompt, and then input the new prompt into the large language model to guide it to correct the integrated code.

[0013] Further, the Q-table is constructed using reinforcement learning. The specific process is as follows: First, initialize the Q-table, including the learning rate, the return values in each error state and the test-pass state, and the initial values of each optimization action in each error state and the test-pass state; When an error occurs during the operation of the integrated code in the test environment, according to the status label of the error s Search the Q-table to find the optimization action with the corresponding maximum value a and input the optimized and updated prompt into the large language model to guide it to correct the integrated code, and then run the corrected integrated code in the test environment to identify the new state x (syntax error, runtime error, logical error, type error, name error, import error, or test pass), and finally update the Q-table through the following formula:

[0014] where: and are the values of action s in the Q-table before and after the update for state a , is the return value for state x , α is the learning rate, γ is the discount factor, is the maximum value of all actions in the Q-table before the update for state x .

[0015] Further, the return values for the six error states are set to -20, and the return value for the test-pass state is set to 100.

[0016] Based on the above technical solutions, the creative contributions and beneficial technical effects of the present invention are mainly reflected in the following aspects: 1. The present invention enables project participants with different backgrounds, knowledge, and development experiences to complete the code development work of the project in a distributed and parallel manner. Through the above process, the integration of sub-codes and the generation of instruction documents under multi-person collaboration can be achieved. This process is automatically carried out based on large models and reinforcement learning algorithms, and the consistent transformation of multi-source heterogeneous mixed codes under unified specifications can be completed with one click, which can greatly improve the algorithm development efficiency for model services. Here, the consistency includes code development styles, code compilation logics, variable and function naming rules, and code annotation habits. Such a mode helps to improve the final integration of the code, reduce the risk probability of overall code bugs caused by inconsistent codes, and lower the human, financial, and time costs for screening and fixing such bugs.

[0017] 2. On this basis, the present invention also helps with the unified maintenance and management of the code in the later stage of the project, reducing costs. Specifically, after the code consistency, it is convenient for code maintainers to read and understand the code, which can reduce the number of maintenance personnel and the depth of their early participation in the project code.

[0018] 3. Through the present invention, an intelligent agent service based on large models can be constructed, which plays a good role for students and teachers in primary schools, junior high schools, senior high schools, universities, etc. in the teaching practice of programming languages. For example, the grading of programming questions in online language exams is no longer rigid, and the effective transformation of the code can be carried out through the method of the present invention, which helps to promote the innovative cultivation of students in the teaching practice process. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic diagram of the overall process of the specific implementation manner of the present invention.

[0020] Figure 2 It is a schematic diagram of the implementation process of code integration in the present invention.

[0021] Figure 3 It is a schematic diagram of the implementation process of backtracking optimization in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the drawings and specific implementation manners.

[0023] The unified specification method for multi-source heterogeneous hybrid code for intelligent model services of the present invention relates to the fields of artificial intelligence technologies such as large language models and reinforcement learning, as well as application fields of large models such as code integration and instruction document generation. It focuses on using large language models and reinforcement learning algorithms to assist in the integration and debugging process of project programs, aiming to handle the unified specification and management of multi-source heterogeneous codes for the same intelligent model service development task. It can be used in application scenarios such as code integration and management in the intelligent model service development scenario of multi-person collaborative development. While ensuring code reliability, it improves the code construction efficiency.

[0024] The overall implementation process of this embodiment is as Figure 1 shown: Step 1: Collect the code blocks of the subtasks and their instruction documents; deploy the code test environment, including completing the configuration of algorithm operation dependencies and the environment; import the Q-table of the previous task as the initial value.

[0025] Step 2: Import the code blocks and instruction documents of the subtasks into the large language model for integration. The integration process is as Figure 2 shown: including the identification of the meanings of variables and methods, the reasoning of the logical relationships between variables and methods in sub-code files, the unified naming of variables and methods, the integration of code blocks, and the output of standardized code and instruction documents; the above process is input into the large language model together with each sub-code block and instruction file in the form of prompt words.

[0026] The prompt word in this step is the task input prompt word, which stipulates the overall task requirements and the sub-requirements corresponding to each subtask. The design of this prompt word is expressed as "The requirements of this task are {...}, which are divided into multiple subtasks, and the corresponding code blocks are respectively. The association relationships between the subtasks, sub-code blocks, and sub-code instruction documents are {subtask name: [sub-code file name, sub-code instruction file name],...}. Please follow the following implementation process {Step ① unify the naming of variables with the same meaning; Step ② reason about the association relationships between variables; Step ③ reason about the relationships between sub-codes according to the previous requirements; Step ④ implement the integration of the code; Step ⑤ generate the task instruction document} to complete the integration of the code and the task instruction document", {...} is the content that needs to be added according to the actual situation in the prompt word framework.

[0027] Step 3: Output the integrated code and the general instruction document.

[0028] Step 4: Run the integrated code in the test environment to obtain the running result. If the test fails, further input the error message combined with the corresponding prompt word into the large language model for status label mapping to obtain the status label of this error.

[0029] After the test environment inputs the test data for the model service, it runs the integrated code and outputs the code running results and the code running situation, such as error messages, etc.; the error messages include the error type, an explanation of the misaligned content combined with specific variable or method names, and the specific location where the error occurs.

[0030] The prompt design in this step is expressed as "An error occurred in the code, and the error message is {...}, please map it to one of the six status labels: syntax error, runtime error, logical error, type error, name error, import error", where {...} is filled with the error message given by the running environment.

[0031] Step 5: Obtain the status label output by the large language model, select the optimization action with the maximum current state value based on the reinforcement learning-based traceability optimization strategy, and merge the action name, error message, and status label into the prompt.

[0032] The specific process is as Figure 3 shown: First, use reinforcement learning to construct the Q-table in the code optimization strategy. The Q-table records the values of different actions in different states, where the actions include four categories: unified variable method naming, logical connection, code combination, and no operation; then, search for the action corresponding to the maximum value in the Q-table according to the status label of this error, and merge the action name, error message, and status label into the prompt to obtain the backtracking optimization prompt.

[0033] Step 6: Input the code block of the subtask and its description document together with the backtracking optimization prompt into the large language model. The backtracking optimization prompt is designed as "... (task input prompt), the following error message appears after the code runs {...}, it is recommended to optimize in the {action} link", where {action} is filled in after being selected through the Q-table, including four types of actions: method naming, logical connection, code combination, and no operation, to guide the large model to make improvements.

[0034] Step 7: Update the Q-table based on the reinforcement learning method.

[0035] (1) Initialize the Q-table, and the initial value is the Q-table saved after the previous task; (2) Set the return values of the six types of error states to -20, and the return value of the passing state to 100; (3) Interact the error code with the running environment. If the test fails, collect the error message; (4) Input the error message into the large model, and the prompt is designed as "An error occurred in the code, and the error message is {...}, please map it to one of the six status labels: syntax error, runtime error, logical error, type error, name error, import error"; (5) The selection of the optimized action is through a greedy strategy, that is, each time the action corresponding to the maximum value of the current state in the Q-table is selected a , and a new state label is obtained.

[0036] (6) Obtain the state label output by the large model, and update the value in the Q-table according to the following formula:

[0037] Where: and are the values of the action s under the state a in the Q-table before and after the update respectively, is the immediate reward, which represents the return value in the next state x , represents the maximum value of all possible actions in the next state x , α represents the learning rate, γ represents the discount factor; the return values of the six types of error reporting states are set to -20, and the return value of the pass state is set to 100.

[0038] (7) Repeat steps (5) and (6) until the code passes the test, and save the Q-table as the initial value for the next task.

[0039] Step 8: Repeat steps 5 to 7 until the code passes the test, and output the integrated code and the overall description document.

[0040] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and apply the present invention. Those who are familiar with the technology in the art can obviously make various modifications to the above embodiments easily, and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A unified standardization method for multi-source heterogeneous mixed codes for intelligent model services, characterized in that: The steps include: (1) Input the code blocks and description documents of each subtask of the intelligent model service into the large language model in combination with the prompt words for code integration, and generate integrated code and overall description documents; (2) Input the integrated code into the test environment and run it to obtain the running results; If an error occurs and the test fails, the error information is further input into the large language model for state label mapping to generate the state label of this error; (3) According to the state label, the prompt word is optimized and updated using the reinforcement learning-based traceability optimization strategy, and the new prompt word is input into the large language model to guide it to correct the integration code. The execution returns to step (2) until the test passes without error, and the final integration code and overall description document are output.

2. According to claim 1, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The prompt word specifies the overall task requirements and the sub-requirements corresponding to each subtask. The optimized and updated prompt word adds error information, status labels and recommended optimization actions on this basis.

3. According to claim 1, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The subtask code block contains subtask-related variables and methods, and the subtask description document contains the meaning of these related variables and methods.

4. According to claim 1, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The specific implementation process of the code integration of the large language model in step (1) is as follows: 1.1 According to the description of the meaning of variables and methods in the subtask description document, identify the similar variables and methods between the subtask code blocks, and unify and replace the names of these variables and methods; 1.2 Infer the logical relationship between each variable, method and code block based on the requirements in the task input prompt and the description document, and integrate the code blocks of each subtask according to the inferred logical relationship to generate integrated code; 1.3 Further integrate the description documents of each subtask based on the logical relationship obtained by reasoning to generate an overall description document.

5. According to claim 4, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The overall description document combines the description documents of each sub-code, including the logical relationships obtained by reasoning, the names and meanings of variables and methods after unified naming.

6. According to claim 1, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: In the step (2), the large language model maps the input error message to a state label according to the corresponding prompt word, and determines whether the state label of the error message belongs to one of the six error states: syntax error, runtime error, logic error, type error, name error, and import error.

7. According to claim 1, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The specific implementation method of step (3) is as follows: obtaining a Q-table, which records the value of each optimization action under different error states and test pass states, where the optimization actions include four categories: unified variable method naming, logical connection, code combination, and no operation; searching the Q-table for the optimization action with the maximum value according to the status label of the current error, and adding the optimization action name, error information, and status label to the prompt word, and then inputting the new prompt word into the large language model to guide it to correct the integrated code.

8. According to claim 7, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The Q-table is constructed by reinforcement learning. The specific process is as follows: first, the Q-table is initialized, including the learning rate, the reward value in each error state and the test pass state, and the initial value of each optimization action in each error state and the test pass state; when the integrated code runs in the test environment and an error occurs, the state label of the error is used to initialize the Q-table. s Look up the Q-table to find the optimal action with the maximum value a The optimized and updated prompt words are input into the large language model to guide it to correct the integration code, and then the corrected integration code is run in the test environment to identify the new state x , and finally update the Q-table using the following formula: ; in: and The status of Q-table before and after the update s Next action a The value of Status x The return value of α is the learning rate, γ is the discount factor, The status of the Q-table before the update x The maximum value of all actions.

9. According to claim 8, a unified standardization method for multi-source heterogeneous mixed codes for intelligent model services is characterized by: The return value for the six types of error status is set to -20, and the return value for the test pass status is set to 100.

Citation Information

Patent Citations

  • Code development auxiliary system and method based on AI code auxiliary large model

    CN119248250A

  • Fuzzy testing method and system for code generation and automatic program repair based on large model

    CN119537254A

  • Application program development method and device, electronic equipment and storage medium

    CN112748914A

  • Handwriting skeleton refinement method, system and equipment based on deep reinforcement learning, and medium

    CN117011856A

  • Hybrid front-end framework migration method based on AST and LLM

    CN117608656A

Cited By

  • Multi-source heterogeneous code conversion method and device for cross-language development

    CN121143797A