Technical solution optimization method, system, device, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP DESIGN INST
- Filing Date
- 2026-03-31
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]然而,这种完全依赖人工的撰写模式存在以下缺陷:其一,受限于撰写人员的个人经验与知识广度,当面对跨学科或复杂度较高的技术问题时,撰写人员易出现认知盲区,导致技术方案在技术逻辑的完整性、创新点的准确性或实施细节的充分公开等方面存在缺陷,不同撰写人员产出的技术方案质量参差不齐;其二,人工反复推敲与优化技术方案的效率较低,一份高质量技术方案的形成往往需要多次内部评审与修改,这一过程耗时费力,难以适应快速迭代的研发节奏
[0028]Therefore, the technical solution provided in this application embodiment, through a central intelligent agent responding to received basic information input by the user, identifies the user intent in the basic information, and responds to the user intent to optimize the technical solution. Based on the user's indication information regarding the user intent, it determines and executes a task planning scheme for the user intent to generate a target technical solution. The task planning scheme includes: calling a technical solution optimization intelligent agent to generate an initial technical solution based on the basic information; calling a technical solution evaluation intelligent agent to evaluate the initial technical solution and generate an evaluation result containing an evaluation score; responding to the evaluation score being less than a preset threshold, the central intelligent agent generates an optimization instruction based on the evaluation result; calling the technical solution optimization intelligent agent to iteratively optimize the initial technical solution based on the optimization instruction until the evaluation score of the optimized technical solution is not less than the preset threshold, and determining the optimized technical solution as the target technical solution.
Smart Images

Figure CN122528898A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, system, device, storage medium, and program product for optimizing a technical solution. Background Technology
[0002] In the process of scientific research innovation and engineering technology development, the writing of technical solutions, such as technical disclosure documents and R&D plan reports, is a crucial step in the transformation of knowledge achievements. Currently, the production of such technical solutions mainly relies on the professional knowledge and personal experience of the relevant writers (researchers or engineers). Writers typically need to consult a large number of documents, conduct rigorous logical deductions, and express the technical solutions in written form according to standardized formats.
[0003] However, this writing model, which relies entirely on manual drafting, has the following drawbacks: First, limited by the personal experience and breadth of knowledge of the writers, when faced with interdisciplinary or highly complex technical problems, writers are prone to cognitive blind spots, resulting in deficiencies in the completeness of technical logic, the accuracy of innovative points, or the full disclosure of implementation details in the technical solutions. The quality of technical solutions produced by different writers varies greatly. Second, the efficiency of repeatedly refining and optimizing technical solutions manually is low. The formation of a high-quality technical solution often requires multiple internal reviews and revisions. This process is time-consuming and labor-intensive, and it is difficult to adapt to the rapid iterative R&D pace.
[0004] To alleviate the aforementioned technical problems, methods for using Artificial Intelligence (AI) tools to assist in writing technical solutions have been disclosed in related technologies. However, existing AI tools still have the following technical problems in practical applications: AI tools focus on "assisting in writing" rather than "assisting in optimization." Most AI tools only output a technical solution once after the user inputs basic information. Their core function is to replace or imitate humans in completing the work of writing a technical solution from scratch, rather than optimizing existing technical solutions. The content of the technical solutions generated by AI tools depends entirely on their internal parameters and training data, lacking human-computer interaction and mechanisms for objective evaluation and quality verification of the generated technical solutions, thus leading to unreliable quality of technical solutions generated by AI tools. Summary of the Invention
[0005] This disclosure aims to at least partially address one of the technical problems in the related art.
[0006] Therefore, one objective of this disclosure is to propose a method for optimizing a technical solution.
[0007] The second objective of this disclosure is to propose a technical solution to optimize the system.
[0008] The third objective of this disclosure is to propose an electronic device.
[0009] The fourth objective of this disclosure is to provide a computer-readable storage medium.
[0010] The fifth objective of this disclosure is to provide a computer program product.
[0011] To achieve the above objectives, the first aspect of this disclosure proposes a method for optimizing a technical solution, the method comprising:
[0012] The central intelligent agent responds to receiving basic information from the user input and identifies the user's intent within that basic information; In response to user intent, the technical solution is optimized by determining and executing a task planning scheme based on the user's indication of the user intent, in order to generate the target technical solution. The task planning scheme includes: calling the technical solution optimization agent to generate an initial technical solution based on basic information; calling the technical solution evaluation agent to evaluate the initial technical solution and generate an evaluation result containing an evaluation score; in response to the evaluation score being less than a preset threshold, the central agent generates an optimization instruction based on the evaluation result; calling the technical solution optimization agent to iteratively optimize the initial technical solution based on the optimization instruction until the evaluation score of the optimized technical solution is not less than the preset threshold, and determining the optimized technical solution as the target technical solution.
[0013] In some embodiments, after determining the optimized technical solution as the target technical solution, the method provided in the first aspect embodiment further includes: The intelligent agent is invoked to verify the technical solution and the target technical solution is verified.
[0014] In some embodiments, the method provided in the first aspect further includes: Through training, a central intelligent agent, a technical solution optimization intelligent agent, a technical solution evaluation intelligent agent, and a technical solution verification intelligent agent are obtained. The process of training the central agent, the technology solution optimization agent, the technology solution evaluation agent, and the technology solution verification agent includes: Obtain the training sample set; For each training sample in the training sample set, the training sample is input into the central intelligent agent, and the central intelligent agent determines the task planning scheme. According to the task planning scheme, the training samples are processed by the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent, and the corresponding reward values of the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent are obtained during the processing. Based on the reward values corresponding to the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent, the cumulative reward of the task planning scheme is determined. Based on the cumulative rewards, the parameters of the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent are optimized until the preset training termination conditions are met, resulting in the trained central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent.
[0015] In some embodiments, the reward value corresponding to the central agent is obtained through the reward function of the central agent; The reward function for the central agent is: ; in, The evaluation score for the target technical solution. To generate the computational resource cost invoked for the target technical solution, It refers to the number of times the central intelligent agent generates optimization instructions during the process of generating the target technical solution. for The weight value, for The weight value, for The weight value.
[0016] In some embodiments, the reward value of the technical solution optimization agent refers to the cumulative sum of reward values obtained in each optimization iteration round of the processing; the reward value obtained in each optimization iteration is obtained through the reward function of the technical solution optimization agent; The reward function for the optimized agent in this technical solution is as follows: ; in, To optimize the reward function value of the intelligent agent in the technical solution, This represents the evaluation score of the optimized technical solution in the current optimization iteration round. It is an internal reward that measures the completeness and logical consistency of the optimized technical solution in the current optimization iteration round. This is a penalty term used to characterize the deviation between the optimized instruction and the optimized technical solution in the current optimization iteration round. for The weight value, for The weight value, for The weight value.
[0017] In some embodiments, the reward value corresponding to the technical solution evaluation agent is obtained through the reward function of the technical solution evaluation agent; The reward function for evaluating the intelligent agent in the technical solution evaluation is as follows: ; in, The reward function value of the agent is used to evaluate the technical solution. The initial technical solution is scored for its novelty. The novelty of the target technical solution is scored. The reward is a correlation factor between the initial technical solution and similar solutions retrieved by the technical solution evaluation agent. This is the weighting value for novelty. for The weight value.
[0018] In some embodiments, the reward function for the technical solution to verify the intelligent agent is:
[0019] in, To verify the reward function value of the intelligent agent in the technical solution, This refers to the i-th specific error instance that is corrected in the target technical solution.
[0020] In some embodiments, a technical solution evaluation agent is invoked to evaluate an initial technical solution and generate an evaluation result containing an evaluation score, including: The intelligent agent that evaluates the technical solutions retrieves at least one similar solution from the initial technical solution. Evaluation results are generated based on similar schemes and initial technical schemes.
[0021] In some embodiments, the technical solution evaluation agent is invoked to retrieve an initial technical solution and obtain at least one similar solution, including: The agent is evaluated by invoking technical solutions, and technical feature vectors are generated based on the initial technical solutions. Calculate the cosine similarity value between the technical feature vector and the candidate feature vectors of each candidate solution in the solution library; Based on the cosine similarity value, at least one similar solution is determined from the candidate solutions in the solution library.
[0022] In some embodiments, evaluation results are generated based on similar solutions and initial technical solutions, including: Compare similar solutions with the initial technical solution, and generate a novelty score and an inventiveness score for the initial technical solution based on the comparison results; An evaluation score is calculated based on the novelty score and the inventiveness score, and an evaluation result is generated based on the novelty score, the inventiveness score, and similar solutions.
[0023] To achieve the above objectives, a second aspect of this disclosure provides a technical solution optimization system, which includes: The central intelligent agent is used to respond to the basic information received from the user input, identify the user intent in the basic information, optimize the technical solution in response to the user intent, determine and execute the task planning scheme based on the user intent, and generate the target technical solution. The technical solution optimizes the intelligent agent, which is used to respond to the call of the central intelligent agent and generate an initial technical solution based on basic information; The technical solution evaluation agent is used to evaluate the initial technical solution in response to the call of the central agent and generate an evaluation result containing an evaluation score. The central intelligent agent is also used to generate optimization instructions based on the evaluation results in response to an evaluation score that is less than a preset threshold. The technical solution optimization agent is also used to respond to the call of the central agent, and iteratively optimize the initial technical solution based on the optimization instructions until the evaluation score of the optimized technical solution is not less than the preset threshold, and then determine the optimized technical solution as the target technical solution.
[0024] In some embodiments, the system provided in the second aspect further includes a technical solution verification agent: The technical solution verification agent is used to verify the target technical solution in response to the call of the central agent.
[0025] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, including: a memory and a processor; The processor reads executable program code stored in memory to run a program corresponding to the executable program code, thereby implementing the technical solution optimization method as described in the first aspect of this disclosure.
[0026] To achieve the above objectives, a fourth aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the technical solution optimization method of the first aspect of this disclosure.
[0027] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product including computer-executable instructions, which, when executed by a processor, implement the technical solution optimization method as described in the first aspect of this disclosure.
[0028] Therefore, the technical solution provided in this application embodiment, through a central intelligent agent responding to received basic information input by the user, identifies the user intent in the basic information, and responds to the user intent to optimize the technical solution. Based on the user's indication information regarding the user intent, it determines and executes a task planning scheme for the user intent to generate a target technical solution. The task planning scheme includes: calling a technical solution optimization intelligent agent to generate an initial technical solution based on the basic information; calling a technical solution evaluation intelligent agent to evaluate the initial technical solution and generate an evaluation result containing an evaluation score; responding to the evaluation score being less than a preset threshold, the central intelligent agent generates an optimization instruction based on the evaluation result; calling the technical solution optimization intelligent agent to iteratively optimize the initial technical solution based on the optimization instruction until the evaluation score of the optimized technical solution is not less than the preset threshold, and determining the optimized technical solution as the target technical solution.
[0029] The technical solution optimization method provided in this application addresses the technical problems of inconsistent quality due to limitations in personal experience and low efficiency of repeated manual optimization in a purely manual writing mode. It also overcomes the shortcomings of existing AI tools that only focus on "assisting writing," lacking human-computer interaction and objective evaluation and quality verification mechanisms, resulting in unreliable quality of the generated technical solutions. This solution introduces a central intelligent agent to identify user intent and formulate task planning accordingly, achieving unified management of the technical solution optimization process. It generates an initial solution by calling the technical solution optimization intelligent agent and further calls the technical solution evaluation intelligent agent to quantitatively evaluate the solution. When the evaluation score does not meet the requirements, the central intelligent agent generates optimization instructions based on the evaluation results and drives the optimization intelligent agent to iteratively optimize until the preset quality standard is reached. This constructs an automated closed loop of "writing-evaluation-optimization," achieving in-depth, objective, and quantifiable iterative optimization of the technical solution while retaining human-computer interaction, significantly improving the quality and reliability of the final target technical solution. Attached Figure Description
[0030] Figure 1 A flowchart illustrating a technical solution optimization method provided in an embodiment of this application; Figure 2 A flowchart illustrating another technical solution optimization method provided in this application embodiment; Figure 3 This application provides a schematic diagram of the structure of a technical solution optimization system according to an embodiment of the present application. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0032] The acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of relevant laws and regulations.
[0033] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0034] To address the issues of inconsistent quality of technical solutions due to limitations in personal experience and low efficiency of repeated manual optimization in the process of optimizing technical solutions, which relies entirely on manual writing, and the shortcomings of existing AI tools that only focus on "assisting writing" and lack human-computer interaction and objective evaluation and quality verification mechanisms, resulting in unreliable quality of generated technical solutions, this application provides a method, system, device, storage medium, and program product for optimizing technical solutions.
[0035] The entity executing the technical solution optimization method can be a technical solution optimization system, which can be installed on an electronic device.
[0036] It should be noted that the technical solution optimization system is an intelligent auxiliary optimization system built on technologies such as artificial intelligence, data analysis, and rule engines. Its core objective is to automatically analyze, evaluate, and propose optimization suggestions for the initially proposed technical solutions, business solutions, engineering designs, or other types of solutions, combined with modification suggestions or other information input by users during human-computer interaction, in order to improve the quality and reliability of the technical solutions.
[0037] The technical solution optimization system constructs an intelligent solution centered on a large language model (LLM) and achieves technical solution evaluation and verification through multi-agent collaboration. A central agent performs task planning for the technical solution optimization task, and then calls a series of related agents, combined with human-computer interaction, to ensure the quality and reliability of the optimized target technical solution. The core of this technical solution optimization system lies in its modular design and the collaborative mechanism between agents.
[0038] Specifically, the technical solution optimization system may include: a central intelligent agent, a technical solution optimization intelligent agent, and a technical solution evaluation intelligent agent; wherein, the central intelligent agent is used to respond to receiving basic information input by the user, identify the user intent in the basic information; responding to the user intent to optimize the technical solution, determine and execute a task planning scheme for the user intent to generate a target technical solution; the technical solution optimization intelligent agent is used to respond to the call of the central intelligent agent, generate an initial technical solution based on the basic information; the technical solution evaluation intelligent agent is used to respond to the call of the central intelligent agent, evaluate the initial technical solution, and generate an evaluation result including an evaluation score; the central intelligent agent is also used to respond to the evaluation score being less than a preset threshold, generate an optimization instruction based on the evaluation result; the technical solution optimization intelligent agent is also used to respond to the call of the central intelligent agent, iteratively optimize the initial technical solution based on the optimization instruction until the evaluation score of the optimized technical solution is not less than the preset threshold, and determine the optimized technical solution as the target technical solution.
[0039] Figure 1 This is a flowchart illustrating a technical solution optimization method provided in an embodiment of this application.
[0040] Applied to the above-mentioned technical solution optimization system, such as Figure 1 As shown, the technical solution optimization method provided in this application includes the following steps: Step 101: The central intelligent agent responds to receiving basic information input by the user and identifies the user's intent in the basic information; In some embodiments, basic information can be input by the user through the human-computer interaction interface of the technical solution optimization system in the form of natural language. This basic information represents the user's intent. For example, if the user inputs the basic information as "I want to optimize the technical solution for drones to avoid obstacles using ultrasonic waves and weather data in adverse weather conditions," the central intelligent agent can perform intent recognition based on this basic information, identifying that the user's intent is not a simple question-and-answer interaction, but rather an optimization of the technical solution.
[0041] In some embodiments, the basic information may also be a semi-structured form, a preliminary technical document, or a draft generated by a large language model.
[0042] In some embodiments, the basic information representing user intent for optimizing technical solutions typically includes core elements such as technical problems, solutions, and expected results.
[0043] In some embodiments, the central intelligent agent performs format standardization, sensitive information desensitization, and key semantic extraction on basic information to ultimately form a structured user intent. For example, the central intelligent agent extracts user intent composed of core technical elements such as "drone," "ultrasonic sensor," "weather data," and "dynamic obstacle avoidance" from the basic information, and displays the identified user intent to the user through technical solutions to optimize the system's human-computer interaction interface.
[0044] In some embodiments, the Central Agent can use LLM to identify intent from basic information input by the user in order to understand the user’s true needs.
[0045] Step 102: In response to user intent, optimize the technical solution by determining and executing a task planning scheme based on the user's instruction information regarding the user intent, so as to generate the target technical solution.
[0046] The task planning scheme includes: Step 1201: Invoke the technical solution to optimize the intelligent agent and generate an initial technical solution based on the basic information; Step 1202: Call the technical solution evaluation agent to evaluate the initial technical solution and generate an evaluation result containing the evaluation score; Step 1203: In response to the evaluation score being less than a preset threshold, the central agent generates optimization instructions based on the evaluation results; Step 1204: Invoke the technical solution optimization agent, and iteratively optimize the initial technical solution based on the optimization instructions until the evaluation score of the optimized technical solution is not less than the preset threshold, and determine the optimized technical solution as the target technical solution.
[0047] In some embodiments, if in step 101 the central agent recognizes that the user's intention is to optimize the technical solution, then task planning is performed, and the technical solution optimization agent and the technical solution evaluation agent in the technical solution optimization system are arranged to finally obtain the optimized target technical solution.
[0048] In some embodiments, after the central intelligence agent displays the user's intent to the user in step 101, the user can input modifications or confirmations to the user's intent through the human-computer interaction interface. For example, if the user believes that the user's intent recognized by the central intelligence agent is accurate, the user can input "OK, please continue"; if the user believes that the user's intent recognized by the central intelligence agent is inaccurate, the user can input modifications to the user's intent, such as "add severe weather to the technical elements".
[0049] In some embodiments, step 1201 includes: The intelligent agent is optimized by invoking technical solutions, generating an initial technical solution based on the technical elements extracted from the basic information. Understandably, if the user inputs modifications to their intent, an initial technical solution also needs to be generated based on these modifications.
[0050] In some embodiments, the process of generating an initial technical solution based on basic information is a process in which the technical solution optimization agent supplements and expands the technical elements in the basic information. For details, please refer to existing AI tools, which will not be elaborated here.
[0051] In some embodiments, after the technical solution optimization agent generates the initial technical solution, it displays the initial technical solution to the user in any way, such as text or a downloadable and previewable document, through a human-computer interaction interface. This allows the user to modify the initial technical solution online or offline, or to make modification suggestions, so that the technical solution optimization agent can further modify the initial technical solution.
[0052] In some embodiments, the process of evaluating the initial technical solution in step 1202 refers to conducting a novelty search evaluation on the initial technical solution generated in step 1201 based on existing technology. For example, if the ultimate purpose of the generated technical solution is to apply for a patent, existing technologies with high similarity to the initial technical solution can be searched, and then the novelty and inventiveness of the initial technical solution generated in step 1201 can be evaluated based on the technologies with high similarity. If the generated technical solution is used to publish papers, etc., existing technologies with high similarity to the initial technical solution can be searched, and then the initial technical solution generated in step 1201 can be evaluated for plagiarism based on the technologies with high similarity.
[0053] Based on this, in some embodiments, step 1202 includes the following steps: The intelligent agent that evaluates the technical solutions retrieves at least one similar solution from the initial technical solution. Evaluation results are generated based on similar schemes and initial technical schemes.
[0054] In some embodiments, the technical solution evaluation agent can connect to a preset database; the retrieval method for solutions in the database can be online or offline. The database may include, but is not limited to: patent literature databases, academic paper databases, technical standard databases, internal enterprise technical document databases, and publicly available technical materials on the Internet.
[0055] In some embodiments, the technical solution evaluation agent may first analyze the initial technical solution and extract key technical features, such as technical problems, technical means, technical effects, application fields, and key parameters. Then, based on the extracted key technical features, a search query is constructed, and a similarity search is performed in the aforementioned database. The retrieved similar solutions may be existing technical solutions that have a high degree of similarity to the initial technical solution in terms of technical field, the technical problem to be solved, the technical means employed, or the technical effects achieved.
[0056] Based on this, the technical solution evaluation agent searches for the initial technical solution and obtains at least one similar solution, including: The agent is evaluated by invoking technical solutions, and technical feature vectors are generated based on the initial technical solutions. Calculate the cosine similarity value between the technical feature vector and the candidate feature vectors of each candidate solution in the solution library; Based on the cosine similarity value, at least one similar solution is determined from the candidate solutions in the solution library.
[0057] In some embodiments, the technical feature vector can be a vector obtained by transforming technical problems, technical means, technical effects, application fields, key parameters, etc.
[0058] In some embodiments, a similar solution refers to the candidate solution whose cosine similarity value with the initial technical solution ranks higher, or whose cosine similarity value is greater than a preset cosine similarity threshold.
[0059] In some embodiments, cosine similarity sim The calculation formula is as follows:
[0060] in, To generate technical feature vectors based on the initial technical solution, These are candidate feature vectors generated based on candidate schemes. Cosine similarity value. sim The range is [-1, 1]), and the closer the value is to 1, the more similar the candidate solution is to the initial technical solution.
[0061] In some embodiments, the technical solution evaluation agent can also display the retrieved similar solutions to the user through a human-computer interaction interface; if the user is satisfied with the retrieved similar solutions, they can enter "OK, please continue"; if the user is not satisfied with the retrieved similar solutions, they can enter "please search again"; In some embodiments, after obtaining at least one similar solution, the technical solution evaluation agent generates an evaluation result based on the similar solution and the initial technical solution. Specifically, the technical solution evaluation agent can perform comparative analysis between the initial technical solution and the retrieved similar solutions, and evaluate the initial technical solution from multiple dimensions.
[0062] In some embodiments, the evaluation score can be a combination of scores from the aforementioned multiple dimensions. For example, multiple scoring dimensions such as novelty and inventiveness can be set, and corresponding weights can be assigned to each scoring dimension. Finally, a comprehensive evaluation score is obtained through weighted calculation.
[0063] Based on this, evaluation results are generated based on similar schemes and initial technical schemes, including: Compare similar solutions with the initial technical solution, and generate a novelty score and an inventiveness score for the initial technical solution based on the comparison results; An evaluation score is calculated based on the novelty score and the inventiveness score, and an evaluation result is generated based on the novelty score, the inventiveness score, and similar solutions.
[0064] In some embodiments, the evaluation results may include not only an evaluation score but also a detailed evaluation report. For example, the evaluation report may point out the main differences between the initial technical solution and similar solutions, list the potential defects or deficiencies of the initial technical solution (e.g., insufficient disclosure of a certain technical feature, lack of data support for a certain technical effect, conflict with similar solutions in a certain statement, etc.), and provide specific modification suggestions for subsequent optimization.
[0065] In some embodiments, in step 1203, in response to the evaluation score being less than a preset threshold, the central agent generates an optimization instruction based on the evaluation result; In some embodiments, the preset threshold in step 1203 is a pre-set boundary line for measuring whether a technical solution meets quality standards. This threshold can be a fixed value or dynamically adjusted according to different technical fields, user needs, or application scenarios. For example, for solutions with high technical complexity, the threshold can be appropriately lowered to increase the space for iterative optimization; for technical solutions with high quality requirements, a higher threshold can be set to ensure the rigor and completeness of the technical solution.
[0066] When the central agent determines that the evaluation score is less than a preset threshold, it confirms that the current technical solution has not yet met the expected quality standard and needs further optimization. At this point, the central agent responds to the determination result and generates optimization instructions based on the evaluation result.
[0067] In some embodiments, the process of generating optimization instructions may include: the central agent first parses the evaluation results and extracts key information. As mentioned earlier, in addition to the evaluation score, the evaluation results may also include a detailed evaluation report, such as pointing out specific defects, shortcomings, or gaps in the current technical solution or compared with similar solutions, as well as possible modification suggestions. Then, based on the parsed information, combined with the original basic information and user intent, the central agent generates targeted optimization instructions. The optimization instructions are used to guide the technical solution optimization agent to perform the next round of optimization operations, and their content may include, but is not limited to, one or more of the following: Indicators for Optimization: Clearly specify the parts of the current technical solution that need to be modified or improved. For example, it could indicate that "the beneficial effects lack data support and experimental data or theoretical derivation evidence needs to be added."
[0068] Optimization Direction Guidance: Provides macro-level optimization directions and strategies for the intelligent agent optimizing technical solutions. For example, it can indicate "enhancing the description of the distinguishing technical features between this solution and similar solution B," etc.
[0069] Optimization constraints: Set the restrictions that must be followed during this optimization process. For example, you can instruct "core technical features must not be changed during the optimization process" or "the terminology used must be consistent with the initial basic information," etc.
[0070] It should be noted that the specific content and form of the above-mentioned optimization instructions are merely illustrative and not intended to limit the scope of this application. In practical applications, different optimization instruction formats and contents can be set according to specific application scenarios and requirements. For example, structured data formats such as JSON or XML can be used to facilitate accurate parsing and execution by the technical solution optimization agent; alternatively, natural language descriptions can be used, which the technical solution optimization agent can understand and execute. All different implementation methods fall within the protection scope of this application.
[0071] In some embodiments, after the central agent generates optimization instructions in step 1023, it calls the technical solution optimization agent to iteratively optimize the initial technical solution based on the optimization instructions, thereby generating the optimized technical solution.
[0072] At this point, the central intelligent agent continues to call the technical solution to optimize the intelligent agent, and uses the optimized technical solution as the initial technical solution. Steps 1023 and 1024 are repeated until the evaluation score of the optimized technical solution is not less than the preset threshold, and the optimized technical solution is determined as the target technical solution.
[0073] During this process, after each optimized technical solution is generated, the intelligent agent can receive modification suggestions from the user through the human-computer interaction interface and optimize the optimized technical solution according to the user's modification suggestions. At the same time, after each optimization instruction is generated for the current optimized technical solution, the central intelligent agent can receive modification suggestions from the user for the optimization instruction through the human-computer interaction interface.
[0074] The technical solution optimization method provided in this application addresses the technical problems of inconsistent quality due to limitations in personal experience and low efficiency of repeated manual optimization in a purely manual writing mode. It also overcomes the shortcomings of existing AI tools that only focus on "assisting writing," lacking human-computer interaction and objective evaluation and quality verification mechanisms, resulting in unreliable quality of the generated technical solutions. This solution introduces a central intelligent agent to identify user intent and formulate task planning accordingly, achieving unified management of the technical solution optimization process. It generates an initial solution by calling the technical solution optimization intelligent agent and further calls the technical solution evaluation intelligent agent to quantitatively evaluate the solution. When the evaluation score does not meet the requirements, the central intelligent agent generates optimization instructions based on the evaluation results and drives the optimization intelligent agent to iteratively optimize until the preset quality standard is reached. This constructs an automated closed loop of "writing-evaluation-optimization," achieving in-depth, objective, and quantifiable iterative optimization of the technical solution while retaining human-computer interaction, significantly improving the quality and reliability of the final target technical solution.
[0075] In some embodiments, the target technical solution also needs to be verified to ensure that the final output of the target technical solution meets the specification requirements in terms of format, syntax, and terminology consistency. For example, verifying whether the content of the target technical solution is complete and whether the numbering is continuous, while correcting inconsistencies in terminology (such as the interchangeable use of "processor" and "processing unit"), correcting format errors such as punctuation, spelling, and paragraph structure, verifying the consistency between the figure labels and text descriptions, and finally outputting a target technical solution with a compliant form.
[0076] Based on this, after step 1024, the technical solution optimization method provided in this application embodiment further includes: The intelligent agent is invoked to verify the technical solution and the target technical solution is verified.
[0077] It is understood that the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent involved in the embodiments of this application can all be obtained through training.
[0078] Based on this, in some embodiments, the optimized method of the technical solution provided in this application further includes the following steps: Through training, a central intelligent agent, a technical solution optimization intelligent agent, a technical solution evaluation intelligent agent, and a technical solution verification intelligent agent are obtained. Specifically, Figure 2 This is a flowchart illustrating another technical solution optimization method provided in an embodiment of this application. Figure 2 As shown, the process of training the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent includes the following steps: Step 201: Obtain the training sample set; In some embodiments, each training sample in the training sample set may include basic information input by the user and its corresponding expected output result (target technical solution).
[0079] In some embodiments, training samples may be derived from historical technical solution writing data. For example, authorized patent technical solutions may be collected as positive samples, while technical solutions that have been identified by examination opinions as having novelty or inventiveness defects may be collected as negative samples.
[0080] In some embodiments, the size of the training sample set can be determined based on actual computing resources and training requirements, and may include tens of thousands to hundreds of thousands of sample data.
[0081] Step 202: For each training sample in the training sample set, input the training sample into the central intelligent agent, and determine the task planning scheme through the central intelligent agent. Specifically, the basic information of user input from the training samples is used as input and fed into the central agent to be trained. Based on the basic information of the current input, the central agent performs intent recognition and generates a corresponding task planning scheme. The task planning scheme can include the types of agents to be invoked (technical solution optimization agent, technical solution evaluation agent, and technical solution verification agent), the invocation order, and the sub-tasks to be performed by each agent. For example, for a certain training sample, the central agent might plan a process of "first invoking the technical solution optimization agent to generate an initial scheme, then invoking the technical solution evaluation agent for evaluation, and if the evaluation score is lower than a threshold, continuing to invoke the optimization agent for iteration."
[0082] Step 203: According to the task planning scheme, the training samples are processed by the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent, and the corresponding reward values of the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent are obtained during the processing. Specifically, the training samples can be processed sequentially by calling each agent according to the task planning scheme determined in step 202. During the processing, each agent will receive a corresponding reward value based on a preset reward function after completing its respective task.
[0083] Step 204: Based on the reward values corresponding to the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent, determine the cumulative reward of the task planning scheme. Specifically, the reward values of each agent obtained in step 203 can be weighted and summed or combined in other ways to obtain the overall cumulative reward for the current task planning scheme. This cumulative reward reflects the overall effect of the task scheme planned by the central agent and the execution of the task scheme by each agent under the current input. The higher the cumulative reward, the better the current agent's processing effect on the training sample, and the closer it is to the expected output result.
[0084] Step 205: Optimize the parameters of the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent based on the cumulative reward until the preset training end conditions are met, and obtain the trained central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent.
[0085] In some embodiments, the reward value corresponding to the central agent is obtained through the reward function of the central agent; The reward function for the central agent is: ; in, The evaluation score for the target technical solution. To generate the computational resource cost invoked for the target technical solution, It refers to the number of times the central intelligent agent generates optimization instructions during the process of generating the target technical solution. for The weight value, for The weight value, for The weight value.
[0086] In some embodiments, , , It can be configured according to the needs of the actual application scenario. For example, in scenarios that pursue high-quality technical solutions, a larger setting can be used. A higher value can be set to encourage the central agent to plan task solutions that generate high-scoring target technical solutions; in cost-sensitive deployment scenarios, a larger value can be set. A value can be set to constrain the central intelligent agent from overusing expensive resources; in scenarios requiring rapid response, a larger value can be set. This value is used to encourage the central intelligent agent to plan task schemes with fewer iterations and higher efficiency.
[0087] In some embodiments, the specific quantification method for calculating resource costs can be determined based on the actual billing model adopted. For example, it can be calculated based on the number of tokens required for each call. The statistical method for counting the number of optimization instructions generated can be the number of optimizations performed from the first generation of optimization instructions to the final output meeting the threshold requirements of the target technical solution.
[0088] In some embodiments, the reward value of the technical solution optimization agent refers to the cumulative sum of reward values obtained in each optimization iteration round of the processing; the reward value obtained in each optimization iteration is obtained through the reward function of the technical solution optimization agent; The reward function for the optimized agent in this technical solution is as follows: ; in, To optimize the reward function value of the intelligent agent in the technical solution, This represents the evaluation score of the optimized technical solution in the current optimization iteration round. It is an internal reward that measures the completeness and logical consistency of the optimized technical solution in the current optimization iteration round. This is a penalty term used to characterize the deviation between the optimized instruction and the optimized technical solution in the current optimization iteration round. for The weight value, for The weight value, for The weight value.
[0089] In some embodiments, the reward value corresponding to the technical solution evaluation agent is obtained through the reward function of the technical solution evaluation agent; The reward function for evaluating the intelligent agent in the technical solution evaluation is as follows: ; in, The reward function value of the agent is used to evaluate the technical solution. The initial technical solution is scored for its novelty. The novelty of the target technical solution is scored. The reward is a correlation factor between the initial technical solution and similar solutions retrieved by the technical solution evaluation agent. This is the weighting value for novelty. for The weight value.
[0090] In some embodiments, the reward function for the technical solution evaluation agent is to maximize the reduction of the novelty score of the initial technical solution. Furthermore, when the agent evaluating the technical solution retrieves a highly relevant similar solution, it receives an additional reward. This drives the technical solution evaluation agent to spare no effort in finding the defects of the initial technical solution or the optimized technical solution, thus demonstrating the "adversarial" characteristics of the technical solution evaluation agent.
[0091] In some embodiments, the reward function for the technical solution to verify the intelligent agent is:
[0092] in, To verify the reward function value of the intelligent agent in the technical solution, This refers to the i-th specific error instance that is corrected in the target technical solution.
[0093] The reward function of the technical solution verification agent shows that it is the reward obtained for each inconsistent format, syntax, or terminology error that the agent finds and corrects, focusing on the formal standardization and rigor of the target technical solution.
[0094] To achieve the technical solution optimization method provided in the embodiments of this application, the embodiments of this application also provide a technical solution optimization system. For example... Figure 3 As shown, Figure 3 A schematic diagram of a technical solution optimization system provided in an embodiment of this application. The technical solution optimization system includes: The central intelligent agent 301 is used to respond to the basic information input by the user, identify the user intent in the basic information, optimize the technical solution in response to the user intent, determine and execute the task planning scheme for the user intent, so as to generate the target technical solution. The technical solution optimizes the intelligent agent 302, which is used to respond to the call of the central intelligent agent and generate an initial technical solution based on basic information; Technical solution evaluation agent 303 is used to evaluate the initial technical solution in response to the call of the central agent and generate an evaluation result containing an evaluation score; The central intelligent agent 301 is also used to generate optimization instructions based on the evaluation results in response to an evaluation score that is less than a preset threshold. The technical solution optimization agent 302 is also used to respond to the call of the central agent, iteratively optimize the initial technical solution based on the optimization instructions, until the evaluation score of the optimized technical solution is not less than the preset threshold, and determine the optimized technical solution as the target technical solution.
[0095] In some embodiments, the technical solution optimization system provided in this application further includes a technical solution verification intelligent agent 304: Technical solution verification agent 304 is used to verify the target technical solution in response to the call of the central agent.
[0096] In some embodiments, the technical solution optimization system provided in this application further includes a training module, which is used for: Through training, a central intelligent agent, a technical solution optimization intelligent agent, a technical solution evaluation intelligent agent, and a technical solution verification intelligent agent are obtained. The process of training the central agent, the technology solution optimization agent, the technology solution evaluation agent, and the technology solution verification agent includes: Obtain the training sample set; For each training sample in the training sample set, the training sample is input into the central intelligent agent, and the central intelligent agent determines the task planning scheme. According to the task planning scheme, the training samples are processed by the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent, and the corresponding reward values of the central intelligent agent, the technical solution optimization intelligent agent, the technical solution evaluation intelligent agent, and the technical solution verification intelligent agent are obtained during the processing. Based on the reward values corresponding to the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent, the cumulative reward of the task planning scheme is determined. Based on the cumulative rewards, the parameters of the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent are optimized until the preset training termination conditions are met, resulting in the trained central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent.
[0097] In some embodiments, the reward value corresponding to the central agent is obtained through the reward function of the central agent; The reward function for the central agent is: ; in, The evaluation score for the target technical solution. To generate the computational resource cost invoked for the target technical solution, It refers to the number of times the central intelligent agent generates optimization instructions during the process of generating the target technical solution. for The weight value, for The weight value, for The weight value.
[0098] In some embodiments, the reward value of the technical solution optimization agent refers to the cumulative sum of reward values obtained in each optimization iteration round of the processing; the reward value obtained in each optimization iteration is obtained through the reward function of the technical solution optimization agent; The reward function for the optimized agent in this technical solution is as follows: ; in, To optimize the reward function value of the intelligent agent in the technical solution, This represents the evaluation score of the optimized technical solution in the current optimization iteration round. It is an internal reward that measures the completeness and logical consistency of the optimized technical solution in the current optimization iteration round. This is a penalty term used to characterize the deviation between the optimized instruction and the optimized technical solution in the current optimization iteration round. for The weight value, for The weight value, for The weight value.
[0099] In some embodiments, the reward value corresponding to the technical solution evaluation agent is obtained through the reward function of the technical solution evaluation agent; The reward function for evaluating the intelligent agent in the technical solution evaluation is as follows: ; in, The reward function value of the agent is used to evaluate the technical solution. The initial technical solution is scored for its novelty. The novelty of the target technical solution is scored. The reward is a correlation factor between the initial technical solution and similar solutions retrieved by the technical solution evaluation agent. This is the weighting value for novelty. for The weight value.
[0100] In some embodiments, the reward function for the technical solution to verify the intelligent agent is:
[0101] in, To verify the reward function value of the intelligent agent in the technical solution, This refers to the i-th specific error instance that is corrected in the target technical solution.
[0102] In some embodiments, the technical solution evaluation agent 303 is specifically used for: The intelligent agent that evaluates the technical solutions retrieves at least one similar solution from the initial technical solution. Evaluation results are generated based on similar schemes and initial technical schemes.
[0103] In some embodiments, the technical solution evaluation agent 303 is further specifically used for: The agent is evaluated by invoking technical solutions, and technical feature vectors are generated based on the initial technical solutions. Calculate the cosine similarity value between the technical feature vector and the candidate feature vectors of each candidate solution in the solution library; Based on the cosine similarity value, at least one similar solution is determined from the candidate solutions in the solution library.
[0104] In some embodiments, the technical solution evaluation agent 303 is further specifically used for: Compare similar solutions with the initial technical solution, and generate a novelty score and an inventiveness score for the initial technical solution based on the comparison results; An evaluation score is calculated based on the novelty score and the inventiveness score, and an evaluation result is generated based on the novelty score, the inventiveness score, and similar solutions.
[0105] The above embodiments of this application describe the system provided in the embodiments of this application. To implement the functions of the methods provided in the above embodiments of this application, the optimized system may include hardware structure and software modules (intelligent agents), implementing the above functions in the form of hardware structure, software modules, or a combination of hardware structure and software modules. One of the above functions may be executed in the form of hardware structure, software module, or a combination of hardware structure and software modules.
[0106] Figure 4 This is a block diagram illustrating an electronic device 400 for implementing the above-described optimized technical solution, according to an exemplary embodiment. For example, the electronic device 400 may be a computer, a personal digital assistant, etc.
[0107] Reference Figure 4 The electronic device 400 may include a communication interface 401, capable of interacting with other devices; a processor 402, connected to the communication interface 401 to enable interaction with other devices, used to execute the methods provided by one or more of the above-described technical solutions when running a computer program; and a memory 403, on which the computer program is stored. Specifically, the specific processing operations of the processor 402 can refer to the optimized methods of the technical solutions described in the above embodiments of this disclosure.
[0108] Of course, in practical applications, the various components in electronic device 400 are coupled together through bus system 404. It can be understood that bus system 404 is used to realize the connection and communication between these components. In addition to a data bus, bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 4 The general designated all buses as Bus System 404.
[0109] The memory 403 in this embodiment is used to store various types of data to support the operation of the electronic device 400. Examples of such data include any computer program used to operate on the electronic device 400.
[0110] The methods disclosed in the embodiments of this application can be applied to processor 402, or implemented by processor 402. Processor 402 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 402 or by instructions in software form. The processor 402 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 402 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 403. Processor 402 reads the information in memory 403 and, in conjunction with its hardware, completes the steps of the aforementioned method.
[0111] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0112] Embodiments of this disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the technical solution optimization method described in the above embodiments of this disclosure.
[0113] The embodiments of this disclosure also propose a computer program product, including a computer program that, when executed by a processor, implements the technical solution optimization method described in the above embodiments of this disclosure.
[0114] Embodiments of this disclosure also propose a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, it causes the electronic device to perform the technical solution optimization method described in the above embodiments of this disclosure.
[0115] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0116] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0117] Any description of operation or method in the flowchart or otherwise described herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or operation, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0118] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0119] It should be understood that various parts of the embodiments of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0120] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by instructing related hardware through an operation sequence, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0121] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc.
[0122] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for optimizing a technical solution, characterized in that, The method includes: The central intelligent agent responds to receiving basic information input by the user and identifies the user's intent in the basic information; In response to the user intent, a technical solution optimization is performed. Based on the user's indication of the user intent, a task planning scheme for the user intent is determined and executed to generate a target technical solution. The task planning scheme includes: invoking a technical solution optimization agent to generate an initial technical solution based on the basic information; invoking a technical solution evaluation agent to evaluate the initial technical solution and generate an evaluation result containing an evaluation score; in response to the evaluation score being less than a preset threshold, the central agent generating an optimization instruction based on the evaluation result; invoking the technical solution optimization agent to iteratively optimize the initial technical solution based on the optimization instruction until the evaluation score of the optimized technical solution is not less than the preset threshold, and determining the optimized technical solution as the target technical solution.
2. The method according to claim 1, characterized in that, After determining the optimized technical solution as the target technical solution, the task planning scheme further includes: The intelligent agent is invoked to verify the target technical solution.
3. The method according to claim 2, characterized in that, The method further includes: Through training, a central intelligent agent, a technical solution optimization intelligent agent, a technical solution evaluation intelligent agent, and a technical solution verification intelligent agent are obtained. The process of training to obtain the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent includes: Obtain the training sample set; For each training sample in the training sample set, the training sample is input into the central intelligent agent, and the central intelligent agent determines the task planning scheme. According to the task planning scheme, the training samples are processed by the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent, and the reward values corresponding to the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent are obtained during the processing. Based on the reward values corresponding to the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent, the cumulative reward of the task planning scheme is determined. Based on the cumulative rewards, optimize the parameters of the central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent until the preset training termination condition is met, and obtain the trained central agent, the technical solution optimization agent, the technical solution evaluation agent, and the technical solution verification agent.
4. The method according to claim 3, characterized in that, The reward value corresponding to the central agent is obtained through the reward function of the central agent; The reward function of the central intelligent agent is as follows: ; in, The evaluation score for the target technical solution is... The computational resource cost invoked to generate the target technical solution It refers to the number of times the central intelligent agent generates the optimization instructions during the process of generating the target technical solution. for The weight value, for The weight value, for The weight value.
5. The method according to claim 3, characterized in that, The reward value of the optimized agent in the technical solution refers to the cumulative sum of the reward values obtained in each optimization iteration round of the processing process; The reward value obtained in each optimization iteration is obtained by optimizing the reward function of the agent using the aforementioned technical solution; The reward function of the optimized agent in the above technical solution is as follows: ; in, To optimize the reward function value of the agent in the described technical solution, This represents the evaluation score of the optimized technical solution in the current optimization iteration round. It is an internal reward that measures the completeness and logical consistency of the optimized technical solution in the current optimization iteration round. This is a penalty term used to characterize the deviation between the optimization instruction and the optimized technical solution in the current optimization iteration round. for The weight value, for The weight value, for The weight value.
6. The method according to claim 3, characterized in that, The reward value corresponding to the evaluation agent of the technical solution is obtained through the reward function of the evaluation agent of the technical solution. The reward function for evaluating the intelligent agent in the aforementioned technical solution is as follows: ; in, To evaluate the reward function value of the agent in the described technical solution, The originality of the initial technical solution is scored. To score the novelty of the target technical solution, the The reward term is the correlation between the initial technical solution and similar solutions retrieved by the evaluation agent. This is the weighting value for novelty. for The weight value.
7. The method according to claim 3, characterized in that, The reward function for the verified intelligent agent in the aforementioned technical solution is: in, To verify the reward function value of the agent in the described technical solution, This refers to the i-th specific error instance that is corrected in the target technical solution.
8. The method according to claim 1, characterized in that, The invoking technology solution evaluation agent evaluates the initial technology solution and generates an evaluation result containing an evaluation score, including: The intelligent agent that evaluates the technical solution is invoked to search for the initial technical solution and obtain at least one similar solution; The evaluation results are generated based on the similar scheme and the initial technical scheme.
9. The method according to claim 8, characterized in that, The process of calling the technical solution evaluation agent to search for the initial technical solution and obtain at least one similar solution includes: The intelligent agent is evaluated by invoking the aforementioned technical solution, and a technical feature vector is generated based on the initial technical solution. Calculate the cosine similarity value between the technical feature vector and the candidate feature vectors of each candidate scheme in the scheme library; Based on the cosine similarity value, at least one similar scheme is determined from the candidate schemes in the scheme library.
10. The method according to claim 8, characterized in that, The process of generating the evaluation result based on the similar scheme and the initial technical scheme includes: The similar solutions and the initial technical solution are compared, and a novelty score and an inventiveness score for the initial technical solution are generated based on the comparison results. The evaluation score is calculated based on the novelty score and the inventiveness score, and the evaluation result is generated based on the novelty score, the inventiveness score, and the similar solutions.
11. A technical solution optimization system, characterized in that, The system includes: A central intelligent agent is used to respond to receiving basic information input by a user, identify the user's intent in the basic information, and, in response to the user's intent for technical solution optimization, determine and execute a task planning scheme for the user's intent to generate a target technical solution. The technical solution optimizes the intelligent agent, which is used to respond to the call of the central intelligent agent and generate an initial technical solution based on the basic information; A technical solution evaluation agent is used to evaluate the initial technical solution in response to a call from the central agent and generate an evaluation result containing an evaluation score. The central intelligent agent is also used to generate an optimization instruction based on the evaluation result in response to the evaluation score being less than a preset threshold. The optimized agent is also used to respond to the call of the central agent and iteratively optimize the initial technical solution based on the optimization instructions until the evaluation score of the optimized technical solution is not less than a preset threshold, and then determine the optimized technical solution as the target technical solution.
12. The system according to claim 11, characterized in that, The system also includes a technical solution verification agent: The technical solution verification agent is used to verify the target technical solution in response to the call of the central agent.
13. An electronic device, characterized in that, Including memory and processor; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, thereby implementing the method as described in any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 10.
15. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the method of any one of claims 1 to 10.