Cooperative interaction strategy optimization method based on multiple agents
Through the multi-agent-based collaborative interaction strategy optimization method, the high-quality parent characteristics of large language model instructions are systematically integrated, and the problem of the failure to effectively integrate the advantageous characteristics of multiple parent instructions in the existing technology is solved, and the efficiency and quality of instruction optimization are improved.
Patent Information
- Application Number
- CN202510837485.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When the existing automation prompt engineering (APE) framework is iteratively optimized to generate large language model instructions, it failed to effectively integrate the advantages of multiple high-quality parent instructions, resulting in insufficient exploration of potential synergistic efficiency value during the iteration proposal process, limiting the continuous improvement of instruction quality and optimization efficiency.
The collaborative interaction strategy optimization method based on multi-agents is adopted, and the collaboration contribution potential index of high-quality parent instructions is obtained, and the advantageous feature fragments of multiple parent instructions are systematically fused, and the proposed model is used to collaborate and generate child instructions, improving the overall performance of instruction optimization.
It improves the efficiency and quality of the instruction optimization process of large language model, and the generated child instructions can better integrate the advantages of multiple sources, improving the overall performance of the automated instruction optimization process.
Smart Images

Figure CN120354878A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of strategy optimization, and particularly to a multi-agent based collaborative interaction strategy optimization method. Background Art
[0002] Large language models (LLMs) demonstrate powerful general computing capabilities by following natural language instructions (prompts), and their performance on various tasks highly depends on the quality of the instructions used. To reduce the cost and difficulty of manually designing and validating high-quality instructions, the Automatic Prompt Engineering (APE) technology has emerged. The process of finding the optimal instructions is formalized as an optimization problem, which includes three core stages: First, use an LLM (as a proposal model) to generate a set of candidate instructions based on the task description and a small number of input-output examples; Second, score each candidate instruction, that is, evaluate its effect of guiding the target LLM to complete the task; Finally, select the instruction with the highest score as the output according to the score.
[0003] To further improve the quality of instructions, the idea of iterative optimization is also introduced in the APE framework. In this iterative mechanism, the system will, based on the relatively better candidate instructions (called parent instructions) selected in the previous round of evaluation, use the proposal model again to generate a new generation of candidate instructions (called child instructions) with expected better performance. Existing mainstream iterative proposal methods usually guide the proposal model to perform independent, mainly semantically similar local rewriting or mutation on each selected parent instruction, in the hope of making small improvements based on the parent instructions, so as to gradually approach better instructions.
[0004] However, when guiding the generation of a new generation of child instructions based on multiple parent instructions selected in the previous round of screening, the commonly adopted strategy is to perform independent, locally semantically similar rewriting on each parent instruction. This strategy essentially treats each parent instruction as an isolated optimization object, and thus lacks an effective mechanism to identify and actively integrate the respective advantageous features or successful instruction construction patterns contained in these different parent instructions. Due to the inability to achieve the complementary advantages between different parent instructions, the newly generated child instructions are only simple improvements on a single parent instruction, and it is difficult to generate instructions that achieve performance breakthroughs. This results in insufficient exploration of the potential synergistic value of multiple parent instructions in this iterative proposal process, limiting the upper limit and optimization efficiency of the APE framework to continuously improve the instruction quality through iteration, and making it prone to prematurely converge to local optimal solutions. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a method for optimizing a collaborative interaction strategy based on multi - agents to solve the problem of how to promote the effective integration of the advantageous features of multiple parent instructions in the generation of child instructions, thereby improving the overall performance of instruction optimization.
[0006] An embodiment of the present invention provides a method for optimizing a collaborative interaction strategy based on multi - agents, and the method includes the following steps: During the process of APE generating instructions through iterative optimization, obtain all high - quality parent instructions in any iteration round; For any high - quality parent instruction, identify and count the text content of any high - quality parent instruction to obtain the number of explicit structural units, and obtain the structured instruction clarity according to the number of explicit structural units; obtain the core instruction semantic intensity according to the occurrence frequency of specific parts of speech in the text content of any high - quality parent instruction, and combine the structured instruction clarity and the core instruction semantic intensity to obtain the collaborative contribution potential index of any high - quality parent instruction; Obtain the collaborative contribution potential index of each high - quality parent instruction, and according to the collaborative contribution potential index of each high - quality parent instruction, sequentially select a core contribution parent instruction from all high - quality parent instructions, extract the dominant advantageous feature segments of each selected core contribution parent instruction, and use a proposal model to perform text recombination on each of the dominant advantageous feature segments to obtain a set of child instructions; Select the high - quality parent instructions for the next iteration round of any iteration round in the set of child instructions, and repeat the method for obtaining the set of child instructions until the iteration round requirements are met to obtain a set of candidate instructions, and select the candidate instruction with the highest quality evaluation result from the set of candidate instructions.
[0007] Preferably, the step of obtaining the structured instruction clarity according to the number of explicit structural units includes: Use the sum of the number of explicit structural units and the constant 1 as the independent variable of the natural logarithm function to obtain the structured instruction clarity.
[0008] Preferably, the step of obtaining the core instruction semantic intensity according to the occurrence frequency of specific parts of speech in the text content of any high - quality parent instruction includes: Respectively count the occurrence frequency of verbs, the occurrence frequency of demonstrative pronouns, and the occurrence frequency of deontic auxiliary verbs in the text content of any high - quality parent instruction, calculate the total frequency of the occurrence frequency of verbs and the occurrence frequency of deontic auxiliary verbs, calculate the product of a preset non - negative penalty coefficient and the occurrence frequency of demonstrative pronouns, use the difference between the total frequency and the product as the numerator, and the instruction length of any high - quality parent instruction as the denominator to obtain the core instruction semantic intensity.
[0009] Preferably, obtaining the collaborative contribution potential index of any high-quality parent instruction by combining the clarity of the structured instruction and the semantic intensity of the core instruction includes: Perform normalization processing on the clarity of the structured instruction and the semantic intensity of the core instruction respectively to obtain the normalized value of the clarity of the structured instruction and the normalized value of the semantic intensity of the core instruction, and record the sum value of the normalized value of the clarity of the structured instruction and the normalized value of the semantic intensity of the core instruction as the instruction feature value of any high-quality parent instruction; Obtain the instruction feature values corresponding to all high-quality parent instructions, use the mean value of all instruction feature values as the bias term, and use the difference between the instruction feature value of any high-quality parent instruction and the bias term as the independent variable of the Sigmoid activation function to obtain the collaborative contribution potential index of any high-quality parent instruction.
[0010] Preferably, according to the collaborative contribution potential index of each high-quality parent instruction, sequentially select a core contribution parent instruction from all high-quality parent instructions, including: According to the collaborative contribution potential index of each high-quality parent instruction, use the roulette selection method to sequentially select a core contribution parent instruction from all high-quality parent instructions.
[0011] Preferably, extracting the dominant advantage feature segment of the core contribution parent instruction selected each time includes: According to the clarity of the structured instruction and the semantic intensity of the core instruction of all high-quality parent instructions, calculate the mean value of the clarity of the structured instruction and the mean value of the semantic intensity of the core instruction respectively. For the core contribution parent instruction selected each time, if the semantic intensity of the core instruction of the core contribution parent instruction selected each time is greater than or equal to the mean value of the semantic intensity of the core instruction, then use the smallest text unit where the verb and deontic verb are located in the text content of the core contribution parent instruction selected each time as the dominant advantage feature segment; If the clarity of the structured instruction of the core contribution parent instruction selected each time is greater than or equal to the mean value of the clarity of the structured instruction, then use all explicit structural units in the text content of the core contribution parent instruction selected each time as the dominant advantage feature segment.
[0012] Preferably, using the proposal model to perform text recombination on each of the dominant advantage feature segments to obtain a set of child instructions, including: Take the dominant advantageous feature segment of any selected core contribution parent instruction as the target segment, use the proposal model to reorganize the text of the target segment to obtain a candidate child instruction, and check the legality and effectiveness of the candidate child instruction. If the check passes, use the candidate child instruction as the child instruction. If the check fails, take the dominant advantageous feature segment of the next selected core contribution parent instruction as the target segment, and repeat the above check method until the number of child instructions meets the quantity requirement to obtain a set of child instructions.
[0013] Preferably, the checking of the legality and effectiveness of the candidate child instruction includes: If the candidate child instruction is empty, or the instruction length is not within the length range, or it contains prohibited words, it is determined that the check fails; if the candidate child instruction is not empty, the instruction length is within the length range, and it does not contain prohibited words, it is determined that the check passes.
[0014] The beneficial effects of the embodiments of the present invention compared with the prior art are: By introducing the collaborative contribution potential index, the present invention characterizes the structural and semantic advantages worthy of reference within each high-quality parent instruction, and then, based on the collaborative contribution potential index, systematically and with emphasis integrates the advantageous feature segments (i.e., the dominant advantageous feature segments) of multiple high-collaborative contribution potential parent instructions. By using the collaborative recombination hint template to guide the proposal model, these advantageous feature segments from different high-quality parents are reconstructed to generate new child instructions. The new child instructions generated by this optimization strategy can more effectively integrate the advantages of multiple sources, making the entire APE framework more efficient when exploring the instruction space and improving the overall performance of the automated instruction optimization process. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of a method for optimizing a collaborative interaction strategy based on multi-agent provided in Embodiment 1 of the present invention. Detailed Embodiments
[0017] The following will describe in detail the embodiments of the present disclosure, and the examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation of the present disclosure.
[0018] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0019] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below. It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device or terminal system capable of implementing the above functions.
[0020] See Figure 1 , which is a method flow chart of a multi-agent-based collaborative interaction strategy optimization method provided in Embodiment 1 of the present invention. As Figure 1 shown, the method may include: Step S101, during the process of the APE generating instructions through iterative optimization, obtain all high-quality parent instructions in any iteration round.
[0021] Large Language Models (LLMs), also known as large language models or large models, first, as the name implies, are large in scale, with the number of network parameters reaching tens of billions, hundreds of billions or even more; second, generality, which means not limited to specific problems or fields; third, emergence, that is, the emergence of unexpected new capabilities. Generative large language models are models or large models with natural language processing, image processing, or speech processing capabilities, and with the ability to generate content such as text, pictures, audio, and video. In many professional fields, such as law, medicine, and finance, large language models are applied to handle complex natural language understanding and generation tasks. In the legal field, the model is required to generate a case summary based on the case file and extract key issues; in the medical field, the model is required to interpret the pathology report and explain it in a way that is easy for patients to understand; in the intelligent customer service scenario, the model needs to conduct multi-round conversations, accurately understand the user's intention, and provide appropriate solutions. For such complex tasks, it is often extremely challenging to manually design prompt words that can comprehensively, accurately, and efficiently guide large language models to produce high-quality outputs, which requires a lot of professional knowledge and repeated trial and error.
[0022] Existing Automated Prompt Engineering (APE) generates large language model instructions through an iterative optimization process, mainly including the following three steps: (1) Initial candidate instruction generation. (2) Iterative optimization, including: selecting a subset of evaluation data, scoring candidate instructions, screening high-quality parent instructions, iterative proposal and child instruction generation (using the proposal model LLM, based on high-quality parent instructions, through semantic similarity rewriting, mutation or resampling of them, to generate a new generation of candidate instructions, that is, child instructions), and checking the iterative termination condition (the number of iterative rounds meets the required number). (3) Selection and output of the final instruction, that is, the candidate instruction with the highest score among candidate instructions. A candidate instruction is a piece of text or a question provided for a generative large language model to guide the generative large language model to generate a specific type of text or answer. A candidate instruction can be a sentence, a question, an article, or a topic, etc., which can guide the generative large language model to generate information or answers related to the candidate instruction. For example, when given the question "What is artificial intelligence?", the generative large language model can generate an article or an answer about artificial intelligence.
[0023] However, when generating large language model instructions through the above iterative optimization process, the core limitation of its iterative optimization step is that the newly generated child instructions are only simple improvements on a single high-quality parent instruction, and fail to fully explore and effectively integrate the unique advantageous features contained in multiple high-quality parent instructions. If only relying on the overall performance score of high-quality parent instructions for subsequent optimization decisions, it is difficult to distinguish whether this score stems from a general excellent instruction design or a coincidence in a specific situation, thus hindering the possibility of inheriting and integrating truly valuable instruction genes into the new generation of instructions. Therefore, the embodiments of the present invention provide an improved iterative proposal method to promote the effective integration of the advantageous features of multiple high-quality parent instructions in the generation of child instructions, thereby improving the overall performance of instruction optimization.
[0024] Since the embodiments of the present invention optimize the step of iterative proposal and child instruction generation based on high-quality parent instructions in the iterative optimization process of step (2), therefore, at the beginning of a typical APE iterative optimization cycle, a set of high-quality parent instructions for guiding the generation of child instructions in this round needs to be determined. Then, in the embodiments of the present invention, taking any iterative round as an example, during the process of APE generating instructions through iterative optimization, according to the candidate instructions (child instructions) generated in the previous iterative round of any iterative round, high-quality parent instructions are screened out from all candidate instructions through candidate instruction scoring for the generation of child instructions in any iterative round. Thus, all high-quality parent instructions in any iterative round can be obtained. It should be noted that in the first iteration, the high-quality parent instructions come from a part of the initial candidate instructions generated by the proposal model based on the initial task description and a small number of examples, which are screened out through candidate instruction scoring.
[0025] Step S102: For any high-quality parent instruction, identify and count the text content of any high-quality parent instruction to obtain the number of explicit structural units. Based on the number of explicit structural units, obtain the clarity of the structured instruction. According to the occurrence frequency of specific parts of speech in the text content of any high-quality parent instruction, obtain the core instruction semantic intensity. Combine the clarity of the structured instruction and the core instruction semantic intensity to obtain the collaborative contribution potential index of any high-quality parent instruction.
[0026] To achieve more efficient iterative proposals, a key prerequisite is to be able to accurately identify and quantify the key structural elements or features in the internal text of each high-quality parent instruction that play a decisive role in its excellent performance and have subsequent migration and combination value. Therefore, the embodiments of the present invention introduce the collaborative contribution potential index. This is used to quantify the potential value of each high-quality parent instruction as a knowledge contributor. This value is specifically reflected in the level of structured organization shown in the text content of the high-quality parent instruction and the clarity and concentration of the core instruction semantics. A high-quality parent instruction with a rich and clear structural hierarchy and capable of accurately and efficiently conveying the core instruction intention has a higher collaborative contribution potential and should be given a higher weight or more preferential consideration in the subsequent process of guiding the generation of better offspring instructions. Therefore, in the embodiments of the present invention, a collaborative contribution potential index is quantified for each high-quality parent instruction, thereby providing a key decision-making basis for the subsequent collaborative integration and innovative combination of high-quality parent instructions.
[0027] Taking any high-quality parent instruction as an example, first, perform standard natural language processing operations on the text content of any high-quality parent instruction, specifically including word segmentation, part-of-speech tagging, and being able to count its total number of words. and the number of explicit structural units . It should be noted that natural language processing operations belong to the prior art and will not be elaborated here. Then, identify and count the number of explicit structural units it contains. Among them, the explicit structural units specifically include: independent list items guided by standard list markers (such as patterns starting with "1.", "a)", "-", "*", etc. at the beginning of the line), and text paragraphs containing substantial instruction content clearly separated by one or more consecutive blank lines. Finally, based on the number of explicit structural units, obtain the structured instruction clarity of any high-quality parent instruction. The specific method is: use the sum of the number of explicit structural units and the constant 1 as the independent variable of the natural logarithm function to obtain the structured instruction clarity.
[0028] In one embodiment, the calculation formula for the structured instruction clarity is:
[0029] Among them, represents the structural instruction clarity of any high-quality parent instruction, represents the natural logarithm function, and 1 represents a constant, represents the number of explicit structural units contained in the text content of any high-quality parent instruction.
[0030] It should be noted that the structural instruction clarity can identify and quantify any high-quality parent instruction in terms of the design advantages and disadvantages in information organization and process guidance. Through the evaluation of the logarithmic transformation of the number of explicit structural units in its text content, it effectively gives a positive evaluation to the high-quality parent instructions that adopt a clear hierarchical and easy-to-follow structure to present the instruction content. This quantification of the instruction structure enables a more basis for reference and retention of the organizational framework of high-quality parent instructions that perform excellently in instruction clarity and standardization during the subsequent collaborative recombination process, thereby improving the understandability and execution accuracy of the child instructions.
[0031] Complementary to the structural level, first, count the occurrence frequencies of verbs , occurrence frequencies of demonstrative pronouns and occurrence frequencies of deontic auxiliary verbs in the text content of any high-quality parent instruction respectively. Among them, verbs include but are not limited to "generate", "analyze", "summarize", "ensure", "compare", "list", etc., especially verbs commonly used to express instructions, commands or requests; deontic auxiliary verbs include but are not limited to modal verbs or auxiliary verbs used to express necessity, mandatory or prohibitive nature, such as: "must", "should", "need", "can", "shall not", "forbid", etc.; demonstrative pronouns refer to words that introduce semantic ambiguity when there is no clear referent, such as: "it", "this", "that", "they", "these", etc. Then, by combining the occurrence frequencies of the above-mentioned parts of speech and considering the instruction length of any high-quality parent instruction (that is, the total number of words ), obtain the core instruction semantic intensity. The specific obtaining method is: Calculate the sum of the occurrence frequencies of verbs and the occurrence frequencies of deontic auxiliary verbs, calculate the product of the preset non-negative penalty coefficient and the occurrence frequency of demonstrative pronouns, take the difference between the sum of frequencies and the product as the numerator, and the instruction length of any high-quality parent instruction as the denominator to obtain the core instruction semantic intensity of any high-quality parent instruction .
[0032] In one embodiment, the calculation formula of the core instruction semantic intensity is:
[0033] Among them, represents the core instruction semantic intensity of any high-quality parent instruction, represents a preset non-negative penalty coefficient, which is set to 0.3.
[0034] It should be noted that focuses on evaluating the instruction concentration and semantic clarity of the text content of any high-quality parent instruction . By positively motivating the density of the core action verbs and restrictive modal words in the instruction, and moderately suppressing the use of demonstrative pronouns that may introduce ambiguity, the efficiency of the instruction text in transmitting key instruction information such as core operation requirements and target constraints per unit length is quantified. A high-quality parent instruction with a higher value usually means that its instruction intention is more direct, the action direction is more clear, and its potential as a high-quality gene contributor is also correspondingly higher. Therefore, the analysis of the core instruction semantic intensity can effectively distinguish between instructions that are concise and to the point and those with redundant content and unclear key points.
[0035] After obtaining the structured instruction clarity and core instruction semantic intensity of any high-quality parent instruction , the structured instruction clarity and core instruction semantic intensity of any high-quality parent instruction are regularly aggregated to obtain the collaborative contribution potential index of any high-quality parent instruction. Among them, the method of regular aggregation is as follows: Normalize the structured instruction clarity and core instruction semantic intensity respectively to obtain the normalized value of the structured instruction clarity and the normalized value of the core instruction semantic intensity. Denote the sum of the normalized value of the structured instruction clarity and the normalized value of the core instruction semantic intensity as the instruction characteristic value of any high-quality parent instruction; Obtain the instruction characteristic values corresponding to all high-quality parent instructions, take the mean of all instruction characteristic values as the bias term, and take the difference between the instruction characteristic value of any high-quality parent instruction and the bias term as the independent variable of the Sigmoid activation function to obtain the collaborative contribution potential index of any high-quality parent instruction.
[0036] In an implementation manner, the calculation formula for the collaborative contribution potential index of any high-quality parent instruction is:
[0037] Among them, represents the collaborative contribution potential index of any high-quality parent instruction, is the standard activation function, and its mathematical expression is , represents the normalization function, represents the bias term.
[0038] It should be noted that after obtaining the structural instruction clarity and core instruction semantic strength of any high-quality parent instruction respectively, they are adjusted to a unified scale through a dynamic min-max normalization function, ensuring the adaptability and comparability of the scores among different iteration stages or different tasks; the two normalized scores are directly added and then subtracted by a bias term dynamically calculated based on all the data of high-quality parent instructions in the current round , forming an original potential value that comprehensively reflects the dual efficacy of instruction structure and content. The setting of this dynamic bias term enables the input of the Sigmoid function to be centered around the average level of the current parent population, thus making more effective use of the sensitive discrimination region of the Sigmoid function. Applying the standard Sigmoid activation function to the original potential value adjusted by the bias term, this non-linear transformation not only normalizes the final collaborative contribution potential index to the interval (0, 1), but more importantly, through its S-shaped curve characteristics, it can effectively amplify the score differences between high-quality parent instructions that are outstanding or poor in both the structural and semantic dimensions, thereby improving the recognition and discrimination of high-quality parent instructions with high collaborative contribution potential.
[0039] Similarly, the collaborative contribution potential index of each high-quality parent instruction in any iteration round is obtained.
[0040] Step S103: Obtain the collaborative contribution potential index of each high-quality parent instruction. According to the collaborative contribution potential index of each high-quality parent instruction, sequentially select a core contribution parent instruction from all high-quality parent instructions, extract the dominant advantage feature segments of each selected core contribution parent instruction, and use the proposal model to perform text recombination on each dominant advantage feature segment to obtain a set of child instructions.
[0041] According to the above method, the collaborative contribution potential index of each high-quality parent instruction is obtained, and this index reveals the internal structural and semantic advantages worthy of reference in each high-quality parent instruction. To overcome the problem that the existing APE iterative proposal method is difficult to effectively integrate the scattered advantages due to the local optimization of high-quality parent instructions, the embodiment of the present invention proposes a child instruction generation strategy based on the guidance of the collaborative contribution potential index to promote the collaborative recombination of the advantage features of multiple high-quality parent instructions. Among them, the specific method of this generation strategy is as follows: (1) According to the collaborative contribution potential index of each high-quality parent instruction, use the Roulette Wheel Selection method to sequentially select a core contribution parent instruction from all high-quality parent instructions.
[0042] Calculate the cumulative value of the collaborative contribution potential index for all high-quality parent instructions ; Calculate the ratio of the collaborative contribution potential index of each high-quality parent instruction to the cumulative value, and record it as the selection probability of the corresponding high-quality parent instruction ; Based on the selection probability of each high-quality parent instruction, conduct a random sampling among all high-quality parent instructions, and the selected high-quality parent instruction is the core contribution parent instruction , this method ensures that the higher the value of the high-quality parent instruction, the more likely its core features and structure will be inherited by the offspring
[0043] (2) Extract the dominant advantageous feature segments of the core contribution parent instruction selected each time
[0044] For any selected core contribution parent instruction , calculate the mean value of the structured instruction clarity and the mean value of the core instruction semantic strength respectively according to the structured instruction clarity and the core instruction semantic strength of all high-quality parent instructions and the mean value of the core instruction semantic strength , if the structured instruction clarity of any selected core contribution parent instruction is greater than or equal to the mean value of the structured instruction clarity , it indicates that the core contribution parent instruction has an advantage in structural organization. At this time, all explicit structural units in the text content of any selected core contribution parent instruction are used as the dominant advantageous feature segments, that is, all explicit structural units identified when calculating the of the core contribution parent instruction are considered as contributing structural feature segments. All explicit structural units in the text content of any selected core contribution parent instruction are used as the dominant advantageous feature segments. Otherwise, no dominant advantageous feature segments are extracted from this part
[0045] If the core instruction semantic strength of any selected core contribution parent instruction is greater than or equal to the mean value of the core instruction semantic strength , it shows that the core contribution parent instruction has an advantage in core semantic expression. At this time, the smallest text unit where the verb and deontic auxiliary verb are located in the text content of any selected core contribution parent instruction is used as the dominant advantageous feature segment. Otherwise, no dominant advantageous feature segments are extracted from this part. It should be noted that if a text unit contains multiple such core words at the same time, it is extracted as a whole segment without double counting. Through this method, it is ensured that the extracted semantic feature segments are the parts in the instruction that actually carry the core operations or constraints
[0046] So far, it is possible to obtain the feature segments with positive contributions in any selected core contributing parent instruction based on the dominant advantageous feature segments extracted from the two parts.
[0047] (3) Use the proposal model to perform text recombination on each dominant advantageous feature segment to obtain a set of child instructions.
[0048] Take the dominant advantageous feature segment of any selected core contributing parent instruction as the target segment , use the proposal model to perform text recombination on the target segment to obtain a candidate child instruction, and perform legality and validity checks on the candidate child instruction. The specific check method is as follows: If the candidate child instruction is empty, or the instruction length is not within the length range, or it contains prohibited words, it is determined that the check fails; if the candidate child instruction is not empty, the instruction length is within the length range, and it does not contain prohibited words, it is determined that the check passes.
[0049] If the check passes, take the candidate child instruction as the child instruction. If the check fails, take the dominant advantageous feature segment of the next selected core contributing parent instruction as the target segment , repeat this check method until the check is passed and the number of obtained child instructions meets the quantity requirement , at this time, form all child instructions into a set of child instructions, and the number of elements in the set of child instructions is . Among them, the quantity requirement is not restricted and can be set according to requirements.
[0050] It should be noted that for performing text recombination on the target segment using the proposal model , a structured prompt template needs to be designed. This template clearly takes the target segment as the main body and guides the proposal model to perform text recombination and generation. Among them, the core composition of the prompt template is: Task context: Clearly describe the target task of the current instruction optimization; Dominant feature instruction: "Please be based on the following core instruction idea and structure: [Organize the segments / patterns in the target segment in a natural language manner]."; Innovation and optimization requirements: "Please integrate, optimize, and innovate the above information to generate a brand-new, complete, and fluent instruction that may perform better in [specific desired improvement directions, such as being clearer, more concise, more guiding, or better able to handle a certain specific situation]. Please note that the core goal of the instruction should be consistent with the task requirements. The generated new instruction should be: xxxx". Submit the constructed collaborative recombination and innovation prompt containing the above guiding content to the preset proposal model (LLM), and the proposal model will process it based on this prompt and generate a new piece of text as a candidate offspring instruction. .
[0051] At this time, the task of generating offspring instructions for high-quality parent instructions in any iteration round is completed. This generation strategy no longer regards high-quality parent instructions as independent optimization starting points, but as a shared pool of high-quality feature fragments. By systematically selecting high-quality parent instructions with high contribution potential (that is, high collaborative contribution potential index), and clearly extracting their most significant dominant feature fragments, and then guiding the proposal model through the collaborative recombination prompt template to reconstruct these dominant feature fragments from different high-quality parent instructions, new offspring instructions are generated.
[0052] Step S104, select the high-quality parent instruction for the next iteration round in any iteration round of the offspring instruction set, and repeat the method for obtaining the offspring instruction set until the iteration round requirement is met, obtaining a set of candidate instruction sets, and select the candidate instruction with the highest quality evaluation result in the candidate instruction set.
[0053] Since the core of the optimization method provided in the embodiments of the present invention lies in improving the iterative proposal link in the automated prompt engineering framework (APE), that is, when generating a new generation of candidate instructions (offspring instructions) based on the high-quality parent instructions evaluated and screened in the previous iteration round, a strategy for generating offspring instructions that promotes the collaborative recombination of the dominant feature fragments of multiple high-quality parent instructions by guiding based on the collaborative contribution potential index as provided above is used to replace the original iterative proposal link in the existing automated prompt engineering (APE). Therefore, after completing the task of generating offspring instructions for high-quality parent instructions in any iteration round and obtaining an offspring instruction set, further evaluate and screen the high-quality parent instruction for the next iteration round in the offspring instruction set, and according to the above method for generating the offspring instruction set, obtain its corresponding offspring instruction set, and repeat the process of obtaining the offspring instruction set until the iteration round meets the iteration round requirement, and output a set of candidate instruction sets, that is, the offspring instruction set obtained in the last iteration round. It should be noted that the iteration round requirement means that the iteration round is between 20 and 50. Preferably, the number of iteration rounds set in the embodiments of the present invention is 30, which is not limited here and can be set according to the actual implementation scenario.
[0054] It is known that in the existing automated prompt engineering (APE), in the step of generating large language model instructions through an iterative optimization process, after completing the iterative optimization to obtain a set of candidate instruction sets, the final instruction can be selected and output from the candidate instruction sets. That is, the candidate instruction with the highest score in the candidate instruction sets is selected as the output of the large language model instruction.
[0055] It should be noted that the focus of the present invention is to provide an improved candidate instruction generation method for the APE framework, that is, by quantifying the contribution potential of each knowledge unit of each high-quality parent instruction and guiding the proposal model to synergistically recombine the advantageous features from these different knowledge units, thereby generating a new generation of candidate instructions (offspring instructions). The specific operation steps are steps S102 - S103. Selecting the candidate instruction with the highest score in the candidate instruction sets belongs to the prior art and will not be elaborated here in detail.
[0056] The automated prompt engineering iterative proposal optimization method proposed by the present invention can play an important role in different application scenarios. By quantifying the collaborative contribution potential of the parent instructions that showed advantages in guiding the model to handle specific sub-aspects (for example, "contract element identification" in legal texts, "key indicator extraction and risk assessment" in medical reports, "context coherence maintenance" in conversations) generated in the previous iteration, and using the synergistic recombination strategy to fuse the advantageous features of these different parent instructions, a series of candidate instructions with better overall performance are automatically generated. Through the continuous iterative optimization of the APE framework, instructions that can guide the model to achieve higher accuracy, stronger logic, better structure, or better meet the requirements of specific fields in specific natural language tasks can finally be screened out.
[0057] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for optimizing collaborative interaction strategies based on multi - agents, characterized in that, The method includes: During the process of the APE generating instructions through iterative optimization, obtaining all high-quality parent instructions in any iteration round; For any high-quality parent instruction, identifying and counting the text content of any high-quality parent instruction to obtain the number of explicit structural units, and obtaining the structured instruction clarity according to the number of explicit structural units; obtaining the core instruction semantic intensity according to the occurrence frequency of specific parts of speech in the text content of any high-quality parent instruction, and combining the structured instruction clarity and the core instruction semantic intensity to obtain the collaborative contribution potential index of any high-quality parent instruction; Obtaining the collaborative contribution potential index of each high-quality parent instruction, and successively selecting a core contribution parent instruction from all high-quality parent instructions according to the collaborative contribution potential index of each high-quality parent instruction, extracting the dominant advantageous feature segment of each selected core contribution parent instruction, and using a proposal model to perform text recombination on each of the dominant advantageous feature segments to obtain a set of child instructions; Selecting the high-quality parent instruction for the next iteration round in the set of child instructions, and repeating the method for obtaining the set of child instructions until the iteration round requirement is met to obtain a set of candidate instruction sets, and selecting the candidate instruction with the highest quality evaluation result in the set of candidate instruction sets.
2. The method for optimizing a collaborative interaction strategy based on multi-agent according to claim 1, wherein The obtaining the structured instruction clarity according to the number of explicit structural units includes: Taking the sum of the number of explicit structural units and the constant 1 as the independent variable of the natural logarithm function to obtain the structured instruction clarity.
3. The collaborative interaction strategy optimization method based on multi-agent according to claim 1, characterized in that The obtaining the core instruction semantic intensity according to the occurrence frequency of specific parts of speech in the text content of any high-quality parent instruction includes: Respectively counting the occurrence frequency of verbs, the occurrence frequency of demonstrative pronouns, and the occurrence frequency of deontic auxiliary verbs in the text content of any high-quality parent instruction, calculating the total frequency of the occurrence frequency of verbs and the occurrence frequency of deontic auxiliary verbs, calculating the product of a preset non-negative penalty coefficient and the occurrence frequency of demonstrative pronouns, and taking the difference between the total frequency and the product as the numerator and the instruction length of any high-quality parent instruction as the denominator to obtain the core instruction semantic intensity.
4. The collaborative interaction strategy optimization method based on multi-agent according to claim 1, wherein The combining the structured instruction clarity and the core instruction semantic intensity to obtain the collaborative contribution potential index of any high-quality parent instruction includes: Respectively performing normalization processing on the structured instruction clarity and the core instruction semantic intensity to obtain the normalized value of the structured instruction clarity and the normalized value of the core instruction semantic intensity, and recording the sum of the normalized value of the structured instruction clarity and the normalized value of the core instruction semantic intensity as the instruction feature value of any high-quality parent instruction; Obtaining the instruction feature values corresponding to all high-quality parent instructions, taking the mean of all instruction feature values as the bias term, and taking the difference between the instruction feature value of any high-quality parent instruction and the bias term as the independent variable of the Sigmoid activation function to obtain the collaborative contribution potential index of any high-quality parent instruction.
5. The method for optimizing a collaborative interaction strategy based on multi-agent according to claim 1, wherein The successively selecting a core contribution parent instruction from all high-quality parent instructions according to the collaborative contribution potential index of each high-quality parent instruction includes: According to the collaborative contribution potential index of each high-quality parent instruction, a core contribution parent instruction is sequentially selected from all high-quality parent instructions using the roulette wheel selection method.
6. The collaborative interaction strategy optimization method based on multi-agent according to claim 3, characterized in that The extraction of the dominant advantage feature segments of the core contribution parent instruction selected each time includes: According to the structured instruction clarity and core instruction semantic intensity of all high-quality parent instructions, the mean value of structured instruction clarity and the mean value of core instruction semantic intensity are calculated respectively. For the core contribution parent instruction selected any time, if the core instruction semantic intensity of the core contribution parent instruction selected any time is greater than or equal to the mean value of the core instruction semantic intensity, then the smallest text unit where the verb and deontic auxiliary verb are located in the text content of the core contribution parent instruction selected any time is used as the dominant advantage feature segment; If the structured instruction clarity of the core contribution parent instruction selected any time is greater than or equal to the mean value of the structured instruction clarity, then all explicit structure units in the text content of the core contribution parent instruction selected any time are used as the dominant advantage feature segment.
7. The collaborative interaction strategy optimization method based on multi-agent according to claim 6, wherein The use of the proposal model to perform text recombination on each of the dominant advantage feature segments to obtain a set of child instructions includes: Taking the dominant advantage feature segment of the core contribution parent instruction selected any time as the target segment, using the proposal model to perform text recombination on the target segment to obtain a candidate child instruction, and performing a legality and validity check on the candidate child instruction. If the check passes, then the candidate child instruction is used as the child instruction. If the check fails, then the dominant advantage feature segment of the core contribution parent instruction selected in the next selection any time is used as the target segment, and the above check method is repeated until the number of child instructions meets the quantity requirement to obtain a set of child instructions.
8. The method for optimizing the collaborative interaction strategy based on multi-agent according to claim 7, characterized in that The check of the legality and validity of the candidate child instruction includes: If the candidate child instruction is empty or the instruction length is not within the length range or contains prohibited words, it is determined that the check fails; if the candidate child instruction is not empty, the instruction length is within the length range and does not contain prohibited words, it is determined that the check passes.
Citation Information
Patent Citations
Strategy processing method and device, equipment, storage medium and program product
CN116943205A
Lead compound discovery method based on large language model
CN120148609A
Methods and apparatus to find optimization opportunities in machine-readable instructions
US20210110308A1