Strategy evaluation method
By generating and optimizing strategy prompt word templates, and utilizing a strategy evaluation agent to evaluate strategy text in wargame simulations from multiple dimensions, the problems of time-consuming, labor-intensive, and resource-wasting methods in existing technologies are solved, achieving efficient and reasonable strategy evaluation results.
Patent Information
- Application Number
- CN202511675864.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-11-17
AI Technical Summary
In existing technologies, strategy evaluation in wargaming relies on human expert referees, which is time-consuming, labor-intensive, and has inconsistent standards. Computer evaluation results lack rationality and have low efficiency in utilizing computer resources.
By generating policy prompt word templates, the policy evaluation agent evaluates the policy text to be evaluated from multiple dimensions, including content accuracy, completeness, logic and compliance. The prompt word templates are optimized until they meet preset conditions. Calibration policy prompt words are generated and evaluated in multiple rounds to obtain consistent and effective policy evaluation results.
It improves the rationality of strategy evaluation and the efficiency of computer processing, reduces the waste of computer resources, enhances the rationality and consistency of evaluation results, and reduces evaluation bias.
Smart Images

Figure CN121118884B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a strategy evaluation method. Background Technology
[0002] Wargaming relies on human participants and expert judges to evaluate strategy texts. This evaluation process is time-consuming, labor-intensive, and involves inconsistent standards that are highly dependent on experience. Evaluating strategy texts using large language models often yields less convincing results compared to expert judges, requiring excessive computer resources for multiple evaluations and consequently reducing computer processing efficiency. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a strategy evaluation method.
[0004] One aspect of this invention provides a strategy evaluation method, comprising: generating strategy prompts based on a strategy text to be evaluated and a strategy prompt template, wherein the strategy prompt template is used to evaluate the strategy text to be evaluated based on at least two of the following: accuracy of strategy text content, completeness of strategy text content, compliance of strategy, logicality of strategy text content, or performance of strategy; evaluating the strategy prompts using a strategy evaluation agent to obtain a strategy evaluation result; wherein the strategy prompt template is a prompt template for a strategy to be calibrated determined under the condition that the prompt validity evaluation result and the prompt consistency evaluation result meet preset conditions, the prompt validity evaluation result is determined based on multiple rounds of respective calibration strategy evaluation results, the prompt consistency evaluation result is determined based on the expert strategy evaluation result of the calibration strategy text and multiple rounds of respective calibration strategy evaluation results, the multiple rounds of respective calibration strategy evaluation results are obtained by evaluating the calibration strategy prompts multiple times using a strategy evaluation model, the calibration strategy prompts are generated by a strategy evaluation agent based on the calibration strategy text and the prompt template for the strategy to be calibrated, and the prompt validity evaluation result and the prompt consistency evaluation result are used to evaluate the calibration strategy prompts.
[0005] According to an embodiment of the present invention, the prompt word validity assessment result and the prompt word consistency assessment result are used to evaluate the calibration strategy prompt words, including: the prompt word validity assessment result is used to evaluate the validity of the calibration strategy prompt words, the prompt word consistency assessment result is used to evaluate the consistency of the calibration strategy prompt words, and the expert strategy assessment result is obtained by experts evaluating the calibration strategy text.
[0006] According to an embodiment of the present invention, the strategy prompt word template is generated by repeatedly performing the following operations until the prompt word effectiveness evaluation result and the prompt word consistency evaluation result meet preset conditions: if the (n-1)th prompt word effectiveness evaluation result and the (n-1)th prompt word consistency evaluation result do not meet preset conditions, the (n-1)th prompt word template to be calibrated is optimized to obtain the nth prompt word template to be calibrated, where n is an integer greater than or equal to 1; based on the calibration strategy text and the nth prompt word template to be calibrated, the nth calibration strategy prompt word is generated; the nth calibration strategy prompt word is evaluated in multiple rounds using a strategy evaluation agent to determine the nth calibration strategy evaluation result for each round; based on the nth calibration strategy evaluation results for each round, the nth prompt word effectiveness evaluation result is determined using the strategy evaluation agent; based on the nth expert strategy evaluation result of the calibration strategy text and the nth calibration strategy evaluation results for each round, the nth prompt word consistency evaluation result is determined using the strategy evaluation agent; the nth prompt word template to be calibrated, determined when the nth prompt word effectiveness evaluation result and the nth prompt word consistency evaluation result meet preset conditions, is used as the strategy prompt template.
[0007] According to an embodiment of the present invention, a strategy evaluation agent is used to evaluate strategy prompts to obtain strategy evaluation results, including: using the strategy evaluation agent to evaluate the strategy text to be evaluated based on strategy prompts related to the accuracy of the strategy text content, to obtain a first strategy evaluation result, wherein the first strategy evaluation result characterizes whether the strategy text to be evaluated satisfies the accuracy of the strategy text content; using the strategy evaluation agent to evaluate text blocks of the strategy text to be evaluated based on strategy prompts related to the completeness and logicality of the strategy text content, to obtain a second strategy evaluation result and a third strategy evaluation result, wherein the text blocks are obtained by segmenting the strategy text to be evaluated, and the second strategy evaluation result characterizes whether the strategy text to be evaluated satisfies the completeness of the strategy text content, and the third ... whether the strategy text to be evaluated satisfies whether the strategy text The strategy evaluation result indicates whether the strategy text to be evaluated meets the logical requirements of the strategy text content; the strategy evaluation agent evaluates the strategy text to be evaluated based on strategy prompts and strategy rules related to strategy compliance, and obtains the fourth strategy evaluation result, which indicates whether the strategy text to be evaluated meets the policy compliance requirements; the strategy evaluation agent evaluates the strategy text to be evaluated based on strategy prompts and preset behavior standards related to policy performance, and obtains the fifth strategy evaluation result, which indicates whether the strategy text to be evaluated meets the policy performance requirements; the strategy evaluation result is obtained based on at least two of the first, second, third, fourth, and fifth strategy evaluation results.
[0008] According to an embodiment of the present invention, a strategy evaluation agent is used to evaluate a strategy text to be evaluated based on strategy prompts related to the accuracy of the strategy text content, and a first strategy evaluation result is obtained. The process includes: using the strategy evaluation agent to extract text from the strategy text to be evaluated to obtain key text, wherein the key text includes text content related to the target object of the strategy in the strategy text to be evaluated; using the strategy evaluation agent to obtain reference text corresponding to the key text from a database; and using the strategy evaluation agent to obtain the first strategy evaluation result based on the key text and the reference text.
[0009] According to embodiments of the present invention, a strategy evaluation agent is used to evaluate text blocks of the strategy text to be evaluated based on strategy cue words related to the completeness and logicality of the strategy text content, to obtain a second strategy evaluation result and a third strategy evaluation result. This includes: using the strategy evaluation agent to evaluate the relevance of multiple text blocks of the strategy text to be evaluated, obtaining the associated text structure of each text block, where the associated text structure represents a text structure with the same strategy category as the text block; using the strategy evaluation agent to obtain a second strategy evaluation result based on the associated text structure of each text block; using the strategy evaluation agent to evaluate the context of multiple text blocks of the strategy text to be evaluated, obtaining the preceding and following text blocks of each text block, where the preceding text block represents the preconditions of the text block, and the text block represents the preconditions of the following text block; and using the strategy evaluation agent to obtain a third strategy evaluation result based on the preceding and following text blocks of each text block.
[0010] According to an embodiment of the present invention, a second policy evaluation result is obtained by using a policy evaluation agent based on the associated text architecture of each of multiple text blocks, including: using the policy evaluation agent to aggregate the associated text architecture of each of the multiple text blocks to generate a expected text architecture; and using the policy evaluation agent to obtain the second policy evaluation result based on the degree of difference between the expected text architecture and the text architecture of the policy text to be evaluated.
[0011] According to an embodiment of the present invention, a third policy evaluation result is obtained by using a policy evaluation agent based on the preceding and following text blocks of each of a plurality of text blocks, including: for any text block among the plurality of text blocks, using the policy evaluation agent to obtain a preceding evaluation result based on the preceding text block and the adjacent preceding text block adjacent to the text block; using the policy evaluation agent to obtain a following evaluation result based on the adjacent following text block and the following text block adjacent to the text block; and using the policy evaluation agent to obtain a third policy evaluation result based on the preceding evaluation result and the following evaluation result corresponding to each of the plurality of text blocks.
[0012] According to an embodiment of the present invention, a strategy evaluation agent is used to evaluate the strategy text to be evaluated based on strategy prompts and strategy rules related to strategy compliance, and to obtain a fourth strategy evaluation result. The evaluation process includes: using the strategy evaluation agent to extract text from the strategy text to be evaluated to obtain behavioral text, which includes text related to the behavior of executing the strategy in the strategy text to be evaluated; and using the strategy evaluation agent to obtain the fourth strategy evaluation result based on the behavioral text and the rule behavior corresponding to the strategy rules.
[0013] According to an embodiment of the present invention, a strategy evaluation agent is used to evaluate the strategy text to be evaluated based on strategy prompts related to strategy performance and preset behavior standards to obtain a fifth strategy evaluation result. The evaluation results include: using the strategy evaluation agent to predict the content of the behavior text based on preset behavior standards to obtain the behavior result of at least one strategy text; using the strategy evaluation agent to determine the behavior value corresponding to each of the at least one behavior result based on preset behavior standards; and obtaining the fifth strategy evaluation result based on the at least one behavior result and the behavior value corresponding to each of the at least one behavior result.
[0014] Another aspect of the present invention provides an electronic device, including: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method described above.
[0015] Another aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the method described above.
[0016] Another aspect of the present invention provides a computer program product including computer-executable instructions, which, when executed, are used to implement the method described above.
[0017] According to embodiments of the present invention, since strategy prompts are generated using strategy prompt templates, different dimensions of evaluation directions can be provided for the strategy evaluation agent, improving the rationality of the strategy evaluation results. Furthermore, the prompt validity evaluation results and prompt consistency evaluation results of the strategy prompt templates meet preset conditions, making the strategy evaluation results consistent and valid with the evaluation results given by expert judges, reducing the evaluation bias of the strategy evaluation agent, avoiding the need for the strategy evaluation agent to conduct multiple evaluations, thereby reducing computer resource consumption while improving computer processing efficiency. Attached Figure Description
[0018] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which the accompanying drawings are shown.
[0019] Figure 1 The diagram illustrates application scenarios of the strategy evaluation method, apparatus, device, medium, and program product according to embodiments of this application.
[0020] Figure 2 A flowchart of a strategy evaluation method according to an embodiment of this application is shown.
[0021] Figure 3 A schematic diagram of a strategy evaluation method according to an embodiment of this application is shown.
[0022] Figure 4 A structural block diagram of a strategy evaluation apparatus according to an embodiment of this application is shown.
[0023] Figure 5 A block diagram of an electronic device suitable for implementing a policy evaluation method according to an embodiment of this application is shown. Detailed Implementation
[0024] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0027] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0028] In the embodiments of this invention, the collection, updating, analysis, processing, use, transmission, provision, invention, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to maintain the security of user personal information and network security.
[0029] In the embodiments of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0030] Figure 1 The diagram illustrates application scenarios for the strategy evaluation method, apparatus, device, medium, and program product according to embodiments of this application. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of the present invention, in order to help those skilled in the art understand the technical content of the present invention, but do not mean that embodiments of the present invention cannot be used in other devices, systems, environments or scenarios.
[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links.
[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] It should be noted that the strategy evaluation method provided in this embodiment of the invention can generally be executed by server 105. Correspondingly, the strategy evaluation system provided in this embodiment of the invention can generally be set up in server 105. The strategy evaluation method provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the strategy evaluation system provided in this embodiment of the invention can also be set up in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the strategy evaluation method provided in this embodiment of the invention can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or it can be executed by other terminal devices that are different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the strategy evaluation system provided in this embodiment of the invention may also be set in the first terminal device 101, the second terminal device 102 or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103.
[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0037] Figure 2 A flowchart illustrating a strategy evaluation method according to an embodiment of the present invention is shown schematically.
[0038] like Figure 2 As shown, the strategy evaluation method includes operations S210 to S220.
[0039] In operation S210, a policy prompt is generated based on the policy text to be evaluated and the policy prompt template.
[0040] The strategy prompt template is used to evaluate the strategy text to be evaluated based on at least two of the following: accuracy of strategy text content, completeness of strategy text content, compliance of strategy, logic of strategy text content, or performance of strategy.
[0041] In operation S220, the policy evaluation agent is used to evaluate the policy prompts and obtain the policy evaluation results.
[0042] The strategy prompt template can be a template for the calibration strategy, determined under the condition that the prompt validity evaluation results and the prompt consistency evaluation results meet preset conditions. The prompt validity evaluation results can be determined based on multiple rounds of individual calibration strategy evaluation results, while the prompt consistency evaluation results are determined based on the expert strategy evaluation results of the calibration strategy text and multiple rounds of individual calibration strategy evaluation results. The multiple rounds of individual calibration strategy evaluation results can be obtained by performing multiple rounds of evaluation on the calibration strategy prompts using a strategy evaluation model. The calibration strategy prompts can be generated by a strategy evaluation agent based on the calibration strategy text and the prompt template for the strategy to be calibrated. The prompt validity evaluation results and the prompt consistency evaluation results can be used to evaluate the calibration strategy prompts.
[0043] In some embodiments, the results of the prompt word effectiveness assessment can be used to evaluate the effectiveness of the calibration strategy prompt words. The results of the prompt word consistency assessment can be used to evaluate the consistency of the calibration strategy prompt words. The expert strategy assessment results can be obtained by experts evaluating the calibration strategy text.
[0044] The policy evaluation agent can analyze input information based on a policy evaluation model (e.g., a large language model) to obtain output results. For example, given the input policy text to be evaluated and policy prompt word templates, the policy evaluation agent uses a large language model to analyze the policy text to obtain the policy evaluation result.
[0045] The policy text to be evaluated can be used in wargaming to simulate and extrapolate real-world scenarios, providing the target object with behavioral strategies that need to be implemented in that scenario. For example, it could describe how the target object should act within a given timeframe and what objective tasks it needs to accomplish. In wargaming, a policy evaluation agent can be used to evaluate this policy text and extrapolate the effects that the behavioral strategies indicated by the policy text would achieve in a real-world scenario.
[0046] The strategy cue word template can be cue words from multiple dimensions, used to assist the large language model in analyzing the text of the strategy to be evaluated.
[0047] Based on the strategy text to be evaluated and the strategy prompt word template, strategy prompt words are generated. These prompt words can be generated in different dimensions to match the strategy text to be evaluated.
[0048] The accuracy of strategy text content can be used to assess whether the content in the strategy text being evaluated conforms to objective facts, such as whether a vehicle has wings.
[0049] The completeness of the strategy text can be used to assess whether the content of the strategy text being evaluated is complete. For example, whether the behavioral strategy for the target object to perform the target task has omitted any behavioral steps.
[0050] Policy compliance can be used to assess whether the content of the policy text being evaluated conforms to the rules. For example, whether a target object's permissions allow it to perform a certain action.
[0051] The logical consistency of strategy text content can be used to assess whether the content of the strategy text being evaluated is logical. For example, whether each step of the target object's execution of the target task has a reasonable sequential relationship.
[0052] Policy performance can be used to evaluate the benefits that the policy can bring after execution in the policy text to be evaluated.
[0053] The calibration strategy prompt template can be an initially set, unadjusted template. The calibration strategy evaluation result can be obtained by evaluating the calibration strategy text based on the calibration strategy prompt template. The calibration strategy text can be preset.
[0054] The validity assessment results of the prompt words can be used to evaluate whether the calibration strategy assessment results are reasonable. For example, whether the assessment results of multiple rounds of individual calibration strategies are stable.
[0055] The consistency assessment results of the prompt words can be used to evaluate whether the calibration strategy assessment results are consistent with the home strategy assessment results. The expert strategy assessment results can be obtained by experts evaluating the calibration strategy text.
[0056] According to embodiments of the present invention, since strategy prompts are generated using strategy prompt templates, different dimensions of evaluation directions can be provided for the strategy evaluation agent, improving the rationality of the strategy evaluation results. Furthermore, the prompt validity evaluation results and prompt consistency evaluation results of the strategy prompt templates meet preset conditions, making the strategy evaluation results consistent and valid with the evaluation results given by expert judges, reducing the evaluation bias of the strategy evaluation agent, avoiding the need for the strategy evaluation agent to conduct multiple evaluations, thereby reducing computer resource consumption while improving computer processing efficiency.
[0057] In some embodiments, the strategy prompt template can be generated by repeatedly performing the following operations until the prompt effectiveness evaluation result and the prompt consistency evaluation result meet preset conditions: if the (n-1)th prompt effectiveness evaluation result and the (n-1)th prompt consistency evaluation result do not meet preset conditions, the (n-1)th prompt template to be calibrated is optimized to obtain the nth prompt template to be calibrated, where n is an integer greater than or equal to 1. Based on the calibration strategy text and the nth prompt template to be calibrated, the nth calibration strategy prompt is generated. The nth calibration strategy prompt is evaluated multiple times using a strategy evaluation agent to determine the nth calibration strategy evaluation result for each round. Based on the nth calibration strategy evaluation results for each round, the strategy evaluation agent determines the nth prompt effectiveness evaluation result. Based on the nth expert strategy evaluation result of the calibration strategy text and the nth calibration strategy evaluation results for each round, the strategy evaluation agent determines the nth prompt consistency evaluation result. The nth prompt template to be calibrated, determined when the nth prompt effectiveness evaluation result and the nth prompt consistency evaluation result meet preset conditions, is used as the strategy prompt template.
[0058] In some embodiments, optimizing the (n-1)th calibration strategy prompt word template to obtain the nth calibration strategy prompt word template may include: optimizing at least one of the following: configuring a preset example of the (n-1)th calibration strategy prompt word template, text indicating the reason for the strategy evaluation, or text indicating the provision of evidence, to obtain the nth calibration strategy prompt word template. The preset example includes sample strategy text and the expected strategy evaluation result corresponding to the sample strategy text.
[0059] The strategy evaluation agent optimizes the (n-1)th policy prompt template. This can be done by randomly adjusting the prompt template or by adjusting it according to requirements. For example, it can adjust the evaluation elements related to the (n-1)th policy prompt template.
[0060] The preset example for the n-1th policy prompt template can be the initially set reference evaluation method. The text indicating the reason for the policy evaluation can be the evaluation reason for the policy evaluation result of the sample policy text provided by the policy evaluation agent. For example, why this evaluation was made. The text indicating the evidence provided can be the basis for the policy evaluation agent's evaluation during the evaluation process. For example, what data was used to evaluate the sample policy text.
[0061] Using a strategy evaluation model to evaluate the nth calibration strategy prompt word in multiple rounds can be based on the nth calibration strategy prompt word, and the calibration strategy text can be evaluated multiple times using a strategy evaluation model (e.g., a large language model) to obtain the nth calibration strategy evaluation results for each round.
[0062] The effectiveness evaluation result of the nth prompt word is determined based on the variation range of the evaluation results of the nth calibration strategy in multiple rounds. For example, the variance of the evaluation results of the nth calibration strategy in multiple rounds is determined. The effectiveness evaluation result of the nth prompt word is obtained based on this variance and a preset variance threshold.
[0063] Determine the correlation value between the evaluation result of the nth expert strategy and the evaluation results of the nth calibration strategy in each round. Based on this correlation value and a preset correlation value, obtain the consistency evaluation result of the nth prompt word.
[0064] If the validity assessment result and the consistency assessment result of the nth prompt word meet the preset conditions, then the nth prompt word template to be calibrated is used as the policy prompt template. If not, further optimization of the prompt word template to be calibrated is performed until the policy prompt template is obtained.
[0065] According to an embodiment of the present invention, a strategy prompt template is obtained by iterating the template of the strategy prompt word to be calibrated multiple times, which makes the strategy evaluation model more stable when evaluating the calibration strategy text, avoiding evaluation instability, and at the same time, it can be consistent with the strategy evaluation results given by experts, avoiding evaluation bias.
[0066] In some embodiments, evaluating policy prompts using a policy evaluation agent to obtain policy evaluation results may include: evaluating the policy text to be evaluated using policy prompts related to the accuracy of policy text content, resulting in a first policy evaluation result. The first policy evaluation result can characterize whether the policy text to be evaluated meets the requirements for policy text content accuracy. Evaluating text blocks of the policy text to be evaluated using policy prompts related to the completeness and logicality of policy text content, resulting in a second and a third policy evaluation result. Text blocks may be obtained by segmenting the policy text to be evaluated. The second policy evaluation result can characterize whether the policy text to be evaluated meets the requirements for policy text content completeness. The third policy evaluation result can characterize whether the policy text to be evaluated meets the requirements for policy text logicality. Evaluating the policy text to be evaluated using policy prompts and policy rules related to policy compliance, resulting in a fourth policy evaluation result. The fourth policy evaluation result can characterize whether the policy text to be evaluated meets the requirements for policy compliance. Evaluating the policy text to be evaluated using policy prompts and preset behavioral standards related to policy performance, resulting in a fifth policy evaluation result. The fifth strategy evaluation result characterizes whether the strategy text being evaluated meets the strategy performance requirements. The strategy evaluation result is obtained based on at least two of the first, second, third, fourth, and fifth strategy evaluation results.
[0067] Policy cue words related to the accuracy of policy text content can assist the policy evaluation model in evaluating the policy text in at least one of the accuracy or authenticity dimensions. The database stores objective facts related to different target objects. The policy text is evaluated based on these objective facts to determine whether it constitutes objective fact, thereby obtaining the first policy evaluation result.
[0068] Policy cue words related to the completeness of policy text content can help policy evaluation models evaluate the policy text to be evaluated in terms of the completeness dimension.
[0069] By segmenting the text to be evaluated into multiple text blocks, which can be words or sentences, different text blocks can be evaluated separately to determine whether the text content of the strategy to be evaluated is complete, whether a certain word or sentence is missing, etc., and thus obtain the evaluation result of the second strategy.
[0070] Policy cue words related to the logical structure of policy text content can help policy evaluation models evaluate policy texts based on their logical structure.
[0071] By evaluating different text blocks separately, determining the relationships between multiple text blocks, and whether the connections between text blocks are logical, the evaluation results of the third strategy can be obtained.
[0072] Policy compliance-related policy prompts can assist policy evaluation models in assessing the compliance dimension of the policy text to be evaluated. Policy rules can be pre-defined rules for different target objects and the corresponding behaviors of the target objects in implementing policies, used to restrict the behavior of the target objects.
[0073] By determining whether the policy text to be evaluated meets the policy rules, the fourth policy evaluation result can be obtained.
[0074] Policy cue words related to policy performance can assist the policy evaluation model in evaluating the value of the policy text to be evaluated, and in assessing the benefits or risks brought about by the results after the policy is executed, i.e., policy performance results. Based on these policy performance results, a matching fifth policy evaluation result can be obtained.
[0075] When evaluating policy prompts using a policy evaluation model configured by a policy evaluation agent, at least two of the above evaluation operations can be performed in parallel. After evaluation in different dimensions, a policy evaluation result can be obtained based on at least two of the first, second, third, fourth, and fifth policy evaluation results. Specifically, the policy evaluation result can be obtained by weighted summation of at least two of the first, second, third, fourth, and fifth policy evaluation results based on the weights of the accuracy, completeness, compliance, logic, and performance of the policy text content in the policy prompt.
[0076] According to embodiments of the present invention, by conducting strategy evaluation operations based on five dimensions—accuracy of strategy text content, completeness of strategy text content, compliance of strategy, logic of strategy text content, and performance of strategy—the accuracy and rationality of the evaluation can be improved.
[0077] In some embodiments, the strategy evaluation agent evaluates the strategy text to be evaluated based on strategy cue words related to the accuracy of the strategy text content to obtain a first strategy evaluation result. This may include: using the strategy evaluation agent to extract text from the strategy text to be evaluated to obtain key text, the key text including text content related to the target object of the strategy in the strategy text to be evaluated; using the strategy evaluation agent to obtain reference text corresponding to the key text from a database; and using the strategy evaluation agent to obtain the first strategy evaluation result based on the key text and the reference text.
[0078] Text extraction of the strategy text to be evaluated can be performed by using a strategy evaluation model to extract text content related to the target object from the strategy text to be evaluated, i.e., key text content, such as at least one of the target object's type, structural information of the target object, or performance configuration information of the target object.
[0079] Based on the type of the target object, retrieve the corresponding reference text content from the database, such as reference structure information and reference performance configuration information related to the target object.
[0080] Comparing key text content with reference text content can determine the similarity between the key text content and the reference text content, and the resulting text content similarity value can be used as the text content comparison result.
[0081] Furthermore, based on the numerical range of the text content similarity value in the text content comparison results, the evaluation value corresponding to the text content similarity value is determined, which is the first strategy evaluation result.
[0082] According to embodiments of the present invention, by utilizing a database to evaluate the strategy text to be evaluated, the reference text content can be accurately extracted from the database, the accuracy of the strategy text to be evaluated can be assessed, and the accuracy of the evaluation of the authenticity of the strategy text to be evaluated can be improved.
[0083] In some embodiments, the strategy evaluation agent evaluates text blocks of the strategy text to be evaluated based on strategy cue words related to the completeness and logicality of the strategy text content, obtaining a second strategy evaluation result and a third strategy evaluation result. This may include: using the strategy evaluation agent to evaluate the relevance of multiple text blocks of the strategy text to be evaluated, obtaining the associated text structure of each text block. The associated text structure can represent the text structure with the same strategy category as the text block. The strategy evaluation agent obtains the second strategy evaluation result based on the associated text structure of each text block. The strategy evaluation agent performs context evaluation on multiple text blocks of the strategy text to be evaluated, obtaining the preceding and following text blocks of each text block. The preceding text block can represent the preconditions of the text block. The text block can represent the preconditions of the following text block. The strategy evaluation agent obtains the third strategy evaluation result based on the preceding and following text blocks of each text block.
[0084] The associated text structure of a text block can be at least one of the following: content related to the text content of the text block, content that needs to be explained based on the text block, or content that may appear when the text content exists. For example, if the text block represents "intermediate stage of strategy A", then the associated text structure of the text block can be "intermediate stage execution content of strategy A", "initial stage of strategy A", etc.
[0085] The preceding and following text blocks of a text block can be text content with a logical order relative to the text block content; that is, the preceding text block must precede the following text block. For example, the text block "Initial Stage" is the preceding text block of the text block "Middle Stage," and the text block "Final Stage" is the following text block of the text block "Middle Stage." The order of the three cannot be changed logically.
[0086] The policy evaluation agent analyzes the preceding and following text blocks of each of the multiple text blocks to determine whether there are logical problems in the arrangement of the multiple text blocks, and obtains the third policy evaluation result.
[0087] According to embodiments of the present invention, by performing relevance prediction on multiple text blocks of the strategy text to be evaluated, it is possible to determine whether the content of the strategy text to be evaluated is complete and logical, thereby further improving the accuracy of the evaluation.
[0088] In some embodiments, obtaining a second policy evaluation result using a policy evaluation agent based on the associated text structures of multiple text blocks may include: aggregating the associated text structures of the multiple text blocks using the policy evaluation agent to generate a desired text structure; and obtaining the second policy evaluation result based on the degree of difference between the desired text structure and the text structure of the policy text to be evaluated using the policy evaluation agent.
[0089] Aggregate the associated text structures of multiple text blocks, remove duplicate associated text structures, and summarize to obtain the complete expected text structure under the strategy category, that is, all text structures that the strategy text of the strategy category should contain.
[0090] The expected text structure and the text structure of the strategy text to be evaluated are compared to determine the amount of text content missing in the current strategy text to be evaluated. This number is used as the degree of difference between the text structures of the expected text structure and the strategy text to be evaluated.
[0091] Furthermore, based on the number of missing text contents, an evaluation value corresponding to that number is determined. Based on the evaluation value, the evaluation result of the second strategy is obtained, which can be used to assess whether the text of the strategy to be evaluated is complete.
[0092] According to embodiments of the present invention, by utilizing the associated text structure of multiple text blocks, it is possible to determine whether the text content of the strategy to be evaluated is complete, thereby improving the evaluation accuracy.
[0093] In some embodiments, the policy evaluation agent obtains a third policy evaluation result based on the preceding and following text blocks of each of the multiple text blocks. This may include: for any text block among the multiple text blocks, the policy evaluation agent obtains a preceding evaluation result based on the preceding text block and the adjacent preceding text block. The policy evaluation agent obtains a following evaluation result based on the adjacent following text block and the adjacent following text block. The policy evaluation agent obtains a third policy evaluation result based on the preceding and following evaluation results corresponding to each of the multiple text blocks.
[0094] Obtain the adjacent preceding and following text blocks of the text block, that is, the context content of the text content in the policy text to be evaluated.
[0095] Comparing adjacent preceding text blocks can be done by calculating the similarity distance between them to obtain a preceding text similarity value. Similarly, the following text blocks can be compared to determine their following text similarity values. The preceding and following text similarity values are then averaged to obtain the third measurement and evaluation result.
[0096] According to embodiments of the present invention, by comparing adjacent preceding text blocks with each other and adjacent following text blocks with each other, the logical relationship between text blocks can be clearly reflected, thereby improving the accuracy of evaluation.
[0097] In some embodiments, the policy evaluation agent evaluates the policy text to be evaluated based on policy prompts and policy rules related to policy compliance, obtaining a fourth policy evaluation result. This may include: using the policy evaluation agent to extract text from the policy text to obtain behavioral text, which includes text related to the actions of executing the policy in the policy text to be evaluated; and using the policy evaluation agent to obtain the fourth policy evaluation result based on the behavioral text and the rule actions corresponding to the policy rules.
[0098] Text extraction is performed on the strategy text to be evaluated to obtain the actions that the target object needs to take to execute the strategy, i.e., the action text. For example, the target object goes to region A at time A, goes to region B at time B, and performs task A in region A, etc.
[0099] Obtaining the rule behavior corresponding to the policy rule can be a constraint on the behavior, which executes a certain behavior or disallows the execution of a certain behavior under certain preset conditions. For example, the rule behavior corresponding to "travel to region A at time A" could be "travel to region A is not allowed at time A" or "region A is only open at time B", etc.
[0100] Comparing behavioral texts with rule behaviors can determine whether the behavioral texts meet the constraints specified by the rule behaviors. The number of behavioral texts that do not conform to the rules is then counted. Furthermore, the evaluation value corresponding to the number of behavioral texts that do not conform to the rules is determined, which is the fourth strategy evaluation result.
[0101] According to embodiments of the present invention, by utilizing policy rules to evaluate the policy text to be evaluated, policy evaluation can be flexibly adapted to different environments, thereby improving the versatility of the policy evaluation agent.
[0102] In some embodiments, the strategy evaluation agent evaluates the policy text to be evaluated based on policy cue words related to policy performance and preset behavior standards to obtain a fifth strategy evaluation result. This may include: using the strategy evaluation agent to predict the behavior text based on preset behavior standards to obtain at least one behavior result for the policy text; using the strategy evaluation agent to determine the behavior value corresponding to each of the at least one behavior result based on preset behavior standards; and obtaining the fifth strategy evaluation result based on the at least one behavior result and the behavior value corresponding to each of the at least one behavior result.
[0103] Preset behavioral standards can be the correspondence between preset behavioral results and preset benefits.
[0104] Based on preset behavioral standards, the content of behavioral text is predicted to determine future possibilities. For example, the risks of implementing the strategy text to be evaluated, whether the current strategy is optimal, etc., are all considered as the behavioral value of the behavioral outcome.
[0105] Furthermore, the assessment value corresponding to the behavioral value is determined, which is the fifth assessment result.
[0106] According to embodiments of the present invention, by using preset behavioral standards to evaluate the text of the strategy to be evaluated, the potential value of different action path strategies can be assessed, thereby improving the flexibility and rationality of the evaluation.
[0107] Figure 3 A schematic diagram of a strategy evaluation method according to an embodiment of this application is shown.
[0108] like Figure 3 As shown, the policy evaluation agent evaluates the policy text to be evaluated based on at least two of the following: the accuracy of the policy text content, the completeness of the policy text content, the compliance of the policy, the logicality of the policy text content, or the performance of the policy, in order to obtain the policy evaluation result.
[0109] Figure 4 A structural block diagram of a strategy evaluation apparatus according to an embodiment of this application is shown.
[0110] like Figure 4 As shown, the strategy evaluation device 400 includes a generation module 410 and an evaluation module 420.
[0111] The generation module 410 is used to generate strategy prompts based on the strategy text to be evaluated and the strategy prompt template. The strategy prompt template is used to evaluate the strategy text to be evaluated based on at least two of the following: the accuracy of the strategy text content, the completeness of the strategy text content, the compliance of the strategy, the logicality of the strategy text content, or the performance of the strategy.
[0112] The evaluation module 420 is used to evaluate the policy prompts using a policy evaluation agent to obtain policy evaluation results. Specifically, the policy prompt template is a template for the policy prompts to be calibrated, determined under preset conditions where the prompt validity evaluation results and prompt consistency evaluation results meet the preset conditions. The prompt validity evaluation results are determined based on multiple rounds of individual calibration policy evaluation results. The prompt consistency evaluation results are determined based on the expert policy evaluation results of the calibration policy text and multiple rounds of individual calibration policy evaluation results. These multiple rounds of individual calibration policy evaluation results are obtained by evaluating the calibration policy prompts using a policy evaluation model. The calibration policy prompts are generated by the policy evaluation agent based on the calibration policy text and the policy prompt template to be calibrated. The prompt validity evaluation results and prompt consistency evaluation results are used to evaluate the calibration policy prompts.
[0113] According to an embodiment of the present invention, strategy prompts are generated based on the strategy text to be evaluated and a strategy prompt template. A strategy evaluation agent evaluates the strategy prompts to obtain a strategy evaluation result. The strategy prompt template is a template to be calibrated, determined when the prompt validity evaluation result and the prompt consistency evaluation result meet preset conditions. Since generating strategy prompts using the strategy prompt template provides the strategy evaluation agent with evaluation directions from different dimensions, improving the rationality of the strategy evaluation result. Furthermore, the prompt validity evaluation result and the prompt consistency evaluation result of the strategy prompt template meet preset conditions, ensuring that the strategy evaluation result is consistent and effective compared to the evaluation result given by expert judges. This reduces the evaluation bias of the strategy evaluation agent and avoids the need for multiple evaluations, thereby improving computer processing efficiency while reducing computer resource consumption.
[0114] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention, or at least part of the functions of any one or more of them, can be implemented in a single module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be implemented by being divided into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, and firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present invention can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0115] For example, any plurality of the generation module 410 and evaluation module 420 can be combined into one module / unit / subunit, or any one of the modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present invention, at least one of the generation module 410 and evaluation module 420 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the generation module 410 and evaluation module 420 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0116] Figure 5 A block diagram of an electronic device suitable for implementing a policy evaluation method according to an embodiment of this application is shown. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0117] like Figure 5 As shown, an electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a ROM 502 (read-only memory) or a program loaded from a storage portion 508 into a RAM 503 (random access memory). The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0118] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 502 and / or RAM 503. It should be noted that programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in one or more memories.
[0119] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The system 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A driver 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the driver 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.
[0120] According to embodiments of the present invention, the method flow according to embodiments of the present invention can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of the embodiments of the present invention. According to embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0121] This invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the invention.
[0122] According to embodiments of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0123] For example, according to embodiments of the present invention, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0124] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of the present invention. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of the present invention.
[0125] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this embodiment of the invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0126] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0127] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not expressly stated in the present invention. In particular, the features described in the various embodiments and / or claims of this invention can be combined and / or combined in various ways without departing from the spirit and teachings of this invention. All such combinations and / or combinations fall within the scope of this invention.
[0129] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of the invention is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. A policy evaluation method, comprising: generating a policy prompt word matched with a policy text to be evaluated based on the policy text to be evaluated and a policy prompt word template, the policy prompt word template being used to evaluate the policy text to be evaluated based on at least two of policy text content accuracy, policy text content integrity, policy compliance, policy text content logic, or policy performance, the policy prompt word template being a prompt word in multiple dimensions for assisting a large language model in analyzing the policy text to be evaluated; evaluating the policy text to be evaluated based on the policy prompt word related to the policy text content accuracy, the policy text content integrity, the policy compliance, the policy text content logic, or the policy performance by a policy evaluation agent, to obtain a policy evaluation result; wherein the policy prompt word template is a to-be-calibrated policy prompt word template determined in a case where a prompt word effectiveness evaluation result and a prompt word consistency evaluation result meet a preset condition, the prompt word effectiveness evaluation result being determined based on respective calibration policy evaluation results of multiple rounds, the prompt word consistency evaluation result being determined based on an expert policy evaluation result of a calibration policy text and the respective calibration policy evaluation results of the multiple rounds, the respective calibration policy evaluation results of the multiple rounds being obtained by evaluating a calibration policy prompt word in multiple rounds by a policy evaluation model, the calibration policy prompt word being generated by the policy evaluation agent based on the calibration policy text and the to-be-calibrated policy prompt word template, the prompt word effectiveness evaluation result and the prompt word consistency evaluation result being used to evaluate the calibration policy prompt word.
2. The method of claim 1, wherein, the prompt word effectiveness evaluation result and the prompt word consistency evaluation result being used to evaluate the calibration policy prompt word, comprising: the prompt word effectiveness evaluation result being used to evaluate effectiveness of the calibration policy prompt word, and the prompt word consistency evaluation result being used to evaluate consistency of the calibration policy prompt word, the expert policy evaluation result being obtained by an expert evaluating the calibration policy text.
3. The method of claim 1 or 2, wherein, the policy prompt word template being obtained by repeatedly performing the following operations until the prompt word effectiveness evaluation result and the prompt word consistency evaluation result meet the preset condition: in a case where an n-1th prompt word effectiveness evaluation result and an n-1th prompt word consistency evaluation result do not meet the preset condition, optimizing an n-1th to-be-calibrated policy prompt word template to obtain an nth to-be-calibrated policy prompt word template, n being an integer greater than or equal to 1; generating an nth calibration policy prompt word based on the calibration policy text and the nth to-be-calibrated policy prompt word template; evaluating the nth calibration policy prompt word in multiple rounds by the policy evaluation agent to determine respective nth calibration policy evaluation results of the multiple rounds; determining an nth prompt word effectiveness evaluation result based on the respective nth calibration policy evaluation results of the multiple rounds by the policy evaluation agent; The strategy evaluation agent evaluates the n-th prompt word based on the n-th expert strategy evaluation result of the agent based on the calibration strategy text and the respective n-th calibration strategy evaluation result of the multiple rounds, to determine an n-th prompt word consistency evaluation result; The n-th strategy prompt word template determined under the condition that the n-th prompt word effectiveness evaluation result and the n-th prompt word consistency evaluation result meet the preset condition is taken as the strategy prompt template.
4. The method of claim 1 or 2, wherein, The strategy evaluation agent evaluates the strategy prompt word to obtain a strategy evaluation result, including: The strategy evaluation agent evaluates the to-be-evaluated strategy text based on the strategy prompt word related to the strategy text content accuracy to obtain a first strategy evaluation result, which represents whether the to-be-evaluated strategy text meets the strategy text content accuracy; The strategy evaluation agent evaluates the text block of the to-be-evaluated strategy text based on the strategy prompt word related to the strategy text content integrity and the strategy text content logic to obtain a second strategy evaluation result and a third strategy evaluation result, the text block being obtained by cutting the to-be-evaluated strategy text, the second strategy evaluation result representing whether the to-be-evaluated strategy text meets the strategy text content integrity, and the third strategy evaluation result representing whether the to-be-evaluated strategy text meets the strategy text content logic; The strategy evaluation agent evaluates the to-be-evaluated strategy text based on the strategy prompt word and the strategy rule related to the strategy compliance to obtain a fourth strategy evaluation result, which represents whether the to-be-evaluated strategy text meets the strategy compliance; The strategy evaluation agent evaluates the to-be-evaluated strategy text based on the strategy prompt word and the preset behavior standard related to the strategy performance to obtain a fifth strategy evaluation result, which represents whether the to-be-evaluated strategy text meets the strategy performance; The strategy evaluation result is obtained based on at least two of the first strategy evaluation result, the second strategy evaluation result, the third strategy evaluation result, the fourth strategy evaluation result, and the fifth strategy evaluation result.
5. The method of claim 4, wherein, The strategy evaluation agent evaluates the to-be-evaluated strategy text based on the strategy prompt word related to the strategy text content accuracy to obtain a first strategy evaluation result, including: The strategy evaluation agent performs text extraction on the to-be-evaluated strategy text to obtain a key text, the key text including text content related to a target object of executing a strategy in the to-be-evaluated strategy text; The strategy evaluation agent obtains a reference text corresponding to the key text from a database; The strategy evaluation agent obtains a first strategy evaluation result based on the key text and the reference text.
6. The method of claim 4, wherein, The second policy evaluation result and the third policy evaluation result are obtained by evaluating the text blocks of the to-be-evaluated policy text based on policy prompt words related to policy text content integrity and policy text content logic by using the policy evaluation agent, and the second policy evaluation result and the third policy evaluation result are obtained by evaluating the text blocks of the to-be-evaluated policy text based on policy prompt words related to policy text content integrity and policy text content logic by using the policy evaluation agent. The correlation of the text blocks of the to-be-evaluated policy text is evaluated by using the policy evaluation agent, and an associated text architecture of each of the text blocks is obtained. The second policy evaluation result is obtained based on the associated text architecture of each of the text blocks by using the policy evaluation agent. The context of the text blocks of the to-be-evaluated policy text is evaluated by using the policy evaluation agent, and an antecedent text block and a consequent text block of each of the text blocks are obtained. The third policy evaluation result is obtained based on the antecedent text block and the consequent text block of each of the text blocks by using the policy evaluation agent.
7. The method of claim 6, wherein, The second policy evaluation result is obtained based on the associated text architecture of each of the text blocks by using the policy evaluation agent. The expected text architecture is generated by aggregating the associated text architecture of each of the text blocks by using the policy evaluation agent. The second policy evaluation result is obtained based on the difference between the expected text architecture and the text architecture of the to-be-evaluated policy text by using the policy evaluation agent.
8. The method of claim 6, wherein, The third policy evaluation result is obtained based on the antecedent text block and the consequent text block of each of the text blocks by using the policy evaluation agent. For any one of the text blocks, The antecedent evaluation result is obtained based on the antecedent text block and a neighboring antecedent text block adjacent to the text block by using the policy evaluation agent. The consequent evaluation result is obtained based on the consequent text block and a neighboring consequent text block adjacent to the text block by using the policy evaluation agent. The third policy evaluation result is obtained based on the antecedent evaluation result and the consequent evaluation result corresponding to each of the text blocks by using the policy evaluation agent.
9. The method of claim 4, wherein, The fourth policy evaluation result is obtained by evaluating the to-be-evaluated policy text based on policy prompt words related to policy compliance and policy rules by using the policy evaluation agent. The behavior text is obtained by extracting the to-be-evaluated policy text by using the policy evaluation agent, and the behavior text includes text related to the behavior of executing the policy in the to-be-evaluated policy text. The fourth policy evaluation result is obtained based on the behavior text and a rule behavior corresponding to the policy rule by using the policy evaluation agent.
10. The method of claim 9, wherein, The fifth policy evaluation result is obtained by evaluating the to-be-evaluated policy text based on policy prompt words related to policy performance and a preset behavior standard by using the policy evaluation agent. The strategy evaluation agent evaluates the agent to predict the behavior text based on the preset behavior standard, and obtain at least one behavior result of the strategy text; The strategy evaluation agent determines the behavior value corresponding to each of the at least one behavior result based on the preset behavior standard; The strategy evaluation agent obtains the fifth strategy evaluation result based on the at least one behavior result and the behavior value corresponding to each of the at least one behavior result.
Citation Information
Patent Citations
Prompt word optimization method, intelligent agent and storage medium
CN120278153A
Strategy generation and evaluation method based on large language model, medium and equipment
CN120929793A