Strategy generation and evaluation method based on large language model, medium and equipment
By selecting appropriate generation methods and model unit architectures, and combining similarity assessment and role-based prompt word weights, the accuracy and multi-dimensional adaptability issues of converting strategy code into natural language text are resolved, ensuring the stability and applicability of strategy text generation and evaluation.
Patent Information
- Application Number
- CN202511461304.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies suffer from inconsistent accuracy when converting policy codes into natural language text, are unable to adapt to multi-angle interpretation and customization needs, and are prone to information truncation and logical confusion when dealing with extremely long policy codes, making it difficult to meet the semantic transmission and evaluation needs in multi-role scenarios.
By selecting between model switching generation and dialog box switching generation methods based on the number of characters in the strategy code and the memory attributes of the large language model, reference strategy codes adapted to different roles are generated using independent or isolated model units. The effectiveness of the strategy text is then evaluated by combining similarity assessment and role-related prompt word weights.
It achieves the integrity and stability of strategy text generation, adapts to semantic transmission in multi-role scenarios, improves the accuracy and reliability of evaluation results, and avoids evaluation bias caused by memory interference and role bias.
Smart Images

Figure CN120929793A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, medium, and device for strategy generation and evaluation based on a large language model. Background Technology
[0002] Policy codes in the form of Open Digital Rights Language (ODRL) specifications are the core carrier for implementing resource access control and permission allocation. However, ODRL explicitly expresses the rules, obligations, and restrictions regarding the use of digital content, data, or services in a machine-readable format, making it difficult for users to understand the raw ODRL code and clearly understand key information such as what is allowed, what is not allowed, and what obligations they are required to fulfill. Therefore, converting policy codes into natural language text to facilitate users' access to key policy information plays an important role in various fields such as digital rights management and data security compliance.
[0003] In existing technologies, when using large language models to convert strategy code into natural language text, the large language model is simply used for one generation. However, large language models may generate strategy texts at different times that are slightly different or even significantly different in terms of accuracy, completeness, and emphasis. Furthermore, a strategy text may need to be reviewed by different roles and needs to adapt to multi-angle and customizable interpretations. The strategy texts converted by the above methods are random and unstable, and are difficult to adapt to multi-dimensional interpretations, thus failing to meet the stability requirements of strategy texts.
[0004] In the exploration of related technologies for policy generation and evaluation based on large language models, the Chinese invention patent "Method and System for Generation and Evaluation of Security Policies for Network / Security Devices Based on Large Language Models" (Publication No. CN120524938A, Publication Date: August 22, 2025) proposes to construct threat assessment datasets and policy generation datasets separately, and use LoRA+ fine-tuning technology based on human feedback reinforcement learning to train the base model, obtaining threat assessment LLM and policy generation LLM. Subsequently, the threat assessment LLM is used to process abnormal logs and output threat information quadruples. The policy generation LLM combines these quadruples with a knowledge base agent to generate device security policies. Finally, the policy is evaluated by evaluating the agent from three dimensions: correctness, effectiveness, and suitability. In-depth analysis of this patent solution reveals that it still has many undeniable shortcomings in addressing the core requirement of converting policy code to natural language text, and fails to effectively solve the key pain points in the current technical field.
[0005] First, this patent fails to consider the multi-role adaptation requirements of the strategy text. In practical applications, the strategy text needs to be interpreted by different roles, such as technical operations personnel, enterprise managers, and compliance review personnel. The focus of each role (e.g., technical personnel focus on configuration details, managers on risk impact, and compliance personnel on compliance clauses) differs significantly. The strategy text generated by this patent uses a uniform output mode, which cannot meet the multi-dimensional and customized interpretation needs, resulting in a significant reduction in the text's practicality. Second, this patent lacks a dynamic adaptation mechanism for the length of the input strategy code and the model's memory capacity. Its threat assessment LLM and strategy generation LLM are based on fine-tuning of a fixed-base model, failing to consider that when the number of characters in the input ODRL strategy code exceeds the model's memory limit, issues such as information truncation and loss of key content will occur, leading to incomplete and inaccurate natural language strategy text. Meanwhile, the solution does not provide a solution for model switching or task splitting. When faced with extremely long policy code, it can only rely on a single model for hard processing. This may affect the generation efficiency due to model overload, and may also cause logical confusion in the policy text due to memory interference. It is difficult to guarantee the stability and reliability of the generated results. It does not take into account the difficulty of understanding the policy text and the usage requirements of different roles, and cannot evaluate the semantic transmission accuracy of the policy text in multi-role scenarios. This leads to one-sided evaluation results and makes it difficult to guide the optimization and iteration of policy text.
[0006] Therefore, how to generate stable strategy texts that can be adapted to multi-dimensional interpretations, while improving the accuracy of evaluation results, has become an urgent problem to be solved. Summary of the Invention
[0007] To address the aforementioned technical problems, the present invention provides a strategy generation and evaluation method based on a large language model, which includes the following steps: S1. Determine the text generation method based on the initial strategy code, the number of characters in the first task prompt and the preset role prompt, as well as the memory attributes of the preset large language model. The text generation method includes model switching generation and dialog box switching generation.
[0008] S2, determine the first model unit and the second model unit according to the text generation method.
[0009] S3. Input the initial strategy code and the first task prompt word into the first model unit to generate the initial strategy text. The first task prompt word is used to assist in the strategy text generation task.
[0010] S4. The initial strategy text and the second task prompt word are combined with each preset role prompt word and then input into the second model unit to obtain the reference strategy code corresponding to each preset role prompt word. The second task prompt word is used to assist in the generation of the strategy code, and the preset role prompt word is used to indicate the role attributes of the receiver of the reference strategy code.
[0011] S5. Based on the first similarity between each reference strategy code and the initial strategy code, and the weight corresponding to each preset role prompt word, the effectiveness of the initial strategy text is evaluated, and the strategy text evaluation result is obtained.
[0012] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described policy generation and evaluation method based on a large language model.
[0013] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0014] This invention has at least the following beneficial effects: By determining the generation method of model switching or dialog box switching based on the number of characters in the initial strategy code and the memory attributes of the preset large language model, the corresponding first and second model units are matched through text generation. This allows the strategy text generation process to adapt to initial strategy codes of different lengths and the memory capacity limitations of the model. While ensuring the independence of the reference strategy code, it optimizes the balance between resources and efficiency, avoids evaluation bias caused by memory interference, and ensures the integrity and stability of the generated initial strategy text. By inputting the initial strategy code and the first task prompt into the first model unit to generate the initial strategy text, and combining the initial strategy text, the second task prompt, and the prompts of each preset role into the second model unit, the corresponding reference strategy code for each role is generated. This allows the same initial strategy text to adapt to the understanding perspectives of different roles, providing multi-dimensional samples for subsequent effectiveness evaluation. By combining the first similarity between each reference strategy code and the initial strategy code, as well as the weight of the corresponding preset role prompt, the effectiveness of the initial strategy text is evaluated. This ensures that the evaluation results reflect the semantic transmission accuracy of the initial strategy text in multi-role scenarios, avoids the one-sidedness of evaluation caused by a single role perspective, and the final output strategy text evaluation results are more in line with the comprehensive needs of actual application scenarios. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of a strategy generation and evaluation method based on a large language model, provided in Embodiment 1 of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0019] Example 1 This first embodiment provides a method for strategy generation and evaluation based on a large language model, such as... Figure 1 As shown, this strategy generation and evaluation method based on a large language model includes the following steps: S1. Determine the text generation method based on the initial strategy code, the number of characters in the first task prompt and the preset role prompt, as well as the memory attributes of the preset large language model. The text generation method includes model switching generation and dialog box switching generation.
[0020] The initial policy code is machine-readable code that conforms to specific specifications (such as ODRL) and is used to define rules for digital rights, access permissions, etc., such as structured code containing elements such as resource subjects, right types, and constraints.
[0021] The pre-selected large language model is a pre-selected artificial intelligence model with natural language understanding and generation capabilities, such as the GPT series and LLaMA. It is used to generate text or code that meets the requirements based on the input content and prompt words, completing the conversion tasks from code to text and from text to code. Those skilled in the art will know that the framework structure, training methods, and applications of existing large language models all fall within the protection scope of this invention, and will not be elaborated further here.
[0022] The first task prompt is a structured instruction designed for the conversion of initial strategy code into natural language text. It includes task objectives (such as fully preserving key elements of the code), output format (such as organizing text according to resource + subject + permission + constraint logic), and quality requirements (such as unambiguity and no missing elements). It provides precise guidance for the preset large language model, ensuring that the generated initial strategy text meets the professional needs of strategy code conversion and reduces irrelevant information or missing elements.
[0023] Predefined role-specific prompts are predefined sets of instructions bound to specific role attributes. They introduce a role's unique perspective and needs, associating the reference strategy code with the role and realistically reflecting how different recipients accept the strategy text. This avoids the limitation of a single reference code failing to cover multiple scenario requirements, providing a multi-dimensional benchmark for subsequent evaluation. For example, a legal role prompt might be: from the perspective of an intellectual property lawyer, prioritize ensuring that the rights and obligations constraints in the code comply with Article X of the Copyright Law, clearly defining infringement liability; a technical role prompt might be: from the perspective of a system architect, optimize code execution efficiency, ensure that constraints can be parsed by the XX system, and reduce redundant parameters.
[0024] The memory properties determine the scale of information that a large language model can remember in context. The memory of a large language model depends on its context window, which can be understood as the short-term memory capacity of the large language model for processing information. When the total amount of input content exceeds this capacity, the large language model will lose some information, especially the content of earlier input.
[0025] Therefore, based on the memory space occupied by the character count of the initial strategy code, the first task prompt word, and the preset role prompt word, and compared with the context window capacity of the large language model, it is determined whether the model can retain the memory of the initial strategy code. Then, it dynamically decides whether to use an independent model or different dialog windows of the same model to handle the subsequent reference strategy code generation task, so as to avoid the memory of the initial strategy code interfering with the generation of the reference strategy code.
[0026] The above-mentioned method quantitatively evaluates the relationship between the number of characters in the initial strategy code, the first task prompt, and the preset role prompt and the memory attributes of the preset large language model, and determines whether the preset large language model can retain the memory of the initial strategy code. This allows for the dynamic selection of subsequent strategy generation methods, ensuring the independence of the reference strategy code while optimizing the balance between resources and efficiency. This avoids evaluation bias caused by memory interference and improves the objectivity and credibility of the initial strategy text validity determination.
[0027] In one specific implementation, the memory attributes of the preset large language model include the context window capacity and the preset ratio threshold. The context window capacity is represented by the number of tokens. S1 includes the following steps: S11, convert the number of characters in the initial strategy code, the number of characters in the first task prompt, and the number of characters in each preset role prompt into the corresponding number of tokens.
[0028] S12, calculate the sum of the number of tokens corresponding to the initial strategy code, the number of tokens corresponding to the first task prompt, and the maximum number of tokens corresponding to all preset role prompts, to obtain the total number of tokens.
[0029] S13, If the ratio of the total number of tokens to the capacity of the context window is greater than the preset ratio threshold, then the text generation method is determined to be dialog box switching generation.
[0030] S14. If the ratio of the total number of tokens to the capacity of the context window is less than or equal to a preset ratio threshold, then the text generation method is determined to be model switching generation.
[0031] The context window capacity is the maximum number of tokens that the preset large language model can process at one time, representing the upper limit of the large language model's memory, and serving as a benchmark for determining whether the model will forget the initial policy code. Correspondingly, when the total amount of input approaches or exceeds this capacity, the preset large language model will prioritize retaining recent inputs and forget earlier information, such as the initial policy code.
[0032] The preset ratio threshold is a pre-set critical value used to distinguish between the state of model memory retention and memory forgetting. For example, 70% means that when the ratio of the total number of tokens to the capacity of the context window exceeds the preset ratio threshold, it is determined that the preset large language model will forget the initial strategy code; otherwise, it is determined that the preset large language model may retain the memory.
[0033] By converting characters to tokens, the unstructured input size (number of characters) is transformed into a quantifiable unit (token) that the model can understand, achieving comparability between the input size and the model's memory attributes, and providing a unified benchmark for subsequent calculations. The conversion ratio from characters to tokens can be determined based on the text type; for example, 1 Chinese character equals 1.5 tokens, and 1 code character equals 2 tokens.
[0034] Since different role prompts will not be input into the same model unit at the same time, among multiple preset role prompts, the prompt with the largest number of tokens is selected to accurately estimate the maximum token usage in a single round of input, avoiding overestimation of the total number of tokens due to repeated calculations.
[0035] The total number of tokens can be used to assess the extent to which input information occupies the memory of the preset large language model. By comparing the ratio of the total number of tokens to the capacity of the context window with a preset ratio threshold, it can be determined whether the preset large language model will forget the initial strategy code.
[0036] Specifically, when the ratio is greater than a preset ratio threshold, it is determined that the preset large language model will forget the initial code due to memory capacity limitations. Different dialog boxes of the same preset large language model can be used safely, reducing the waste of resources caused by using an independent model when a dialog box could have been used.
[0037] When the ratio is less than or equal to the preset ratio threshold, it is determined that the preset large language model may retain the memory of the initial strategy code. An independent large language model is required to ensure that the generation of the reference strategy code is not affected by the potential interference of the initial strategy code. This allows the reference strategy code corresponding to different preset role prompt words to truly reflect the semantic transmission effect of the initial strategy text. Thus, based on the evaluation results of the unbiased reference strategy code, the effectiveness of the initial strategy text can be more objectively reflected, providing an accurate basis for subsequent evaluations.
[0038] As described above, by converting characters to token counts and calculating the total number of tokens, and comparing the ratio of the total number of tokens to the context window capacity with a preset threshold, the precise quantification of input size and the precise division of model memory states are achieved, ensuring the scientific nature of the generation method selection. This balances the improvement of resource utilization with the guarantee of a clean generation environment free from memory pollution for subsequent reference strategy code, thereby improving the reliability of the evaluation results.
[0039] S2, determine the first model unit and the second model unit based on the text generation method.
[0040] In one specific embodiment, S2 includes the following steps: When the text generation method is model switching generation, the first model unit and the second model unit are two independently deployed preset large language models with the same parameter configuration. The input and output data of the two large language models are isolated from each other.
[0041] When the text generation method is dialog box switching generation, the first model unit and the second model unit are two independent dialogue contexts in the same preset large language model process. The two dialogue contexts share model parameters and each maintains an independent input history.
[0042] The first model unit is a large language model processing unit responsible for performing the task of converting the initial policy code to the initial policy text. The second model unit is a large language model processing unit responsible for performing the task of converting the initial policy text to the reference policy code, and it must ensure that it has no interfering information from the initial policy code in its memory.
[0043] When the generation method is model switching generation, physical isolation must be achieved through two independently deployed, pre-configured large language model instances with identical parameters to completely sever the memory transfer of the initial strategy code. These two large language models are built based on the same pre-trained parameters and configurations (such as model structure and inference parameters), but run on different computing resources (such as independent servers or containers), and their input and output data are not shared. This achieves physical-level memory isolation, ensuring that when the first model unit remembers the initial strategy code, the second model unit, due to its independent deployment, has no such memory. This ensures that the reference code generation relies solely on the initial strategy text's second task prompt and the pre-configured role prompt.
[0044] When the generation method is dialog box switching, logical isolation can be achieved through independent dialogue contexts within the same preset large language model, avoiding mutual interference of input history while sharing model parameters. Specifically, the two independent dialogue contexts within the same preset large language model process are two logical dialogue spaces created through technical means (such as session ID isolation) during the execution process of a single preset large language model. They share the model's core computational parameters (such as the weight matrix). By using independent dialogue contexts to prevent the input history of the first model unit from affecting the second model unit, model computational resources are reused, improving efficiency.
[0045] The above describes how the model unit architecture is selected by matching the text generation method to achieve precise adaptation of memory isolation, ensuring that the generation of reference strategy code is not affected by the memory of the initial strategy code, and optimizing the balance between resources and efficiency while ensuring the isolation effect.
[0046] S3. Input the initial strategy code and the first task prompt word into the first model unit to generate the initial strategy text. The first task prompt word is used to assist in the strategy text generation task.
[0047] In one specific embodiment, S3 includes the following steps: S31, the initial strategy code and the first task prompt are combined and input into the first model unit to obtain the original strategy text corresponding to the initial strategy code.
[0048] S32 extracts key elements from the original policy text using an entity recognition algorithm. These key elements include resource identifiers, subject information, scope of permissions, time constraints, geographical constraints, and obligation constraints.
[0049] S33 performs structured verification on key elements and obtains the verification result, which is either pass or fail.
[0050] S34. If the verification result is not passed, a supplementary instruction is sent to the first model unit to obtain the original strategy text regenerated by the first model unit.
[0051] S35, Repeat step S32 until the verification result is passed, and obtain the initial strategy text.
[0052] The original strategy text is an unverified natural language text directly generated by the first model unit based on the initial strategy code and the first task prompt words. It may have problems such as missing key elements and ambiguous logical expressions. Therefore, the key elements in the original strategy text are extracted by entity recognition algorithm. The key elements are then structured and verified to obtain the verification results. When the key elements are missing or incorrect, that is, when the verification fails, supplementary instructions are passed to the first model unit to correct the generation deviations. An iterative process of generation, verification, and optimization is executed until the generated strategy text meets the requirements, forming a quality control mechanism of generation, verification, and optimization, thereby improving the generation quality and reliability of the initial strategy text.
[0053] Specifically, entity recognition algorithms can automatically identify entities with specific meanings in the original strategy text, such as resource identifiers, time constraints, and other preset key elements, and mark their categories and locations. For example, December 31, 2025 can be marked as a time constraint, thereby transforming the unstructured original strategy text into a verifiable set of structured elements, which makes it easier to accurately locate missing or incorrect elements in the original strategy text.
[0054] Further, a verification rule base is constructed based on the key elements of the initial strategy code. The elements extracted by the entity recognition algorithm are compared with the requirements of the verification rule base to determine whether the original strategy text completely contains all necessary elements and whether the element description is consistent with the semantics of the code. For example, "allow commercial use" in the original strategy text should match "commercialuse=yes" in the initial strategy code. This quantitatively evaluates the completeness and accuracy of the original strategy text, avoids the problem of seemingly fluent text but missing key information, and provides a clear basis for feedback optimization.
[0055] Furthermore, when the structured validation fails, targeted supplementary instructions are sent to the first model unit, including missing element prompts (such as the need to supplement regional constraint information) and correction requirements (such as explicitly marking the resource identifier as image001), thereby helping the first model unit to accurately locate its own generation deviation, reduce invalid iterations, and improve optimization efficiency.
[0056] As described above, the precise guidance of the first task prompt reduces the randomness of the first model unit generation, ensuring that the original strategy text meets the professional conversion requirements. Through the combined operation of entity recognition algorithm and structured verification, the elements in the original strategy text are accurately checked. Through the iterative closed loop of verification failure → feedback supplementary instructions → regeneration, the strategy text is dynamically optimized, ensuring the integrity and accuracy of the final initial strategy text.
[0057] S4. The initial strategy text and the second task prompt word are combined with each preset role prompt word and then input into the second model unit to obtain the reference strategy code corresponding to each preset role prompt word. The second task prompt word is used to assist in the generation of the strategy code, and the preset role prompt word is used to indicate the role attributes of the receiver of the reference strategy code.
[0058] The second task prompt is a general instruction designed for the reverse conversion of the initial strategy text to strategy code. It includes basic specifications for code generation, such as converting to machine-readable code that conforms to ODRL syntax, including complete entity definitions, constraints and action instructions, and ensuring that the syntax is error-free. This provides a unified technical benchmark for the generation of reference code for all roles, avoids code format chaos and missing elements caused by the free play of the second model unit, and ensures that the reference strategy code has basic standardization and executability.
[0059] The reference strategy code serves as a mirror reference for evaluating the effectiveness of the initial strategy text. If the reference strategy codes corresponding to different role prompts can maintain a high degree of similarity with the initial strategy code, it indicates that the initial strategy text has strong semantic transmission ability and high effectiveness; otherwise, it indicates that the initial strategy text has ambiguity or missing elements and low effectiveness.
[0060] As described above, by combining the input of the second task prompt and the preset role prompt, a unique reference strategy code is generated for each preset role, constructing a multi-dimensional evaluation benchmark, avoiding the one-sidedness of the initial strategy text effectiveness evaluation, and improving the objectivity of the initial strategy text effectiveness evaluation.
[0061] S5. Based on the first similarity between each reference strategy code and the initial strategy code, and the weight corresponding to each preset role prompt word, the effectiveness of the initial strategy text is evaluated, and the strategy text evaluation result is obtained.
[0062] The first similarity measure assesses the consistency between a single reference strategy code and the initial strategy code across semantic, structural, and elemental dimensions, quantifying the transformation effect from the initial strategy text to the reference strategy code. The calculation method can include a fusion of multiple dimensions such as semantic vector cosine similarity and abstract syntax tree structure similarity. The value can be between 0 and 1, with values closer to 1 indicating higher consistency.
[0063] The weights corresponding to the preset role prompts are weighted coefficients set according to the importance of different roles in actual applications. These can be determined based on historical data or business needs, reflecting the priority of different roles' demands for the initial strategy text and avoiding evaluation bias caused by all roles being equally important. For example, because the compliance risk impact of legal roles is greater, their weight is usually higher than that of ordinary business roles. The weight corresponding to the prompts for legal roles is 0.4, the weight corresponding to the prompts for technical roles is 0.3, and the weight corresponding to the prompts for business roles is 0.3, with a total weight of 1.
[0064] The above-mentioned method achieves an objective assessment of the effectiveness of the initial strategy text by calculating the consistency between the first similarity metric reference strategy code and the initial strategy code, avoiding the subjectivity and bias of manual assessment based on experience, and making the assessment results more convincing. By pre-setting the weight of role prompts, the assessment results are made to match the differences in the importance of roles in actual business, improving the scenario adaptability of the assessment. At the same time, it comprehensively examines the semantic transmission capability of the initial strategy text in multi-role scenarios, avoiding the one-sidedness of single-role assessment, and improving the comprehensiveness and reliability of the assessment.
[0065] In one specific embodiment, S5 includes the following steps: S511 performs syntax parsing on each reference strategy code and initial strategy code to obtain the abstract syntax tree corresponding to each reference strategy code and initial strategy code, as well as the code elements in each reference strategy code and initial strategy code. The code elements include rights entities, constraint parameters, and action instructions.
[0066] S512 converts each code element into a code feature vector, where the dimensions of the code feature vector include the type of right action, the number of constraints, the complexity of subject relationships, and the completeness of the syntactic structure.
[0067] S513, calculate the semantic similarity between the code feature vector corresponding to each reference strategy code and the code feature vector corresponding to the initial strategy code.
[0068] S514, calculate the structural similarity between the abstract syntax tree corresponding to each reference strategy code and the abstract syntax tree corresponding to the initial strategy code.
[0069] S515, calculate the first similarity between each reference strategy code and the initial strategy code based on the semantic similarity and structural similarity corresponding to each reference strategy code and the initial strategy code.
[0070] This involves using existing parsing tools (such as ANTLR and Lark) to generate a parser that transforms the input strategy code into a structured abstract syntax tree. Each node in the abstract syntax tree corresponds to a grammatical component in the code, such as a tag, attribute, or text.
[0071] Each node of each abstract syntax tree (AST) is traversed using either a depth-first or breadth-first search algorithm. The node type, hierarchical relationship, and connection method are recorded to form a structural description of the AST, such as a node tree diagram or a hierarchy list, for subsequent structural similarity calculation. Nodes are matched according to element extraction rules, and all elements conforming to the rules are collected to form a code element set. This process removes irrelevant formatting information, focuses on core semantic comparison, and lays the foundation for subsequent feature quantification.
[0072] The specific content of the code elements can be set by the implementer according to the actual situation. In this embodiment, the rights entity can be a participating entity such as a resource owner or user; the constraint parameters can be conditions such as time range or geographical restrictions; and the action instructions can be rights actions such as allowing use or prohibiting modification.
[0073] Correspondingly, when constructing the code element collection, if an asset node is encountered, its id attribute is extracted as the rights entity, and if... <asset> / <party>The node extracts its type and value attributes as constraint parameters, and when it encounters an action node, it extracts its text content as an action instruction.
[0074] Furthermore, unstructured code elements are transformed into multi-dimensional code feature vectors that can be processed by machine learning. Each dimension corresponds to a quantifiable code feature, and the feature transformation form for each dimension can be set according to the actual situation of the dimension. For example, for the type of right action, one-hot encoding is assigned to each action, and 0-1 encoded feature values are used to represent actions such as use / modification / distribution. For the number of constraints, the total number of extracted constraint parameters is counted (e.g., time constraint + geographical constraint = 2), and a maximum value is set (e.g., 5). The feature values in the range [0, 1] are obtained by normalizing the actual number / maximum value. For the complexity of subject relationships, the association method of right entities is scored. For example, 1-to-1 relationship = 1, 1-to-many = 2, nested relationship = 3. The feature values in the range [0, 1] are obtained by normalizing the score / maximum score. For the completeness of the syntactic structure, several basic syntactic check items are set, such as whether necessary tags are included, whether the logic is closed, etc. The proportion of passing items is calculated and normalized to obtain the feature values in the range [0, 1].
[0075] Then, in a fixed order, such as the type of right action, the number of constraints, the complexity of subject relationships, and the completeness of syntactic structure, the feature values of each dimension are concatenated to form the final code feature vector.
[0076] Furthermore, the structural descriptions of the abstract syntax trees (abstract syntax trees) of each reference strategy code (denoted as T1) and the initial strategy code (abstract syntax tree) are extracted. These structural descriptions include node types, hierarchical relationships, and connection methods. Three basic operations are defined: node insertion, node deletion, and node replacement. The minimum number of operations required to transform T1 into T2 is calculated, i.e., the tree edit distance M. The total number of nodes in the two abstract syntax trees (denoted as N) is counted, and the structural similarity X = 1 - (M ÷ N), where X ranges from [0, 1], where 1 indicates completely identical structures and 0 indicates completely different structures.
[0077] Those skilled in the art will recognize that any method for calculating semantic similarity in the prior art falls within the protection scope of this invention, such as cosine similarity, which will not be elaborated upon here.
[0078] For each reference strategy code, the average of its semantic similarity and structural similarity with the initial strategy code is determined as its first similarity with the initial strategy code.
[0079] In one specific implementation, the weights of semantic similarity and structural similarity can be adjusted according to actual needs, thereby performing a weighted average of semantic similarity and structural similarity, making the evaluation results more in line with the needs of the scenario, improving the universality of the evaluation method, and adapting to the differentiated requirements of strategy code in different fields.
[0080] As described above, parallel extraction of abstract syntax tree code elements through syntax parsing improves processing efficiency while ensuring data consistency. By evaluating semantic similarity and structural similarity in two dimensions, a comprehensive characterization of code matching is achieved, improving the accuracy of the first similarity measurement.
[0081] In one specific embodiment, S5 further includes the following steps: S521, Obtain the historical database, which stores historical strategy codes, preset role prompts, intermediate strategy codes generated by each historical strategy code for each preset role prompt, and matching score corresponding to each intermediate strategy code.
[0082] S522, For each preset role prompt word, calculate the average matching score of all intermediate strategy codes generated corresponding to the current preset role prompt word.
[0083] S523, normalize the average matching score of all preset character prompts to obtain the weight corresponding to each preset character prompt.
[0084] Among them, the historical strategy code is the initial strategy code that has been processed in the past, the intermediate strategy code is the reference code generated by the historical strategy code for each preset role prompt, and the matching score is the similarity score between the intermediate strategy code and the historical strategy code that has been manually or automatically evaluated in the past, ranging from 0 to 100 points.
[0085] The arithmetic mean of the matching scores of all intermediate strategy codes corresponding to the same preset role prompt is taken to obtain the average matching score of each preset role prompt, thus quantifying the overall performance of each preset role prompt in historical scenarios.
[0086] The average matching score of all preset role prompts is normalized, and the normalization result is determined as the weight of each preset role prompt. The weight is positively correlated with the historical performance of the preset role prompt, ensuring that preset role prompts with high importance have a higher proportion in the evaluation.
[0087] As mentioned above, by introducing empirical data from historical databases, the weight of preset role prompts is transformed from subjective setting to data-driven, thereby improving the objectivity of weight allocation.
[0088] In one specific embodiment, S5 further includes the following steps: S531, based on the weight corresponding to each preset role prompt word, the first similarity corresponding to the reference strategy code corresponding to each preset role prompt word is weighted and averaged to obtain the comprehensive similarity.
[0089] S532, if the overall similarity is greater than or equal to the preset similarity threshold, the strategy text evaluation result is determined to be valid.
[0090] S533, if the overall similarity is less than the preset similarity threshold, the strategy text evaluation result is determined to be invalid.
[0091] The preset similarity threshold is a pre-defined critical value for judging the validity of text, avoiding subjectivity based on intuition and ensuring a consistent decision-making basis for evaluations across different batches and scenarios. It can be determined based on the comprehensive similarity distribution of valid strategy texts in historical data; for example, the lowest comprehensive similarity of 0.7 among historically valid strategy texts can be used as the threshold.
[0092] A higher overall similarity indicates a stronger consistency between the reference strategy code generated under different preset role prompts and the initial strategy code. This means the initial strategy text can accurately and unambiguously convey the core semantics of the initial code to each role, resulting in high effectiveness. Conversely, a lower overall similarity indicates a greater deviation between the reference strategy code and the initial strategy code. This suggests that the initial strategy text suffers from semantic ambiguity, missing elements, or other problems, resulting in low effectiveness.
[0093] In one specific implementation, the strategy generation and evaluation method based on a large language model further includes the following steps: S6. When the strategy text evaluation result is invalid, input all reference strategy codes, initial strategy codes and third task prompt words into the preset large language model to obtain a difference analysis report. The difference analysis report includes the differences between each reference strategy code and the initial strategy code, key conflict elements and optimization suggestions.
[0094] S7. Adjust the first task prompt words according to the difference analysis report. The adjustment methods include adding domain qualifiers and / or supplementing contextual constraints.
[0095] S8, after adjustment, return to step S3 until the strategy text evaluation result is valid or the preset adjustment limit is reached.
[0096] The third task prompt is an instruction designed to drive the difference analysis. It includes the analysis objective (such as comparing the semantic differences between the reference code and the initial code), the output format (such as presenting a list of differences, conflict levels, and optimization suggestions), and the professional requirements (such as focusing on the consistency of rights actions and constraints). It is used to guide the preset large language model to focus on the core differences of the strategy code, generate a structured and targeted analysis report, and avoid the generalization of analysis results or the omission of key conflicts.
[0097] Guided by the prompts in the third task, the pre-defined large language model analyzes the differences between the reference strategy code and the initial strategy code, generating a difference analysis report that includes differences, conflicting elements, and optimization suggestions. This transforms the fuzzy results with low overall similarity into an actionable list of issues, providing precise basis for adjustments, avoiding blind adjustments, and improving optimization efficiency.
[0098] The differences are the specific distinctions between the reference policy code and the initial policy code in elements such as rights entities and constraint parameters. For example, the reference code omits the time constraint 2025-12-31. The key conflict elements are the core differences that may lead to serious deviations. For example, the initial policy code allows use, while the reference policy code prohibits use. The optimization suggestions are the directions for adjusting the prompts based on the differences. For example, suggestions such as "clearly require the complete retention of time constraints in the first task prompt" and "this policy is applicable to internal data sharing scenarios within the enterprise, and the confidentiality obligations after data use should be clearly defined" are provided.
[0099] Based on the optimization suggestions in the difference analysis report, the first task prompt words were adjusted to optimize the conversion quality from initial strategy code to initial strategy text at the source. The adjustments included adding domain qualifiers and / or supplementing contextual constraints. Domain qualifiers are specific terms, industry rules, or scenario tags that clearly define the professional domain or application scenario to which the initial strategy text belongs. They are used to define the scope of strategy generation, avoiding the generation of generalized content beyond the target domain by a pre-set broad language model, and ensuring that the strategy text highly matches the actual business scenario. Supplementing contextual constraints involves adding strategy-related background information to the first task prompt words, addressing the omission of strategy text elements caused by missing domain qualifiers and / or missing context, making the generated text more closely aligned with the actual application scenario, and reducing deviations in subsequent code conversion.
[0100] The above-mentioned method generates a difference analysis report by pre-setting a large language model, thereby automating and accurately locating code differences. Based on the difference analysis report, the first task prompt words are adjusted to solve text quality problems from the source of conversion, avoid repeated and ineffective iterations, and help generate stable, high-quality strategy text that can adapt to multi-dimensional interpretation.
[0101] The above describes a process where the generation method for model switching or dialog box switching is determined by the number of characters in the initial strategy code and the memory attributes of the preset large language model. The text generation method matches the corresponding first and second model units, allowing the strategy text generation process to adapt to initial strategy codes of varying lengths and the model's memory limitations. This ensures the independence of the reference strategy code while optimizing resource and efficiency balance, avoiding evaluation bias caused by memory interference, and guaranteeing the integrity and stability of the generated initial strategy text. The initial strategy text is generated by inputting the initial strategy code and the first task prompt into the first model unit. The initial strategy text, the second task prompt, and the prompts for each preset role are then input into the second model unit to generate reference strategy codes for the corresponding roles. This allows the same initial strategy text to adapt to the understanding perspectives of different roles, providing multi-dimensional samples for subsequent effectiveness evaluation. Finally, the effectiveness of the initial strategy text is evaluated by combining the first similarity between each reference strategy code and the initial strategy code, as well as the weight of the corresponding preset role prompts. This ensures that the evaluation results reflect the semantic accuracy of the initial strategy text in multi-role scenarios, avoiding the bias caused by a single role's perspective. The final output strategy text evaluation results are more in line with the comprehensive needs of actual application scenarios.
[0102] Example 2 Embodiment 2 of the present invention provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the policy generation and evaluation method based on a large language model provided in the above embodiment.
[0103] Example 3 Embodiment 3 of the present invention provides an electronic device, which includes a processor and the non-transitory computer-readable storage medium of Embodiment 2 of the present invention.
[0104] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.< / party> < / asset>
Claims
1. A strategy generation and evaluation method based on a large language model, characterized in that, The method includes the following steps: S1. The text generation method is determined based on the number of characters in the initial strategy code, the first task prompt word, and the preset role prompt word, as well as the memory attributes of the preset large language model. The text generation method includes model switching generation and dialog box switching generation. S2, determine the first model unit and the second model unit according to the text generation method; S3, input the initial strategy code and the first task prompt word into the first model unit to generate the initial strategy text, wherein the first task prompt word is used to assist in the strategy text generation task; S4, the initial strategy text and the second task prompt word are combined with each preset role prompt word and then input into the second model unit to obtain the reference strategy code corresponding to each preset role prompt word. The second task prompt word is used to assist in the generation of the strategy code, and the preset role prompt word is used to indicate the role attribute of the receiver of the reference strategy code. S5. Based on the first similarity between each reference strategy code and the initial strategy code, and the weight corresponding to each preset role prompt word, the effectiveness of the initial strategy text is evaluated to obtain the strategy text evaluation result.
2. The strategy generation and evaluation method based on a large language model according to claim 1, characterized in that, The memory attributes of the preset large language model include the context window capacity and the preset ratio threshold. The context window capacity is represented by the number of tokens. S1 includes the following steps: S11, convert the number of characters in the initial strategy code, the number of characters in the first task prompt word, and the number of characters in each preset role prompt word into the corresponding number of tokens; S12, calculate the sum of the number of tokens corresponding to the initial strategy code, the number of tokens corresponding to the first task prompt word, and the maximum number of tokens corresponding to all preset role prompt words to obtain the total number of tokens; S13, if the ratio of the total number of tokens to the capacity of the context window is greater than the preset ratio threshold, then the text generation method is determined to be dialog box switching generation; S14, if the ratio of the total number of tokens to the capacity of the context window is less than or equal to the preset ratio threshold, then the text generation method is determined to be model switching generation.
3. The strategy generation and evaluation method based on a large language model according to claim 2, characterized in that, S2 includes the following steps: When the text generation method is model switching generation, the first model unit and the second model unit are two independently deployed preset large language models with the same parameter configuration, wherein the input and output data of the two large language models are isolated from each other; When the text generation method is dialog box switching generation, the first model unit and the second model unit are two independent dialogue contexts in the same preset large language model process. The two dialogue contexts share model parameters and each maintains an independent input history.
4. The strategy generation and evaluation method based on a large language model according to claim 1, characterized in that, S3 includes the following steps: S31, the initial strategy code and the first task prompt word are combined and input into the first model unit to obtain the original strategy text corresponding to the initial strategy code; S32, extract key elements from the original policy text using an entity recognition algorithm, wherein the key elements include resource identifier, subject information, scope of authority, time constraint, geographical constraint, and obligation constraint; S33, perform structured verification on the key elements and obtain the verification result, wherein the verification result is either pass or fail; S34, if the verification result is not passed, a supplementary instruction is sent to the first model unit to obtain the original strategy text regenerated by the first model unit; S35, Repeat step S32 until the verification result is passed, and obtain the initial strategy text.
5. The strategy generation and evaluation method based on a large language model according to claim 1, characterized in that, S5 includes the following steps: S511, perform syntax parsing on each reference strategy code and the initial strategy code respectively to obtain the abstract syntax tree corresponding to each reference strategy code and the initial strategy code respectively, as well as the code elements in each reference strategy code and the initial strategy code; S512, each code element is converted into a code feature vector, wherein the dimensions of the code feature vector include the type of right action, the number of constraints, the complexity of subject relationships, and the completeness of the syntactic structure; S513, calculate the semantic similarity between the code feature vector corresponding to each reference strategy code and the code feature vector corresponding to the initial strategy code; S514, calculate the structural similarity between the abstract syntax tree corresponding to each reference strategy code and the abstract syntax tree corresponding to the initial strategy code; S515, calculate the first similarity between each reference strategy code and the initial strategy code based on the semantic similarity and structural similarity corresponding to each reference strategy code and the initial strategy code.
6. The strategy generation and evaluation method based on a large language model according to claim 5, characterized in that, S5 also includes the following steps: S521, Obtain the historical database, wherein the historical database stores historical strategy codes, preset role prompt words, intermediate strategy codes generated by each historical strategy code for each preset role prompt word, and matching degree scores corresponding to each intermediate strategy code; S522, For each preset role prompt word, calculate the average matching score of all intermediate strategy codes generated corresponding to the current preset role prompt word; S523, normalize the average matching score of all preset character prompts to obtain the weight corresponding to each preset character prompt.
7. The strategy generation and evaluation method based on a large language model according to claim 6, characterized in that, S5 also includes the following steps: S531, based on the weight corresponding to each preset role prompt word, the first similarity corresponding to the reference strategy code corresponding to each preset role prompt word is weighted and averaged to obtain the comprehensive similarity; S532, if the overall similarity is greater than or equal to the preset similarity threshold, then the strategy text evaluation result is determined to be valid; S533, if the overall similarity is less than the preset similarity threshold, the strategy text evaluation result is determined to be invalid.
8. The strategy generation and evaluation method based on a large language model according to claim 7, characterized in that, The method further includes the following steps: S6, when the strategy text evaluation result is invalid, input all reference strategy codes, the initial strategy code and the third task prompt word into the preset large language model to obtain a difference analysis report, wherein the difference analysis report includes the differences between each reference strategy code and the initial strategy code, key conflict elements and optimization suggestions; S7. Adjust the first task prompt word according to the difference analysis report, wherein the adjustment method includes adding domain qualifiers and / or supplementing contextual constraints; S8, after adjustment, return to step S3 until the strategy text evaluation result is valid or the preset adjustment limit is reached.
9. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the policy generation and evaluation method based on a large language model as described in any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 9.
Citation Information
Patent Citations
Quantitative strategy natural language construction and interpretability auxiliary system based on multi-agent collaborative architecture
CN120524938A
Code generation model fine tuning method and device based on clustering and natural language strategy optimization algorithm
CN118468982A
Large language model discrete cue word searching method and device
CN120296148A
Remote sensing satellite group adaptive task planning method and system based on large language model
CN120430582A
Method and device for carrying out resource configuration on multi-mode network, equipment and medium
CN120768759A
Cited By
Code automatic generation method and system based on algorithm model
CN121116382A
Strategy evaluation method
CN121118884A
Strategy evaluation method
CN121118884B
Active optical alignment method based on local lightweight language model
CN121455222A
Energy storage abnormity early warning and risk grading system based on large language model
CN121599461A