Policy field-oriented reasoning generation method, device and equipment and storage medium
By identifying task types and combining semantic direction vectors and generation constraint mechanisms, the problems of non-standard and inaccurate generated results in the policy field are solved, ensuring the standardization of output format and the authenticity of content, and improving the parsability and credibility of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-17
AI Technical Summary
When existing large language models generate inference results in the policy domain, they lack fine-grained control over the output format, resulting in non-standard text format, low content accuracy, and a tendency to produce fictitious content and misleading conclusions, which affects the credibility and usability of the system.
By identifying the input question and determining the task type, structured generation constraints based on context-free grammar CFG and classification output constraints based on closed option sets are adopted. Combined with pre-generated semantic direction vectors, these constraints are injected into the hidden layer of the inference model to ensure that the generated results conform to the standard format and are authentic.
It has achieved standardized format and authenticity of policy reasoning results, improved the parsability and credibility of generated content, reduced the risk of generating fictitious content, and improved the reliability and efficiency of the system.
Smart Images

Figure CN121279459B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of response reasoning technology in the policy domain, and in particular to a reasoning generation method, apparatus, device and storage medium for the policy domain. Background Technology
[0002] In practical applications in the policy field, large language models are widely used to generate policy reasoning results, such as policy question answering, clause analysis, and decision support. These applications impose two stringent requirements on the quality of the generated content: on the one hand, the factual accuracy of the content must be ensured to avoid information bias; on the other hand, it must strictly adhere to pre-defined structured output specifications, including standardized policy citation formats, standardized clause citation methods, and structured comparative analysis frameworks, to ensure that downstream systems can efficiently parse and automate the process.
[0003] However, current large language models generally lack fine-grained control mechanisms over output format during the generation process. This results in generated text that, while semantically plausible, often suffers from formatting issues such as non-standard citations, disorganized wording, or loose analytical structures, severely weakening the usability of the results and the efficiency of system integration. A more prominent problem is that, because models excessively pursue fluency while neglecting the constraint of content authenticity, they are highly susceptible to external illusions—fabricating non-existent policy provisions, fabricating invalid citations, or generating misleading conclusions. This not only damages the credibility of the system but may also lead to policy implementation risks and legal compliance hazards. Summary of the Invention
[0004] The purpose of this application is to provide a reasoning generation method, apparatus, device, and storage medium for the policy field, so as to solve the technical problems of non-standard reasoning generation and low accuracy and authenticity of content in related technologies, and to ensure that the output format of policy reasoning results is standardized and the authenticity of content is improved.
[0005] In a first aspect, embodiments of this application provide a reasoning generation method for the policy domain, including:
[0006] The input policy domain problem is identified to determine the task type of the policy domain problem;
[0007] Based on the task type, a generation constraint mechanism is determined, wherein the generation constraint mechanism includes at least structured generation constraints based on context-free grammar (CFG) and classification output constraints based on closed option sets.
[0008] Obtain the pre-generated semantic direction vector based on the task type;
[0009] The semantic direction vector is injected into the hidden layer of the pre-trained inference model, and the policy domain problem and the generation constraint mechanism are input into the inference model injected with the semantic direction vector. The inference result corresponding to the policy domain problem is output. The step of injecting the semantic direction vector into the hidden layer of the pre-trained inference model is to perform vector operation on the output result of the hidden layer based on the semantic direction vector, and use the result of the operation as the input of the next adjacent layer. The hidden layer is several intermediate layers in the inference model.
[0010] Secondly, embodiments of this application provide a reasoning generation apparatus for the policy domain, comprising:
[0011] The identification and processing module is used to identify the input policy domain problem and obtain the task type of the policy domain problem;
[0012] The constraint determination module is used to determine the generation constraint mechanism according to the task type, wherein the generation constraint mechanism includes at least structured generation constraints based on context-free grammar (CFG) and classification output constraints based on closed option sets.
[0013] The vector acquisition module is used to acquire a pre-generated semantic direction vector according to the task type;
[0014] The inference processing module is used to inject the semantic direction vector into the hidden layer of a pre-trained inference model, and input the policy domain problem and the generation constraint mechanism into the inference model injected with the semantic direction vector, and output the inference result corresponding to the policy domain problem. The step of injecting the semantic direction vector into the hidden layer of the pre-trained inference model is to perform vector operation on the output result of the hidden layer based on the semantic direction vector, and use the result of the operation as the input of the next adjacent layer. The hidden layer is a number of intermediate layers in the inference model.
[0015] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the policy domain-oriented reasoning generation method described above.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the policy-oriented reasoning generation method described above.
[0017] This application provides a method, apparatus, device, and storage medium for generating inference in the policy domain. During inference generation, the input policy domain question is identified to determine its task type. Based on the task type, a generation constraint mechanism is determined, which includes at least structured generation constraints based on context-free grammar (CFG) and classification output constraints based on closed option sets. A pre-generated semantic direction vector is obtained based on the task type. This semantic direction vector is injected into the hidden layer of a pre-trained inference model. The policy domain question and the generation constraint mechanism are then input into the inference model with the injected semantic direction vector, outputting the inference result corresponding to the policy domain question. By identifying the task type and determining the generation constraint mechanism, which includes structured generation constraints and classification output constraints, and combining this with the injection of the semantic direction vector into the inference model, precise control of the policy inference result format and assurance of content authenticity are achieved. This approach ensures standardized output format and enhances content authenticity. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating one step of the reasoning generation method for the policy domain provided in this application embodiment;
[0019] Figure 2 This is a flowchart illustrating one step of obtaining the task type provided in an embodiment of this application;
[0020] Figure 3 This is a flowchart illustrating one step in obtaining a semantic direction vector according to an embodiment of this application;
[0021] Figure 4 This is another flowchart illustrating the steps for obtaining the semantic direction vector provided in an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of a policy-oriented reasoning generation device provided in an embodiment of this application;
[0023] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0024] Figure 7 This is another structural schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] It should be understood that the steps described in the method embodiments disclosed in this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] In related technologies, when dealing with policy-related issues, it is often impossible to ensure that the output conforms to preset structured format requirements, such as standardized policy citation formats or structured comparative analysis frameworks. This makes it difficult for downstream systems to automatically parse the generated content. Furthermore, the model may fabricate non-existent policy provisions or citations during the generation process due to a lack of effective constraints, creating an external illusion phenomenon. These problems directly weaken the parsability and credibility of the generated content, thereby affecting the overall operational efficiency and decision-making reliability of policy support systems and hindering automated processing.
[0029] To address the technical problems existing in related technologies, this application provides a reasoning generation method for the policy domain. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart illustrating the steps of a policy-oriented reasoning generation method provided in this application embodiment, which includes steps 101 to 105.
[0030] Step 101: Identify the input policy domain problem to obtain the task type of the policy domain problem.
[0031] In one embodiment, when generating reasoning for policy domain questions, the system first receives the policy domain question input by the user, and then performs type identification processing on the policy domain question to obtain the task type corresponding to the received policy domain question.
[0032] For example, the task types include at least policy reference generation, cross-regional policy comparison and analysis, description of clause differences, compliance / applicability judgment, process guidance and interpretation of sensitive policies. For different task types, multiple predefined grammar templates (CFG Schema) can be maintained, and each template corresponds to a specific output format (such as standardized policy reference JSON, clause comparison list, comparison dimension description template, etc.).
[0033] In practical applications, when identifying and processing policy issues to determine their corresponding task types, semantic analysis and feature extraction are performed on the policy issues, followed by appropriate analysis to determine the corresponding task type. Therefore, one can refer to... Figure 2 , Figure 2 This is a flowchart illustrating a step for obtaining a task type according to an embodiment of this application, wherein the step includes steps 201 to 203.
[0034] Step 201: Receive the input policy domain question, and perform text cleaning and feature extraction on the policy domain question to obtain the feature type corresponding to the policy domain question;
[0035] Step 202: Perform semantic recognition processing on policy domain issues to obtain the type recognition results corresponding to the policy domain issues;
[0036] Step 203: Perform type analysis based on feature type and type identification results to obtain the task type corresponding to the policy domain issue.
[0037] Among them, text cleaning and feature extraction refer to removing noise from the input text and extracting key attributes. This can be achieved by using regular expressions to filter special characters and stop words, and by using the TF-IDF algorithm to extract keywords. Semantic recognition can be understood as parsing the deep intent and contextual relationships of a question. This can be achieved by fine-tuning a pre-trained language model such as BERT to generate semantic vectors. Type analysis refers to the comprehensive evaluation by integrating explicit features and implicit semantics. This can be achieved by using decision trees or rule engines to combine feature types and type recognition results for cross-validation.
[0038] Specifically, the proposed solution first performs text cleaning and feature extraction on the input question to obtain structured feature types, providing a quantifiable preliminary basis for task type identification. Based on this, semantic recognition processing is used to deeply analyze the semantic essence of the question, generating type identification results that reflect deeper intentions. Finally, type analysis is performed based on the feature types and type identification results to achieve cross-validation of explicit features and implicit semantics, thereby comprehensively determining the task type. This multi-dimensional identification mechanism ensures the accuracy of task type classification and lays the foundation for the correct matching of subsequent constraint mechanisms.
[0039] In other words, upon receiving a policy domain question as input, the system first performs text cleaning to remove irrelevant characters and noise, and then performs feature extraction operations such as word segmentation, part-of-speech tagging, and named entity recognition to obtain the key elements of the question. Subsequently, a semantic recognition model is used to determine the intended meaning and extract contextual semantic features. Finally, the feature types and semantic recognition results are fused, and a pre-trained classifier is used to determine the task type to which the question belongs, providing a basis for subsequently calling the appropriate grammar template and reasoning logic.
[0040] In practical applications, the processing may include the following steps:
[0041] Step 1: Problem Preprocessing and Feature Extraction
[0042] 1. Text standardization: Clean user input, correct spelling errors, and standardize expression.
[0043] 2. Word segmentation and part-of-speech tagging: Use a policy domain dictionary for word segmentation and identify the part of speech of keywords (such as verbs "compare", "judge", "basis").
[0044] 3. Key Feature Extraction:
[0045] Intent keyword matching: Matches the core words in the question with the preset "intent keyword library".
[0046] Example: Compare / compare → implies "policy comparison"; whether it meets the requirements or whether it can be applied for → implies "compliance assessment".
[0047] Policy Entity Identification: Identify and label named entities in the question, such as policy name, region, institution, clause number, etc. Entity type and quantity are important clues (for example, the appearance of two regional entities is likely a comparison task).
[0048] Step Two: Semantic Understanding and Classification Reasoning
[0049] 1. Semantic Vectorization: Convert the entire question sentence into a high-dimensional semantic vector (sentence embedding) to capture its overall meaning.
[0050] 2. Classification model prediction:
[0051] Input the semantic vector into a pre-trained task classification model (e.g., a classifier fine-tuned based on architectures such as BERT).
[0052] The model has been trained on a large number of labeled policy questions and is able to understand the semantic patterns of various tasks.
[0053] The model outputs the probability that the problem belongs to each preset task type.
[0054] 3. Confidence Assessment: Check the confidence score of the highest probability. If the score is below the threshold (e.g., 0.8), the question is marked as uncertain, which may trigger a clarification process or be downgraded to a general question-and-answer strategy.
[0055] Step 3: Contextual Relationship and Final Ruling
[0056] 1. Context integration: In multi-turn dialogues, combine previous dialogue history to understand the true intent of the current question and ensure the consistency of judgment.
[0057] 2. Feature and Model Result Fusion: The regularized features (such as keywords and entities) extracted in step one are combined with the model prediction probabilities in step two and weighted to make a final decision. For example, even if the model's prediction probability for "policy comparison" is slightly lower, but two clear "regional entities" are identified in the question, the system may still ultimately determine it as a comparison task.
[0058] 3. Output structured task identifiers: Generate the final task type label (e.g., policy_comparison).
[0059] Step 102: Determine the generation constraint mechanism according to the task type. The generation constraint mechanism includes at least structured generation constraints based on context-free grammar (CFG) and classification output constraints based on closed option sets.
[0060] In one embodiment, after determining the task type corresponding to the policy domain issue, the generation constraint mechanism for the current reasoning generation can be determined according to the task type, so that more accurate reasoning results can be output when generating reasoning.
[0061] For example, in the process of generating policy domain reasoning results, after identifying the input policy domain question to determine its task type, a generation constraint mechanism is dynamically selected based on the determined task type. This generation constraint mechanism includes two core forms: structured generation constraints based on context-free grammars (CFG) and classification output constraints based on closed option sets. When the task type requires strictly formatted output, the structured generation constraint based on CFG is activated. This constraint ensures that the generated content conforms to a specific structure through preset grammatical rules, such as mandating that policy citations must include the policy name, clause number, and specific clause information. The content is as follows: When the task type involves classification decisions, a classification output constraint based on a closed option set is adopted. This constraint limits the output to a predefined set of legal options, thereby avoiding the generation of invalid or fictitious content by the model. The set of legal options varies for different tasks. For example, for the "applicability judgment" task, {"applicable", "not applicable"} is loaded; for the "pilot area classification" task, {"core area", "expanded area", "non-pilot area"} is loaded; and for the "compliance judgment" task, {"compliant", "non-compliant", "supplementary materials required"} is loaded. At the same time, the set of legal options can be dynamically expanded and customized.
[0062] In other words, determining the generation constraint mechanism includes: obtaining a preset control strategy matrix, and querying and matching the control strategy matrix based on the task type to obtain configuration information that matches the task type. The control strategy matrix records the association between the task type and the configuration information. The generation constraint mechanism corresponding to the task type is generated based on the configuration blueprint. When the generation constraint mechanism is a structured generation constraint based on the context-free grammar CFG, the corresponding specific context-free posting template is loaded according to the task type. When the generation constraint mechanism is a classification output constraint based on a closed option set, the set of legal options associated with the task type is loaded.
[0063] Among them, the control strategy matrix refers to the data structure that stores the mapping relationship between task types and configuration information. It can be implemented using relational database tables or configuration files (such as JSON format). Its purpose is to provide a queryable global mapping mechanism, avoiding the tediousness and error risk of manual configuration at runtime. Configuration information can be understood as a set of configuration parameters associated with task types. It can include constraint mechanism type identifiers, which serve as the input basis for generating constraint mechanisms. The configuration blueprint specifically refers to the rule template used to dynamically generate constraint mechanisms. The context-free posting template refers to the predefined CFG rule set for task types. It can be implemented using syntax rules described by BNF paradigm, with the purpose of forcing the output to strictly follow the predefined structure. The set of legal options is a list of predefined options associated with task types. It can be implemented using an enumerated value set, with the purpose of limiting the output range to eliminate the possibility of generating invalid content.
[0064] Specifically, during the processing, a dynamic adaptation from task type to generation constraint mechanism is achieved through the linkage mechanism between the control strategy matrix and the configuration blueprint. First, a preset control strategy matrix is obtained, which establishes a global mapping relationship between task type and configuration information. Then, a query and matching operation is performed in the matrix based on the task type, using the task type as a unique index key to ensure the targeted and efficient nature of configuration retrieval. Next, based on the matched configuration information, the corresponding generation constraint mechanism is dynamically generated through the configuration blueprint. Finally, according to the specific type of the generation constraint mechanism, a corresponding context-independent document template or set of legal options is loaded, ensuring that the constraint mechanism highly matches the policy and task requirements, thus forming a complete task-constraint matching process.
[0065] The control strategy matrix can be shown in the following table:
[0066]
[0067] For example, when the task type is policy reference, the system queries and matches the configuration information in the control strategy matrix. This configuration information indicates that the generation constraint mechanism is a structured generation constraint based on context-free grammar (CFG). Subsequently, the constraint mechanism is generated according to the configuration blueprint, and a specific context-free document template for the policy reference task is loaded. This template defines the standard format structure of the policy reference.
[0068] When the task type is policy classification, the system queries and matches the configuration information indicating that the constraint generation mechanism is a classification output constraint based on a closed option set. Then, it loads the set of legal options associated with the policy classification, which contains a list of predefined policy category names.
[0069] When the task type is compliance judgment, the system matches the guided_choice constraint mechanism according to the control strategy matrix and loads three types of limiting options: "compliant", "non-compliant", and "requires supplementary materials" to ensure that the output is strictly limited to the preset compliance conclusion. At the same time, it combines the semantic direction vector v_truth (α=2.2) to strengthen the model's weight tilt towards factual accuracy and avoid ambiguity or subjective judgment.
[0070] Based on the processing method described above, by realizing the automatic matching of task type and generation constraint mechanism, manual configuration errors are avoided and matching efficiency is improved; at the same time, the loading of the task type-driven constraint mechanism ensures the standardization of output format and the compliance of content, effectively preventing format deviation and content illusion, and ensuring the parsability of downstream systems.
[0071] Step 103: Obtain the pre-generated semantic direction vector according to the task type.
[0072] The semantic direction vector is generated through pre-calculation and other processing, and then stored accordingly for later use.
[0073] In one embodiment, to improve the accuracy of subsequent reasoning, after identifying the task type, a pre-generated semantic direction vector is obtained based on the identified task type to guide the semantic orientation of the generated content. For example, in a compliance judgment task, v_truth (α=2.2) is invoked to enhance the reliance on factual evidence; in sensitive consultation, v_safe (α=2.0), v_fact (α=1.8), and v_neutral (α=1.5) are integrated to ensure that the response is safe, objective, and neutral, thereby achieving precise control and risk avoidance of semantic output.
[0074] For example, when obtaining the pre-generated semantic direction vector based on the task type, it is also determined by matching in the control policy matrix. Specifically, determining the semantic direction vector includes: performing a configuration query in the control policy matrix based on the task type to obtain a vector recipe that matches the task type; parsing the vector recipe to obtain the semantic direction vector and intervention intensity parameter corresponding to the task type.
[0075] In other words, during the processing, by deeply integrating the acquisition of semantic direction vectors into the task type management framework, the configuration is first queried in the control strategy matrix according to the task type to obtain vector recipes. This fully utilizes the established association between task types and configuration information in the control strategy matrix. Subsequently, the vector recipes are parsed to obtain semantic direction vectors and intervention intensity parameters, transforming structured configurations into operable parameters. These parameters are used to inject into the hidden layer of the inference model to ensure that the generation process can dynamically adapt to the guidance needs of different task types, thereby strengthening the model's ability to follow the normative format of the policy domain.
[0076] When policy issues involve structured generation tasks, the system queries the control policy matrix for a matching vector recipe based on the task type. This vector recipe can be stored in JSON format. The parser module reads the JSON data and extracts the semantic direction vector (represented as a 128-dimensional floating-point array) and the intervention strength parameter (scalar value 0.5). These parameters are then injected into the hidden layer of the inference model to guide the model in generating outputs that conform to the policy citation format.
[0077] Based on the above processing method, the vector recipe is accurately matched with the task requirements, system redundancy and parameter configuration inconsistencies are reduced, the standardization and reliability of the inference results are improved, and format errors and content illusions are effectively reduced.
[0078] It should be noted that the semantic direction vector is pre-calculated and stored. Obtaining the semantic direction vector includes: acquiring a sample set, inputting the sample set into the inference model, and calculating the semantic direction vector based on the output of the hidden layer in the inference model.
[0079] Reference Figure 3 , Figure 3 This is a flowchart illustrating a step for obtaining a semantic direction vector according to an embodiment of this application, wherein the step includes steps 301 to 304.
[0080] Step 301: Construct a comparison dataset based on the sample set, wherein the comparison dataset includes a positive sample set and a negative sample set;
[0081] Step 302: Input the comparison dataset into the hidden layer of the inference model, and record the hidden state vector of each data in the comparison dataset in all layers;
[0082] Step 303: The hidden state vector is processed based on the positive and negative classification of the samples and the mean calculation to obtain the first state vector corresponding to the positive sample set and the second state vector corresponding to the negative sample set.
[0083] Step 304: Calculate the vector difference between the first state vector and the second state vector, and normalize the vector obtained by the difference calculation to obtain the semantic direction vector.
[0084] Here, the sample set refers to the basic data source collection used to generate semantic direction vectors, which can be implemented using historical question-and-answer records or manually annotated policy document fragments in the policy field; the positive and negative sample sets obtained by classifying the sample sets when comparing datasets, the positive sample set contains correct output examples that conform to policy specifications, and the negative sample set contains erroneous output examples with formatting errors or fictitious content, which can be constructed through rule-based filtering or expert review; the hidden state vector specifically refers to the internal representation vector of the input data in the hidden layer of the inference model, which can be obtained by recording the activation values of the model in a specific layer; the first state vector and the second state vector refer to... The representative vector is obtained by aggregating the hidden state vectors of the positive and negative sample sets; the vector difference calculation can be understood as performing a subtraction operation on the first and second state vectors to quantify the semantic offset direction, which can be implemented using the vector subtraction formula; the pairing process specifically involves matching positive and negative samples one-to-one according to semantic relevance; the hidden state difference refers to the vector difference value output by the data pair in the hidden layer, which can be obtained by calculating the difference between the hidden state vectors of the paired samples; principal component analysis is a statistical method for dimensionality reduction and feature extraction of the hidden state difference, which can be implemented using the PCA algorithm to prioritize the retention of the semantic dimensions that have the most significant impact on policy generation.
[0085] In fact, when obtaining the semantic direction vector, an appropriate method can be selected for processing according to the actual situation and needs. The methods used include the difference mean method and the principal component analysis method. However, both methods involve obtaining a sample set and then using the corresponding inference model to take the sample set as input, and then processing it according to the corresponding calculation and determination methods based on the output of the hidden layer.
[0086] When using the difference mean method, the sample set is first divided into positive and negative subsets, and the hidden state vectors of each subset are obtained through forward propagation of the inference model. Then, the hidden state vectors of each subset at each layer of the inference model are averaged, resulting in the first and second state vectors for the positive and negative sample sets, respectively. Finally, vector subtraction is performed to obtain the semantic direction vector pointing to the direction of semantic enhancement. The resulting vectors are then normalized. This method is computationally simple, highly interpretable, and suitable for scenarios with relatively concentrated sample distributions and clear policy orientations.
[0087] For example, when processing based on the difference mean method, semantic direction is extracted by calculating the average activation difference between the two types of behavioral samples. One type of sample represents the "target behavior" (such as answering a fact-based question), and the other represents the "non-target behavior" (such as hallucinations or fabricated content). The difference between the average hidden states of the two types of samples at a certain layer of the model is the guiding vector.
[0088]
[0089] in, This represents the average activation of the target sample. This represents the average activation of the non-target samples. This vector can be understood as the semantic transfer direction "from illusion to reality".
[0090] Additionally, when obtaining the semantic direction vector, one can also refer to... Figure 4 , Figure 4 This is another flowchart illustrating the steps for obtaining a semantic direction vector provided in this application embodiment, wherein the steps include steps 401 to 404.
[0091] Step 401: Construct a comparison dataset based on the sample set and perform pairing processing to obtain several data pairs. The comparison dataset includes a positive sample set and a negative sample set, with one positive sample paired with one negative sample.
[0092] Step 402: Input the data pairs into the hidden layer of the inference model to obtain the hidden state difference corresponding to the data pairs;
[0093] Step 403: Perform principal component analysis on the hidden state differences and select the hidden state difference with the largest variance as the semantic direction vector.
[0094] When using Principal Component Analysis (PCA), the positive and negative sample sets are first paired to form several semantically related data pairs. Each data pair is then input into the hidden layer of the inference model to obtain its corresponding hidden state difference. Subsequently, PCA is performed on the matrix formed by all hidden state differences, and the direction with the largest variance is extracted as the semantic direction vector. This method effectively captures the dominant change direction in complex semantic spaces, is suitable for scenarios with diverse sample differences and sensitivity to policy adjustments, and improves the robustness and generalization ability of the direction vector.
[0095] The two methods described above can be selected or combined depending on the specific task requirements and data characteristics. The difference-means method is suitable for situations with clear policy guidance and concentrated sample distribution, offering advantages such as high computational efficiency and strong interpretability. Principal component analysis (PCA) is more suitable for handling high-dimensional, complex semantic changes, effectively extracting the most discriminative direction vectors. In practical applications, PCA can be used to first determine the main semantic dimensions, followed by direction calibration using the difference-means method, to balance accuracy and stability.
[0096] Therefore, when obtaining the semantic direction vector, a dual generation mechanism works collaboratively. First, based on a global statistical path, a comparative dataset containing positive and negative sample sets is constructed using the sample set. After inputting it into the inference model, the hidden state vector is recorded. The first and second state vectors are obtained by calculating the mean of the positive and negative sample sets respectively. Then, the semantic direction vector is directly extracted by calculating the vector difference. This path enhances the robustness of the vector by aggregating features of similar samples. Simultaneously, based on a local optimization path, the comparative dataset is paired to form data pairs. After inputting it into the model, the hidden state difference is obtained. Then, principal component analysis is used to select the direction with the largest variance as the semantic direction vector. This path focuses on subtle changes in a single sample pair to improve the sensitivity of the vector. The two paths complement each other. The global path ensures that the vector stably reflects the essence of the task, while the local path accurately captures key semantic dimensions. Together, they enable the semantic direction vector to accurately represent the difference direction between normative outputs and erroneous outputs in the policy domain, thereby providing a quantifiable target for subsequent constraint injection.
[0097] In practical applications, the sample set for generating policy citation formats uses historical policy consultation records. The positive sample set includes policy citation examples that conform to national standards, while the negative sample set includes examples containing fictitious provisions or format errors. After constructing the comparison dataset, the positive and negative sample sets are input into the hidden layer of the inference model, respectively. The hidden state vectors of each data set are recorded, and the mean of the positive sample set is calculated to obtain the first state vector, while the mean of the negative sample set is calculated to obtain the second state vector. The difference between these two vectors is the semantic direction vector. Simultaneously, positive samples can be paired with semantically similar negative samples to form data pairs. After inputting these pairs into the model, the hidden state difference of each pair is calculated, and principal component analysis is used to select the direction with the largest variance as a supplementary semantic direction vector. Finally, the vectors generated by the two paths are fused to constrain the model output.
[0098] When processing data using principal component analysis (PCA), PCA further extracts the most representative directions based on the difference mean. This is achieved by analyzing the activation differences of multiple samples. Principal component analysis was performed, and the principal direction that best distinguishes between target and non-target behaviors was selected as the guiding vector:
[0099]
[0100] This method is more robust with large datasets and diverse samples, avoiding noise caused by differences in a single mean. It can capture multi-dimensional behavioral differences, resulting in a more robust direction vector.
[0101] Furthermore, for semantic direction vectors, the intensity parameter α can control the intervention magnitude. During model inference, direction vectors can be added to intermediate layers of the model to guide the generation behavior. Let the hidden state at position i in layer l be... , where d represents the hidden dimension. The activation intervention involves injecting a direction vector into this layer. And through strength parameters By controlling the magnitude of intervention, the behavior of the model during the generation phase can be altered. The basic formula is as follows:
[0102] ;
[0103] when When the model is in the "target direction," it will be generated, for example, more cautious and more fact-based; while when When this happens, the model will shift in the opposite direction.
[0104] By adjusting the value of α, the semantic bias of the model's output can be flexibly controlled during the generation process, enabling fine-grained regulation in scenarios such as fact-following, creative expression, or safety constraints, thereby improving the quality and controllability of the generated content. This intervention mechanism not only significantly improves the quality of generation but also effectively mitigates behavioral biases of the model in complex tasks, demonstrating good adaptability, especially in open-domain question answering and dialogue systems.
[0105] Step 104: Inject the semantic direction vector into the hidden layer of the pre-trained inference model, and input the policy domain problem and the generation constraint mechanism into the inference model with injected semantic direction vector, and output the inference result corresponding to the policy domain problem. Injecting the semantic direction vector into the hidden layer of the pre-trained inference model means performing vector operation on the output result of the hidden layer based on the semantic direction vector, and using the result of the operation as the input of the next adjacent layer. The hidden layer is several intermediate layers in the inference model.
[0106] In one embodiment, after obtaining the semantic direction vector, the semantic direction vector is injected into a designated hidden layer of the inference model to guide the model to reason along the true semantic direction during the generation process. The injection method can employ vector addition or attention mechanisms to control the process, making the model more inclined to activate paths consistent with the facts when dealing with policy domain issues. Simultaneously, combined with a generation constraint mechanism, the output content is validated in real time to ensure it conforms to the standardized expression of policy provisions. This method effectively suppresses the generation of fictitious content and improves the accuracy and reliability of the answers. Specifically, the hidden layers into which the semantic direction vector is injected are several intermediate layers of the inference model. These can be a continuous number of intermediate layers; for example, if the inference model has 30 layers, the semantic direction vector can be injected into layers 10-26. Alternatively, several non-continuous intermediate layers can be used as hidden layers; for example, if the inference model has 30 layers, the semantic direction vector can be injected into layers 10-15 and 20-26.
[0107] For example, when handling social security policy consultations, the model significantly reduces the tendency to speculate on unfamiliar policy provisions by injecting a "authenticity" semantic direction vector. Combined with keyword matching and logical consistency checks in the constraint mechanism, it ensures that the output content is consistent with official statements.
[0108] Furthermore, when injecting semantic direction vectors into the hidden layer of the inference model, the output of the hidden layer is adjusted by performing vector operations on the output based on the semantic direction vectors to obtain the final output of the hidden layer. The specific details can be found in the basic formula described above. However, in practical applications, there may be multiple vector controls, meaning the task type may have multiple behaviors. In this case, the basic formula can be adjusted during injection to obtain the following formula:
[0109] ;
[0110] in, The hidden state after injecting the semantic direction vector at the i-th position in the l-th layer. This represents the original hidden state at the i-th position of the l-th layer when no semantic vector is injected. Let j be the semantic direction vector. for The corresponding intervention intensity parameter, J, is the total number of semantic direction vectors.
[0111] In practice, multiple semantic direction vectors can be used during inference processing. When multiple behaviors need to be controlled, multiple semantic direction vectors can be determined based on the identified task type. Then, by weighted fusion of the semantic direction vectors, a composite guiding direction is generated and injected into the corresponding hidden layer in the inference model. After weighted fusion of multiple semantic direction vectors, the hidden layer adjusts the output based on the above formula to ensure that the output meets the actual requirements.
[0112] Furthermore, after obtaining the corresponding inference results based on the above processing method, it can be determined whether verification processing is required based on the actual situation to ensure that the output inference results meet the actual needs. Since the inference results are obtained based on the determined generation constraint mechanism, the inference results need to meet the requirements set by the generation constraint mechanism, such as whether the inference results match the corresponding template or whether they conform to the preset logical structure and keyword specifications. If the verification fails, a correction mechanism is triggered to readjust the weights of the semantic direction vector or strengthen the constraint conditions until the output meets the consistency and accuracy requirements of the policy text, ensuring that the final result is applicable in real-world scenarios.
[0113] In practical applications, for structured generation constraints based on Context-Free Grammar (CFG), the inference model performs restricted decoding in guided_grammar mode to ensure that the output strictly follows the selected grammar structure. Subsequently, the inference result is verified by a format validation module to ensure that the generated result fully matches the template definition. If a mismatch is found, the model can be rolled back or resampled to ensure that the referenced fields are complete and the format is standardized, while significantly reducing the risk of illusions in complex structured tasks.
[0114] Furthermore, the semantic direction vector and related intervention intensity parameters can be dynamically updated. Based on actual task requirements and feedback signals, the parameter configuration is continuously optimized using an online learning mechanism, improving the model's adaptability to different semantic objectives. Specifically, this includes: receiving feedback data on the inference results, including user evaluations and automatic evaluations; associating the feedback data with policy domain issues and the semantic direction vector, and storing it in the corresponding feedback dataset; and updating the semantic direction vector and the associated intervention intensity parameters based on the feedback sample set when an update instruction is triggered.
[0115] Feedback data refers to multi-dimensional quality evaluation information of the reasoning results. It can be achieved through user-submitted ratings or error markers via an interactive interface, as well as compliance checks automatically executed by the system based on preset verification rules. This integrates subjective judgments with objective indicators, avoids bias from a single feedback source, and ensures the reliability and representativeness of the data. Association storage refers to establishing a structured mapping relationship between feedback data and policy domain issues and semantic direction vectors. This can be achieved by using a relational database to bind and store feedback data, issue text features, and vector identifiers, preserving the correlation between the issue context and the vector state, and providing a scenario-specific basis for subsequent updates. Update processing refers to iteratively optimizing the semantic direction vector and its intervention intensity parameters. This can be achieved based on the feedback sample set through vector difference calculation or principal component analysis algorithms, to specifically correct historical errors and adapt to new policy scenarios, rather than simply resetting parameters.
[0116] Specifically, the receiving stage synchronously integrates user evaluation and automatic evaluation results to ensure the semantic accuracy and format standardization of the feedback data in the inference results. The associated storage stage binds the feedback data to specific policy domain issues and the currently used semantic direction vector, avoiding noise introduced by generalization adjustments. The update processing stage recalculates the vector parameters based on the feedback sample set when an instruction is triggered, keeping the core components of the constraint generation mechanism synchronized with actual needs, thus forming a closed-loop feedback mechanism and achieving dynamic adaptability of the semantic direction vector. During update processing, the vector can be recalculated based on the aforementioned calculation and processing methods to obtain a new semantic direction vector.
[0117] For example, when a user rates the policy Q&A results and marks them as formatted incorrectly, the system automatically records the feedback data and associates it with the current policy domain question text and the semantic direction vector used. This data is structured and stored in the feedback database. When the system detects that the accumulated feedback samples have reached a preset threshold, it triggers an update process, using positive and negative sample pairs from the feedback samples to recalculate the semantic direction vector, for example by analyzing the principal components of the hidden state differences, thereby optimizing the output quality of the subsequent inference process.
[0118] In summary, the above embodiments provide a reasoning generation method for the policy domain. During reasoning generation, the input policy domain question is identified to determine its task type. Based on the task type, a generation constraint mechanism is determined, which includes at least structured generation constraints based on context-free grammar (CFG) and classification output constraints based on closed option sets. A pre-generated semantic direction vector is obtained based on the task type. This semantic direction vector is injected into the hidden layer of a pre-trained reasoning model. The policy domain question and the generation constraint mechanism are then input into the inference model with the injected semantic direction vector, outputting the reasoning result corresponding to the policy domain question. By identifying the task type and determining the generation constraint mechanism, which includes structured generation constraints and classification output constraints, and combining this with the injection of the semantic direction vector into the reasoning model, precise control of the policy reasoning result format and assurance of content authenticity are achieved. This method has the advantages of ensuring standardized output format and improving content authenticity.
[0119] Based on the method described in the above embodiments, this embodiment will further describe it from the perspective of a policy-oriented reasoning generation device. This policy-oriented reasoning generation device can be implemented as an independent entity or integrated into an electronic device, such as a terminal, which may include a mobile phone, a tablet computer, etc.
[0120] Please see Figure 5 , Figure 5 This is a schematic diagram of a policy-oriented reasoning generation device provided in an embodiment of this application, such as... Figure 5 As shown in the embodiment of this application, the policy-oriented reasoning generation apparatus 500 includes:
[0121] The identification and processing module 501 is used to identify the input policy domain issues and obtain the task type of the policy domain issues;
[0122] The constraint determination module 502 is used to determine the generation constraint mechanism according to the task type, wherein the generation constraint mechanism includes at least structured generation constraints based on context-free grammar CFG and classification output constraints based on closed option sets;
[0123] The vector acquisition module 503 is used to acquire pre-generated semantic direction vectors according to the task type;
[0124] The inference processing module 504 is used to inject semantic direction vectors into the hidden layer of a pre-trained inference model, and input the policy domain problem and the generation constraint mechanism into the inference model with injected semantic direction vectors, and output the inference result corresponding to the policy domain problem. Injecting semantic direction vectors into the hidden layer of the pre-trained inference model involves performing vector operations on the output result of the hidden layer based on the semantic direction vectors, and using the result of the operation as the input of the next adjacent layer. The hidden layer consists of several intermediate layers in the inference model.
[0125] In one embodiment, the identification processing module 501 is further configured to:
[0126] The system receives input policy domain questions and performs text cleaning and feature extraction on the policy domain questions to obtain the feature types corresponding to the policy domain questions.
[0127] Semantic recognition processing is performed on policy-related issues to obtain the type identification results corresponding to the policy-related issues;
[0128] Type analysis is performed based on feature types and type identification results to obtain the task types corresponding to policy domain issues.
[0129] In one embodiment, the constraint determination module 502 is further configured to:
[0130] Obtain the preset control strategy matrix, and perform a query and match in the control strategy matrix based on the task type to obtain the configuration information that matches the task type. The control strategy matrix records the association between the task type and the configuration information.
[0131] The generation constraint mechanism corresponding to the task type is generated based on the configuration blueprint;
[0132] When the constraint generation mechanism is a structured constraint generation mechanism based on context-free grammar CFG, the corresponding specific context-free posting template is loaded according to the task type.
[0133] When generating constraint mechanisms based on categorized output constraints with closed option sets, load the set of valid options associated with the task type.
[0134] In one embodiment, the vector acquisition module 503 is further configured to:
[0135] Based on the task type, a configuration query is performed in the control strategy matrix to obtain a vector recipe that matches the task type;
[0136] The vector formula is parsed to obtain the semantic direction vector and intervention intensity parameter corresponding to the task type.
[0137] In one embodiment, the policy-oriented reasoning generation apparatus 500 further includes a vector preprocessing module for:
[0138] Obtain a sample set and input it into the inference model. Calculate the semantic direction vector based on the output results of all layers in the inference model.
[0139] The vector preprocessing module is also used for:
[0140] A comparison dataset is constructed based on the sample set, which includes a positive sample set and a negative sample set;
[0141] The comparison dataset is input into the hidden layer of the inference model, and the hidden state vector of each data point in the comparison dataset is recorded in each layer.
[0142] The hidden state vectors are classified and mean-calculated based on the positive or negative nature of the samples to obtain the first state vector corresponding to the positive sample set and the second state vector corresponding to the negative sample set.
[0143] The vector difference between the first state vector and the second state vector is calculated, and the vector obtained by the difference calculation is normalized to obtain the semantic direction vector.
[0144] The vector preprocessing module is also used for:
[0145] A comparison dataset is constructed based on the sample set, and several data pairs are obtained by pairing. The comparison dataset includes a positive sample set and a negative sample set, with one positive sample paired with one negative sample.
[0146] The data pairs are input into the hidden layer of the inference model to obtain the hidden state difference corresponding to the data pairs;
[0147] Principal component analysis is performed on the hidden state differences, and the hidden state difference with the largest variance is selected as the semantic direction vector.
[0148] In one embodiment, the inference processing module 504 injects the semantic direction vector into the hidden layer based on the following formula, which includes:
[0149] ;
[0150] in, The hidden state after injecting the semantic direction vector at the i-th position in the l-th layer. This represents the original hidden state at the i-th position of the l-th layer when no semantic vector is injected. Let j be the semantic direction vector. for The corresponding intervention intensity parameter, J, is the total number of semantic direction vectors.
[0151] In one embodiment, the policy-oriented reasoning generation apparatus 500 further includes an update processing module for:
[0152] Receive feedback data on the inference results, including user evaluations and automatic evaluations;
[0153] The feedback data is associated with policy area issues and semantic direction vectors, and stored in the corresponding feedback dataset;
[0154] When an update instruction is triggered, the semantic direction vector and the intervention intensity parameter associated with the semantic direction vector are updated based on the feedback sample set.
[0155] Additionally, please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a mobile terminal such as a smartphone or tablet computer. Figure 6 As shown, the electronic device 600 includes a processor 601 and a memory 602. The processor 601 and the memory 602 are electrically connected.
[0156] The processor 601 is the control center of the electronic device 600. It connects various parts of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 602 and calling the data stored in the memory 602, it performs various functions of the electronic device 600 and processes data, thereby monitoring the electronic device 600 as a whole.
[0157] In this embodiment, the processor 601 in the electronic device 600 loads the instructions corresponding to the processes of one or more applications into the memory 602 according to the following steps, and the processor 601 runs the applications stored in the memory 602, thereby realizing any step in the policy domain-oriented reasoning generation method provided in the above embodiment.
[0158] The electronic device 600 can implement the steps of any embodiment of the policy-oriented reasoning generation method provided in the embodiments of this application. Therefore, it can achieve the beneficial effects that any policy-oriented reasoning generation method provided in the embodiments of this application can achieve. For details, please refer to the previous embodiments, which will not be repeated here.
[0159] Please see Figure 7 , Figure 7 This is another structural schematic diagram of the electronic device provided in the embodiments of this application, such as... Figure 7 As shown, Figure 7A specific structural block diagram of an electronic device provided in an embodiment of this application is shown. This electronic device can be used to implement the policy-oriented reasoning generation method provided in the above embodiments. The electronic device 700 can be a mobile terminal such as a smartphone or a laptop computer.
[0160] RF circuit 710 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 710 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity modules (SIM cards), memory, etc. RF circuit 710 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.
[0161] The memory 720 can be used to store software programs and modules, such as the program instructions / modules corresponding to the policy domain-oriented reasoning generation method in the above embodiment. The processor 780 executes various functional applications and the policy domain-oriented reasoning generation method by running the software programs and modules stored in the memory 720.
[0162] Memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, memory 720 may further include memory remotely located relative to processor 780, which can be connected to electronic device 700 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0163] The input unit 730 can be used to receive uploaded digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 780, and can receive and execute commands sent by the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0164] Display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic device 700. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 740 may include display panel 741, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms. Further, touch-sensitive surface 731 may cover display panel 741. When touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to processor 780 to determine the type of touch event. Subsequently, processor 780 provides corresponding visual output on display panel 741 according to the type of touch event. Although in the figures, touch-sensitive surface 731 and display panel 741 are implemented as two separate components to achieve input and output functions, in some embodiments, touch-sensitive surface 731 and display panel 741 can be integrated to achieve input and output functions.
[0165] The electronic device 700 may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can generate an interruption when the flip is closed or shut down. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the electronic device 700, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0166] Audio circuitry 760, speaker 761, and microphone 762 provide an audio interface between the user and electronic device 700. Audio circuitry 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. Conversely, microphone 762 converts collected sound signals into electrical signals, which are then received by audio circuitry 760, converted back into audio data, and processed by processor 780. The audio data is then transmitted via RF circuitry 710 to, for example, another terminal, or output to memory 720 for further processing. Audio circuitry 760 may also include an earphone jack to facilitate communication between peripheral headphones and electronic device 700.
[0167] Electronic device 700, through transmission module 770 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 770 is shown in the figure, it is understood that it is not an essential component of electronic device 700 and can be omitted as needed without changing the essence of the invention.
[0168] The processor 780 is the control center of the electronic device 700. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 700 by running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, thereby providing overall monitoring of the electronic device. Optionally, the processor 780 may include one or more processing cores; in some embodiments, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.
[0169] The electronic device 700 also includes a power supply 790 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to the processor 780 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 790 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0170] Although not shown, the electronic device 700 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors to implement any step of the policy domain-oriented reasoning generation method provided in the above embodiments.
[0171] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0172] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of this application provide a storage medium storing multiple instructions that, when executed by a processor, can implement any step in the policy-oriented reasoning generation method provided in the above embodiments.
[0173] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0174] Since the instructions stored in the storage medium can execute the steps in any embodiment of the policy-oriented reasoning generation method provided in this application, the beneficial effects that any policy-oriented reasoning generation method provided in this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0175] The foregoing has provided a detailed description of a reasoning generation method, apparatus, electronic device, and storage medium for the policy field provided by embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application. Moreover, those skilled in the art can make several improvements and modifications without departing from the principles of this application, and these improvements and modifications are also considered within the scope of protection of this application.
Claims
1. A policy domain-oriented reasoning generation method, characterized in that, The method comprises the following steps: identifying an input policy field question to obtain a task type of the policy field question; determining a generation constraint mechanism according to the task type, wherein the generation constraint mechanism at least includes a structured generation constraint based on a context-free grammar (CFG) and a classification output constraint based on a closed option set; obtaining a pre-generated semantic direction vector according to the task type, wherein the step of pre-generating the semantic direction vector comprises the following steps: obtaining a sample set, inputting the sample set into an inference model, and calculating a semantic direction vector according to the output results of all layers in the inference model, wherein the semantic direction vector is used to guide the semantic tendency when generating content based on the policy field question; injecting the semantic direction vector into a hidden layer of a pre-trained inference model, inputting the policy field question and the generation constraint mechanism into the inference model with the injected semantic direction vector, and outputting an inference result corresponding to the policy field question, wherein the injecting of the semantic direction vector into the hidden layer of the pre-trained inference model is a vector operation on the output result of the hidden layer based on the semantic direction vector, and the operation result is used as the input of the adjacent next layer, and the hidden layer is a plurality of intermediate layers in the inference model.
2. The policy domain oriented inference generation method of claim 1, wherein, The identification of the input policy field question to obtain the task type of the policy field question comprises the following steps: receiving an input policy field question, performing text cleaning processing and feature extraction on the policy field question, and obtaining a feature type corresponding to the policy field question; performing semantic recognition processing on the policy field question to obtain a type recognition result corresponding to the policy field question; performing type analysis according to the feature type and the type recognition result to obtain a task type corresponding to the policy field question.
3. The policy domain oriented inference generation method of claim 1, wherein, The determination of the generation constraint mechanism according to the task type comprises the following steps: obtaining a pre-set control strategy matrix, and performing query matching in the control strategy matrix based on the task type to obtain configuration information matched with the task type, wherein the control strategy matrix records the association between the task type and the configuration information; generating the generation constraint mechanism corresponding to the task type according to the configuration information; when the generation constraint mechanism is a structured generation constraint based on a context-free grammar (CFG), loading a corresponding specific context-free grammar template according to the task type; when the generation constraint mechanism is a classification output constraint based on a closed option set, loading a legal option set associated with the task type.
4. The policy domain oriented inference generation method of claim 3, wherein, The obtaining of the pre-generated semantic direction vector according to the task type comprises the following steps: performing configuration query in the control strategy matrix according to the task type to obtain a vector formula matched with the task type; performing analysis processing on the vector formula to obtain a semantic direction vector and an intervention intensity parameter corresponding to the task type.
5. The policy domain oriented inference generation method of claim 1, wherein, The method further comprises the following steps: acquire a sample set, input the sample set into the inference model, and calculate a semantic direction vector according to output results of all layers in the inference model; The acquiring of the sample set and the inputting of the sample set into the inference model and the calculating of the semantic direction vector according to the output results of all layers in the inference model comprise: constructing a comparison data set according to the sample set, wherein the comparison data set comprises a positive sample set and a negative sample set; inputting the comparison data set into the hidden layer of the inference model, and recording a hidden state vector of each data in the comparison data set at each layer; classifying and calculating the mean of the hidden state vector based on sample positivity, to obtain a first state vector corresponding to the positive sample set and a second state vector corresponding to the negative sample set; performing vector difference calculation on the first state vector and the second state vector, and performing normalization processing on the difference value to obtain a semantic direction vector; The acquiring of the sample set and the inputting of the sample set into the inference model and the calculating of the semantic direction vector according to the output results of all layers in the inference model comprise: constructing a comparison data set according to the sample set, and performing pairing processing to obtain a plurality of data pairs, wherein the comparison data set comprises a positive sample set and a negative sample set, and one positive sample is paired with one negative sample; inputting the data pairs into the hidden layer of the inference model to obtain hidden state differences corresponding to the data pairs; performing principal component analysis on the hidden state differences, and selecting a hidden state difference with the largest variance as the semantic direction vector.
6. The policy domain oriented inference generation method of claim 1, wherein, The injecting of the semantic direction vector into the hidden layer of the pre-trained inference model is implemented through a preset formula, and the formula comprises: ; wherein, is the hidden state after injecting the semantic direction vector for the i-th position of the l-th layer, is the original hidden state without injecting the semantic vector for the i-th position of the l-th layer, is the j-th semantic direction vector, is the semantic direction vector for the i-th position of the l-th layer, is the corresponding intervention intensity parameter, and J is the total number of semantic direction vectors.
7. The policy domain oriented inference generation method of claim 1, wherein, The method further comprises: receiving feedback data for the inference result, wherein the feedback data comprises user evaluation and automatic evaluation; associating the feedback data with the policy field problem and the semantic direction vector, and storing into a corresponding feedback data set; when it is determined that an update instruction is triggered, updating the semantic direction vector and an intervention intensity parameter associated with the semantic direction vector according to the feedback data set.
8. A policy domain-oriented reasoning generation apparatus characterized by comprising: comprise: an identification processing module configured to identify an input policy field problem to obtain a task type of the policy field problem; a constraint determination module configured to determine a generation constraint mechanism according to the task type, wherein the generation constraint mechanism at least comprises a structured generation constraint based on a context-free grammar (CFG) and a classification output constraint based on a closed option set; a vector acquisition module configured to acquire a pre-generated semantic direction vector according to the task type, wherein the pre-generation of the semantic direction vector comprises: acquiring a sample set, inputting the sample set into an inference model, and calculating a semantic direction vector according to output results of all layers in the inference model, and the semantic direction vector is used to guide the semantic tendency when generating content based on the policy field problem. The inference processing module is configured to inject the semantic direction vector into a hidden layer of a pre-trained inference model, and input the policy field question and the generated constraint mechanism into the inference model to which the semantic direction vector is injected, and output an inference result corresponding to the policy field question.
9. An electronic device, comprising: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the policy field-oriented inference generation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the policy field-oriented inference generation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Machine learning task execution and feedback method and device, equipment and medium
CN120542570A
Civil administration service question and answer method based on large model and knowledge graph retrieval enhancement
CN120780798A