Method for guaranteeing security of large model generation content through alignment mechanism

By integrating security specification inspection and alignment strategies in the inference stage of the big model, predicting and avoiding potential violations, and negotiating the alignment between the model and user needs, the problem of insufficient security risks and flexibility in generating content in the big model is solved, and higher security, reliability and user satisfaction are achieved.

CN120220696AInactive Publication Date: 2025-06-27INSPUR SOFTWARE TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510695414.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art poses security risks when generating content in large models and lacks flexible responses, making it difficult to balance meeting user needs and complying with safety specifications.

Method used

By actively integrating security specification inspection and alignment strategies in the large-scale model inference stage, explicit inference pre-analysis and prediction are adopted to avoid potential violations, and coordinate between the model and user needs through a negotiated alignment strategy to ensure that the final output content complies with security specifications and meets user intentions as much as possible.

Benefits of technology

It significantly improves the security and reliability of large-scale applications, reduces excessive rejection, improves user satisfaction, and enhances interpretability and auditability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220696A_ABST
    Figure CN120220696A_ABST
Patent Text Reader

Abstract

The invention provides a method for guaranteeing the safety of large model generation content through an alignment mechanism, and belongs to the technical field of artificial intelligence and content safeguard.The method comprises the steps that firstly, explicit reasoning analysis is conducted before a user request is answered, pre-stored safety specifications are retrieved to obtain guidance, and the compliance of the user request is judged; for possibly non-compliant requests, a user request or an answer scheme is adjusted through a negotiation type alignment strategy; and then, the large model generates contents conforming to safety specifications, compliance verification is carried out on the generated contents through the safety verification subsystem, and finally a safe answer is output. According to the method and the device, the risk of generating harmful contents by a large model is effectively reduced, the safety and the reliability of content generation are improved, and meanwhile, user requirements and use experience are considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and content security, and particularly to a method for ensuring the security of content generated by large models through an alignment mechanism. Background Art

[0002] With the development of artificial intelligence technology, large models have been widely used in the fields of content generation such as text and images. However, while large models generate rich content, they may also produce harmful information that does not conform to ethical norms or laws and regulations. The spread of such insecure content not only has a negative impact on users but may also lead to legal liabilities and social problems. Therefore, how to effectively restrict the output of large models to make it conform to safety and ethical norms has become a key issue that needs to be solved urgently in current artificial intelligence applications. The existing content security guarantee means mainly include two categories: one is to perform security alignment during the model training stage, such as fine-tuning the output tendency of the model through a harmful content filtering dataset and human feedback reinforcement learning (RLHF); the other is to perform result filtering or rule verification during the model inference stage, such as using keyword masking, blacklist / whitelist rules, or an independent content review model to check the generated results. However, these existing methods still have many deficiencies: the alignment method that simply relies on offline training lacks flexibility and is difficult to update security policies in a timely manner; and only performing passive filtering after generation often cannot prevent risks during the model generation process, and sometimes intercepts after the content has been generated, which may cause delays or even omissions. In addition, simple and rough keyword filtering is prone to false positives or bypasses, and directly refusing to answer will affect the user experience and cannot achieve a balance between meeting user needs and complying with safety regulations. Therefore, there is an urgent need for an improved method that can actively incorporate safety regulations during the content generation process of large models, can dynamically identify and avoid potential insecure outputs, and can negotiate and align with user needs through flexible strategies, so as to provide useful information as much as possible on the premise of ensuring safety. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a method for ensuring the security of content generated by large models through an alignment mechanism, which solves the problem that the content generated by large models in the prior art has security risks and lacks flexible countermeasures. This method actively integrates security specification inspection and alignment strategies during the large model inference stage, anticipates and avoids potential illegal content through explicit pre-inference analysis, and uses a negotiation-based alignment strategy to coordinate between the model and user needs, ensuring that the final output content not only conforms to security specifications but also meets user intentions as much as possible. Compared with the existing methods that only rely on post-filtration or single rejection, the present invention can more actively and efficiently prevent the generation of harmful content, and greatly improve the security and reliability of large model applications.

[0004] The technical solution of the present invention is as follows: A method for ensuring the security of the content generated by a large model through an alignment mechanism, which performs explicit reasoning and analysis before answering the user's request, retrieves the pre-stored security specifications for guidance, and judges the compliance of the user's request; for compliant requests, directly generate an answer; for non-compliant but adjustable requests, adjust the user's request or the answer plan; for non-compliant and non-providable requests, output a rejection or warning message; subsequently, the large model generates content that complies with the security specifications, and the security verification subsystem conducts compliance verification on the generated content, and finally provides the generated secure and compliant answer to the user Specifically, it includes the following steps: Receive the user's question: Obtain the natural language question input by the user; Explicit security reasoning: Before generating an answer, conduct security analysis and reasoning on the user's question, retrieve the pre-stored content security specifications, and judge item by item the semantics and intentions of the user's question based on the security specifications to obtain a reasoning conclusion on whether the question is compliant or not; Alignment decision: Determine the answer strategy according to the reasoning conclusion, specifically including: when the user's question does not violate any security specifications, generate a compliant answer corresponding to the user's question; when the user's question violates the security specifications but there are secure alternative answer plans, adjust the user's question or its answer to generate a compliant answer to meet the user's needs while complying with the specifications; when the user's question does not meet the security specifications and there is no secure answer plan, generate a rejection answer or a warning message; Answer generation and output: Call the large language model to generate the final response content according to the determined answer strategy and output it to the user, where the output content meets the requirements of the security specifications.

[0005] Furthermore, the explicit security reasoning is carried out in a chain reasoning manner. During the process, the large language model generates an interpretable reasoning link, quotes one or more security rule clauses related to the user's question as the basis, and judges the degree of compliance between the request content of the user's question and the security rules.

[0006] Furthermore, there is a security specification storage unit for storing pre-defined content security policies and rule sets. The explicit security reasoning step includes retrieving the specification clauses related to the user's question from the security specification storage unit and using them as the reasoning basis to evaluate the compliance of the user's question.

[0007] Furthermore, when it is detected that the user question involves illegal content, the method further includes: reformulating the user question or partially satisfying the user's needs without violating the security regulations to generate safe answer content; considering the answer adjustment as a constrained optimization problem: minimizing the difference between the answer content and the original answer under the condition of satisfying the security constraints. Suppose the answer initially generated by the large model is , the final answer after alignment adjustment is A. Define the answer deviation measurement function Δ(A, ) to measure the adjusted answer A relative to the initial answer The degree of deviation (e.g., quantified by semantic change, information loss, or modification amplitude). The optimization goal of the negotiated alignment process can be expressed as minimizing Δ under safety constraints, that is: ; The optimization variable A represents the adjustable answer content. is the optimal answer after optimization; constraint condition S(A) θ requires that the final output answer A must meet the security threshold requirement of risk assessment (that is, the content risk value does not exceed θ). The objective function Δ(A, ) describes the relationship between A and The larger the value, the more the adjusted answer deviates from the original answer. On the one hand, the security of output A is strictly guaranteed by constraining S(A) <=θ, and on the other hand, by minimizing Δ(A, ) Minimize the extent of modification to the original answer as much as possible to ensure that the model's answer is both safe and not distorted, and achieve negotiated alignment so that the final answer meets the user's reasonable needs without violating safety regulations.

[0008] Furthermore, the large language model used to implement the method is pre-trained or fine-tuned to include security specification reasoning tasks, so that it can make inferences and judgments based on the input security specification content before generating an answer; the pre-training process explicitly instills security policy knowledge into the model and guides the model to form an alignment strategy of reasoning first and then answering, thereby giving the model the inherent ability to generate content in accordance with security specifications.

[0009] A security verification step is included before outputting the final answer: the generated answer content is reviewed by the security verification subsystem, and the answer is compared and verified with the security specifications to ensure that the output content does not contain any components that violate the security policy; the answer is sent to the user only when it passes the security verification. The security verification subsystem reviews and evaluates the generated answer content through rule matching and / or machine learning models; when it is detected that the answer content contains prohibited information, it triggers content regeneration or a security prompt process to ensure the compliance and security of the final output content.

[0010] The beneficial effects of the present invention are Improved security compliance and robustness: The model performs explicit security specification reasoning before answering, ensuring that its decisions are based on clear security guidelines, eliminating the generation of content that violates policies from the source. Experimental analysis shows that this mechanism enables the model to resist various jailbreak attacks and adversarial prompts, significantly improving the accuracy of identifying bad requests. Even when faced with complex or unseen harmful questions, the model can rationally follow the specifications and not easily produce illegal outputs.

[0011] Reduce excessive rejections and improve user satisfaction: Adopt a negotiated alignment strategy to try to meet users' reasonable needs while ensuring safety. When user requests involve borderline content, the model will not simply reject them, but will adjust the answer strategy to provide alternative safety information or suggestions. This flexible response mechanism effectively reduces excessive rejections of benign requests, making the model's responses both safe and helpful, and improving the user experience.

[0012] Enhanced explainability and auditability: The decision-making process of the model contains an analyzable explicit reasoning chain, which records its consideration and judgment basis for security specifications. Compared with a pure end-to-end black box model, the method of the present invention makes the logic behind each answer more transparent and available for developers or auditors to check. This explainability helps debug and improve the model, and is also very friendly to regulatory compliance: if a content security incident occurs, the reasoning process of the model can be traced back to locate the problem.

[0013] Versatility and flexible expansion: The present invention achieves the decoupling of security knowledge and model reasoning through an independent security specification storage unit and reasoning module. When the security policy is updated or needs to adapt to new application scenarios, only the specification library or rule set needs to be updated, and the model can adapt to the new requirements through explicit reasoning without the need for complete retraining, which greatly improves the flexibility and maintainability of the system. At the same time, this method can be combined with existing model training paradigms (such as SFT, RLHF, etc.) to significantly enhance the security alignment capability while retaining the original effectiveness of the model.

[0014] In summary, the alignment mechanism provided by the present invention greatly improves the effect of large model content security control. The model can strictly comply with the predetermined security specifications and resist malicious attacks, and can also show sufficient flexibility and intelligence in human-computer dialogue to avoid unnecessary rejection. Through the innovative combination of explicit reasoning pre-position and negotiated alignment, the present invention achieves a good balance between security, robustness and usability, and can be widely used in artificial intelligence services that have high requirements for dialogue content security. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the workflow of the present invention. Detailed implementation manners

[0016] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] As Figure 1 shown, the present invention provides a method for ensuring the security of the content generated by a large model through an alignment mechanism. The flowchart shows the whole process from receiving a user's question to the model performing explicit security reasoning and making a decision to generate the final output. First, the model obtains the question input by the user, and then enters the security reasoning stage, retrieves the security specifications and analyzes the compliance of the user's request. Then, the response strategy is determined through a decision branch: for compliant requests, an answer is directly generated; for requests involving inappropriate content but can be adjusted, the security of the answer content is modified; for requests that are non-compliant and cannot be provided, a rejection or warning message is output. Finally, the generated secure and compliant answer is provided to the user.

[0018] The specific steps are as follows: Receiving user input: Receive a content generation request from the user, including the prompt or question input by the user. Analyze the information input by the user to obtain its intention and requirements, and submit the request to the subsequent security check process.

[0019] Pre-analysis of explicit reasoning: Before generating the final answer, trigger the explicit reasoning pre-mechanism. This mechanism calls the large model or its auxiliary module to perform internal reasoning and deduction on the user's request, generates an intermediate chain of thought to analyze the meaning of the request and the possible security issues it may cause. At the same time, the system retrieves the security specification terms or policy guidelines related to this request from the security specification storage and retrieval module and integrates them into the reasoning process. In this step, the large model reviews the user's request according to the retrieved security rules, identifies potential inappropriate requirements, sensitive topics or violation risks, and outputs a preliminary secure and compliant answer plan or processing strategy. Through explicit chain reasoning analysis, the system can clearly evaluate the compliance of the request and provide a decision basis for the subsequent steps.

[0020] Request compliance judgment: Based on the above reasoning analysis results and the extracted security specifications, the user request is judged for compliance. Specifically, the system compares the user's request intent with the security specification requirements to determine whether the request is within the permitted range. If the request content does not violate any security regulations, it is judged as a compliant request; if the request involves potential illegal information (such as a request to generate illegal and harmful content), it is judged as non-compliant. In the judgment process, the analysis conclusions output by the explicit reasoning stage can be combined, such as marking which parts may violate the rules and what rules have been violated. The judgment result will guide the selection of subsequent answer strategies.

[0021] Negotiated alignment adjustment: When it is detected that the user request is non-compliant or there is a security risk, the negotiated alignment strategy is initiated to adjust the request. The negotiated alignment strategy is an interactive or adaptive content guidance mechanism: on the one hand, the system can conduct limited interactive "negotiations" with the user, such as giving security warnings or suggesting that the user modify the wording and scope of the request; on the other hand, the system can also automatically adjust the answer plan by rewriting or partially filtering sensitive requirements in the request to generate an alternative plan that meets user needs as much as possible without violating regulations. For example, for requests involving confidential or harmful information, the system will explain that it cannot be provided directly and propose to provide relevant security public information, or treat the answer adjustment as a constrained optimization problem, and adjust the output expectations while complying with security regulations: that is, minimize the difference between the answer content and the original answer while satisfying security constraints. Suppose the answer initially generated by the large model is , the final answer after alignment adjustment is . Define the answer deviation measurement function Δ(A, ) to measure the adjusted answer A relative to the initial answer The degree of deviation (e.g., quantified by semantic change, information loss, or modification amplitude). The optimization goal of the negotiated alignment process can be expressed as minimizing Δ under safety constraints, that is: ; Among them, the optimization variable A represents the adjustable answer content, is the optimal answer after optimization; constraint S(A) θ requires that the final output answer A must meet the security threshold requirement of risk assessment (that is, the content risk value does not exceed θ). The objective function Δ(A, ) describes the relationship between A and The degree of deviation between them. The larger the value, the more the adjusted answer deviates from the original answer. By minimizing this deviation function, the system ensures that while meeting security requirements, the answer content retains as much valid information and tone style from the original answer as possible to match the user's question intention. In other words, the negotiation-based alignment module achieves a balance between "answer accuracy / completeness" and "content security compliance" through the above optimization formula: on the one hand, it strictly ensures the security of the output A by constraining S(A) θ, and on the other hand, it minimizes Δ(A, ) to minimize the modification to the original answer as much as possible, thus ensuring that the model's answer result is both secure and distortion-free. In actual implementation, the deviation metric function Δ(A, ) can be selected in an appropriate form according to specific requirements. For example, Δ can be defined as the minimum edit distance of content changes ; or calculate the semantic difference between A and based on semantic embeddings. Regardless of the measurement method used, the above optimization objective modeling clearly describes the negotiation-based alignment process: that is, through optimization, the final output meets both security rules and maximally maintains consistency with the original answer. Negotiation-based alignment ensures a balance between security and usability: it not only avoids the generation of illegal content but also provides acceptable alternative answers or suggestions for users. When the request meets the compliance requirements after adjustment, the process proceeds to the next step; if the request still cannot meet the compliance requirements after multiple negotiations, it may ultimately politely refuse to provide unsafe content.

[0022] Content generation: Based on the original compliant request or the compliant request adjusted through negotiation, the large model performs content generation. At this time, under the guidance of the secure and compliant answer plan formed in the explicit reasoning stage, the large model outputs the final answer content. During the generation process, the large model will follow the guidance provided by the secure specification storage and retrieval module, such as avoiding the use of inappropriate words, avoiding taboo topics, or adopting predefined secure response templates, etc., so as to strictly comply with security requirements at the output level. Since the previous explicit reasoning analysis has planned a secure answer idea, the content generation of the model will be based on this idea, significantly reducing the probability of outputting illegal content.

[0023] Security Verification: After the large model generates a preliminary answer, it enters the security verification subsystem for review. The security verification subsystem conducts a comprehensive compliance check and risk assessment on the generated content. The implementation method can be rule-based content scanning, model-based discriminant analysis, or a combination of both. Specifically, the verification subsystem will retrieve the security specification storage module again, extract the rules related to the current answer content, and then compare item by item whether the answer violates any prohibitions, such as containing sensitive information, privacy data, or inappropriate remarks. If the security verification fails (it is detected that the content does not meet the specifications), the system can prevent direct output and feedback the problem to the explicit reasoning pre-mechanism or the negotiation alignment module for secondary adjustment to regenerate content that meets the requirements. Through this closed-loop verification mechanism, it is ensured that even if there are omissions in the previous steps, the content finally output to the user is still secure and compliant. If the verification passes, the answer content is confirmed to be secure.

[0024] Output Result: The final content that has passed the security verification is output to the user. The output stage also includes logging and audit tracking, saving the relevant information of this alignment process (including reasoning analysis and adjustment) for subsequent improvement of the model or providing compliance certification. Thus, the entire process is completed. Through the above steps, what the user obtains is an answer that has undergone multiple security guarantee processes, maximizing the avoidance of bad content.

[0025] Each module cooperates with each other to ensure content security. Among them, the explicit reasoning pre-mechanism is used to conduct chained reasoning and security assessment on the user request before answer generation; the negotiation alignment strategy is used to align security requirements through human-machine collaboration or adaptive adjustment when the request is non-compliant; the security specification storage and retrieval module provides the basis for security policies, including pre-stored content security rules, specification terms, and corresponding processing strategies, which can be dynamically retrieved and called during the reasoning and verification processes; the security verification subsystem is responsible for conducting the final content compliance review after answer generation. The interaction logic of each module is as follows: The explicit reasoning mechanism analyzes the input according to the specifications provided by the retrieval module, and the judgment result triggers the negotiation alignment module to adjust the input or output plan. The adjusted plan guides the content generation module to generate an answer, and the answer is then submitted to the security verification subsystem for review. The rules required for verification are still provided by the specification storage module. In this way, a closed-loop alignment mechanism is formed to ensure that the entire process from request parsing to content output runs under the guidance and constraints of security specifications. The technical solution of the present invention can be implemented in software form (such as a security plug-in module of a conversational AI system), or can be implemented in the form of a system combination in the model architecture (such as a complete content generation system including a rule library and an audit component). Through the above modular design, the update of security specifications and policy adjustment can be independent of the core architecture of the large model, enabling the system to quickly adapt to new content security requirements.

[0026] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for ensuring the security of the content generated by a large model through an alignment mechanism, characterized in that, before answering the user's request, perform explicit reasoning analysis, retrieve the pre-stored security specifications to obtain guidance, and judge the compliance of the user's request; for compliant requests, directly generate an answer; for non-compliant but adjustable requests, adjust the user's request or the answer plan; for non-compliant and non-providable requests, output a rejection or warning message; subsequently, the large model generates content that complies with the security specifications, and the security verification subsystem conducts compliance verification on the generated content, and finally provides the generated secure and compliant answer to the user.

2. The method according to claim 1, characterized in that, specifically includes the following steps: Receive the user's question: Obtain the natural language question input by the user; Explicit security reasoning: Before generating an answer, conduct security analysis and reasoning on the user's question, retrieve the pre-stored content security specifications, and judge item by item the semantics and intentions of the user's question based on the security specifications to obtain an inference conclusion on whether the question is compliant or not; Alignment decision: Determine the answer strategy according to the inference conclusion, including: when the user's question does not violate any security specifications, generate a compliant answer corresponding to the user's question; when the user's question violates the security specifications but there are safe alternative answer plans, adjust the user's question or its answer to generate a compliant answer to meet the user's needs while complying with the specifications; when the user's question does not conform to the security specifications and there is no safe answer plan, generate a rejection answer or a warning message; Answer generation and output: Call the large language model to generate the final response content according to the determined answer strategy and output it to the user, where the output content meets the requirements of the security specifications.

3. The method according to claim 2, characterized in that, the explicit security reasoning is carried out in a chained reasoning manner, and during the process, the large language model generates an interpretable reasoning link, quotes more than one security rule clause related to the user's question as a basis, and judges the degree of compliance of the request content of the user's question with the security rules.

4. The method according to claim 3, characterized in that, set up a security specification storage unit to store the predefined content security policies and rule sets, and the explicit security reasoning step includes retrieving the specification clauses related to the user's question from the security specification storage unit and using them as the basis for reasoning to evaluate the compliance of the user's question.

5. The method according to claim 2, characterized in that, when it is detected that the user's question involves illegal content, without violating the security specifications, rephrase the user's question or partially meet the user's needs to generate a secure answer content; regard the answer adjustment as a constrained optimization problem, that is, minimize the difference between the answer content and the original answer under the condition of meeting the security constraints.

6. The method according to claim 5, characterized in that, Let the answer initially generated by the large model be , and the final answer after alignment adjustment be A; define the answer deviation metric function Δ(A, ) to measure the degree of deviation of the adjusted answer A from the initial answer ; the optimization goal of the negotiation-based alignment process is expressed as minimizing Δ under the security constraints, that is: ; The optimized variable A represents the adjustable response content, is the optimal response after optimization; the constraint condition S(A) <= θ requires that the finally output response A must meet the safety threshold requirement of risk assessment, that is, the content risk value does not exceed θ; the objective function Δ(A, ) then characterizes the deviation degree between A and The larger the value, the more the adjusted response deviates from the original response.

7. The method according to claim 6, characterized in that, Strictly ensure the security of the output A by constraining S(A) <= θ; minimize the modification to the original answer as much as possible by minimizing Δ(A, ) to ensure that the model's answer is both secure and distortion-free, thus achieving negotiated alignment.

8. The method according to claim 1, characterized in that, Large language models are pre-trained or fine-tuned with security specification reasoning tasks, enabling them to make reasoning judgments based on the input security specification content before generating answers; the pre-training process explicitly instills security policy knowledge into the model and guides the model to form an alignment strategy of reasoning first and then answering, thus endowing the model with the inherent ability to generate content in accordance with security specifications.

9. The method according to claim 1, wherein before outputting the final answer, it includes a security verification step: the generated answer content is reviewed by the security verification subsystem, and the answer is compared and verified with the security specifications to ensure that the output content does not contain components that violate the security policy; only when the answer passes the security verification is it sent to the user.

10. The method according to claim 9, wherein the security verification subsystem reviews and evaluates the generated answer content through rule matching and / or machine learning models; when prohibited information is detected in the answer content, a process of regenerating the content or providing a security prompt is triggered to ensure the compliance and security of the final output content.

Citation Information

Patent Citations

  • Double-record quality inspection method and device, computer device and storage medium

    CN109767335A

  • Intelligent customer service question and answer method based on large language model technology

    CN118364084A

  • Enterprise financial knowledge question answering method and system based on large model

    CN119046428A

  • Large language model question and answer rule packaging method, medium and system

    CN119149703A

  • Access control method and device and computer readable medium

    CN119728272A