Large model self-reflection safety knowledge condensation method and system

Through the self-reflection of large-model security knowledge condensation method, the problem of lack of lightweight and highly versatile security knowledge filtering and condensation methods in the existing technology is solved, and effective security knowledge filtering and condensation in the generation process of large language models is realized to ensure the security and versatility of the output without the need for an external knowledge base.

CN120217358APending Publication Date: 2025-06-27PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510176741.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art lacks a lightweight and versatile method that can effectively filter and condense safety knowledge in the generation process of large language models, avoid the introduction of irrelevant information or noise, and eliminate the need for an external knowledge base.

Method used

The security knowledge condensation method of large-scale self-reflection is adopted to generate initial output by receiving prompt words input by users, and self-reflection is carried out to determine whether there are potential defects or security vulnerabilities. If present, iterative optimization is performed until the output meets all safety requirements. Ultimately, the optimized and improved responses and task processes are invested in the safety knowledge condensation mechanism, and the safety knowledge and security concepts of safety specifications are condensed, and stored in the self-generated safety knowledge base.

Benefits of technology

It realizes the construction of a security mechanism based on self-reflection on a large model, ensures the security of the output, and is highly versatile, without the need for an external knowledge base, and has the characteristics of self-iteration optimization and scalability, which improves the versatility and efficiency of security knowledge condensation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217358A_ABST
    Figure CN120217358A_ABST
Patent Text Reader

Abstract

The invention provides a large model self-reflection safety knowledge condensation method and system. The method comprises the steps that S1, initial output is generated; s2, carrying out self-reflection, judging whether the output has potential defects or security vulnerabilities or not, and if yes, carrying out iterative optimization; s3, in the iterative optimization process, firstly, based on defects or security vulnerabilities judged by self-reflection and additional security specifications, improving measures are put forward, and optimization and improvement are carried out; then performing self-reflection again, and repeating continuously until the safety requirement is met; and S4, inputting the final reply and the task process into a safety knowledge condensation mechanism, condensing the safety knowledge and the safety concept of the safety specification, and storing the safety knowledge and the safety concept in a self-generated safety knowledge base in a unified preset format. According to the method, the universality of safety knowledge congealing can be effectively improved, the thinking depth of large model generation can be improved, safety knowledge can be used for the generation process in an extensible mode, the safety requirement is met, and the method does not depend on an external knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a security knowledge condensation solution for large language models, in particular to a method for condensing security knowledge through self-reflection of a large model, and further relates to a system adopting the method for condensing security knowledge through self-reflection of the large model. Background Art

[0002] Currently, large language models (LLMs), also known as large models for short, such as ChatGPT and Claude, have shown great potential in various fields such as reasoning, programming, and scientific research. LLMs are widely adopted in various application scenarios. However, LLMs are not always reliable. They can generate toxic or unsafe content and are vulnerable to "hallucination", resulting in incorrect or unsafe outputs. In artificial intelligence application scenarios such as government and enterprises, data security is of utmost importance. Especially when using the output content of large models to participate in the work process, how to ensure that the text, code, and other content generated by large models do not contain risky content and security vulnerabilities has become one of the key technical issues in the application of large models.

[0003] In June 2024, Tian, a doctoral student from a university who was interning in an enterprise, took advantage of a vulnerability in Hugging Face (HF) to write malicious code into the shared model of his internship company. As a result, the training effect of the model fluctuated, but the team could not verify the reason, affecting the model training task of the team. Later, his internship company sued the intern, requesting the court to order him to compensate the company for the infringement losses and publicly apologize.

[0004] In November 2024, the first real case of "corrupted corpus" of a large model occurred. A user used ChatGPT to program to build an auxiliary trading robot, but the code generated by ChatGPT called a malicious API address and directly provided the private key in plain text to the malicious API for processing. After the code ran, the user's wallet was stolen with a loss of $2.5k. It can be seen that when the model uses the search method for output and does not review the search results, using malicious reference materials may lead to the generation of malicious content.

[0005] With the development of large model applications, various protection measures can be adopted to protect the output and reasoning process of large models, including: 1. Strictly review the user input on the user input side to ensure its compliance and security. When the system detects that the user input contains malicious content, the current session should be immediately terminated to avoid potential security risks.

[0006] 2. Continuously detect and review the content output by the large model, with a focus on detecting whether there are any issues of illegal content, infringement content, and privacy leakage in the output content, to ensure that users can obtain safe, reliable, and legal information.

[0007] 3. Strengthen management on the model side: Monitor the entire process of user interaction with the large model, promptly detect and prevent Prompt injection attacks by malicious users, and avoid guiding the large model to generate illegal content. In the case of sensitive issues, do not answer or use a safe response to prevent the large model from misleading users or spreading non-compliant information.

[0008] 4. For the hallucination problem, adopt an external knowledge base to improve its accuracy during generation, or adopt a method of involving human experts in the review to reduce hallucinations.

[0009] However, in the existing technology, there is a lack of a lightweight and highly versatile RAG security knowledge base that can be used in the generation process. Using an external knowledge base may introduce irrelevant information or noise, thereby misleading the large model. Moreover, the process of the large model obtaining knowledge from lengthy information will also introduce additional overhead. At this time, how to filter and condense the obtained knowledge content is particularly important. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to provide a method for condensing security knowledge through self-reflection of a large model. By adopting a technical solution for condensing security knowledge generated by the large model itself, a security mechanism based on self-reflection can be implemented on any large model, ensuring the security of the output of the large model and having strong versatility; without relying on an external knowledge base, and having the characteristics of self-iterative optimization and scalability. On this basis, a system adopting the method for condensing security knowledge through self-reflection of the large model is further provided.

[0011] To this end, the present invention provides a method for condensing security knowledge through self-reflection of a large model, including the following steps: Step S1, receive the prompt word input by the user, and generate an initial output through the large language model; Step S2, conduct self-reflection through the large language model, and determine whether there are potential defects or security vulnerabilities in the output. If so, jump to Step S3 for iterative optimization; if not, display the output result; Step S3, in the iterative optimization process, first, based on the defects or security vulnerabilities determined by self-reflection, and the security specifications appended to the output, propose improvement measures through the large language model and conduct optimization and improvement; then, conduct self-reflection again after the optimization and improvement. Repeat the process of self-reflection and iterative optimization in this way until the optimized and improved reply meets all security requirements. Use this reply as the final reply, and jump to Step S4; Step S4: Input the final reply and the task process into the safety knowledge refinement mechanism to refine safety knowledge and safety concepts that meet safety specifications, and store them in the self-generated safety knowledge base in a unified preset format.

[0012] A further improvement of the present invention lies in that, in step S2, when judging whether there are potential defects or security vulnerabilities in the output, it is judged whether there is already safety knowledge and safety specifications in the current safety knowledge base. If so, the knowledge base is used to review the output; if not, the large language model is used for judgment.

[0013] A further improvement of the present invention lies in that the implementation process of self-reflection in step S2 is as follows: Use the output and the input of the task to perform a retrieval-augmented generation (RAG) query to obtain relevant safety knowledge. The relevant safety knowledge includes safety specifications for secure coding, safety reply specifications for sensitive issues, and standard replies in the historical records. When applicable safety knowledge is queried, the corresponding safety knowledge is concatenated with the output.

[0014] A further improvement of the present invention lies in that the process of concatenating the corresponding safety knowledge with the output is as follows: After querying applicable safety knowledge, append the safety specification corresponding to the safety knowledge to the output, and jump to step S3 for iterative optimization.

[0015] A further improvement of the present invention lies in that, in step S3, if applicable safety knowledge has been queried, the large language model is required to optimize and improve the initial reply according to the corresponding safety specification, and self-reflection is performed again after the optimization and improvement.

[0016] A further improvement of the present invention lies in that step S4 uses the Milvus vector database to store the self-generated safety knowledge, and the preset storage format is: ID; DATE; CHUNKS["INDEX", "CHUNK_CONTENT"], where ID represents a unique identifier used to distinguish knowledge entries; DATE represents the date of record generation or update; CHUNKS represents a list of knowledge fragments; INDEX represents the index number of the knowledge fragment; CHUNK_CONTENT represents the content of the knowledge fragment.

[0017] A further improvement of the present invention lies in that step S4 includes the following sub-steps: Step S401: For the final reply that meets safety requirements and the text generation task in the corresponding context, according to the background of the task, compare it with the initial reply with defects, and point out that in the process of self-reflection and iterative optimization, the realization of the transformation from existing defects or security vulnerabilities to safety expressions that meet safety specifications is completed to complete the refinement at the safety expression level. Step S402: For the transformation of the identified security expressions and in-depth reflection on the reasons for such transformation, present all the thoughts in the reflection process in the form of Chain of Thought (COT), and resend these thoughts to the large language model. Summarize them in a concise, essential, and noise-free manner, extract security knowledge from them, obtain the deep-seated reasons for the security expressions, and store them in the self-generated security knowledge base to complete the refinement of the security knowledge hierarchy. Step S403: For the extracted security knowledge, extract the security concepts of this security knowledge. The security concepts refer to the core points involved in the security knowledge, including one or more of risk functions such as strcpy(), overflow vulnerabilities, boundary review coding specifications, and toxic substances. When any one security concept is extracted, record another security concept and its corresponding interpretation for use as the core concept related to the subsequent Retrieval-Augmented Generation (RAG) query to complete the refinement of the security concept hierarchy and accelerate the speed of self-reflection and Retrieval-Augmented Generation (RAG) query.

[0018] A further improvement of the present invention is that in step S401, the security expressions include the repair code corresponding to the generation of secure code and / or the modified noun or sentence expression corresponding to compliance with security specifications.

[0019] A further improvement of the present invention is that in step S402, the process of summarizing in a concise, essential, and noise-free manner includes: reviewing for overflow vulnerabilities before code output and paying attention to the use of the strcpy() risk function; or, replacing the strcpy() risk function with the strcpy_s() secure function for such security repair; or, not answering if the generated content involves the configuration process information of toxic substances.

[0020] The present invention also provides a system for refining security knowledge through self-reflection of a large model, which adopts the method for refining security knowledge through self-reflection of a large model as described above, and includes: An initial reply generation module that receives the prompt words input by the user and generates an initial output through the large language model; A self-reflection module that conducts self-reflection through the large language model and determines whether there are potential defects or security vulnerabilities in the output. If so, it jumps to the iterative optimization module for iterative optimization; if not, it displays the output result. Iterative optimization module. During the iterative optimization process, first, based on the defects or security vulnerabilities identified through self-reflection, as well as the security specifications appended to the output, improvement measures are proposed through the large language model and optimization and improvement are carried out; then, self-reflection is carried out again after the optimization and improvement, and in this way, the process of self-reflection and iterative optimization is continuously repeated until the optimized and improved response meets all security requirements. This response is used as the final response and jumps to the security knowledge refinement module; Security knowledge refinement module. The final response and the task process are input into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications and store them in the self-generated security knowledge base in a unified preset format.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: after generating the initial output through the large language model, self-reflection is first carried out, and it is judged whether there are potential defects or security vulnerabilities in the output. If so, iterative optimization is carried out; during the iterative optimization process, first, based on the defects or security vulnerabilities identified through self-reflection, as well as the security specifications appended to the output, improvement measures are proposed through the large language model and optimization and improvement are carried out; then, self-reflection is carried out again after the optimization and improvement, and in this way, the process of self-reflection and iterative optimization is continuously repeated until the optimized and improved response meets all security requirements; finally, the optimized and improved response and the task process are input into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications and store them in the self-generated security knowledge base in a unified preset format, thereby forming a complete self-reflection mechanism.

[0022] Therefore, the present invention can effectively improve the generality of security knowledge refinement, is applicable to various large models, and can achieve the self-optimizing security gain effect without fine-tuning training, achieving the purpose of being ready to use out of the box; on this basis, the thinking depth of the large model generation is also improved through the self-reflection mechanism including the self-reflection and iterative optimization process, and the security knowledge can be expandably used in the generation process to ensure that the content such as the text and code generated by the large model can meet the security requirements, and the security coding, security specifications, security knowledge, and security concepts therein are refined to form a self-generated security knowledge base; it does not rely on an external knowledge base, effectively avoids the risk of introducing irrelevant information or noise, and has the characteristics of self-iterative optimization and scalability, providing a better foundation for data security. Description of the Drawings

[0023] Figure 1 It is a schematic flow diagram of the self-reflection mechanism according to an embodiment of the present invention; Figure 2 It is a schematic work flow diagram according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the security knowledge refinement mechanism according to an embodiment of the present invention; Figure 4 is a system architecture diagram of a related technology of the present invention; Figure 5 is a system architecture diagram of another related technology of the present invention; Figure 6 is a model architecture diagram of another related technology of the present invention. Specific Embodiments

[0024] In the description of the present invention, if it involves "several", its meaning is more than one; if it involves "multiple", its meaning is more than two; if it involves "greater than", "less than", "exceeding", it should be understood as not including the present number; if it involves "above", "below", "within", it should be understood as including the present number. If it involves "first", "second", etc., it should be understood as only used for the distinction of the same or similar technical feature names, and cannot be understood as implying / specifying the relative importance of technical features, cannot be understood as implying / specifying the quantity of technical features, nor can it be understood as implying / specifying the sequence relationship of technical features.

[0025] Before describing in detail the preferred embodiments of the present invention, the related technologies are described first. Among them, the abbreviations and key term definitions involved include: LLM, which refers to Large Language Model, that is, a large language model, also known as a large model, is an artificial intelligence model based on deep learning technology, with a huge number of parameters and powerful natural language processing capabilities. Transformer, a machine learning model based on the Self-Attention mechanism proposed in the paper "Attention is All You Need". Self-reflection, which refers to the ability of a large language model to evaluate and improve its own generated output, aiming to improve the reliability of the model. Knowledge Condensation, the ability of a large language model to refine, compress and integrate massive and complex knowledge information to form a more structured, efficient and accurate knowledge representation; RAG, which refers to Retrieval-augmented Generation, that is, retrieval-enhanced generation.

[0026] A technical solution related to the present invention is an interactive reflection method. Ziwei Ji et al. proposed an interactive reflection method, which combines knowledge acquisition and answer generation. The factuality, consistency and connotation of the generated answers are steadily improved through the feedback process. Utilizing the interactivity and multitasking capabilities of LLMs, more accurate answers are gradually generated.

[0027] This technical solution focuses on the medical field, involving uncommon professional concepts and potential social risks, making the challenges posed by hallucinations particularly critical because inaccurate or misleading information can have serious consequences for patient care. Its system designs an iterative and introspective process that utilizes the multi-round interactivity and multi-tasking capabilities of LLMs. Self-reflection methods are involved, first generating relevant background knowledge for a given question and then conducting a factual assessment. Once a discrepancy is detected, the model is urged to self-correct, leveraging its inherent ability to refine knowledge. This cyclic process repeats until a satisfactory level of factuality is achieved. During the answer stage, a similar generate-score-refine strategy is adopted to ensure the consistency between the generated answer and the background knowledge. Additionally, an entailment assessment is conducted between the answer and the question. If the generated answer does not meet the criteria, the process returns to the initial stage and repeats the cycle.

[0028] Specifically, as Figure 4 shown, its system includes three loops: the factual knowledge acquisition loop, the knowledge consistency answer loop, and the question entailment answer loop.

[0029] (1) Factual knowledge acquisition loop: First, the model generates background knowledge based on the provided question. This step utilizes the inherent ability of LLMs to understand context. Then, a scorer is used to conduct a factual assessment of the generated knowledge. If the factual score is below the threshold set in the assessment stage, the model is required to self-reflect and is asked to "please refine the knowledge to improve its truthfulness".

[0030] This generate-score-refine strategy is repeated interactively until the generated knowledge reaches a satisfactory level of factuality. This iterative process promotes the dynamic and iterative interaction between the system and the knowledge it generates. And it ensures that the model gradually refines the generated background knowledge, integrating it with established facts.

[0031] (2) Knowledge consistency answer loop: The model continues to generate an answer based on the provided question and the generated knowledge. If the consistency score of the generated answer is below the threshold, the system prompts the model to introspect and self-correct, asking it to "please optimize the response to improve its consistency.", repeating this generate-score-fine strategy until the generated answer is consistent. This iterative process ensures that the model gradually refines the generated answer based on the vetted background knowledge, thus maintaining its integrity.

[0032] (3) Question-Answer Implication Loop: Evaluate the implication relationship of the generated answer through the similarity score to determine whether there is a reasonable implication relationship between this answer and the background knowledge, expected semantics, etc. involved in the entire task, so as to ensure the rationality of the answer and that it is indeed recognizable and can effectively respond to the question (i.e., has answerability). If it does not meet the corresponding implication requirements, the entire process has to be looped again to optimize the answer.

[0033] The main work of this interactive reflection method is to reduce hallucinations in the generative QA scenario in the medical field. This technical solution is still in its early stage and not ready for direct practical deployment in the security field or when the model lacks professional knowledge, especially in complex or ambiguous application scenarios. Moreover, this technical solution completely relies on the judgment and scoring of the large model itself, has extremely high requirements for the capabilities of the large model itself, has great generality and usage limitations, and cannot harvest knowledge from historical records to improve its capabilities.

[0034] Another technical solution related to the present invention is the collaborative model. As Figure 5 and Figure 6 shown, Dongze Hao et al. proposed two collaborative models: the knowledge condensation model and the knowledge reasoning model. First, utilize the multi-modal perception and reasoning capabilities of the vision-language model to extract concise knowledge concepts from the retrieved long paragraphs to ensure relevance to the visual content and the question. Second, utilize the text understanding capabilities of the large language model to summarize and condense the paragraphs into the knowledge essence that helps answer the question. Then integrate these two types of condensed knowledge into the knowledge reasoning model, which wisely browses the merged information to draw a conclusive answer.

[0035] The work focuses on knowledge-based visual question answering (KB-VQA), which requires the model to utilize external knowledge to understand and answer questions based on visual content. Recent research retrieves knowledge paragraphs from external knowledge bases and then uses them to answer questions. However, these retrieved knowledge paragraphs usually contain irrelevant or noisy information. For example, convert the image into visual context and send them together with the question and the retrieved knowledge paragraphs to the LLM to generate an answer. Since the retrieved knowledge paragraphs contain a lot of noisy information, it will mislead the model to predict the wrong answer.

[0036] The method proposed by the collaborative model consists of two models: the knowledge condensation model and the knowledge reasoning model.

[0037] The knowledge condensation model uses BLIP as the vision-language model and the open-source large language model Vicuna as the knowledge concentrator to extract useful information from the retrieved knowledge. The image is first input into the image encoder to extract visual features, and then the visual features and text embeddings are input into the LLM to generate text. The retrieved knowledge paragraphs are condensed into knowledge concepts and the essence of knowledge.

[0038] The knowledge reasoning model, after obtaining the condensed knowledge concepts and the essence of knowledge, uses an encoder-decoder architecture to reason about this knowledge to predict the answer. Two types of knowledge reasoning methods are used to generate the answer. For the concatenated knowledge pattern, the visual context, question, knowledge concepts, and essence, etc. are concatenated into a sentence, and then the sentence is input into a series of encoder layers to jointly encode this text information. Then the embedding vector is obtained and passed to a series of decoder layers to generate the answer. For the concatenated embedding pattern, the visual context and question are connected to different types of knowledge as different sentences. These sentences are input into a series of encoder layers to encode different information separately. Then these embeddings are connected together and passed to the decoder.

[0039] The work of this collaborative model focuses on the knowledge-based visual question answering (KB-VQA) scenario, aiming to improve the performance of the model by condensing and retaining more effective information. However, it still highly depends on an additional knowledge base and correctly obtaining knowledge, and due to the lack of in-depth thinking and self-reflection of the model, situations may occur where the knowledge condensation model converts all knowledge paragraphs into useless information.

[0040] Therefore, the above two related technical solutions cannot meet the actual application requirements of high output security, strong generality, and no need for an external knowledge base.

[0041] Next, in conjunction with the accompanying drawings, a more optimal embodiment of the present invention will be further described in detail.

[0042] As Figure 1 and Figure 2 shown, this embodiment provides a method for safe knowledge refinement with self-reflection of a large model, including the following steps: Step S1, receive the prompt word input by the user and generate an initial output through the large language model; Step S2, conduct self-reflection through the large language model and determine whether there are potential defects or security vulnerabilities in the output. If so, jump to step S3 for iterative optimization. If not, display the output result; Step S3, during the iterative optimization process, first, based on the defects or security vulnerabilities identified through self-reflection, and the security specifications appended to the output, propose improvement measures through the large language model and conduct optimization improvements; then, conduct self-reflection again after the optimization improvements, and repeat the self-reflection and iterative optimization process in this way until the optimized and improved response meets all the security requirements determined by the large model during the self-reflection and iterative optimization process, and use this as the final response, then jump to Step S4; Step S4, input the final response and the task process into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications, and store them in the self-generated security knowledge base in a unified preset format.

[0043] It should be noted that although there are also measures / schemes for the security scenarios of large model applications in the existing technologies, most of the current existing schemes focus on the fine-tuning and security alignment of specific models, or the filtering mechanism on the input and output sides. Such existing schemes cannot achieve their general performance for various large models.

[0044] Different from the existing technologies, the present embodiment provides a security mechanism technical solution that is universal for various large models, does not require fine-tuning training, and can self-optimize, achieving the ability of out-of-the-box use.

[0045] In this embodiment, after generating the initial output through the large language model, first conduct self-reflection, that is, reflect on the security scenarios, security vulnerabilities, and security knowledge in the information generated by the large model itself; and judge whether there are potential defects or security vulnerabilities in the output, and if so, conduct iterative optimization; during the iterative optimization process, first, based on the defects or security vulnerabilities identified through self-reflection, and the security specifications appended to the output, propose improvement measures through the large language model and conduct optimization improvements; then, conduct self-reflection again after the optimization improvements, and repeat the self-reflection and iterative optimization process in this way until the optimized and improved response meets all the specified security requirements; all the specified security requirements refer to all the security requirements determined by the large model during the self-reflection and iterative optimization process. Finally, input the optimized and improved response and the task process into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications, refine the security coding, security specifications, security knowledge, and security concepts, etc., to form a self-generated security knowledge base of the large model. This security knowledge base is also called the RAG knowledge base or the RAG security knowledge base, which can be continuously expanded, and is stored in the self-generated security knowledge base in a unified preset format, thus forming a complete self-reflection mechanism. In subsequent generation tasks, use the security knowledge therein for appending to improve the security of the large model output.

[0046] To further enhance the security of large model applications, in this embodiment, security knowledge is added during the generation process, and a self-reflection mechanism is adopted to further enhance the reliability and robustness of the generated content, and the information content generated by reflection is condensed and summarized. Moreover, by adopting a method for condensing security knowledge generated by the large model itself to generate security knowledge, a security mechanism that can think and construct itself can be implemented on any large model without relying on an external knowledge base. Therefore, this embodiment can effectively improve the generality of security knowledge condensation, is applicable to various large models, and can achieve a self-optimizing security gain effect without fine-tuning training, achieving the purpose of being ready to use out of the box.

[0047] On this basis, this embodiment also improves the thinking depth of large model generation through a self-reflection mechanism including a self-reflection and iterative optimization process, can extend the use of security knowledge in the generation process, ensure that the content such as text and code generated by the large model can meet security requirements, and condense out security coding, security specifications, security knowledge, security concepts, etc. therein to form a self-generated security knowledge base, which does not depend on an external knowledge base and has the characteristics of self-iterative optimization and scalability, providing a better foundation for data security.

[0048] Therefore, the method and system for condensing security knowledge in this embodiment can provide a better foundation for application scenarios that pay special attention to data security in artificial intelligence application scenarios such as governments and enterprises. It has practical significance for ensuring the security and reliability of the content such as text and code generated by large models, can effectively avoid outputting content with risks or security vulnerabilities, has strong generality and does not depend on an external knowledge base, and effectively achieves the effect of being ready to use out of the box.

[0049] Step S1 described in this embodiment is used to implement the generation of an initial reply. In step S1, the user inputs a prompt word Prompt, and the generation model of the large language model will generate an initial output, also called an initial reply. Although this initial output usually meets the basic requirements in the input, there are still some security problems or non-compliances that need to be further improved. The generation model can select existing commercial large models, such as GPT-4o or Claude-3.5-Sonnet; or open-source large models, such as Llama3.3 and Qwen 2.5, etc. Any of the above large models, whether closed-source or open-source, can be used as the base for this embodiment.

[0050] Step S2 in this embodiment is used to achieve self-reflection. In step S2, a large model is used to self-reflect and judge / determine whether there are any potential defects and / or security vulnerabilities in the output. When judging whether there are potential defects or security vulnerabilities in the output, it is judged whether there is already security knowledge and security specifications in the current security knowledge base. If so, the knowledge base is used to review the output; if not, the judgment is made through a large language model.

[0051] If the output has no potential defects and security vulnerabilities, the output result will be normally displayed. If a defect or a situation not in line with the security specifications (there is a security vulnerability) is found, it will transition to an iterative optimization process, that is, step S3.

[0052] The implementation process of self-reflection in step S2 of this embodiment is as follows: Use the output and the input of the task to perform a Retrieval-Augmented Generation (RAG) query to obtain relevant security knowledge. The relevant security knowledge includes security specifications for secure coding, security response specifications for sensitive issues, and standard responses in the historical records, etc.; when applicable security knowledge is queried, the corresponding security knowledge is concatenated with the output (initial output). For example, when generating C++ code, if the Retrieval-Augmented Generation (RAG) query retrieves security specifications related to the request, the security specifications that should be complied with are appended after the initial output. It should be noted that in this step S2 of this embodiment, no modification is made, but it enters the subsequent process for iterative optimization. After that, whether relevant knowledge is obtained from the Retrieval-Augmented Generation (RAG) query or not, it will enter the next iterative optimization process.

[0053] The process of concatenating the corresponding security knowledge with the output in this embodiment is as follows: After querying applicable security knowledge, the security specifications corresponding to this security knowledge are appended after the output, and it jumps to step S3 for iterative optimization.

[0054] Step S3 in this embodiment is used to achieve iterative optimization. Based on the defects or security vulnerabilities judged through self-reflection, and the security specifications appended after the initial output, the large language models (LLMs) will propose evaluations and measures for potential improvement directions and make modifications. For example, when generating C++ code, if the Retrieval-Augmented Generation (RAG) query retrieves security specifications related to the request and they have been appended after the initial output (i.e., the initial response), the large model will be required to modify the initial response according to this specification to ensure that the new version of the output no longer has defects. Then, after the modification, it will enter the self-reflection operation process again, using the improved response as the content to be judged, and reflect again whether there are any potential defects and / or security vulnerabilities; that is, continuously repeat the process of self-reflection → iterative optimization → self-reflection... until the improved response meets all security requirements.

[0055] Therefore, in step S3 of this embodiment, if applicable security knowledge has been queried, the large language model is required to optimize and improve the initial response according to the corresponding security specifications, and perform self-reflection again after the optimization and improvement. For the case where this instance is entered for the first time and there is no refined knowledge base yet, when retrieving and enhancing generation (RAG) search cannot query applicable security knowledge, only the capabilities of the large model itself will be used to reflect on and optimize the content at this time. As the system is used more, the system will have a self-generated security knowledge base obtained through security knowledge refinement, and historical experience can be reused for new generation tasks. Therefore, the security of the system will continue to improve as it is used. The large model will continuously repeat the process of self-reflection and iterative optimization until the improved response meets all specified security requirements. At this time, the improved final response and task process will be input into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications and store them in the security knowledge base in a unified format.

[0056] In step S4 of this embodiment, it is preferable to use the Milvus vector database to store the self-generated security knowledge, and the preset storage format is: ID; DATE; CHUNKS["INDEX", "CHUNK_CONTENT"], where ID represents the unique identifier used to distinguish knowledge entries; DATE represents the date of record generation or update; CHUNKS represents the list of knowledge fragments; INDEX represents the index number of the knowledge fragment; CHUNK_CONTENT represents the content of the knowledge fragment. Of course, the preset format described in this embodiment refers to the default pre-set unified format to facilitate the expansion of the self-generated security knowledge base. In actual applications, this preset format can be adjusted according to actual situations and requirements.

[0057] Through the self-reflection mechanism of this embodiment, the unsafe parts can be repaired in the generation task. The optimized response not only meets the initial functional requirements but also significantly improves its security and standardization. And each iterative optimization is based on the results of the previous iterative optimization, and thus can dynamically update the self-generated security knowledge base by combining insights from previous outputs. This continuously evolving self-generated security knowledge base enables the system to more effectively identify and address potential defects and security vulnerabilities over time.

[0058] Step S4 of this embodiment is used to implement the security knowledge refinement mechanism based on the self-reflection of the large model. In the iterative optimization process (also known as the iterative improvement process), for the case where the improved final response meets all specified security requirements, the improved final response and task process will be input into the security knowledge refinement mechanism to refine the security knowledge and security concepts of the security specifications.

[0059] It should be noted that the safety knowledge refinement mechanism described in this embodiment will use a large model for progressive reflection, and in combination with the COT (Chain of Thought) process, perform progressive refinement from the safety expression level to the safety knowledge level and then to the safety concept level.

[0060] Specifically, as Figure 3 shown, step S4 described in this embodiment includes steps S401 to S403.

[0061] In step S401, first, for the final reply that meets the safety requirements and the text generation task in the corresponding context, according to the background of the task, for the already safe final reply and the targeted question, compare it with the defective initial reply, and point out that in the process of self-reflection and iterative optimization, the realization of the defective or security vulnerability is transformed into the safety expression corresponding to the safety specification, that is, point out from which expressions in this self-reflection and iterative optimization process the transformation from defective to safe is achieved, so as to complete the refinement of the safety expression level. For example, which lines of code are repaired to achieve safe code generation, or which nouns or sentence patterns are modified to achieve an expression that conforms to the safety specification. Thus, the refinement of the first level, that is, the safety expression level, is completed.

[0062] Therefore, in step S401 described in this embodiment, the safety expression includes the repaired code corresponding to the realization of safe code generation and / or the modified nouns or sentence patterns corresponding to the realization of compliance with the safety specification. In step S401, the input is the final reply that meets the safety requirements and the text generation task in the corresponding context, and the output is the refinement result of the safety expression level.

[0063] In step S402, for the identified transformations of multiple security statements and an in-depth reflection on the reasons for such transformations, that is, the input is the transformation of the security statement condensed in step S401 and its reasons; then, all the thinking contents of the reflection process are presented in the form of Chain of Thought (COT), and these thinking contents are sent back to the large language model for summarization in a concise, essential, and noise-free manner, from which security knowledge is extracted to obtain the deep-seated reasons for the security statement. For example, security review knowledge such as "Before code output, it is necessary to review for overflow vulnerabilities, especially pay attention to the use of risk functions such as strcpy()", or security repair knowledge such as "For the strcpy() risk function, it should be replaced with the safer strcpy_s() function", or security specification knowledge such as "Do not answer the configuration process information of toxic substances involved in the generated content". These concise and clear security knowledge will be refined from the context thinking process and stored in the self-generated security knowledge base to complete the condensation of the security knowledge level; that is, the condensation result of the security knowledge level is output and stored in the self-generated security knowledge base.

[0064] Therefore, in step S402 of this embodiment, the process of summarization in a concise, essential, and noise-free manner includes, but is not limited to: security review knowledge such as reviewing for overflow vulnerabilities before code output and paying attention to the use of the strcpy() function; or, security repair knowledge such as replacing the strcpy() risk function with the strcpy_s() security function; or, security specification knowledge such as not answering the configuration process information of toxic substances involved in the generated content.

[0065] Finally, in step S403, for the extracted security knowledge (that is, the input is the security knowledge extracted in step S402), further extract its security concepts from the security knowledge. The security concept refers to the core points involved in the security knowledge, including one or more of risk functions such as strcpy(), overflow vulnerabilities, boundary review, and toxic substances; when any security concept is extracted, another record of the security concept and its corresponding interpretation is made, which is used as the core concept related to the subsequent Retrieval-Augmented Generation (RAG) query to complete the condensation of the security concept level and accelerate the speed of self-reflection and Retrieval-Augmented Generation (RAG) query.

[0066] In this embodiment, if the transition of the security statement cannot be recognized in step S401, the security statement will be extracted only from the final reply at this time; or if the security knowledge cannot be extracted in step S402, the security statement will be text summarized at this time, and the generated text summary will be quality evaluated and only high-quality summaries will be recorded, such as only recording text summaries that exceed a preset threshold. The preset threshold refers to the quality evaluation threshold set in advance, and this preset threshold can be set and adjusted according to the actual situation and requirements; or if the security concept cannot be extracted in step S403, the system will record an error message and issue a warning, such as recording the occurrence time, location, and relevant context information corresponding to the error message for subsequent troubleshooting and analysis, and by performing word segmentation on the text, only the keywords will be recorded after the word segmentation, and common words (such as common words like "of", "is", "in", etc.) will be filtered.

[0067] This embodiment also provides a system for refining security knowledge through self-reflection of a large model, which adopts the method for refining security knowledge through self-reflection of a large model as described above, and includes: An initial reply generation module that receives the prompt words input by the user and generates an initial output through a large language model; A self-reflection module that performs self-reflection through a large language model and determines whether there are potential defects or security vulnerabilities in the output. If so, it jumps to the iterative optimization module for iterative optimization. If not, it displays the output result; An iterative optimization module. During the iterative optimization process, first, based on the defects or security vulnerabilities determined by self-reflection, and the security specifications appended to the output, improvement measures are proposed and optimized through a large language model; then, self-reflection is performed again after the optimization improvement. In this way, the self-reflection and iterative optimization processes are continuously repeated until the reply after optimization improvement meets all specified security requirements. This reply is used as the final reply and jumps to the security knowledge refinement module; A security knowledge refinement module that inputs the final reply and the task process into the security knowledge refinement mechanism, refines the security knowledge and security concepts of the security specifications, and stores them in the self-generated security knowledge base in a unified preset format. In summary, this embodiment proposes a method and system for refining security knowledge through self-reflection of a large model, which can bring general, non-fine-tuning training required, and self-iterative optimization security gain effects to various large models, realizing the ability to be used out of the box. Moreover, this embodiment effectively improves the thinking depth of the large model generated through the self-reflection mechanism, and can expandably use security knowledge in the generation process to ensure that the text, code, etc. generated by the large model no longer contain potential defects (such as risky content) and security vulnerabilities.

[0068] On this basis, this embodiment also conducts self-reflection on security scenarios, security vulnerabilities, and security knowledge in the information generated by the large model itself, and conducts progressive refinement and summary on the information content generated by self-reflection, and condenses security coding, security specifications, security knowledge, and security concepts, etc. to form a self-generated and continuously expandable security knowledge base (also known as the RAG knowledge base) of the large model. In subsequent generation tasks, the security knowledge and security specifications therein are used for addition, so as to improve the security of the large model output. At the same time, it no longer relies on external knowledge bases to reduce the limitations on application scenarios and effectively avoid the risk of introducing irrelevant information or noise.

[0069] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for condensing safety knowledge through self-reflection of a large model, characterized in that: The following steps are involved: Step S1, receiving a prompt word input by a user and generating an initial output through a large language model; Step S2, self-reflection is performed through the large language model, and whether the output has potential defects or security vulnerabilities is determined. If so, jump to step S3 for iterative optimization, otherwise the output result is displayed; Step S3, in the iterative optimization process, based on the defects or security vulnerabilities determined by self-reflection and the security specifications attached to the output, improvement measures are proposed and optimized through the large language model; then, self-reflection is performed again after the optimization and improvement, and the self-reflection and iterative optimization process is repeated in this way until the optimized and improved response meets all security requirements, and this response is used as the final response, and the process jumps to step S4; In step S4, the final response and task flow are put into the security knowledge condensation mechanism to condense the security knowledge and security concepts of the security specifications, and are stored in a self-generated security knowledge base in a unified preset format.

2. The method for condensing safety knowledge through self-reflection of a large model according to claim 1 is characterized in that: In step S2, when judging whether the output has potential defects or security vulnerabilities, it is judged whether there is already security knowledge and security specifications in the current security knowledge base. If so, the output is reviewed using the knowledge base; if not, the judgment is made through the large language model.

3. The method for condensing safety knowledge through self-reflection of a large model according to claim 1 or 2, characterized in that: The implementation process of the self-reflection in step S2 is: using the output and the input of the task to perform retrieval enhancement to generate RAG queries, and obtain relevant security knowledge, the relevant security knowledge includes security specifications for security coding, security response specifications for sensitive issues, and standard responses in historical records; When applicable security knowledge is found, the corresponding security knowledge is concatenated with the output.

4. The method for condensing safety knowledge through self-reflection of a large model according to claim 3 is characterized in that: The process of splicing the corresponding security knowledge with the output is as follows: after the applicable security knowledge is queried, the security specification corresponding to the security knowledge is appended to the output, and then jump to step S3 for iterative optimization.

5. The method for condensing safety knowledge by self-reflection of a large model according to claim 3 is characterized in that: In step S3, if applicable security knowledge has been found, the large language model is required to optimize and improve the initial response according to the corresponding security specifications, and to perform self-reflection again after the optimization and improvement.

6. The method for condensing safety knowledge by self-reflection of a large model according to claim 1 or 2, characterized in that: The step S4 uses the Milvus vector database to store self-generated security knowledge, and the preset storage format is: ID; DATE; CHUNKS["INDEX", "CHUNK_CONTENT"], where ID represents a unique identifier used to distinguish knowledge entries; DATE represents the date when the record is generated or updated; CHUNKS represents a list of knowledge fragments; INDEX represents the index number of the knowledge fragment; CHUNK_CONTENT represents the content of the knowledge fragment.

7. The method for condensing safety knowledge through self-reflection of a large model according to claim 1 or 2, characterized in that: The step S4 comprises the following sub-steps: Step S401: For the final response that meets the security requirements and the corresponding text generation task in the context, compare the initial response with defects according to the task background, point out that in the process of self-reflection and iterative optimization, the defects or security vulnerabilities are transformed into security statements that meet the security specifications, so as to complete the condensation of the security statement level; Step S402: In-depth reflection is conducted on the changes in the identified security expressions and the reasons for such changes. All the thinking contents of the reflection process are presented in the form of a thinking chain COT, and these thinking contents are sent back to the large language model to be summarized in a concise, essential and noise-free manner, from which security knowledge is extracted, the deep-seated reasons for the security expressions are obtained, and the results are stored in the self-generated security knowledge base to complete the condensation of the security knowledge level; Step S403, extracting the security concepts of the security knowledge obtained through extraction, wherein the security concepts refer to the core points involved in the security knowledge, including one or more of risk functions such as strcpy(), overflow vulnerabilities, boundary review coding specifications, and toxic substances; when any security concept is extracted, another security concept and its corresponding interpretation will be recorded as the core concept related to the subsequent retrieval enhancement generation of RAG queries, so as to complete the condensation of the security concept level and speed up the self-reflection and retrieval enhancement generation of RAG queries.

8. The method for condensing safety knowledge through self-reflection of a large model according to claim 7 is characterized in that: In the step S401, the security expression includes a repair code corresponding to the generation of the security code, and / or a modified noun or sentence expression corresponding to the compliance with the security specification.

9. The method for condensing safety knowledge through self-reflection of a large model according to claim 7 is characterized in that: In step S402, the process of summarizing in a concise, essential and noise-free manner includes: reviewing overflow vulnerabilities before code output and paying attention to the use of risk functions such as strcpy(); or, replacing the strcpy() risk function with the strcpy_s() safe function for security repair; or, not answering the configuration process information involving toxic substances in the generated content.

10. A system for condensing safety knowledge through self-reflection of a large model, characterized in that: The method for condensing safety knowledge by self-reflection of a large model as claimed in any one of claims 1 to 9 is adopted, and includes: The initial response generation module receives the prompt word input by the user and generates the initial output through the large language model; The self-reflection module performs self-reflection through the large language model and determines whether the output has potential defects or security vulnerabilities. If so, it jumps to the iterative optimization module for iterative optimization. If not, it displays the output results. Iterative optimization module: In the iterative optimization process, based on the defects or security vulnerabilities identified by self-reflection and the security specifications attached to the output, improvement measures are proposed and optimized through the large language model. Then, after the optimization and improvement, self-reflection is performed again. In this way, the self-reflection and iterative optimization process is repeated continuously until the optimized and improved response meets all security requirements. This response is used as the final response, and the module jumps to the security knowledge condensation module. The safety knowledge condensation module puts the final response and task process into the safety knowledge condensation mechanism, condenses the safety knowledge and safety concepts of safety specifications, and stores them in a self-generated safety knowledge base in a unified preset format.

Citation Information

Cited By

  • Automatic heuristic algorithm planning method based on large language model

    CN120832937A

  • Automatic heuristic algorithm planning method based on large language model

    CN120832937B