Rust language system security enhancement method and device based on large language model

By applying large language models and Miri tests in the Rust language system, combined with adaptive rollback mechanism and knowledge base support, the undefined behavior problems caused by the interaction between secure Rust and unsafe Rust in the Rust language system are solved, and the security and development efficiency of the Rust system are improved.

CN120068077AActive Publication Date: 2025-05-30NAT UNIV OF DEFENSE TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411954495.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing technology has limitations in semantic understanding, complex scenario reasoning and dynamic problem repair, which leads to huge challenges in improving the security of the Rust language system.

Method used

The Rust language system security enhancement method based on large language model is adopted, and the code structure is optimized and security is improved through Miri testing and large language model-driven assertions and modifications, combined with adaptive rollback mechanism and knowledge base support.

Benefits of technology

It effectively solves the undefined behavior problems caused by the interaction between secure Rust and unsafe Rust in the Rust language, reduces the uncertainty of unsafe Rust, and improves the overall security, reliability and development efficiency of the Rust system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068077A_ABST
    Figure CN120068077A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for enhancing the security of a Rust language system based on a large language model, and the method comprises the steps: testing an input Rust system code snippet, adding assertion and modification to the Rust system code snippet in combination with the large language model and a cue word if the test is not passed, and carrying out the test again; after repeated iteration for multiple times, if the test is not passed, rolling back to the optimal code with the least error frequency; on the basis of an abstract syntax tree AST and a knowledge base enhanced cue word extracted by a Miri test, continuing iteration to add assertion and modify; and if the test is passed, testing the semantic acceptability of the code snippets of the Rust system. The method aims at solving the system safety problem introduced by undefined behaviors generated by interaction of safe Rusts and unsafe Rusts in the Rust language, the uncertainty of the unsafe Rusts is reduced, and therefore the overall safety, reliability and development efficiency of a Rust system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer operating systems, and particularly to a method and device for enhancing the security of the Rust language system based on large language models. Background Art

[0002] In a computing system, the security of the operating system is of utmost importance. The Rust language provides a new paradigm for building a secure operating system with comprehensive protection. Code written solely in safe Rust has no memory safety issues. However, even though the Rust language has significant advantages in terms of security, there are still potential risk areas, namely the unsafe Rust parts. These operations are inevitable in system development, thus reducing the stability of the operating system based on the Rust language. Therefore, there are still huge challenges in the aspect of its security enhancement technology.

[0003] Firstly, the unique design philosophy of the Rust language, including the ownership mechanism, borrowing rules, and lifetimes, enables developers to explicitly manage memory and reduce common memory leakage and data race problems in traditional programming languages. However, this design also increases the learning curve of the language and may lead to misuse of the unsafe Rust parts. Although the types of unsafe Rust operations are limited (five unsafe operations), minor semantic changes can significantly affect the way the code is modified, resulting in the fact that similar undefined behaviors may not produce comparable solutions. In this case, developers must deeply understand the semantics of the code and design reliable verification mechanisms. Therefore, understanding semantics and having professional knowledge are crucial for solving Rust security problems.

[0004] Secondly, the current methods for enhancing the security of the Rust language mainly rely on code analysis tools, static checkers, and runtime validators to discover and fix potential problems in unsafe code. These tools usually combine rule templates or automated models to detect common undefined behavior patterns and generate repair suggestions. However, these methods often fail to fully capture the complex semantic relationships and context dependencies implicit in the code. For example, for dependencies on underlying hardware interactions, these tools may be difficult to provide accurate and efficient solutions, requiring engineers to apply professional knowledge along a strict logical reasoning, judgment, and verification process, where the efficiency of the solution is closely related to experience and professional knowledge. Therefore, it is necessary to explore more efficient, flexible, and intelligent technical means to reduce the dependence on manual intervention by engineers and improve the accuracy and speed of problem discovery and repair.

[0005] Finally, the powerful self-reasoning and semantic understanding capabilities of large language models (LLMs) provide new opportunities for enhancing the security of Rust programs. LLMs can learn complex syntax and semantic patterns from vast amounts of code data, assisting developers in more efficiently identifying potential risk points, generating repair suggestions, and optimizing code structures. However, due to high training costs and the lack of high-quality data in specific domains, traditional LLMs struggle to deeply understand complex relational reasoning when dealing with professional languages such as Rust. Additionally, LLMs may encounter the "hallucination" problem when generating suggestions, i.e., outputting content that does not conform to the actual semantics or context, further limiting the effectiveness of their direct application. Therefore, optimizing the architecture and training methods of LLMs for the Rust domain, combined with domain knowledge and semantic constraint mechanisms, has become a new direction for enhancing Rust security.

[0006] Generally speaking, due to the limitations of existing technologies in semantic understanding, complex scenario reasoning, and dynamic problem repair, the overall improvement of Rust language security still faces significant challenges. Summary of the Invention

[0007] The technical problem to be solved by the present invention: In view of the above problems of the prior art, a method and device for enhancing the security of a Rust language system based on a large language model are provided. The present invention aims to solve the system security problems introduced by undefined behaviors resulting from the interaction between safe Rust and unsafe Rust in the Rust language from a systematic perspective, reduce the uncertainty of unsafe Rust, and improve the overall security, reliability, and development efficiency of the Rust system.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for enhancing the security of a Rust language system based on a large language model, comprising the following steps: S1, obtaining an input Rust system code snippet; S2, performing Miri testing on the Rust system code snippet. If the test passes, jump to step S8; otherwise, jump to step S3; S3, adding assertions and modifications to the Rust system code snippet in combination with a large language model and prompting words; S4, performing Miri testing on the Rust system code snippet after adding assertions and modifications. If the test passes, jump to step S8; otherwise, jump to step S5; S5, determining whether the iteration maximum value has been reached. If the iteration maximum value has not been reached, jump to step S3 to continue the iteration; otherwise, jump to step S6; S6. Select the Rust system code snippet with the fewest error occurrences among the original Rust system code snippet, the Rust system code snippet with added assertions, and the modified Rust system code snippet as the best code, and roll back the Rust system code snippet to the best code; S7. Based on the abstract syntax tree (AST) extracted from the Miri test and the knowledge base enhanced prompt words, jump to step S3; S8. Test the semantic acceptability of the Rust system code snippet.

[0009] Optionally, step S2 includes: S2.1. Conduct a Miri test on the Rust system code snippet to determine whether there is an error message of undefined behavior in the Rust system code snippet. If there is no error message of undefined behavior, it is determined that the test passes and jump to step S8; otherwise, jump to step S2.2; S2.2. Determine the unsafe Rust type that causes undefined behavior. If the unsafe Rust type that causes undefined behavior is the target unsafe Rust type, it is determined that the test fails and jump to step S3. The target unsafe Rust types include five types: dereferencing a raw pointer, calling an unsafe function or method, accessing or modifying a mutable static variable, implementing an unsafe trait, and accessing union fields; otherwise, it is determined that the test passes and jump to step S8.

[0010] Optionally, step S3 includes: S3.1. Combine the large language model and prompt words to add assertions to the Rust code that causes undefined behavior of the target unsafe Rust type in the Rust system code snippet for automatic repair; S3.2. Combine the large language model and prompt words to modify the Rust code that causes undefined behavior of the target unsafe Rust type in the Rust system code snippet to provide an alternative safe implementation or modify it to reduce undefined behavior in the Rust code without compromising the code semantics.

[0011] Optionally, the prompt words used in step S3.1 include: " ... The following is the error message... {log}, the idea of adding assertions is as follows: 1. Add an assertion before the undefined behavior becomes possible to prevent it from occurring; 2. This is important and necessary; Only add assertions without adjusting or deleting other code modules; " Among them, "..." represents the omitted content, {log} is the log of the Miri test, and " " represents emphasis; Optionally, the prompt words used in step S3.2 include: " Undefined behavior occurred when running the Miri test according to the code. The following are the ideas for fixing it: 1. Identify undefined behavior; 2. The code for adding assertions cannot be modified as there is a problem with the logic itself; 3. To maintain the functionality and semantics of the source code and avoid significantly changing the logical structure; 4. Design a safe alternative. Refactor the code according to the safe alternative to ensure that the modified code not only avoids undefined behavior but also maintains the original functional logic and performance standards; …Here is the code…{code}" where "…" represents the omitted content, {code} is the modified Rust system code snippet, and " " represents emphasis.

[0012] Optionally, the abstract syntax tree AST and knowledge base enhancement prompt words extracted based on the Miri test in step S7 include: S7.1. Generate an abstract syntax tree AST for the Rust system code snippet that rolls back to the best code. The abstract syntax tree AST is a tree - shaped data structure representing the program's syntax structure, where each node represents a syntax element and the edges represent the parent - child relationships between nodes, connecting the hierarchical structure between syntax elements; S7.2. Prune and simplify the abstract syntax tree AST using a pruning function; S7.3. Calculate the similarity between the pruned and simplified abstract syntax tree AST and the correct abstract syntax tree AST stored in the knowledge base. If there is no correct abstract syntax tree AST with a similarity exceeding the preset threshold, directly end and jump to step S3; otherwise, use the correct abstract syntax tree AST with the best similarity as the target correct abstract syntax tree AST; S7.4. Extract knowledge from the target correct abstract syntax tree AST, including: inputting the pruned and simplified abstract syntax tree AST and the target correct abstract syntax tree AST into a pre - trained large - language model to obtain repair suggestions for the unsafe Rust types and their repair paths for undefined behavior in the pruned and simplified abstract syntax tree AST. The repair suggestions include indicating which nodes need to be adjusted, which edges need to be replaced, and which structures need to add specific assertions to ensure code security. The large - language model has established a mapping relationship between the abstract syntax tree AST extracted from the Miri test, the correct abstract syntax tree AST stored in the knowledge base, and the repair suggestions for the unsafe Rust types and their repair paths for undefined behavior in the abstract syntax tree AST extracted from the Miri test; S7.5 enhances the prompt with the knowledge extracted from the target correct Abstract Syntax Tree (AST), including adding the unsafe Rust types with undefined behavior and their repair suggestions for the repair paths to the prompt to achieve the enhancement of the prompt.

[0013] Optionally, before step S7.4, it also includes: collecting the Abstract Syntax Trees (ASTs) extracted from Miri tests, the correct ASTs stored in the knowledge base, and the repair methods for the unsafe Rust types with undefined behavior and their repair paths in the ASTs extracted from Miri tests, and constructing a training dataset; using the training dataset to train a large language model to establish the mapping relationship between the ASTs extracted from Miri tests, the correct ASTs stored in the knowledge base, and the repair suggestions for the unsafe Rust types with undefined behavior and their repair paths in the ASTs extracted from Miri tests.

[0014] In addition, the present invention also provides a Rust language system security enhancement device based on a large language model, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the Rust language system security enhancement method based on the large language model.

[0015] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the Rust language system security enhancement method based on the large language model through a processor.

[0016] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the Rust language system security enhancement method based on the large language model through a processor.

[0017] Compared with the prior art, the present invention mainly has the following advantages: aiming at the system security problem introduced by the undefined behavior generated by the interaction between safe Rust and unsafe Rust in the Rust language, the present invention constructs a Rust language system security enhancement method driven by large language model (LLM) inference. Through security enhancement based on domain knowledge, semantic reasoning and automated verification, combined with repair optimized by an adaptive rollback mechanism, knowledge base support and large language model (LLM), it can effectively solve the system security problem introduced by the undefined behavior generated by the interaction between safe Rust and unsafe Rust in the Rust language, reduce the uncertainty of unsafe Rust, and improve the overall security, reliability and development efficiency of the Rust system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of the basic process of the method according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the two-stage repair process in an embodiment of the present invention.

[0020] Figure 3 This is a comparison schematic diagram of the adaptive rollback mechanism (b) and the rollback mechanism (a) in an embodiment of the present invention.

[0021] Figure 4 This is a schematic diagram of the process based on AST and knowledge base enhanced prompt words in an embodiment of the present invention.

[0022] Figure 5 This is the percentage of different types of error repairs passing the Miri test in an embodiment of the present invention.

[0023] Figure 6 This is the percentage of different types of error repairs passing semantic acceptability in an embodiment of the present invention.

[0024] Figure 7 This is the difference ratio of different types of error repairs passing the Miri test and semantic acceptability in an embodiment of the present invention.

[0025] Specific implementation manners In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] As Figure 1 shown, the method for enhancing the security of the Rust language system based on the large language model in this embodiment includes the following steps: S1. Obtain the input Rust system code snippet; S2. Perform a Miri test on the Rust system code snippet. If the test passes, jump to step S8; otherwise, jump to step S3; S3. Combine the large language model and prompt words to add assertions and modifications to the Rust system code snippet; S4. Perform a Miri test on the Rust system code snippet after adding assertions and modifications. If the test passes, jump to step S8; otherwise, jump to step S5; S5. Determine whether the iteration maximum value is reached. If the iteration maximum value has not been reached, jump to step S3 to continue the iteration; otherwise, jump to step S6; S6. Select the Rust system code snippet with the fewest error counts among the original Rust system code snippet, the Rust system code snippet with added assertions, and the modified Rust system code snippet as the best code, and roll back the Rust system code snippet to the best code; S7. Based on the abstract syntax tree (AST) extracted from the Miri test and the knowledge base enhanced prompt words, jump to step S3; S8. Test the semantic acceptability of the Rust system code snippet.

[0027] The Miri compiler is a tool specifically used to interpret the mid-level intermediate representation (MIR) of Rust. It can execute Rust programs by interpreting this intermediate representation to perform static and dynamic analysis to detect error messages of undefined behavior (UB) in Rust programs. These error messages include potential problems triggered during code execution, memory access errors, errors related to ownership or lifetime, etc. Based on these logs, the system can identify the code paths and locations that may cause UB. Step S2 in this embodiment includes: S2.1. Conduct a Miri test on the Rust system code snippet to determine whether there are error messages of undefined behavior in the Rust system code snippet. If there are no error messages of undefined behavior, it is determined that the test passes and jumps to step S8; otherwise, it jumps to step S2.2; S2.2. Determine the unsafe Rust types that cause undefined behavior. If the unsafe Rust type that causes undefined behavior is the target unsafe Rust type, it is determined that the test fails and jumps to step S3. The target unsafe Rust types include dereferencing raw pointers, calling unsafe functions or methods, accessing or modifying mutable static variables, implementing unsafe traits, and accessing union fields; otherwise, it is determined that the test passes and jumps to step S8.

[0028] For the Rust language, unsafe Rust is limited and only exists in five categories: dereferencing raw pointers; calling unsafe functions or methods; accessing or modifying mutable static variables; implementing unsafe traits, and accessing union fields. The limited types of undefined behavior caused by unsafe Rust can be further divided into types that can be replaced by safe Rust APIs and types that require semantic modification. A key point in this embodiment lies in the two-stage repair design driven by the large language model (LLM) in step S3. Specifically, as Figure 2 shown, step S3 in this embodiment includes: S3.1. Combine the large language model and the prompt words to automatically fix the Rust code that causes undefined behavior of the target unsafe Rust type in the Rust system code snippet by adding assertions. S3.2. Combine the large language model and the prompt words to modify the Rust code that causes undefined behavior of the target unsafe Rust type in the Rust system code snippet to provide an alternative safe implementation or modify it to reduce the undefined behavior in the Rust code without compromising the code semantics. For the Rust language, unsafe Rust is limited and only exists in five categories: dereferencing raw pointers; calling unsafe functions or methods; accessing or modifying mutable static variables; implementing unsafe traits, and accessing union fields. The limited types of undefined behavior caused by unsafe Rust can be further divided into types that can be replaced by safe Rust APIs and types that require semantic modification.

[0029] In this embodiment, the prompt words used in step S3.1 include: " ... The following is the error message... {log}, the idea of adding assertions is as follows: 1. Add assertions before the undefined behavior becomes possible to prevent it from occurring; 2. This is important and necessary; Only add assertions without adjusting or deleting other code modules; " Among them, "..." represents the omitted content, {log} is the log of the Miri test, " " represents emphasis. Each time step S3.1 is iteratively executed, the large language model will first try to add assertions at appropriate places to prevent undefined behavior, ensuring that any potential undefined behavior can be captured at runtime before it spreads. The advantage of this method is that it allows potential undefined behavior to be detected and identified in a timely manner during code execution, preventing them from causing more serious system errors or security vulnerabilities. The corresponding prompt word engineering is as shown in the above part. However, in some complex scenarios, the undefined behavior involves deeper logical errors or dependencies on external states, and assertions alone cannot directly solve these problems. To solve this problem, a modification stage is introduced, in which it adjusts the error semantics in the Rust code to ensure that the function of the code remains unchanged while effectively reducing the risk of undefined behavior. These suggestions provide alternative safe implementations or modifications, enabling developers to maintain the correctness of the code without compromising security.

[0030] In this embodiment, the prompt words used in step S3.2 include: " Undefined behavior occurred when running the Miri test on the code. The repair idea is as follows: 1. Identify undefined behavior; 2. The code with added assertions cannot be modified as there are problems with the logic itself; 3. To maintain the functionality and semantics of the source code, avoid significantly changing the logical structure; 4. Design a safe alternative. Refactor the code according to the safe alternative to ensure that the modified code not only avoids undefined behavior but also maintains the original functional logic and performance standards; … Here is the code …{code}” where “…” represents the omitted content, {code} is the modified Rust system code snippet, “ ” indicates emphasis.

[0031] When using a large language model framework to fix undefined behavior in Rust code, most of the errors will be significantly reduced. However, there are also some errors that increase. This phenomenon is called the “large model hallucination” because the fixes generated by the model inadvertently introduce new problems, and the errors accumulate over time. As shown in Figure 3 (a), without a rollback mechanism, the process will enter the next stage from the final state of the assertion stage, exacerbating the impact of the hallucination. Current solutions usually involve a direct rollback mechanism (rolling back to the initial state, resetting the code to its initial state to eliminate the accumulated errors, thus reducing the spread of incorrect fixes. However, this method sometimes causes valuable partial corrections to be lost during the iteration. For example, during the repair process, the number of errors increases during the iteration, but the overall trend may show a fluctuating decline. This behavior is analogous to humans often providing frequent and random errors during the reasoning process, just like large language models, but still being able to solve challenging problems in the end. Therefore, this embodiment believes that large language models have a certain self - correction ability, allowing them to gradually converge to the correct answer through multiple iterations.

[0032] To both avoid the large model hallucination and utilize the adaptive problem - solving ability of large language models, this embodiment introduces an adaptive rollback mechanism. Specifically, in step S6 of this embodiment, the Rust system code snippet with the fewest number of errors is selected as the best code from the original Rust system code snippet, the code with added assertions, and the modified Rust system code snippet, and the Rust system code snippet is rolled back to the best code, which is to introduce an adaptive rollback mechanism in the embodiment. As shown in Figure 3As shown in (b) of , this adaptive rollback mechanism does not simply roll back to the initial state, but instead intelligently selects an intermediate state based on the observed error reduction pattern, ensuring a more effective and refined correction process. After each phase ends, the system rolls back to the best code state before entering the next phase, where the best code refers to the state with the fewest detected errors. For example, during the assertion phase, two ideas and are generated, resulting in the sequence . Then the system rolls back from this sequence to the best code state with the fewest errors and then enters the next modification phase. This ensures that subsequent iterations are based on the most refined and stable version of the code, minimizing error propagation and improving the accuracy of the overall repair.

[0033] To enhance the correction process and support more specialized Rust code repair, this embodiment integrates a knowledge base centered on Abstract Syntax Trees (AST) and semantic analysis.

[0034] See Figure 4, in this embodiment, in order to enhance the correction process and support more professional Rust code repair, this embodiment integrates a knowledge base centered on Abstract Syntax Trees (ASTs) and semantic analysis. The specific implementation process of this part includes: First, generate an AST tree from the Rust system code snippet and perform pruning optimization on it. Then, compare the similarity with the database. If there is a similarity, extract the knowledge base content to modify the prompt. Finally, make a more accurate modification to the code. First, parse the Rust system code snippet to generate the corresponding abstract syntax tree. The abstract syntax tree (AST) represents each syntax element (such as variables, functions, operators, etc.) and its hierarchical structure in the source code in a tree structure, which can clearly reflect the syntax and structural characteristics of the code. This embodiment constructs a prompt engineering. However, in order to better extract effective information, this article prunes the output AST text to retain the part related to code security and function repair. For the Rust language, all unsafe Rust operations are marked with the "unsafe" keyword. Therefore, by locating the keyword, prune the AST to retain the code block related to "unsafe". Although this method cannot accurately find the exact problem statement, it greatly reduces the scope of code that needs to be analyzed and reduces the interference of irrelevant code to the large language model (LLM), thereby improving the efficiency and accuracy of the repair process. Second, compare the similarity between the incorrect AST and the existing ASTs in the database. This embodiment constructs a database containing a large number of historical code snippets and their repair records. The system calculates the similarity between the AST of the current code and the existing ASTs in the database to identify code snippets with similar error types and structures. If the system detects a high-similarity match, it will automatically extract the relevant repair knowledge base content. These knowledge base contents include common error types, repair strategies, and repair suggestions for specific problems. Finally, combine the incorrect AST structure and the correct AST structure extracted from the knowledge base with the prompt in the modification stage as the training data for few-shot learning to guide the large language model to generate a more accurate repair strategy. Specifically, in step S7 of this embodiment, the AST and the knowledge base enhanced prompt extracted based on the Miri test include: S7.1, Generate an abstract syntax tree (AST) for the Rust system code snippet that rolls back to the best code using a large language model. The AST is a tree-shaped data structure that represents the program's syntax structure, where each node represents a syntax element and the edges represent the parent-child relationship between the nodes, connecting the hierarchical structure between the syntax elements; S7.2, Use a pruning function to simplify the AST by pruning; S7.3. Calculate the similarity between the pruned and simplified Abstract Syntax Tree (AST) and the correct ASTs stored in the knowledge base (stored in SQL in this embodiment). If there is no correct AST with a similarity exceeding the preset threshold, directly end and jump to step S3; otherwise, use the correct AST with the best similarity as the target correct AST. S7.4. Extract knowledge from the target correct AST, including: inputting the pruned and simplified AST and the target correct AST into a pre-trained large language model to obtain repair suggestions for the unsafe Rust types and their repair paths for the undefined behaviors in the pruned and simplified AST. The repair suggestions include indicating which nodes need to be adjusted, which edges need to be replaced, and which structures need to add specific assertions to ensure code security. The large language model has established a mapping relationship between the ASTs extracted from the Miri test, the correct ASTs stored in the knowledge base, and the repair suggestions for the unsafe Rust types and their repair paths for the undefined behaviors in the ASTs extracted from the Miri test. S7.5. Enhance the prompt with the knowledge extracted from the target correct AST, including adding the repair suggestions for the unsafe Rust types and their repair paths for the undefined behaviors to the prompt to enhance the prompt, so as to dynamically adjust the content of the prompt according to the needs by combining the code context information and the structural features of the AST, ensuring that the generated prompt can fully express the logical and semantic requirements of the code repair process.

[0035] In this embodiment, when generating the AST from the Rust system code snippet that rolls back to the best code in step S7.1, the required prompt can be used as needed. For example, the prompt used in this embodiment is: You are an AST extractor, and you will extract the AST from the requested code I send; 1. Your answer should only include the AST, without any additional comments or explanations; 2. Ensure that the AST output maintains the correct hierarchical structure, and appropriately use indentation or parentheses; …This is the code…{code} Let's think step by step; where "…" represents the omitted content, {code} is the Rust system code snippet that rolls back to the best code, and " " represents emphasis.

[0036] Before step S7.4 in this embodiment, the following steps are also included: collecting the abstract syntax tree (AST) extracted by Miri testing, the correct AST stored in the knowledge base, and the repair methods for the unsafe Rust types with undefined behavior in the AST extracted by Miri testing and their repair paths, and constructing a training dataset; using the training dataset to train a large language model to establish a mapping relationship between the AST extracted by Miri testing, the correct AST stored in the knowledge base, and the repair suggestions for the unsafe Rust types with undefined behavior in the AST extracted by Miri testing and their repair paths.

[0037] Rust is a systems programming language that emphasizes safety, speed, and concurrency. In Rust development, testing is a very important part of the development process because one of the design philosophies of the Rust language is to ensure the safety and correctness of code through compile-time checks. The test semantics of Rust system code snippets refer to the rules and conventions followed when writing test code to ensure the acceptability and effectiveness of the tests. Step S8 for the acceptability of the test semantics of Rust system code snippets is an existing method, and the required test method can be adopted according to needs, so it will not be elaborated here.

[0038] To verify the method for enhancing the security of the Rust language system based on the large language model in this embodiment, in this embodiment, a code test set with undefined behavior provided by Miri official is selected for repair as an example. The test set includes different types of errors, such as memory errors, concurrency errors, dangling pointers, borrowing errors, and data race errors, etc. In this embodiment, the repair situations for different types of undefined behavior using only GPT4 and GPT4 combined with this embodiment will be compared. This embodiment is compared based on passing the Miri test (pass) and maintaining the semantic requirements (exec).

[0039] Figure 5 The percentage of passing the Miri test after repairing different types of errors in the embodiment. From Figure 5 it can be seen that compared with using only GPT4 for repair, the repair accuracy rate of the method for enhancing the security of the Rust language system based on the large language model in this embodiment has increased by about 30%-40% on average. For example, for the alloc errors in the memory allocation series, GPT4 can only repair 53% of the undefined behavior, while the repair rate of this project reaches a passing rate of 93%, an increase of 40%. It can be seen that compared with using only GPT4, the method for enhancing the security of the Rust language system based on the large language model in this embodiment can significantly reduce the possibility of the code having undefined behavior and has good adaptability to multiple types of errors.

[0040] To better ensure the equivalence of the semantic functions of the modified code, this embodiment further judges its semantic equivalence, and the results are as Figure 6 and Figure 7 shown. Figure 6 This is the percentage of semantic acceptability after fixing different types of errors in this embodiment, Figure 7 This is the difference ratio between passing the Miri test and semantic acceptability after fixing different types of errors in this embodiment. From Figure 6 and Figure 7 it can be seen that compared with passing the Miri test, the passing rate has decreased. This is because the randomness and hallucinations generated by the large model lack precise descriptions of the context for code snippets, so the passing rate of semantic acceptability is about 50%. However, the semantic acceptability success rate of the code repaired by GPT4 has been significantly improved, about 25%-35%. Generally speaking, after using the method of this embodiment to repair the undefined behavior code, the memory safety and runtime stability of the code have been significantly improved, and potential undefined behaviors have been effectively eliminated. At the same time, the automated repair mechanism of this embodiment can also reduce the manual intervention of developers, improve the repair efficiency and reduce the risk of human errors, providing a reliable guarantee for building a highly secure Rust system.

[0041] In summary, the method for enhancing the security of the Rust language system based on the large language model in this embodiment includes performing Miri detection on the input Rust system code snippet. The code snippets that pass the Miri detection directly enter the semantic acceptability test. Otherwise, first add assertions or modify the code, and then use Miri to detect whether the code repaired by the LLM is incorrect. If it is incorrect, check whether the maximum iteration value has been reached. If the maximum iteration number has not been reached, repeat the process of repairing the code by the LLM. Otherwise, directly roll back to the code in the best state, and update the prompt words of the large model based on the knowledge base of the abstract syntax tree (AST). If the best code still cannot pass the Miri test, directly exit, indicating that the repair fails. Otherwise, enter the stage of testing semantic acceptability and then directly exit. This embodiment utilizes the self-reasoning and semantic understanding capabilities of the large language model, combined with static analysis tools, to input the code with undefined behavior and its reasons into the large model. The repair process is divided into two stages: assertion and modification. First, try to add assertions at appropriate places to prevent undefined behavior, ensuring that any potential undefined behavior can be captured at runtime before it spreads. For some complex scenarios involving deeper logical errors or dependencies on external states, a modification stage is introduced. In this stage, it adjusts the incorrect semantics in the Rust code to ensure that the functionality of the code remains unchanged while effectively reducing the risk of undefined behavior. To avoid the problem of large model hallucinations during the repair process, which may lead to a decrease in the repair accuracy rate, this embodiment designs an adaptive rollback mechanism to prevent error accumulation and roll back to the best state in a more efficient manner. In addition, this embodiment provides knowledge base support based on the abstract syntax tree, allowing for more general repair based on the code structure, thereby having the ability to provide a generalization solution and improving the repair accuracy rate. Finally, code correctness and semantic acceptability are detected. Based on these methods, this embodiment efficiently and intelligently reduces the undefined behavior in insecure Rust and ensures the security of the Rust system. The method of this embodiment can accurately identify potential risks in insecure Rust code and provide intelligent automatic repair, which is beneficial to reducing insecure operations and undefined behavior in the Rust operating system, enhancing the security of the Rust system, and improving the development efficiency, providing technical support for building a highly reliable and secure operating system.

[0042] In addition, this embodiment also provides a device for enhancing the security of the Rust language system based on the large language model, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the method for enhancing the security of the Rust language system based on the large language model.

[0043] In addition, this embodiment also provides a computer-readable storage medium storing a computer program or instructions, which are programmed or configured to execute the method for enhancing the security of the Rust language system based on the large language model through a processor.

[0044] In addition, this embodiment also provides a computer program product including a computer program or instructions, which are programmed or configured to execute the method for enhancing the security of the Rust language system based on the large language model through a processor.

[0045] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application can be in the form of a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0046] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A Rust language system security enhancement method based on a large language model, characterized in that: The steps include: S1, obtain the input Rust system code snippet; S2, perform Miri test on the Rust system code snippet, if the test passes, jump to step S8; Otherwise, jump to step S3; S3, combining the large language model and prompt words to add assertions and modifications to Rust system code snippets; S4, perform Miri test on the Rust system code snippet with added assertions and modifications, and jump to step S8 if the test passes; Otherwise, jump to step S5; S5, judging whether the maximum iteration value has been reached, if not, jumping to step S3 to continue iterating; otherwise, jumping to step S6; S6, select the Rust system code snippet with the least number of errors among the original Rust system code snippet, the Rust system code snippet with added assertions, and the modified Rust system code snippet as the best code, and roll back the Rust system code snippet to the best code; S7, based on the abstract syntax tree AST extracted by Miri test and the knowledge base enhanced prompt words, jump to step S3; S8, testing semantic acceptability of Rust system code snippets.

2. The Rust language system security enhancement method based on a large language model according to claim 1 is characterized in that: Step S2 includes: S2.1, perform Miri test on the Rust system code snippet to determine whether there is an error message of undefined behavior in the Rust system code snippet. If there is no error message of undefined behavior, the test is determined to be passed and jump to step S8; otherwise, jump to step S2.2; S2.2, determine the unsafe Rust type that causes undefined behavior. If the unsafe Rust type that causes undefined behavior is the target unsafe Rust type, determine that the test fails and jump to step S3. The target unsafe Rust type includes five types: dereferencing raw pointers, calling unsafe functions or methods, accessing or modifying mutable static variables, implementing unsafe features, and accessing union fields; otherwise, determine that the test passes and jump to step S8.

3. The Rust language system security enhancement method based on a large language model according to claim 1 is characterized in that: Step S3 includes: S3.1, combining the large language model and prompt words, adds assertions to automatically repair Rust code in Rust system code snippets that cause undefined behavior of target unsafe Rust types; S3.2, in combination with the large language model and prompt words, modify the Rust code in the Rust system code snippet that causes undefined behavior of the target unsafe Rust type to provide an alternative safe implementation or modify it to reduce the undefined behavior in the Rust code without compromising the semantics of the code.

4. The Rust language system security enhancement method based on a large language model according to claim 3 is characterized in that: The prompt words used in step S3.1 include: "…The following is the error message…{log}", the ideas for adding assertions are as follows:

1. Add assertions before undefined behavior becomes possible to prevent it from happening; 2. Doing so is important and necessary; Only add assertions, do not adjust or delete other code modules; "where "..." indicates omitted content, {log} is the log of Miri test, " ” for emphasis.

5. The Rust language system security enhancement method based on a large language model according to claim 3 is characterized in that: The prompt words used in step S3.2 include: "Undefined behavior occurs when running the Miri test according to the code. The repair ideas are as follows:

1. Identify undefined behavior; 2. The code that adds assertions cannot be modified, and there are problems with the logic itself; 3. In order to maintain the functionality and semantics of the source code, avoid drastically changing the logical structure; 4. Design safe alternatives; refactor the code according to the safe alternatives to ensure that the modified code not only avoids undefined behavior but also maintains the original functional logic and performance standards; ...Here is the code...{code}", where "..." indicates omitted content, and {code} is the modified Rust system code snippet, " ” for emphasis.

6. The Rust language system security enhancement method based on a large language model according to claim 1 is characterized in that: The abstract syntax tree AST and knowledge base enhanced prompt words extracted based on the Miri test in step S7 include: S7.1, using the large language model to generate an abstract syntax tree AST for the Rust system code snippet rolled back to the best code, wherein the abstract syntax tree AST is a tree data structure that expresses the syntax structure of the program, wherein each node represents a syntax element, and the edge represents the parent-child relationship between the nodes, connecting the hierarchical structure between the syntax elements; S7.2, use the pruning function to prune and simplify the abstract syntax tree AST; S7.3, calculate the similarity between the pruned and simplified abstract syntax tree AST and the correct abstract syntax tree AST stored in the knowledge base. If there is no correct abstract syntax tree AST whose similarity exceeds a preset threshold, then directly end and jump to step S3; otherwise, the correct abstract syntax tree AST with the best similarity is used as the target correct abstract syntax tree AST; S7.4, extracting knowledge from the target correct abstract syntax tree AST, including: inputting the pruned and simplified abstract syntax tree AST and the target correct abstract syntax tree AST into a pre-trained large language model, obtaining repair suggestions for unsafe Rust types with undefined behavior and their repair paths in the pruned and simplified abstract syntax tree AST, wherein the repair suggestions include indicating which nodes need to be adjusted, which edges need to be replaced, and which structures need to add specific assertions to ensure code security, and the large language model is pre-trained to establish a mapping relationship between the abstract syntax tree AST extracted by the Miri test, the correct abstract syntax tree AST stored in the knowledge base, and the repair suggestions for unsafe Rust types with undefined behavior and their repair paths in the abstract syntax tree AST extracted by the Miri test; S7.5, enhances the hint words with the knowledge extracted from the target correct abstract syntax tree AST, including adding repair suggestions for unsafe Rust types with undefined behavior and their repair paths to the hint words to enhance the hint words.

7. The Rust language system security enhancement method based on a large language model according to claim 6 is characterized in that: Before step S7.4, it also includes: collecting the abstract syntax tree AST extracted by the Miri test, the correct abstract syntax tree AST stored in the knowledge base, and the repair methods for the unsafe Rust types with undefined behavior and their repair paths in the abstract syntax tree AST extracted by the Miri test and constructing a training data set; using the training data set to train the large language model to establish a mapping relationship between the abstract syntax tree AST extracted by the Miri test, the correct abstract syntax tree AST stored in the knowledge base, and the repair suggestions for the unsafe Rust types with undefined behavior and their repair paths in the abstract syntax tree AST extracted by the Miri test.

8. A Rust language system security enhancement device based on a large language model, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the Rust language system security enhancement method based on a large language model as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the Rust language system security enhancement method based on a large language model as described in any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the Rust language system security enhancement method based on a large language model as described in any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Rust language-oriented vulnerability automatic positioning and analysis method

    CN116305163A

  • Security enhancement model development method and system based on Rust language

    CN116484439A

  • Intelligent contract cross-contract detection method and system based on abstract syntax tree

    CN116610561A

  • Code annotation generation method and device

    CN116661855A

  • Rust language document test automatic generation method and device based on large code model

    CN117951038A