Rust imperfect packaging detection method and device based on large language model
Through the detection method based on the large language model, decomposing and analyzing the security description of Rust insecure functions, the problem of difficult to detect insane unsecure call packaging in the prior art is solved, and higher detection accuracy and software security are achieved.
Patent Information
- Application Number
- CN202510652149.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art is difficult to effectively detect unsolicited unsafe call encapsulation in the Rust language, resulting in latent vulnerabilities and undefined behavior, affecting the security and reliability of the software.
The detection method based on the large language model is adopted, and the context of the target encapsulation is obtained through static analysis tools. The security description of the unsafe function is decomposed into fine-grained contracts, and the contract type and guarantee mode are analyzed to determine whether the encapsulation provides guarantees for each contract.
It significantly improves the accuracy and applicability of detection of imperfect packaging in the Rust code base, assists developers in reviewing codes, and improves the security and reliability of the software.
Smart Images

Figure CN120179529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer program analysis based on large language models, and particularly to a method and device for detecting Rust unsound encapsulation based on large language models. Background Art
[0002] In the software field, system software is computer software that directly executes or controls hardware, and supports the development and operation of upper-layer application software as underlying software; ensuring the stability and security of system software is of great significance. In the past, C / C++ was the mainstream language for developing system software; however, the C / C++ language does not restrict the use of memory and pointers, resulting in software developed in the C / C++ language being prone to hidden memory leaks and pointer usage security issues, bringing great maintenance pressure to software developers and maintainers.
[0003] Rust is a new general-purpose system-level programming language. Through unique ownership and lifetime mechanisms, it can effectively avoid introducing memory problems during programming while maintaining high performance similar to the C / C++ language. Rust provides a security guarantee through strict checks during the compilation stage: programs written in safe Rust are guaranteed to be memory-safe as long as they can pass compilation verification. However, in actual development, in order to optimize performance or call underlying system functions, developers usually require higher flexibility, and these operations are prohibited in safe Rust. For this reason, Rust provides explicit unsafe blocks that allow developers to bypass compiler checks and perform unsafe operations. Existing research has shown that calling unsafe functions is the main purpose of using unsafe code. Usually, a Rust function is declared as unsafe to emphasize that it has additional security requirements for the function caller, and these requirements are called contracts by the community. The caller must ensure that all these contracts are satisfied; otherwise, calling an unsafe function may lead to undefined behavior and introduce latent vulnerabilities.
[0004] The official "Unsafe Code Guide" advocates encapsulating unsafe code into safe functions, enabling users to directly use the safe encapsulation without concerning about the underlying security details. The call correctness of a sound unsafe call encapsulation should be verifiable by the Rust compiler, meaning it should not impose any additional requirements on the caller other than the parameter types. To achieve this, the encapsulation itself must ensure that all contracts of its internal unsafe functions are guaranteed. If an unsafe call is encapsulated in a safe function that does not guarantee all contracts, it will introduce unsoundness. Unsound encapsulation means that the caller can input specific values to this encapsulation in safe Rust and trigger undefined behavior, thus undermining Rust's security promise. Additionally, the unsoundness of the encapsulation may spread through function calls and data flows, leading to complex bugs and compromising Rust's advantages. Therefore, unsoundness is intolerable in the Rust community, and some partially unsound unsafe call encapsulations have even been disclosed as security vulnerabilities. In summary, detecting unsound unsafe encapsulations is of great importance.
[0005] The contracts of unsafe functions are described in an unstructured and flexible natural language form in the security notes section of the documentation. An unsafe function usually contains multiple contracts, making the security notes lengthy and complex. Additionally, this section may include extended descriptions (such as examples and consequences), making the content overly cumbersome. Therefore, traditional natural language processing techniques are difficult to effectively analyze various security notes. Moreover, some contracts that need to be independently checked may be mixed in a single sentence, increasing the risk of omission in subsequent checks. Besides the complexity of the security notes, verifying these contracts themselves is also challenging. First, this task requires familiarity with Rust features and an in-depth understanding of the contracts of unsafe functions. Checking contracts may involve reasoning in complex contexts, involving numerous structs, functions, traits, and variable types. Overall, verifying the soundness of an unsafe call by checking whether all contracts are guaranteed is a tedious and error-prone task. Summary of the Invention
[0006] The objective of the present invention is to provide a method and device for detecting unsound encapsulations in Rust based on large language models in view of the deficiencies of the prior art.
[0007] To achieve the above objective, the present invention provides a method for detecting unsound encapsulations in Rust based on large language models, including the following steps:
[0008] (1) By analyzing the documentation and code of unsafe calls in the standard library and popular open-source projects, summarize the types of contracts and the corresponding guarantee modes for each contract type, and design examples for each guarantee mode;
[0009] (2) Obtain the relevant context of the target encapsulation through a static analysis tool, including code hints and reference information;
[0010] (3) Decompose the original security description of the insecure functions within the encapsulation into multiple fine-grained contracts through a large language model, and classify each contract into the types defined in step (1);
[0011] (4) Use the large language model to analyze each fine-grained contract, and provide corresponding examples for the large language model in combination with the contract types classified in step (3) and the corresponding safeguard modes to analyze whether the encapsulation provides safeguards for the target fine-grained contract; the examples include requests and reference answers, and the requests include reference information, the code of the encapsulation, and the fine-grained contract to be inspected;
[0012] (5) Summarize the analysis results of all fine-grained contracts. If there are contracts that are not safeguarded, determine that the insecure call encapsulation is an unsound encapsulation; otherwise, determine it as a sound encapsulation.
[0013] Further, the reference information includes the called functions, structures, and their documentation; the code hints include parameter name hints and variable type hints.
[0014] Further, it also includes trimming the reference information:
[0015] (1) Retain the code snippets of the structure;
[0016] (2) In the code of the called functions and macros, the implementation details will be omitted, and only their function signatures will be retained;
[0017] (3) Retain the description part of the documentation, which provides a brief overview of the functions of the elements.
[0018] Further, step (3) also includes: improving the quality of contract splitting and classification through the self-review and optimization process iteration of the large language model. Specifically:
[0019] The large language model reviews whether the currently decomposed fine-grained contracts meet all the requirements of consistency, non-overlap, atomicity, unique classification, and clarity, provides detailed review opinions for each requirement, and finally decides whether to optimize the decomposition results; if optimization is required, the large language model optimizes the decomposition results according to the original security description, the current decomposition and classification results, and the review opinions, and outputs the optimized fine-grained contracts and corresponding types.
[0020] Further, in step (4), the large language model will check each fine-grained contract alone; when checking a certain fine-grained contract, it will conduct multiple rounds of checks according to the type of the fine-grained contract, with each round corresponding to a safeguard mode of the contract type, and provide examples of the corresponding safeguard mode when using the large language model for checking.
[0021] Furthermore, it also includes a pre - constructed example library; the examples in the example library are divided into two categories: positive examples and negative examples of the safeguard mode, which are used as demonstrations to be added to the prompts; the positive examples of the safeguard mode correspond to a certain safeguard mode and describe how the encapsulation safeguards the establishment of the contract in that safeguard mode; the negative examples are used to explain the reasons why a certain fine - grained contract is considered to be unguaranteed.
[0022] Furthermore, in step (5), first summarize the analysis results corresponding to the safeguard mode in step (4). As long as a fine - grained contract is safeguarded by any one mode, it is considered that the contract is guaranteed; otherwise, it is unguaranteed. Then summarize the results of each fine - grained contract to obtain the soundness of the encapsulation. If any one fine - grained contract is unguaranteed, then the encapsulation is unsound.
[0023] Furthermore, when the model does not output as expected, corresponding measures are taken to improve the usability of the system, including:
[0024] For the unknown classification results in step (3), determine the contract type through vector similarity;
[0025] For the case where the large - language model in step (4) does not give an affirmative answer, it is determined to be unguaranteed.
[0026] To achieve the above - mentioned purpose, the present invention also provides a Rust unsound encapsulation detection device based on a large - language model, including one or more processors for implementing the above - mentioned method.
[0027] To achieve the above - mentioned purpose, the present invention also provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above - mentioned Rust unsound encapsulation detection method based on a large - language model.
[0028] To achieve the above - mentioned purpose, the present invention also provides a computer - readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above - mentioned Rust unsound encapsulation detection method based on a large - language model.
[0029] The beneficial effects of the present invention are as follows: Existing methods have great limitations in detecting unsound encapsulations of unsafe calls in the Rust language, while the present invention uses a large - language model combined with static analysis technology to achieve intelligent detection of the soundness of encapsulation code. Through fine - grained contract splitting, classification, and multi - round safeguard mode analysis, the present invention significantly improves the accuracy and applicability of detection, can effectively detect unsound encapsulations in the Rust code library, assist developers in reviewing code, and improve software security and reliability. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0031] Figure 1 is the overall framework diagram of the method of the present invention;
[0032] Figure 2 is the safety description decomposition classification flowchart including iterative self-check in the present invention;
[0033] Figure 3 is the structural schematic diagram of the device of the present invention;
[0034] Figure 4 is the schematic diagram of an electronic device of the present invention. Specific embodiments
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0036] A Rust unsound encapsulation detection method based on a large language model of the present invention, see Figure 1 , includes the following steps:
[0037] (1) By analyzing the documentation and code of unsafe calls in the standard library and popular open-source projects, summarize the types of contracts and the possible guarantee modes corresponding to each contract type, and design examples for each guarantee mode;
[0038] In one embodiment, open-source Rust code is first crawled from the GitHub open-source code hosting platform as an experimental data set. Specifically, the present invention first analyzed the unsafe call encapsulations in the Rust standard library and the top 500 packages with the most downloads on the central repository. All code libraries were cloned from the latest commits on GitHub on May 25, 2024. These packages cover multiple fields and have been extensively reviewed to ensure the quality and comprehensiveness of the preliminary research.
[0039] After collecting the documentation of all insecure call wrappers and the insecure functions they reference, the present invention introduces two rounds of manual analysis. The goal of the first round of analysis is to classify the content in the security specifications into different contract types, while the second round of analysis is dedicated to refining the safeguard patterns for each contract type. In the first round of analysis, the researchers read through the documentation of the involved insecure APIs one by one, summarized the contract types, and classified the contracts in the documentation. In the second round of analysis, based on the results of the first round, the researchers refined the corresponding safeguard patterns for each contract type. For a specific contract type, we manually examined the code and documentation of all relevant insecure call wrappers and grouped these wrappers into several subgroups according to the way the safeguard pattern holds.
[0040] Referring to Table 1 and Table 2, the embodiments of the present invention finally define 16 contract types and their corresponding 34 safeguard patterns. These 16 contract types focus on different aspects of security requirements, including memory and pointers, values, concurrency, lifecycle, ownership, data flow, and environment.
[0041] Table 1: Some contract types and their corresponding safeguard patterns
[0042] Table 2: The remaining contract types and their corresponding safeguard patterns
[0043] To fully activate the context learning ability of the large language model, the present invention pre-constructs an example library, and the examples in the example library can be added to the prompt as demonstrations. These examples are divided into two categories: positive examples and negative examples of safeguard patterns. The positive examples of safeguard patterns correspond to specific safeguard patterns and describe how the wrapper safeguards the establishment of the contract in that safeguard pattern. On the contrary, the negative examples are used to explain why a certain contract is considered unprotected. According to the previously summarized 16 contract types and 34 safeguard patterns, the example library contains a total of 34 positive examples of patterns and 16 negative examples. Both the positive examples and negative examples of safeguard patterns consist of requests and reference answers. Among them, the requests include reference information, the code of the wrapper, and the fine-grained contract to be inspected. The reference answers are step-by-step analyses written by humans, implicitly demonstrating a specific safeguard pattern or explaining why the contract cannot be protected.
[0044] (2) Obtain the relevant context of the target wrapper through a static analysis tool, including code hints and reference information;
[0045] In one embodiment, for the encapsulation of unsafe calls to be inspected, the embodiments of the present invention first use a static analysis tool to retrieve relevant context information, including code hints and reference information. Code hints are code snippets that can be attached to the original code to provide richer information, such as the deduced variable types and the parameter names shown in function calls. Reference information includes relevant information about elements such as structures, functions, and features involved in the target unsafe call encapsulation.
[0046] Rust is an implicitly statically typed language, meaning that variable types can be deduced based on the code context without explicit annotation. In addition, the contracts of unsafe functions are usually described in terms of parameter names, so when checking whether the contracts are guaranteed, it is necessary to match the function formal parameters with specific variables. However, this matching process is error-prone for large language models, resulting in inaccurate generated answers. For example, assume an unsafe function " " stipulates that "a must be greater than b", and the call method is " ", then the large language model may incorrectly map variables b and a to parameters a and b, making it difficult to correctly check this unsafe call. In addition, the type of the new variable c is also unclear, and variable types contain some semantic information that is crucial when providing guarantees for the contract. For example, variables of type must be non-negative. After adding code hints, the call code will become " ", making the variable types clearer and the correspondence between variables and parameters more distinct.
[0047] To accurately and efficiently obtain variable types, the present invention includes a retriever that can extract the type deduction results from the compiler and attach the deduced types to the original code in a format conforming to Rust syntax. Specifically, for a new variable declared by the statement " ", its type will be attached after the variable in the format " "; for the return type of a closure, the format is " ". In addition, to establish the matching relationship between function formal parameters and specific incoming variables, the present invention analyzes the parameter list declared in the function signature and prefixes the variable name with the parameter name in the format " ".
[0048] In one embodiment, considering that the warehouse to be inspected is a real project with a large amount of code and is relatively complex, the target non-safe call encapsulation usually involves a large number of external elements, such as structures and functions. Without providing the reference information of these elements, the large language model often gives incorrect answers containing serious hallucinations. To retrieve the reference information, the present invention extends the Rust language server Rust Analyzer. First, by analyzing the abstract syntax tree of the target non-safe call encapsulation, the present invention extracts all the referenced elements. Subsequently, Rust Analyzer combines the analysis results of the entire project to retrieve the documentation and code of these elements.
[0049] However, since the reference information may grow exponentially, the present invention only includes directly referenced elements. Even so, the retrieved information may still be too redundant and not conducive to the large language model's retrieval of information. Therefore, the present invention uses the following rules to trim the reference information:
[0050] (a) The code snippet of the structure (i.e., the structure definition) will be completely retained because its field information is very important and can provide information such as field types and the association relationships between attributes.
[0051] (b) The overly long implementation details in the code of the referenced functions and macros will be omitted, and only their function signatures will be retained.
[0052] (c) Although the documentation information is valuable, it is usually too verbose to be directly used. To balance the amount of information, the present invention retains the description part of the documentation, which provides a brief overview of the function of the element.
[0053] Additionally, for the used unsafe functions, the present invention extracts the security instructions in their documentation for further inspection.
[0054] (3) Decompose the original security instructions of the unsafe functions within the encapsulation into multiple fine-grained contracts by the large language model, and classify each contract into the types defined in step (1);
[0055] In one embodiment, the embodiment of the present invention uses the large language model to decompose the original security instructions into fine-grained contracts so that each contract can be independently inspected. This decomposition process needs to meet the following requirements:
[0056] Consistency: The decomposed contracts must be derived from the original security instructions and cover all the described security requirements.
[0057] Non-overlap: The decomposed contracts cannot overlap with each other.
[0058] Atomicity: The decomposed contracts cannot be further disassembled.
[0059] Unique Classification: Each decomposed contract can only belong to one contract category.
[0060] Clarity: The decomposed contracts should be concise but clearly expressed.
[0061] To meet the unique classification requirement and facilitate subsequent analysis, the present invention adds the names and definitions of all contract types to the prompt information for the large language model and requires the large language model to complete classification while decomposing the contracts. The prompt information also contains seven examples to fully stimulate the context learning ability of the large language model to further improve the accuracy of classification. By decomposing and classifying the security statements, the present invention can obtain multiple fine-grained contracts and their corresponding types.
[0062] See Figure 2 Since the quality of decomposition and classification directly affects the accuracy of subsequent inspections, the present invention further uses the large language model to self-check and optimize the results of decomposition and classification. Specifically, the present invention uses specific prompt words to let the large language model review whether the currently decomposed contracts meet all the requirements of consistency, non-overlap, atomicity, unique classification, and clarity:
[0063] """
[0064] ## Role Description:
[0065] You are a senior software engineer proficient in Rust.
[0066] ## Background:
[0067] {Omit the specific description of the background}
[0068] ## Task Definition: Your task is to strictly check whether the "Safety" section of the unsafe Rust function is correctly and faithfully decomposed into fine-grained contracts. Be skeptical of the given contracts and adopt the strictest and most conservative analysis method. Assume that each contract may imply problems and analyze it critically. Do not accept any contract without fully verifying its correctness. You must traverse all the decomposed contracts and analyze whether each contract meets the following criteria:
[0069] {Omit the specific description of the decomposition requirements}
[0070] Finally, you must give a summary and a final judgment, indicating in bold "Yes" or "No" whether the decomposed contracts are perfect and require no modification.
[0071] ## Contract Type Definition:
[0072] {Omit the names and definitions of the contract types}
[0073] """
[0074] The large language model will provide detailed review comments for each requirement and finally decide whether the decomposition result needs to be optimized. If optimization is required, the present invention will further let the large language model optimize the decomposition result according to the original security description, the current decomposition and classification results, and the review comments given by the large language model, and output the optimized fine-grained contracts and corresponding types. The specific prompt words are as follows:
[0075] """
[0076] ## Role description:
[0077] You are a senior software engineer proficient in Rust
[0078] ## Background:
[0079] In Rust, unsafe functions always declare their contract requirements to the caller through the "Safety" section. Programmers must satisfy all the required contracts, otherwise unsafe calls may lead to undefined behavior. However, the content of the "Safety" section may be redundant and not easy to read. To improve clarity, we need to decompose the "Safety" section into fine-grained contracts and classify them according to clearly defined contract types. The decomposition process must meet the following criteria:
[0080] {Omit the specific description of the decomposition criteria}
[0081] ## Task definition:
[0082] You will receive: 1) information about an unsafe Rust function, including the original "Safety" section; 2) the decomposed fine-grained contracts and their types; 3) an analysis of the decomposed contracts based on the decomposition criteria. Your task is to improve these contract decompositions based on all the information provided (especially the analysis comments). You must 1) carefully analyze each analysis point and determine the appropriate modification plan; 2) modify the contract accordingly if the analysis is correct; 3) clarify the contract under the premise of ensuring correctness if the analysis is incorrect or the expression is ambiguous.
[0083] ## Response format:
[0084] The response content should only contain the list of refined fine-grained contracts in the following format: "- bytes must contain valid UTF-8 encoding (Encoding)". Each entry corresponds to a specific contract and its type.
[0085] ## Contract type definition:
[0086] {Omit the name and definition of the contract type}
[0087] """
[0088] The above process will be iteratively executed until the large language model determines that no further optimization is required. To avoid self-checking for infinite loops, the present invention sets a limit on the maximum number of iterations to ensure that the decomposition and classification processes are completed within a reasonable range.
[0089] (4) Use the large language model to analyze each fine-grained contract, and provide corresponding examples for the large language model in combination with the contract types and corresponding safeguard modes classified in step (3) to analyze whether the encapsulation provides safeguards for the target fine-grained contract.
[0090] In one embodiment, the embodiment of the present invention checks whether each fine-grained contract is safeguarded within the encapsulation by sending a request to the large language model. This request is generated by filling in a template, and the filled content includes the cropped reference information, the code with code hints, and the description of the target fine-grained contract.
[0091] To provide accurate domain knowledge to the large language model, the present invention selects corresponding examples from the example library according to the type of the target contract. However, due to the significant differences between different safeguard modes, it is difficult for the large language model to learn the domain knowledge of all safeguard modes simultaneously from the context. Therefore, the present invention designs multi-round mode-oriented checks, and each round of check corresponds to a single safeguard mode related to the target contract type. Specifically, in each round of check, the present invention selects a mode example related to this safeguard mode and a counterexample corresponding to the target contract type. In this round of check, the large language model will ultimately make a judgment on whether the contract is safeguarded by the non-secure call encapsulation under this safeguard mode.
[0092] To effectively stimulate the reasoning ability of the large language model, the present invention instructs it to use the chain-of-thought method for checking, rather than simply requiring it to analyze step by step. The present invention requires the large language model to follow the following steps for reasoning:
[0093] (I) Locate the code snippet of the non-secure call.
[0094] (II) List the variables related to the contract.
[0095] (III) Analyze step by step around the contract.
[0096] (IV) Judge whether the contract is safeguarded.
[0097] In the first step, the large language model clarifies its analysis objective and repeats relevant code snippets to reduce hallucination phenomena caused by input conflicts. The second step extracts key expressions in the code to facilitate subsequent reasoning steps. The third step is the core analysis phase, where the large language model uses its inherent capabilities and domain knowledge learned from examples to perform detailed step-by-step reasoning. Finally, the large language model needs to give a clear "yes" or "no" as the final judgment to automate the output processing. This chain of checks is not only specified in the system prompt of the request but also demonstrated through few-shot examples. The difficulty of the four steps increases gradually, and each step depends on the output of the previous step. In addition to enhancing the reasoning ability of the large language model, the chain-of-thought method also enhances the interpretability of the output, enabling human reviewers to more easily determine whether it is a false positive by reading the output of the large language model.
[0098] After all the pattern-oriented checks, the judgment results for each safeguard pattern can be obtained. Through the "OR" operation, these judgment results are combined into the overall judgment result for the contract. Finally, considering that a sound non-secure call encapsulation must provide safeguards for all contracts, the present invention further aggregates the judgment results of all contracts through the "AND" operation as the final judgment on whether the non-secure call encapsulation is sound. If the final judgment result of a certain contract is unprotected, it means that the contract is not safeguarded by the non-secure call encapsulation and corresponding repairs are required. The present invention defines such unprotected fine-grained contracts as contract-level unsoundness.
[0099] In addition to the above three parts, namely the encapsulation context acquirer, the security description splitter, and the fine-grained contract discriminator, the present invention also implements solutions for abnormal situations that may affect automation in actual scenarios, thereby improving the usability of the present invention.
[0100] First, since the present invention checks reliability based on security instructions, insecure functions lacking documentation or security instructions will be directly skipped. During the decomposition and classification phase, the large language model may misclassify the contract as an undefined type or directly output "unknown". In response to these situations, the present invention will obtain the vector representation of the contract through the encoding model and match the contract most similar to it in the example library through cosine similarity to determine its type. During the contract checking process, although the large language model ultimately needs to answer "yes" or "no", it may output "unknown" in uncertain situations. During result aggregation, the present invention treats "unknown" as a variant of "not guaranteed". For insecure call encapsulations involving multiple non-secure calls, the present invention will separately decompose the security instructions of each call and perform pattern-based checks on all fine-grained contracts. Similarly, only when all contracts are determined to be "guaranteed" will this insecure call encapsulation be considered "reliable". Additionally, if any contract of a non-secure call is determined to be "not guaranteed", then this non-secure call will be considered function-level unreliable.
[0101] Corresponding to the foregoing embodiment of a method for detecting Rust unsound encapsulations based on a large language model, the present invention also provides an embodiment of a device for detecting Rust unsound encapsulations based on a large language model.
[0102] See Figure 3 , an embodiment of a device for detecting Rust unsound encapsulations based on a large language model provided by an embodiment of the present invention includes one or more processors for implementing a method for detecting Rust unsound encapsulations based on a large language model in the foregoing embodiment.
[0103] An embodiment of a device for detecting Rust unsound encapsulations based on a large language model of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the device for detecting unsound encapsulations of insecure calls in the Rust language based on a large language model of the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0104] The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0105] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0106] Corresponding to the foregoing embodiment of a method for detecting unsound encapsulation of Rust based on a large language model, an embodiment of the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for detecting unsound encapsulation of Rust based on a large language model as described above. As Figure 4 shown, it is a hardware structure diagram of a device with any data processing ability where the method for detecting unsound encapsulation of Rust based on a large language model provided by an embodiment of the present application is located. In addition to Figure 4 the processors, memory, DMA controller, disk, and non-volatile memory shown, any device with data processing ability where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with any data processing ability, which will not be elaborated here.
[0107] Corresponding to the foregoing embodiment of the method for detecting unsound encapsulation of unsafe calls in the Rust language based on a large language model, an embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for detecting unsound encapsulation of unsafe calls in the Rust language based on a large language model in the above embodiment.
[0108] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0109] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
[0110] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design concepts disclosed by the present invention are within the scope of protection of the present invention.
Claims
1. A Rust incomplete encapsulation detection method based on a large language model, characterized in that: The following steps are involved: (1) By analyzing the documents and codes of unsafe calls in standard libraries and popular open source projects, we summarize the types of contracts and the corresponding protection modes for each contract type, and design examples for each protection mode; (2) Obtain relevant context of the target package through static analysis tools, including code hints and reference information; (3) Decompose the original security description of the unsafe function in the package into multiple fine-grained contracts through the large language model, and classify each contract into the type defined in step (1); (4) Analyze each fine-grained contract using the large language model, and provide corresponding examples for the large language model based on the contract types and corresponding protection modes classified in step (3), and analyze whether the encapsulation provides protection for the target fine-grained contract; the examples include requests and reference answers, and the requests include reference information, encapsulated code, and the fine-grained contract to be checked; (5) Summarize the analysis results of all fine-grained contracts. If there are unsecured contracts, the unsafe call encapsulation is determined to be an incomplete encapsulation; otherwise, it is determined to be a complete encapsulation.
2. The method according to claim 1, characterized in that Reference information includes called functions, structures and their documents; code hints include parameter name hints and variable type hints.
3. The method according to claim 1, characterized in that It also includes the citation information clipping: (1) Code snippets that retain the structure; (2) The implementation details of the referenced functions and macros will be omitted, leaving only their function signatures; (3) Keep the description section of the document, which provides a brief overview of the element's functionality.
4. The method according to claim 1, characterized in that: The step (3) also includes: improving the quality of contract splitting and classification through self-review of the large language model and iterative optimization process, specifically: The large language model reviews whether the currently decomposed fine-grained contract meets all consistency, non-overlapping, atomicity, unique classification, and clarity requirements, and provides detailed review opinions for each requirement, and ultimately decides whether the decomposition results need to be optimized; if optimization is required, the large language model optimizes the decomposition results based on the original security description, the current decomposition and classification results, and the review opinions, and outputs the optimized fine-grained contracts and corresponding types.
5. The method according to claim 1, characterized in that In step (4), the large language model will check each fine-grained contract independently; when checking a fine-grained contract, it will perform multiple rounds of checks based on the type of the fine-grained contract, each round corresponding to a protection mode of the contract type, and provide examples of the corresponding protection modes when using the large language model for checking.
6. The method according to claim 1 or 5, characterized in that: It also includes a pre-built example library; the examples in the example library are divided into two categories: positive examples and negative examples of guarantee patterns, which are used to add to prompts as demonstrations; positive examples of guarantee patterns correspond to a certain guarantee pattern, describing how the package guarantees the establishment of the contract with the guarantee pattern; Counterexamples are used to explain why a fine-grained contract is considered not guaranteed.
7. The method according to claim 1, characterized in that In the step (5), the analysis results of the corresponding guarantee modes in step (4) are first summarized. As long as the fine-grained contract is guaranteed by any mode, it is considered that the contract is guaranteed, otherwise it is not guaranteed. Then, the results of each fine-grained contract are summarized to obtain the soundness of the encapsulation. If any fine-grained contract is not guaranteed, the encapsulation is not sound.
8. The method according to claim 1, characterized in that When the model does not output as expected, take appropriate measures to improve the availability of the system, including: For unknown classification results in step (3), the contract type is determined by vector similarity; If the large language model does not give a positive answer in step (4), it is considered to be unsecured.
9. A Rust insane encapsulation detection device based on a large language model, characterized in that: It includes one or more processors, used to implement a Rust unsound encapsulation detection method based on a large language model as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the Rust insound encapsulation detection method based on a large language model described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Automatic Rust program defect detection method and system based on feature extraction
CN116680705A
Secure transport software update
US20240004639A1