A Cross-Language Compiler Vulnerability Mining Method and Device Based on Transfer Learning

The pre-trained model is corrected and fine-tuned through the transfer learning method, and the test samples of the target language are generated, which solves the problems of large quality and time overhead of the training set in the existing technology, and realizes efficient compiler vulnerability mining.

CN115758379BActive Publication Date: 2025-07-11INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211460458.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-07-11
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

The existing learning-based fuzz testing method is greatly affected by the quality of the training set in the mining of compiler vulnerability, has a large training time overhead, and is inefficient in generating test cases.

Method used

Using the transfer learning method, the model is corrected and fine-tuned by selecting the pre-trained model, calculating the sequence distance difference and correcting regular terms, and a test sample of the target language is generated to explore loopholes.

Benefits of technology

It improves the efficiency and effectiveness of compiler vulnerability mining, reduces the requirements for the quality of the target language training set, and has cross-language universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115758379B_ABST
    Figure CN115758379B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross - language compiler vulnerability mining method and device based on transfer learning. The steps of the method include: 1) calculating the distribution difference D and the corrected regularization term L between the source programming language data D S and the target programming language data D T ; 2) correcting the pre - trained model M S according to the difference between the corrected regularization term L and the source language sequence S T and the target language sequence S S to obtain the corrected model M S ’ ; 3) fine - tuning and training M S ’ using the target language sequence S T to obtain the model M T ; 4) generating a target - language program as a sample according to the model M T and performing fuzz testing to mine vulnerabilities. The present invention proposes an optimization and reuse technology of the pre - trained model and a test sample generation method to solve the timeliness and effectiveness problems of test sample generation in compiler fuzz testing. The present invention can improve the speed and the number of samples of vulnerability miners when generating the target language as test samples, and thus improve the vulnerability mining ability for compilers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and particularly to a cross-language compiler vulnerability mining method and device based on transfer learning. Background Art

[0002] With the development of the computer industry, computer software has increasingly become an important part and asset in social life and industrial production. However, software security incidents emerge in an endless stream, posing a serious threat to social security and stability. Therefore, the security of software has become even more important. As a core basic software, the compiler supports the stable use of most other software, and its security is particularly important. The "Trojan Source" vulnerability discovered by Nicholas et al. is a typical case where compiler vulnerabilities threaten global software code. This vulnerability exists in Unicode and can affect the logical order of most compilers by rearranging characters, thus affecting almost all software written in computer languages (Boucher N, Anderson R. Trojan Source: Invisible Vulnerabilities[J]. arXiv preprint arXiv:2111.00169, 2021).

[0003] Currently, learning-based fuzz testing is one of the common methods for mining compiler vulnerabilities. Due to the large code size and high complexity of compilers, fuzz testing has the advantages of being fast and effective when mining compiler vulnerabilities. Currently, fuzz testing usually adopts a learning-based test case generation method. This technology continuously generates new program test cases by training a neural network and uses them as inputs for compiler fuzz testing to observe whether the program test cases cause compiler exceptions. Godefroid et al. trained an RNN to generate well-formatted PDF files and used them as test cases to test the PDF parser on the edge (Godefroid, P.; Peleg, H.; and Singh, R. 2017. Learn&Fuzz: Machine learning for input fuzzing. In ASE’17). Liu et al. proposed a fuzz testing tool DEEPFUZZ by training a SequenceToSequence model. This tool can automatically and continuously generate well-formatted C programs and successfully detected 8 vulnerabilities existing in GCC.

[0004] Although the learning-based test case generation method is effective, there are still two problems as follows: First, the training effect of the model is affected by the quality of the training set. Godefroid's work was supported by the official developers of Edge, so a large number of examples provided by official personnel were obtained for training. The C language programs used by Liu for training were also collected from the GCC test suite. However, not all compiler or virtual machine engine program test cases have high-quality test suites as training sets, which will undoubtedly greatly reduce the generation effect. Second, the training time cost of the model is very high. To obtain a better model, it takes more than dozens of hours of training time, and a lot of energy is also spent on optimizing the parameters of the model. In this case, the time cost for generating test cases is much higher than the cost of testing and vulnerability mining.

[0005] The characteristics of transfer learning can exactly make up for the problems existing in traditional learning-based fuzz testing methods. Transfer learning can still achieve good training effects even when the data quality of the target data set is not high, and using the pre-trained model can shorten the training duration. Using transfer learning to quickly train a generation model can improve the efficiency and effectiveness of test case generation and vulnerability mining. Summary of the Invention

[0006] The purpose of the present invention is to provide a cross-language compiler vulnerability mining method and device based on transfer learning for existing problems.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A cross-language compiler vulnerability mining method based on transfer learning, which includes the following steps:

[0009] 1) Select a model for generating fuzz test cases based on learning as the pre-trained model, called M S , and on this basis, select the source language data set and the target language data set, serialize them into the source language sequence S S and the target language sequence S T , and calculate the distance difference D between the two sequences and the correction regularization term L of the model;

[0010] 2) For the pre-trained model M S of the source language, according to the obtained correction regularization term L, and the difference between the sequences S S and S T , correct the pre-trained model M S to obtain the corrected model M S ';

[0011] 3) For the corrected M S ', use the target language sequence S TPerform fine-tuning training to finally obtain the generation model M T ;

[0012] 4) Use the model M T to generate code, combine it with the seed program to generate a target language program as a test case, and use this test case for fuzz testing to detect vulnerabilities.

[0013] Furthermore, in step 1) above, select a publicly available model for learning to generate fuzz test cases on an open-source platform as the pre-training model. The model selected on the open-source platform is called the pre-training model M S , and the source code programming language generated by this model is called the "source language". Determine the target compiler to be detected, and the programming language compiled by this compiler is called the "target language". Select a certain number of source language programs from the vulnerability information release platform or the official test suite as the source language dataset D S , and target language programs as the target language dataset D T . Consider each program in D S and D T as a sequence of tokens (lemmas), and concatenate all programs into a large sequence S S and S T . And calculate the distance difference D and the regularization term L between the two sequences.

[0014] Furthermore, the calculation of the distance difference D and the regularization term L between the two sequences includes the following steps:

[0015] a) Use the maximum mean discrepancy to describe the distance difference D between the two sequences:

[0016]

[0017] where the maximum mean discrepancy (MMD) is used to measure the distance between two different but related random variable distributions, that is, the distance between elements s S , s T in sequences S s and s t . In the formula, the function is the projection function that maps samples s s and s t to a unified space, and then calculates the sum of the means of the samples of the two distributions on the function, and then takes the difference.

[0018] b) Use the distance difference D to calculate the regularization term L of the loss function:

[0019] L = Lc + 0.25 * MMD 2 (S S , S T )

[0020] Among them, L C is the loss function of the pre-trained model itself, representing the difference between the output sample and the true sample, while MMD represents the difference between the source domain and target domain samples.

[0021] Furthermore, step 2) above corrects the pre-trained model, and the corrected model M S ' has some common characteristics of programming languages and can better fit the target language. Among them, the pre-trained model Ms is a SequenceToSequence model. This model consists of two variant long short-term memory artificial neural networks (LSTM) of recurrent neural networks (RNN), namely the encoder and the decoder. Among them, the encoder processes an input sequence into a vector c representation of a fixed dimension, which hides the memory information of this input sequence, and the decoder uses the vector c to generate an output sequence. Each LSTM is a neural network composed of a hidden state h and an optional output y, which is applied to a variable sequence X = <x1, x2, x3,..., x T >. The model M S , M S ', and M T mentioned in the present invention all belong to this kind of model, and the difference between the models lies in the different internal activation functions.

[0022] Furthermore, the correction of the pre-trained model M S includes the following steps:

[0023] a) Assume that for each time step t of this LSTM, the update function of the hidden state h t is: h t = f(h t-1 , x t ), and the update function of the optional output y t is y t = Ф(h t ) = g(h t -1, x t ). Among them, f and g are the non-linear activation functions of the specific LSTM. Among them, x t represents the input obtained by the LSTM model at time step t.

[0024] b) For the sequence S T manually select tokens containing different code semantics as the target language semantic standard set set. Among them, tokens with different semantics refer to tokens with different semantics such as macro definitions, reserved words, operators, user variables, comments, etc. in programming languages.

[0025] c) For S SAs model M S For the input sequence, for each x in the input sequence X = <x1, x2, x3, …, x t > i , calculate the distance d between x i and the target language semantic standard set, and modify the update function of h t to so as to change its weight. The calculation formula of the distance d is as follows:

[0026]

[0027] where N refers to all token sets in the language semantic standard set set, and n represents the specific token selected by traversing the set. In the formula the function is the projection function that maps the sample n and x i to the unified space, and then calculates the sum of the means of the samples of the two distributions on the function, and then finds the difference.

[0028] d) Use the regularization term L obtained in step 1) to modify the update function of y t to y t = Ф(h t ) = g(h t -1, x t ) + L.

[0029] e) In the training of the pre-trained model M S using the source language sequence S S , apply the above modifications to obtain M S ’.

[0030] Furthermore, the above step 3) is to train the final generation model using the target data. The training effect of multiple groups of shorter input sequences is better than that of a single input sequence during the training process. Therefore, S T is sliced into multiple short sequences of the same length. The training of the generation model includes the following steps:

[0031] a) Divide the target language sequence S T into multiple training sequences of a fixed size p as the input sequence of the model. Through sequence division, the i-th training sequence is X i = S T [i*p:(i + 1)*p], where S[i:j] refers to the subsequence of S between indices i and j.

[0032] b) The output sequence of the model is the sequence obtained by shifting the training sequence one position to the left. That is, the output sequence O i = S T [i*p + 1:(i + 1)*p + 1].

[0033] c) Fix the parameters of the bottom - most network of the pre - trained model M S ’, and keep the parameters of the remaining layers of the network as the initial parameters for training.

[0034] d) Use the training sequence X i and the output sequence O i to train the activation functions f and g of the model, and obtain the model M T , so that the input sequence can reduce the cross - threshold with the output sequence after being processed by the model. The actual learning objective is to obtain the probability distribution of the next token based on the prefix sequence.

[0035] Furthermore, in step 4), select a program sequence in the target language sequence S T as the seed program for fuzz testing. And select a prefix sequence in the seed program, use the encoder part of the model M T to process this prefix, and use the M T decoder and sampling strategy to continuously generate a token sequence until a termination condition is encountered, or the length of the token sequence exceeds the upper limit. These tokens finally form a new program code snippet, and this code snippet will generate a new program based on the replacement strategy in combination with the seed program. Use this program as the input for fuzz testing.

[0036] Furthermore, generating the target - language program as a fuzz - testing example in step 4) includes the following steps:

[0037] a) Select a program from the test set D T as the seed program, randomly select an integer bg as the starting line number, and len as the number of lines of the prefix sequence. Take the lines [bg:bg + len] in the program code as the prefix sequence S p for generating a new code snippet. Use S p as the input sequence of the encoder of the model M T and encode it into a vector c that hides the information of this sequence. The decoder of the model M T uses c and the initial y0 = <bos>Decode. Among them <bos>The start symbol representing the target language program.

[0038] b) During the decoding process, each LSTM at time step t outputs a probability distribution p(y t |y t+1 ,…,y1,c) = g(h t ,…,y1,c) based on c, h, and y, which represents the probability distribution of generating the next token based on the sequence S t and the sequence <y t ,…,y1>. p and the sequence <y t ,…,y1>.

[0039] c) Perform top-3 sampling on this probability distribution, i.e., randomly select one token from the top three tokens with the highest distribution probabilities each time as y t+1 , and use it as the input to the LSTM at the next time step t + 1.

[0040] d) Until generated in this way <eos>, as the end symbol of the sequence, or when the length of the generated token sequence exceeds a threshold, for example, the threshold can be set to 3000. Take the generated token sequence as the code snippet, and record the number of lines of the generated code snippet as len2.

[0041] e) Delete the last len2 lines of code after the prefix sequence of the original seed program, and insert the newly generated code snippet at the deleted code position to replace the original seed program, obtaining a new test program.

[0042] f) Use this new test program as the input of the target compiler for fuzz testing.

[0043] A cross - language compiler vulnerability mining device based on transfer learning, which includes:

[0044] A pre - processing module for selecting a model that generates fuzz test cases based on learning as the pre - trained model M S , and on this basis, select a source - language dataset and a target - language dataset, and serialize them into a source - language sequence S S and a target - language sequence S T , and calculate the distance difference D between the two sequences and the correction regularization term L of the model;

[0045] A model correction module for the pre - trained model M of the source language S , according to the obtained correction regularization term L, and the sequences S S and S T between the differences, correct the pre - trained model M S to obtain the corrected model M S ';

[0046] A model fine - tuning module for the corrected M S ', using the target - language sequence S T generated based on the target dataset D T to perform fine - tuning training, and finally obtain the generation model M T ;

[0047] A test case generation module for generating a target - language program as a test case according to the generation model M T and using the test case for fuzz testing to mine vulnerabilities.

[0048] The beneficial effects of the present invention are:

[0049] By adjusting the existing pre - trained generation model and then training with the target language to obtain the generation model of the target language, the present invention can generate program test cases and mine vulnerabilities, which can effectively improve the efficiency and effectiveness of security analysts in vulnerability detection.

[0050] Through the model fine-tuning module, the present invention can quickly train a generation model for a target language program, and this training process reduces the requirements for the quality of the training set of the target language program. Therefore, it can be effectively applied to the generation of fuzz testing cases for compilers in different programming languages, improving language-independent generality and having cross-language characteristics. Description of the Drawings

[0051] Figure 1 is a flowchart of a cross-language compiler vulnerability mining method based on transfer learning;

[0052] Figure 2 is a flowchart for correcting a pre-trained model;

[0053] Figure 3 is a flowchart for obtaining a generation model using a pre-trained model;

[0054] Figure 4 is a flowchart for generating a target language program using a generation model for fuzz testing. Detailed Embodiments

[0055] The present invention will be further described below in conjunction with the drawings through embodiments, but the scope of the present invention is not limited in any way.

[0056] The overall process of the vulnerability mining method of the present invention based on transfer learning and fuzz testing is as Figure 1 shown. Taking the fuzz testing of a C++ compiler g++ to mine vulnerabilities as an example, it specifically includes:

[0057] 1) Select source language and target language data sets, serialize them, and calculate the distance difference D and the correction regularization term L between the two sequences. The specific description is as follows:

[0058] 1a) Select the publicly available model "DeepFuzz" as the pre-trained model M S on the open source platform, C language as the source language, and C++ language as the target language. Collect the test suite programs of C language and C++ language respectively in the SARD test library, and splice the test programs into large sequences S S and S T . Go to 1b).

[0059] 1b) Use the maximum mean discrepancy (MMD) to calculate the distance difference D between the two sequences. The formula is as follows:

[0060]

[0061] And use the loss function to calculate the correction regularization term L. The formula is as follows:

[0062] L = Lc + 0.25 * MMD 2 (S S , S T )

[0063] 2) For the pre - trained model M of the source language S Make corrections to obtain M S ’. Its flowchart is as Figure 2 shown and is specifically described as follows:

[0064] 2a) For the sequence S T Manually select tokens with different code semantics as the target - language semantic standard set set. Go to 2b).

[0065] 2b) Take S S as the input sequence of the model M S For each xi in the input sequence X = <x1, x2, x3, …, x t >>, calculate the distance d between xi and the target - language semantic standard set, and modify the update function of h t to be thus changing its weight. Go to 2c). The calculation formula for the distance d is as follows:

[0066]

[0067] where N refers to all token sets within the language semantic standard set set, and n represents a specific token selected by traversing within the set. In the formula the function is the projection function that maps the sample n and x i to a unified space, then calculates the sum of the means of the samples of the two distributions on the function and then takes the difference.

[0068] 2c) Use the regularization term L obtained in step 1) to modify the update function of y t to be y t = Ф(h t ) = g(h t - 1, x t ) + L. Go to 2d).

[0069] 2d) In the training of the pre - trained model M S using the source - language sequence S S apply the above modifications to obtain M S ’.

[0070] 3) For the corrected pre - trained model M S ’, use the target - language data for fine - tuning training to obtain the model M T . Its process is as Figure 3 As shown below, the specific description is as follows:

[0071] 3a) Divide the target language sequence S T into multiple training sequences of a fixed size p as the input sequences of the model. The i-th training sequence is X i = S T [i*p:(i + 1)*p]. Go to 3b).

[0072] 3b) The output sequence O obtained by shifting the elements of the training sequence X one position to the left, where O i = S T [i*p + 1:(i + 1)*p + 1]. Go to 3c).

[0073] 3c) Fix the parameters of the lowest layer network of the pre-trained model M S ', and use the parameters of the remaining layer networks as the initial parameters for training. Use the training sequence X i and the output sequence O i , to train the modified pre-trained model M S ' to obtain the model M T .

[0074] 4) According to the model M T , generate the target language program as a fuzzing test case and perform fuzzing testing to discover vulnerabilities.

[0075] Its process is as Figure 4 shown below, and the specific description is as follows:

[0076] 4a) Select a program from the C++ language test suite as the seed program. Randomly select an integer bg as the starting line number and len as the number of lines of the prefix sequence. Take the lines [bg:bg + len] in the program code as the prefix sequence S P . Go to 4b).

[0077] 4b) Use the prefix sequence S P as the input sequence of the encoder of the model M T and encode it into a vector c. The decoder of the model M T utilizes the vector c and the initial y0 = <bos>Decode. Go to 4c).

[0078] 4c) During the decoding process, each LSTM outputs a probability distribution p(y t |y t+1 ,…,y1,c) = g(h t ,y t ,c) at time step t based on c, h, and y. Sample the top-3 from this probability distribution, i.e., randomly select one token from the three tokens with the highest occurrence probabilities each time as y t , and use it as the input to the LSTM at the next time step t+1. Go to 4d). t+1

[0079] 4d) Repeat the decoding process in 4c) until generated in this way <eos>As the end symbol of the sequence, or when the length of the generated token sequence exceeds the threshold of 3000. Use the generated token sequence as a code snippet, and record the number of lines of the code snippet as len2. Go to 4e).

[0080] 4e) Delete the len2 lines of code after the prefix sequence of the original seed program, and insert the newly generated code snippet into the deleted code position to replace the original seed program, obtaining a new test program. Go to 4f).

[0081] 4f) Use this new test program as the input to the target compiler g++ for fuzz testing. Continuously repeat the steps of selecting the seed program in 4a) for repeated input testing until the fuzz testing process is manually stopped.

[0082] Another embodiment of the present invention provides a cross-language compiler vulnerability mining device based on transfer learning, which includes:

[0083] A preprocessing module for selecting a model for generating fuzz test cases based on learning as the pre-trained model M S , and on this basis, selecting a source language dataset and a target language dataset, serializing them into a source language sequence S S and a target language sequence S T , and calculating the distance difference D between the two sequences and the correction regularization term L of the model;

[0084] A model correction module for the pre-trained model M of the source language S , according to the obtained correction regularization term L, and the sequences S S and S T to correct the pre-trained model M S to obtain the corrected model M S ';

[0085] A model fine-tuning module for fine-tuning and training the corrected M S ' using the target language sequence S T generated based on the target dataset D T to finally obtain the generation model M T ;

[0086] A test case generation module for generating a target language program as a test case according to the generation model M T and using the test case for fuzz testing to mine vulnerabilities.

[0087] For the specific implementation process of each module, refer to the description of the method of the present invention above.

[0088] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for performing the steps in the method of the present invention.

[0089] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disk). When the computer program stored in the computer-readable storage medium is executed by a computer, the various steps of the method of the present invention are implemented.

[0090] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.< / eos> ​< / bos> < / eos> < / bos> < / bos>

Claims

1. A cross - language compiler vulnerability mining method based on transfer learning, the steps of which include: Select the model for learning-based generation of fuzz test cases as the pre-trained model M S , and on this basis, select the source language dataset and the target language dataset, and serialize them into the source language sequence S S and the target language sequence S T , and calculate the distance difference D between the two sequences and the correction regularization term L of the model; For the pre-trained model M S , according to the obtained corrected regularization term L, and the sequences S S and S T the difference between them corrects the pre-trained model M S to obtain the corrected model M S ’ ; For the corrected M S ’ , using the target language sequence S T generated based on the target language dataset D T to perform fine-tuning training, and finally obtaining the generation model M T ; Generate a target language program according to the generation model M T Generate a target language program as a test case, and use the test case for fuzz testing to discover vulnerabilities; The distance difference D is calculated using the maximum mean discrepancy, and the corrected regularization term L is calculated using a loss function; the calculation formula of D is as follows: Among them, the maximum mean discrepancy (MMD) is used to measure the distance between the distributions of two different but related random variables, that is, for the sequences S S , S T the distance between elements s s and s t ; The function is a projection function that maps the samples s s and s t to a unified space; The calculation formula of L is as follows: L = Lc + 0.25 * MMD 2 (S S ,S T ) Where L is composed of the loss function Lc of the pre - trained model itself and MMD. The loss function of itself represents the difference between the output sample and the real sample, and MMD represents the difference between the source - domain and target - domain samples; The pre-trained model M S is corrected by modifying the activation function of each neural network in the pre-trained model through the token difference between sequences and the correction regularization term L, including the following steps: For the target language sequence S T Manually select tokens with different code semantics as the target language semantic standard set; calculate the distance d between each element in the input sequence and the target language semantic standard set, and use this distance to modify the hidden state h t The update function of Modify y using the corrected regularization term L t The update function of y t = Ф(h t ) = g(h t -1, x t ) + L Among them, x t represents the input obtained by the LSTM model at time sequence t, and f and g are the non-linear activation functions of the LSTM.

2. The method according to claim 1, wherein The serialization method of the source - language dataset and the target - language dataset is: regarding each program constituting the dataset as a sequence composed of tokens, and directly splicing multiple programs in the dataset into an entire sequence.

3. The method according to claim 1, characterized in that The corrected M S ’ , and use the target language sequence S T generated based on the target language dataset D T for fine-tuning training, including: using the target language sequence S T to generate input sequences and output sequences for training. During the training process, fix the parameters of the lowest layer neural network of the original model and use the parameters of the remaining layers of the neural network as the initial parameters.

4. The method according to claim 1, wherein The generation model M T generates a target language program as a test example, including: randomly selecting a seed program and a prefix sequence therein from a target language dataset, and using the prefix sequence and the generation model M T to obtain a code snippet, and using the code snippet to replace the subsequent code segment of the prefix sequence in the original seed program.

5. The method according to claim 4, wherein The said according to the generation model M T Generating a target language program as a test example, including: a) Select a program from the target language dataset D T as the seed program, randomly select an integer bg as the starting line number and len as the number of lines of the prefix sequence, and take the [bg:bg+len] lines in the program code as the prefix sequence S p , and take S p as the input sequence of the encoder of the model M T , and encode it into a vector c that hides the sequence information. The decoder of the model M T uses c and the initial y0 = <bos>Decode; among which <bos>Represents the start symbol of the target - language program; < / bos> < / bos> b) During the decoding process, each LSTM at time t is based on c, h and y t Output a probability distribution p(y t+1 |y t ,…,y1,c)=g t (h t ,y t ,c), which represents the sequence S p and sequence <y t ,…,y1> and generate the probability distribution of the next token; c) Perform top-3 sampling on this probability distribution, that is, randomly select one token from the top three tokens with the highest distribution probabilities each time as y t+1 , and use it as the input to the LSTM at the next time step t+1; d) until generated by this method <eos>As the end symbol of the sequence, or when the length of the generated token sequence exceeds the threshold, taking the generated token sequence as a code snippet, and recording the number of lines of the generated code snippet as len2; < / eos> e) Deleting the last len2 lines of code after the prefix sequence of the original seed program, and inserting the newly generated code snippet into the deleted code position to replace the original seed program, obtaining a new test program; f) Using this new test program as the input of the target compiler for fuzz testing.

6. A cross-language compiler vulnerability mining device based on transfer learning using the method according to any one of claims 1 to 5, characterized in that, Includes: A preprocessing module for selecting a model that generates fuzz test cases based on learning as a pre-trained model M S , and on this basis, selecting a source language dataset and a target language dataset, and serializing them into a source language sequence S S and a target language sequence S T , and calculating the distance difference D between the two sequences and the correction regularization term L of the model; A model correction module for the pre-trained model M S , according to the obtained correction regularization term L, and the sequence S S and S T to correct the pre-trained model M S to obtain the corrected model M S ’ ; A model fine-tuning module for fine-tuning the corrected M S ’ , using the target language sequence S T generated based on the target language dataset D T for fine-tuning training to finally obtain the generation model M T ; A test case generation module, which is used to generate a target language program as a test case according to the generation model M T and perform fuzz testing using the test case to discover vulnerabilities.

7. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores a computer program, and when the computer program is executed by a computer, it implements the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Joint information extraction method based on weak supervised learning

    CN110826303A

  • Small target identification method and system based on space-time neural network

    CN113160050A