Software unit test code automatic generation method based on retrieval and editing combination

CN116820484BActive Publication Date: 2026-09-29CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310860338.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2026-09-29
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

通过调研发现,①Integration方法的整体性能主要归功于它在检索断言方面的成功;②Integration方法很难理解检索到的焦点测试和输入的焦点测试之间的语义差异,导致许多标记被错误地修改;③Integration方法仅限于特定类型的编辑操作(即替换),并且不能处理令牌添加或删除

Benefits of technology

[0051]本发明通过检索输入焦点测试的相似焦点测试的断言视为原型,并将原型与输入焦点测试和类似焦点测试之间的语义差异所反映断言编辑模式相结合来生成目标测试断言。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116820484B_ABST
    Figure CN116820484B_ABST
Patent Text Reader

Abstract

The application relates to a software unit test code automatic generation method based on retrieval and editing combination. Based on a given input focus test and a corpus, a retrieval component is used to calculate the similarity of each focus test in the input focus test and the corpus based on a Jaccard similarity algorithm, to obtain a similar focus test with the highest similarity value in the corpus and a corresponding similar test assertion; an editing-based component is used to learn the editing mode of the input focus test and the similar focus test instance, and the editing mode is applied to editing of the similar test assertion, so that a target test assertion is generated. The method is much better than the most advanced baseline, and the method can be applied to actual working scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to software test assertion generation, and more particularly to a method for automatically generating software unit test code based on a combination of retrieval and editing. Background Technology

[0002] Unit testing is a crucial activity in software development, involving testing individual units of a software application, such as methods, classes, or modules. While integration and system testing evaluate the overall performance of a system, unit testing focuses on verifying that each unit of code works as the developers expected and envisioned, detecting and diagnosing faults before they propagate to the entire system, and preventing regressions. Therefore, effective unit testing can improve software quality, reduce the incidence and cost of software failures, and improve the overall software development process. Unit tests consist of test prefixes and test assertions. A test prefix is ​​a series of statements that operate on the unit under test to obtain a specific state, while test assertions typically include assertions specifying the expected behavior in that state.

[0003] While testing offers significant benefits, creating effective unit tests is a non-trivial and time-consuming task. Previous research has shown that developers can spend over 15% of their time on test generation. To simplify unit test generation, various automated testing tools, such as Randoop and EvoSuite, have been proposed. However, these test generation tools prioritize generating high-coverage tests rather than meaningful assertions and remain difficult to interpret in terms of expected program behavior, thus failing to replace the need for manual unit testing.

[0004] To overcome the challenges of assertion generation, numerous test assertion generation methods have been proposed. Meanwhile, with the development of deep learning technology and the ever-increasing volume of source code data, automatically learning code summaries from a large number of test assertions using deep learning models has become a very popular research topic. Recently, the deep learning-based test assertion generation method ATLAS circumvents the problem of low scenario applicability of traditional rule-based generation methods. However, test assertions generated from scratch typically favor high-frequency words in the corpus and may encounter problems with low-frequency words, such as item-specific identifiers, and perform poorly when generating long sequences of test assertions. Currently, a state-of-the-art ensemble method (called Integration) has been proposed, which combines information retrieval (IR) with deep learning-based methods to generate assertions for unit tests. The integration method verifies the compatibility between the retrieved assertions and the current focus test (focal-test). If the compatibility exceeds a threshold, the retrieved assertion is returned as the final result. Otherwise, the deep learning-based method generates the assertion. Research revealed that: ① the overall performance of the integration method is primarily attributed to its success in retrieving assertions; ② the integration method struggles to understand the semantic differences between the retrieved focus test and the input focus test, leading to many tokens being incorrectly modified; ③ the integration method is limited to specific types of edit operations (i.e., replacement) and cannot handle token addition or deletion. Summary of the Invention

[0005] To alleviate the above limitations and improve the effectiveness of assertion generation, this invention proposes an automatic software unit test code generation method based on a combination of retrieval and editing.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: a method for automatically generating software unit test code based on a combination of retrieval and editing, comprising component one and component two, wherein component one includes the following steps:

[0007] S101: Input Focus Test Each input focus test is preceded by an input test prefix. and input test method composition.

[0008] S102: Regarding Calculate using the Jaccard similarity algorithm and every focus test in the corpus similarity Finally, similarity vectors are obtained.

[0009] S103: For each The similarity focus test with the highest similarity value in the corpus was obtained. and its corresponding similar test assertions Among them, the similarity focus test is prefixed with similarity test. and similar test methods composition.

[0010] Component 2 includes the following steps:

[0011] S104: Use the diff tool to... and Perform comparisons and create edit sequences based on the comparison results.

[0012] S105: respectively for and Perform word embedding and access contextual information. Each editor in E ij Convert to text embedding vector h′ j , Each token x in h Convert to text embedding vector h h The final result is the edit sequence text embedding vector H′=[h′1,h′2,…,h′] s The similarity test assertion text embedding vector H = [h1, h2, ..., h] is used. t ]. Where s represents the edit sequence. The number of edits, where t represents the similarity test assertion. The number of tokens.

[0013] S106: Construction and Shared attention layer for fusion and Information, and capture and The relationship between them.

[0014] S107: Using two different bidirectional LSTM pairs Each token x h and Each editor in E ij The corresponding generated final representation vector z h and z′ j The final representation matrix is ​​obtained as Z = [z1, z2, ..., z]. t ],Z′=[z′1,z′2,…,z′ s ].

[0015] S108: Using the outputs Z and Z′ of the two encoders as input, generate the target test assertion using an LSTM-based decoder. output .

[0016] Preferably, the similarity matrix in S102 The calculation steps are as follows: Calculate The formula is:

[0017]

[0018] Where |·| represents the number of elements in the set.

[0019] For each Obtain similar vectors Where n represents the total number of data in the corpus.

[0020] Preferably, S104 creates an edit sequence. The specific steps are as follows: Each element in the array is a triple. This is called editing. Among them, yes One of the tokens, yes One of the tokens, a j yes Convert to Editing operations.

[0021] There are four editing operations: insertion, deletion, equality, or replacement. When a... j When it is an insert or delete operation, it means or It will be an empty token.

[0022] Preferably, the step of S105 to obtain the edit sequence text embedding vector H′ and the similarity test assertion text embedding vector H is as follows: First, using a pre-trained model, obtain the edit sequence text embedding vector H′ and the similarity test assertion text embedding vector H′ respectively. and Word embedding sequence.

[0023] Then, a bidirectional long short-term memory (Bi-LSTM) is used to process the word embedding sequence to access contextual information. For editing sequences... Each editor in First, horizontal connections and a j Then input it into the Bi-LSTM, as shown below:

[0024]

[0025]

[0026]

[0027] Where, h′ j Editor E ij The context vector, It is a concatenation operation; similarly, the assertion encoder obtains each assertion. Token x h context vector h h The final result is the edit sequence text embedding vector H′=[h′1,h′2,…,h′] s The similarity test assertion text embedding vector H = [h1, h2, ..., h] is used. t ].

[0028] Preferably, S106 includes the following step: the attention layer sets H′=[h′1,h′2,…,h′ s ] and H = [h1,h2,…,h t As input, they correspond to edit sequences respectively. Similar test assertions And for each h′ j Output the corresponding feature vector g′ j and the original context vector h′ j , for each h h Output the corresponding feature vector g h and the original context vector h h .

[0029] g′ j The formula is shown below, where attention weight α′ j Measure each assertion token x j Compared to editor E ij Importance:

[0030]

[0031] Among them, H T This represents the transpose of the similarity test assertion text embedding vector H. This represents the transpose of the trainable parameters.

[0032] eigenvector g h The formula is shown below, where attention weight α h Measuring each editor E ij Compared to assertion token x h importance:

[0033] g h =H′αh ;α h =softmax(H′) T W α h h (6);

[0034] Among them, H′ T W represents the transpose of the edit sequence text embedding vector H′. α This represents the trainable parameters.

[0035] Preferably, S107 calculates z h and z′ j The steps are as follows: For Each token x h z h By h h and g h The calculation is as follows:

[0036]

[0037] for Each editor in E ij z′ j From h′ j and g′ j The calculation is as follows:

[0038]

[0039] Preferably, step S108 includes the following steps:

[0040] During decoding step k, the decoder embeds the k-th word based on the basic fact assertion. Previously hidden state s k-1 and the previous output vector o k-1 To calculate the hidden state s k As shown below:

[0041]

[0042] Then, the context vector for each time step is calculated as the representation of the encoder input for the dot product attention mechanism, as shown in Equation (6). Given two encoders, the decoder obtains two context vectors: ck from the retrieved assertions and c′ from the focus test edit sequence. k Using C k c′ k and s k Calculate the output vector o k The corresponding vocabulary distribution is obtained using a softmax layer.

[0043]

[0044] Where V c and V′ c These are trainable parameters. It records the probability of each token being generated, and the token with the highest probability will be output in decoding step k.

[0045] The token is copied from the retrieved similar test assertions and input focus tests using a pointer generator:

[0046]

[0047] in and These are copies of y from the retrieved assertions and input focus tests, respectively. k The probability, y k This represents the token output in decoding step k. β kl and β′ kl It is y l and E l Attention weights at time step k, y l E represents a token indicating a similarity test assertion. l This represents the token used for the input focus test. k The conditional probability at time step k is and The combination, that is,

[0048]

[0049] Where γ k and θ k These represent generating y by selecting from the vocabulary and copying from the retrieved assertions, respectively. k The probability of.

[0050] Compared with the prior art, the present invention has at least the following advantages:

[0051] This invention generates target test assertions by retrieving assertions from similar focus tests of the input focus test as prototypes and combining the prototypes with assertion editing patterns reflected by semantic differences between the input focus test and similar focus tests.

[0052] Experimental results demonstrate that the method of this invention significantly outperforms state-of-the-art baselines. For assertion generation tasks, retrieving similar assertions and learning to modify the retrieved assertions by applying a set of editing operations yields satisfactory performance. Furthermore, this invention can be applied to real-world work scenarios. Attached Figure Description

[0053] Figure 1 This is a simplified flowchart of the method of the present invention. Detailed Implementation

[0054] The present invention will now be described in further detail.

[0055] The core idea of ​​this invention is to treat the assertions of similar focus tests as prototypes and utilize neural sequence-to-sequence models to learn assertion editing patterns for modifying the prototypes. The motivation for this invention is that the retrieved assertions guide the neural model on "how to assert," while the assertion editing patterns emphasize "what to assert" to the neural model. This invention can: (1) comprehensively understand the semantic differences between the input and similar focus tests; (2) flexibly apply appropriate assertion editing patterns; and (3) generate diverse editing operations.

[0056] This invention consists of two main components: a retrieval component and an editing component. The retrieval component retrieves similar focus tests from a corpus for a given input focus test and uses the test assertions of the retrieved similar focus tests as prototypes. The editing component trains a sequence-to-sequence neural network to learn editing patterns for a given input focus test and similar focus tests, and edits and modifies the prototypes to generate test assertions.

[0057] Specifically, in component one, given input focus test Input focus test based on Jaccard similarity algorithm and every focus test in the corpus similarity Finally, the similarity focus test with the highest similarity value in the corpus was obtained. and its corresponding similar test assertions In component two, for a given input focus test and similar focus test examples and their corresponding assertions and The neural editing model aims to find a function f such that Therefore, for an input focus test, a target test assertion is generated. output .

[0058] See Figure 1 A method for automatically generating software unit test code based on a combination of retrieval and editing includes component one and component two. Component one obtains similar focus tests and corresponding test assertions based on a similarity retrieval corpus of input focus tests. Component two takes the test assertions corresponding to similar focus tests as prototypes, combines them with the assertion editing mode reflected by the semantic differences between input focus tests and similar focus tests, and edits the prototypes to generate target test assertions.

[0059] A method for automatically generating software unit test code based on a combination of retrieval and editing, comprising component one and component two, wherein component one includes the following steps:

[0060] S101: Input Focus Test Each input focus test is preceded by an input test prefix. and input test method composition.

[0061] S102: Regarding Calculate using the Jaccard similarity algorithm and every focus test in the corpus similarity Finally, similarity vectors are obtained.

[0062] S103: For each The similarity focus test with the highest similarity value in the corpus was obtained. and its corresponding similar test assertions Among them, the similarity focus test is prefixed with similarity test. and similar test methods Composition. For each one The highest similarity value of Jaccard in the corpus was obtained. And obtain the corresponding similarity focus test with index x. Similar test assertions

[0063] Component 2 includes the following steps;

[0064] S104: Use the diff tool to... and Perform comparisons and create edit sequences based on the comparison results.

[0065] S105: respectively for and Perform word embedding and access contextual information. Each editor in E ij Convert to text embedding vector h′ j , Each token x in h Convert to text embedding vector h h The final result is the edit sequence text embedding vector H′=[h′1,h′2,…,h′] s The similarity test assertion text embedding vector H = [h1, h2, ..., h] is used. t ]. Where s represents the edit sequence. The number of edits, where t represents the similarity test assertion. The number of tokens.

[0066] S106: Construction and Shared attention layer for fusion and Information, and capture and The relationship between them.

[0067] S107: Using two different bidirectional LSTM pairs Each token x h and Each editor in E ij The corresponding generated final representation vector z h and z′ j The final representation matrix is ​​obtained as Z = [z1, z2, ..., z]. t ],Z′=[z′1,z′2,…,z′ s ].

[0068] S108: Using the outputs Z and Z′ of the two encoders as input, generate the target test assertion using an LSTM-based decoder. output .

[0069] Specifically, the similarity matrix in S102 The calculation steps are as follows:

[0070] calculate The formula is:

[0071]

[0072] Where |·| represents the number of elements in the set.

[0073] For each Obtain similar vectors Where n represents the total number of data in the corpus.

[0074] Specifically, S104 creates an edit sequence. The specific steps are as follows:

[0075] Each element in the array is a triple. This is called editing. Among them, yes One of the tokens, yes One of the tokens, a j yes Convert to Editing operations.

[0076] There are four editing operations: insertion, deletion, equality, or replacement. When a... j When it is an insert or delete operation, it means or It will be an empty token. Constructing such an edit sequence not only preserves focus tests (i.e.) and ) information, and can be obtained through a j Highlight their fine-grained differences.

[0077] Specifically, the steps in S105 to obtain the edit sequence text embedding vector H′ and the similarity test assertion text embedding vector H are as follows:

[0078] To capture syntactic and semantic information, a pre-trained model, such as fastText, is first used to acquire them separately. and Word embedding sequence.

[0079] Then, a bidirectional long short-term memory (Bi-LSTM) is used to process the word embedding sequence to access contextual information. For editing sequences... Each editor in First, horizontal connections and a j Then input it into the Bi-LSTM, as shown below:

[0080]

[0081]

[0082]

[0083] Where, h′ j Editor E ij The context vector, It is a concatenation operation; similarly, the assertion encoder obtains each assertion. Token x h context vector h h The final result is the edit sequence text embedding vector H′=[h′1,h′2,…,h′] s The similarity test assertion text embedding vector H = [h1, h2, ..., h] is used. t ].

[0084] Specifically, step S106 includes the following steps:

[0085] The attention layer will have H′=[h′1,h′2,…,h′]s ] and H = [h1,h2,…,h t As input, they correspond to edit sequences respectively. Similar test assertions And for each h′ j Output the corresponding feature vector g′ j and the original context vector h′ j , for each h h Output the corresponding feature vector g h and the original context vector h h ;

[0086] g′ j The formula is shown below, where attention weight α′ j Measure each assertion token x j Compared to editor E ij Importance:

[0087]

[0088] Among them, H T This represents the transpose of the similarity test assertion text embedding vector H. This represents the transpose of the trainable parameters.

[0089] eigenvector g h The formula is shown below, where attention weight α h Measuring each editor E ij Compared to assertion token x h importance:

[0090] g h =H′α h ;α h =softmax(H′) T W α h h (6);

[0091] Among them, H′ T W represents the transpose of the edit sequence text embedding vector H′. α This represents the trainable parameters.

[0092] Specifically, S107 calculates z h and z′ j The steps are as follows:

[0093] for Each token x h z h By h h and g h The calculation is as follows:

[0094]

[0095] for Each editor in E ij z′ j From h′ j and g′ j The calculation is as follows:

[0096]

[0097] Specifically, step S108 includes the following steps:

[0098] To construct the initial state s0 of the LSTM, the final representation matrix Z = [z1, z2, ..., z] output by the encoder is used. t ] and Z′=[z′1,z′2,…,z′ s Connect them. During decoding step k, the decoder embeds the k-th word based on the basic fact assertion. Previously hidden state s k-1 and the previous output vector o k-1 To calculate the hidden state s k As shown below:

[0099]

[0100] Then, the context vector for each time step is calculated as the representation of the encoder input for the dot product attention mechanism, as shown in Equation (6). Given two encoders, the decoder obtains two context vectors, namely c from the retrieved assertions. k and c′ from the focus test edit sequence k Using C k c′ k and s k Calculate the output vector o k The corresponding vocabulary distribution is obtained using a softmax layer.

[0101]

[0102] Where V c and V′ c These are trainable parameters. It records the probability of each token being generated, and the token with the highest probability will be output in decoding step k.

[0103] Due to the similarity of the focus tests, it is reasonable to assume that some tokens in the new assertion should also appear in the retrieved assertions, while other tokens not present in the retrieved assertions should be included in the input focus test. Therefore, a pointer generator is used to copy tokens from the retrieved similar test assertions and the input focus test:

[0104]

[0105] in and These are copies of y from the retrieved assertions and input focus tests, respectively. k The probability, y k This represents the token output in decoding step k. β kl and β′ kl It is y l and E l Attention weights at time step k, y l E represents a token indicating a similarity test assertion. l This represents the token used for the input focus test. k The conditional probability at time step k is and The combination, that is,

[0106]

[0107] Where γ k and θ k These represent generating y by selecting from the vocabulary and copying from the retrieved assertions, respectively. k The probability of.

[0108] The data used in this invention comes from two publicly available datasets provided by Yu et al. in their paper, namely Data old and Data new .

[0109] (1) Data old :Data old This data originates from the original dataset used by the ATLAS method. Initially, Data... old It was extracted from a pool of 2.5 million test methods on GitHub, including test prefixes and their corresponding assertion statements. For each test method, Data... old All include focus methods. Then, for Data... old Preprocessing is performed to exclude test methods with tag lengths exceeding 1K, and assertions containing focus tests and unknown tags not present in the vocabulary are filtered out, following established natural language processing practices. After removing duplicates, the Data... oldWe obtained 156,760 data points, which were further divided into training, validation, and test sets in an 8:1:1 ratio.

[0110] (2) Data new Excluding assertions with unknown tokens may oversimplify the assertion generation problem, making Data... old This is unsuitable for representing the true data distribution. This, in turn, poses a significant threat to the validity of the experimental conclusions. Therefore, Yu et al. added those data points that were not representative of the true data distribution. old The samples with unknown labels that were excluded from the dataset were used to construct an expanded dataset, denoted as Data. new Besides Data old In addition to the existing data items, Data new It also includes 108,660 samples with unknown tokens, forming a total of 265,420 data sets, which are further divided into training, validation, and test sets in an 8:1:1 ratio.

[0111] To verify the effectiveness of this invention, we compared it with five baselines. We first chose ATLAS, the first and classic neural network-based assertion generation method. ATLAS utilizes a sequence-to-sequence model to generate assertions from scratch. Given that EDITAS aims to re-examine and improve retrieval-enhanced software assertion generation methods, we adopted three state-of-the-art retrieval methods, including IR. ar , and And integration, a method for combining retrieval with deep learning. ar Using the same input as ATLAS, assertions most similar to a given focus test are retrieved based on the Jaccard similarity coefficient. Then, The tokens in the retrieved assertions are further adjusted based on the context. Furthermore, Integration combines IR-based and DL-based methods to improve assertion generation capabilities. The Integration method first verifies the compatibility between the retrieved assertions and the current focus test; if the compatibility exceeds a threshold, the retrieved assertion is returned as the final result. Otherwise, the DL-based method generates the assertion. The method proposed in this invention is called EDITAS.

[0112] This invention uses Accuracy and multi-BLEU scores as evaluation metrics. (1) Accuracy: A generated assertion is considered accurate if and only if it perfectly matches the ground truth. Accuracy determines the percentage of samples in which the generated output matches the expected output grammatically. (2) Multi-BLEU: BLEU calculates the corrected n-gram accuracy of the candidate sequence (i.e., the generated assertion) against the reference sequence (i.e., the ground truth), where n ranges from 1 to 4. The corrected n-gram accuracy values ​​are then averaged, and sentences that are too short are penalized.

[0113] This invention calculates the accuracy and BLEU score between assertions generated by different methods and manually written assertions. The experimental results are shown in Table 1. It can be seen that ATLAS performs the worst among all methods. This is mainly attributed to two reasons: 1) As a typical sequence-to-sequence DL model, ATLAS suffers from exposure bias and gradient vanishing, resulting in poor effectiveness of generating long sequence tokens as assertions. 2) ATLAS has a weak ability to generate statements containing unknown tokens, which significantly reduces its overall performance. ar Retrieving assertions from a corpus and using them as output yields better performance than ATLAS. This demonstrates that assertions based on focus-tests contain valuable and reusable information, justifying our use of similar focus-tests as prototype assertions. and Further adjustments were made to the retrieved assertions to enhance the assertion generation capability of the IR-based method. However, as shown in Table 1, and The performance of adaptive operations is limited, especially for complex datasets. For example, compared to IR... ar compared to, In Data old This can improve accuracy by 20.33%, while in Data... new The improved accuracy was only 6.94%. Integration combines IR and DL techniques and achieves better accuracy and BLEU scores than assertion generation methods based on ATLAS and IR.

[0114] As shown in Table 1, EDITAS achieves a significant performance improvement compared to ATLAS, with an average accuracy increase of 87.48% and a BLEU score increase of 42.65% across both datasets. This is because EDITAS utilizes the rich semantic information from retrieved assertions, rather than generating assertions from scratch. The proposed method, EDITAS, outperforms IR-based baseline methods and ensembles on all evaluation metrics. Specifically, compared to IR… ar , Compared to Integration, EDITAS achieved average accuracy improvements of 32.24%, 21.19%, 15.99%, and 10.00%, respectively, demonstrating the effectiveness of the editing module of this invention. Compared to an IR-based baseline, EDITAS uses retrieved assertions as a prototype and modifies them by considering semantic differences between the input and similar focus tests. By combining the advantages of neural networks and IR-based methods, EDITAS achieves optimal performance.

[0115] Table 1. Evaluation results of EDITAS on two datasets compared to five state-of-the-art test assertion generation methods.

[0116]

[0117] We further compared the effectiveness of EDITAS and the baseline method for different types of assertions. Table 2 shows the results on the Data dataset. old and Data new This section provides detailed statistics for each assertion type. Each column represents an assertion type, and the brackets in each cell show the number of assertions in the dataset and their corresponding percentage.

[0118] Table 2 Data old and Data new Detailed statistics for each type

[0119]

[0120] Tables 3 and 4 show the baselines for each dataset. old and Data new The above shows the effectiveness of EDITAS for each assertion type. Each column represents an assertion type, and the brackets in each cell show the number of correctly generated assertions and their corresponding ratio. The results show that EDITAS outperforms the baseline method for almost all assertion types, especially for standard JUnit assertion types. Overall, the experimental results demonstrate the versatility of EDITAS in generating different types of assertions.

[0121] Table 3 EDITAS and each baseline in Data old Validity of each assertion type on the dataset

[0122]

[0123] Table 4 EDITAS and each baseline in Data new Validity of each assertion type on the dataset

[0124]

[0125]

[0126] EDITAS has the following advantages: 1) EDITAS can learn and apply different assertion editing modes, and and Unable to process token addition or deletion operations. 2) and An assertion is modified only if the retrieved assertion contains at least one token that is not present in the input focus test. However, even if all tokens in the retrieved assertion appear in the input focus test, it may still need to be modified due to semantic differences between focus tests. Instead, EDITAS uses a probabilistic model to learn common patterns of assertion editing from the semantic differences between existing focus tests. Overall, the editing patterns learned by EDITAS are more diverse and can cover a wider range of samples.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for automatically generating software unit test code based on a combination of retrieval and editing, characterized in that: It includes component one and component two, wherein component one includes the following steps: S101: Input Focus Test Each input focus test is preceded by an input test prefix. and input test method composition; S102: Regarding Calculate using the Jaccard similarity algorithm and every focus test in the corpus similarity Ultimately, similar results were obtained. ; S103: For each The similarity focus test with the highest similarity value in the corpus was obtained. and its corresponding similar test assertions Among them, the similarity focus test is prefixed with similarity test. and similar test methods composition; Component 2 includes the following steps; S104: Use tool pair and Perform comparisons and create edit sequences based on the comparison results. ; S105: respectively for and Perform word embedding and access contextual information to... Each editor in Convert to text embedding vector , Each token in Convert to text embedding vector The final result is the edit sequence text embedding vector. Similar test assertion text embedding vector ,in Represents the edit sequence Number of edits Representative similarity test assertion The number of tokens; S106: Construction and Shared attention layer for fusion and Information, and capture and The relationship between them; S107: Using two different bidirectional LSTM pairs Each token and Each editor in The corresponding generated final representation vector and Finally, the final representation matrix is ​​obtained. , ; S108: Converts the outputs of the two encoders and As input, a target test assertion is generated using an LSTM-based decoder. ; S108 includes the following steps: In the decoding step During this period, the decoder asserts the first based on basic facts. Embedded characters Previously hidden state and the previous output vector To calculate the hidden state As shown below: (9); Then, the context vector for each time step is calculated as the representation of the encoder input for the dot product attention mechanism, as shown in Equation (6). Given two encoders, the decoder obtains two context vectors, namely, the context vectors from the retrieved assertions. and from the focus test edit sequence ,use , and Calculate the output vector and use The layer obtains the corresponding vocabulary distribution. ; (10); in and These are trainable parameters. It records the probability of each token being generated, with the token having the highest probability being used in the decoding step. The following will be output; The token is copied from the retrieved similar test assertions and input focus tests using a pointer generator: (11); in and These are copied from the retrieved assertions and input focus tests, respectively. The probability, Indicating in the decoding step The token that is output below. and yes and time step Attention weights Tokens representing similar test assertions, The token indicating the input focus test. time step The conditional probability is , and The combination is: (12); in and These respectively represent generation by selecting from the vocabulary and copying from the retrieved assertions. The probability of.

2. The method for automatically generating software unit test code based on a combination of retrieval and editing as described in claim 1, characterized in that: The similarity matrix in S102 The calculation steps are as follows: calculate The formula is: (1); in, Indicates the number of elements in the set; For each Obtain similar vectors , , where n represents the total number of data in the corpus.

3. The method for automatically generating software unit test code based on a combination of retrieval and editing as described in claim 2, characterized in that: S104 creates an edit sequence The specific steps are as follows: Each element in the array is a triple. = This is referred to as editing, where, yes One of the tokens, yes One of the tokens, yes Convert to Editing operations; There are four editing operations: insertion, deletion, equality, or replacement. When it is an insert or delete operation, it means or It will be an empty token. .

4. The method for automatically generating software unit test code based on a combination of retrieval and editing as described in claim 3, characterized in that: S105 obtains the edit sequence text embedding vector. Similar test assertion text embedding vectors The steps are as follows: First, using a pre-trained model, obtain... and Word embedding sequence; Then use bidirectional long short-term memory. To process word embedding sequences to access contextual information, for editing sequences Each editor in First, horizontal connection as well as Then enter As shown below: (2); ; (3); (4); in, Editor The context vector, It is a concatenation operation; similarly, the assertion encoder obtains each assertion. Token context vector The final result is the edit sequence text embedding vector. Similar test assertion text embedding vectors .

5. The method for automatically generating software unit test code based on a combination of retrieval and editing as described in claim 4, characterized in that: S106 includes the following steps: The attention layer will and As input, they correspond to edit sequences respectively. Similar test assertions and for each Output the corresponding feature vector and the original context vector , for each Output the corresponding feature vector and the original context vector ; The formula is shown below, attention weights Measure each assertion token Compared to editing Importance: (5); in, Represents similarity test assertion text embedding vectors transpose, This represents the transpose of the trainable parameters; Feature vector The formula is shown below, attention weights Measure each editor Compared to assertion tokens importance: (6); in, Represents the text embedding vector of the edit sequence. transpose, This represents the trainable parameters.

6. The method for automatically generating software unit test code based on a combination of retrieval and editing as described in claim 3, characterized in that: The calculation in S107 and The steps are as follows: for Each token of Depend on and The calculation is as follows: (7); for Each editor in of Depend on and The calculation is as follows: (8)。