Software regression testing method and device, equipment and storage medium

Through the combination of vector model and large language model, relevant test cases are automatically filtered and executed, the problem of inefficient regression testing of existing software is solved, and efficient and accurate software regression testing is achieved.

CN120029901APending Publication Date: 2025-05-23PCI TECH GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411893363.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing software regression testing methods are inefficient and have poor testing results. Especially in an environment where software is rapidly iterated, it is difficult to meet the needs of manual regression use cases.

Method used

Using a method of combining vector model and large language model, we obtain the software function descriptions corresponding to the test cases to be executed, vectorized processing and search, automatically filter out the software function descriptions most relevant to the question, and obtain the associated test scripts to perform regression testing.

Benefits of technology

It improves the efficiency and accuracy of test case selection, improves the quality and efficiency of software regression testing, and realizes a fully automated regression testing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029901A_ABST
    Figure CN120029901A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software testing, and discloses a software regression testing method and device, equipment and a storage medium. The software regression test method comprises the following steps: acquiring first software function description corresponding to a to-be-executed test case; inputting the first software function description as a question into a preset vector model for retrieval, and outputting a plurality of retrieval results related to the question and retrieved from a preset vector database by the vector model; the retrieval result serves as a cue word and is input into a preset large language model together with the question for semantic transformation, a second software function description most relevant to the question in the retrieval result is output, and the second software function description carries a test case identifier; and obtaining a test script associated with the test case identifier and executing a software regression test. According to the method and the device, the test case selection efficiency is improved, meanwhile, the accuracy of the test case is ensured, and the test quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software testing, and in particular to a software regression testing method, device, equipment and storage medium. Background Art

[0002] Software regression testing refers to retesting after the software system has modified the code (usually to fix software defects) to confirm that the modification has not introduced new errors or caused errors in other codes. As an integral part of the software life cycle, regression testing accounts for a large proportion of the workload in the entire software testing process. Multiple regression tests are performed at each stage of software development. In agile development, the continuous release of new versions makes regression testing more frequent, and some even require regression testing several times a day. Therefore, it is very necessary to improve the efficiency of regression testing.

[0003] At present, regression testing mainly relies on testers to judge the possible impact of code modification on software system functions based on their understanding of the software system business. Manually selecting relevant test cases for regression testing not only has poor test results, but also affects test efficiency. The quality of regression case selection often depends on the tester's experience and the depth of understanding of the business, and it is easy to miss selections, resulting in insufficient test coverage. In addition, with the iteration of software system functions, the test case library continues to accumulate, and sometimes there are as many as thousands of test cases. It becomes very difficult and inefficient to manually select regression cases from them, and it cannot meet the needs of rapid software iteration. Summary of the invention

[0004] The main purpose of the present invention is to provide a software regression testing method, device, equipment and storage medium, aiming to solve the technical problems of poor testing effect and low testing efficiency in existing software regression testing.

[0005] A first aspect of the present invention provides a software regression testing method, the software regression testing method comprising:

[0006] Obtaining a first software function description corresponding to the test case to be executed;

[0007] Input the first software function description as a question into a preset vector model for retrieval, and output a number of retrieval results related to the question retrieved by the vector model from a preset vector database;

[0008] Input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output a second software function description in the search result that is most relevant to the question, wherein the second software function description carries a test case identifier;

[0009] Obtain the test script associated with the test case identifier and perform software regression testing.

[0010] Optionally, in a first implementation of the first aspect of the present invention, before obtaining the first software function description corresponding to the test case to be executed, the method further includes:

[0011] Obtain software function description document;

[0012] Dividing the software function description document into function points to obtain multiple software function points and function descriptions of each software function point;

[0013] Associating each test case identifier with the functional description of the corresponding software function point;

[0014] Vectorize the function description of each software function point with the test case identifier to obtain the function description text vector corresponding to each software function point;

[0015] The function description text vector corresponding to each software function point is stored in the vector database.

[0016] Optionally, in a second implementation of the first aspect of the present invention, dividing the software function description document into function points to obtain multiple software function points and a function description of each software function point includes:

[0017] Based on the correlation between the text contents in the software function description document, the software function description document is divided into functional points to obtain a plurality of document blocks;

[0018] Marking the software function points of each document block respectively, and obtaining multiple secondary software function points and the function description of each secondary software function point;

[0019] The second-level software function points are subject-clustered to obtain a plurality of first-level software function points, wherein each first-level software function point contains one or more second-level software function points.

[0020] Optionally, in a third implementation of the first aspect of the present invention, the functional description of each secondary software functional point is accompanied by several associated test case identifiers, and different test case identifiers have different priorities.

[0021] Optionally, in a fourth implementation of the first aspect of the present invention, vectorizing the function description of each software function point with a test case identifier to obtain a function description text vector corresponding to each software function point includes:

[0022] Convert the functional description of each software function point with the test case identifier into block data of Json structure, wherein the block data includes the document path, block title, block content and the position of the block content in the document;

[0023] The block data is input into the vector model for vectorization processing and dense vector and sparse vector calculation to obtain the function description text vector corresponding to each software function point.

[0024] Optionally, in a fifth implementation of the first aspect of the present invention, inputting the first software function description as a question into a preset vector model for retrieval, and outputting a number of retrieval results related to the question retrieved by the vector model from a preset vector database includes:

[0025] Input the first software function description as a question into a preset vector model for vectorization processing to obtain an input vector;

[0026] Calculating the similarity score between the input vector and each function description text vector in the vector database respectively through the vector model;

[0027] Based on the similarity score, each function description text vector in the vector database is sorted by the vector model, and several function description text vectors with the highest similarity scores are selected as retrieval results related to the question.

[0028] Optionally, in a sixth implementation of the first aspect of the present invention, the inputting the search result as a prompt word together with the question into a preset large language model for semantic conversion, and outputting the second software function description in the search result that is most relevant to the question includes:

[0029] Combining the search results as prompt words with the question to construct an input sequence;

[0030] Inputting the input sequence into a preset large language model, and performing preliminary parsing of the input sequence by the large language model to identify words, phrases, and sentence structures in the input sequence;

[0031] Based on the vocabulary, phrases and sentence structures in the input sequence, mining the deep semantic information of the input sequence through the large language model;

[0032] Based on the deep semantic information of the input sequence, the second software function description most relevant to the question in the retrieval results is outputted through the large language model.

[0033] A second aspect of the present invention further provides a software regression testing device, the software regression testing device comprising:

[0034] An acquisition module, used to acquire a first software function description corresponding to a test case to be executed;

[0035] A vector retrieval module, configured to input the first software function description as a question into a preset vector model for retrieval, and output a number of retrieval results related to the question retrieved by the vector model from a preset vector database;

[0036] A semantic understanding module, used to input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output a second software function description in the search result that is most relevant to the question, wherein the second software function description carries a test case identifier;

[0037] The execution module is used to obtain the test script associated with the test case identifier and execute software regression testing.

[0038] Optionally, in a first implementation of the second aspect of the present invention, the software regression test further comprises: a partitioning module, an association module, a vectorization module and a storage module;

[0039] The acquisition module is also used to: acquire a software function description document;

[0040] The division module is used to: divide the software function description document into function points to obtain multiple software function points and function descriptions of each software function point;

[0041] The association module is used to: associate each test case identifier with a functional description of a corresponding software function point;

[0042] The vectorization module is used to: perform vectorization processing on the function description of each software function point with the test case identifier to obtain the function description text vector corresponding to each software function point;

[0043] The storage module is used to store the function description text vector corresponding to each software function point into the vector database.

[0044] Optionally, in a second implementation of the second aspect of the present invention, the division module is specifically used to:

[0045] Based on the correlation between the text contents in the software function description document, the software function description document is divided into functional points to obtain a plurality of document blocks;

[0046] Marking the software function points of each document block respectively, and obtaining multiple secondary software function points and the function description of each secondary software function point;

[0047] The second-level software function points are subject-clustered to obtain a plurality of first-level software function points, wherein each first-level software function point contains one or more second-level software function points.

[0048] Optionally, in a third implementation of the second aspect of the present invention, the functional description of each secondary software functional point is accompanied by several associated test case identifiers, and different test case identifiers have different priorities.

[0049] Optionally, in a fourth implementation manner of the second aspect of the present invention, the vectorization module is specifically used to:

[0050] Convert the functional description of each software function point with the test case identifier into block data of Json structure, wherein the block data includes the document path, block title, block content and the position of the block content in the document;

[0051] The block data is input into the vector model for vectorization processing and dense vector and sparse vector calculation to obtain the function description text vector corresponding to each software function point.

[0052] Optionally, in a fifth implementation manner of the second aspect of the present invention, the vector retrieval module is specifically used to:

[0053] Input the first software function description as a question into a preset vector model for vectorization processing to obtain an input vector;

[0054] Calculating the similarity score between the input vector and each function description text vector in the vector database respectively through the vector model;

[0055] Based on the similarity score, each function description text vector in the vector database is sorted by the vector model, and several function description text vectors with the highest similarity scores are selected as retrieval results related to the question.

[0056] Optionally, in a sixth implementation of the second aspect of the present invention, the semantic understanding module is specifically used to:

[0057] Combining the search results as prompt words with the question to construct an input sequence;

[0058] Inputting the input sequence into a preset large language model, and performing preliminary parsing of the input sequence by the large language model to identify words, phrases, and sentence structures in the input sequence;

[0059] Based on the vocabulary, phrases and sentence structures in the input sequence, mining the deep semantic information of the input sequence through the large language model;

[0060] Based on the deep semantic information of the input sequence, the second software function description most relevant to the question in the retrieval results is outputted through the large language model.

[0061] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned software regression testing method.

[0062] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned software regression testing method.

[0063] The software regression testing method provided by the present invention introduces a vector model and a large language model. When the tester performs software regression testing, he only needs to input the software function description corresponding to the test case to be executed as a question into the vector model. The vector model will automatically find the search results related to the question (including several approximate software function descriptions). In order to improve the accuracy of the search results, the search results are further input into the large language model together with the question as prompt words for semantic conversion. Through the semantic understanding and analysis of the large language model, the software function description most relevant to the question is screened out from multiple search results. The software function description is accompanied by a test case identifier, and finally the associated test script is obtained and executed according to the test case identifier, thereby fully automating the complete software regression test process. The present invention not only improves the efficiency of test case selection, but also ensures the accuracy of test cases and improves the test quality. At the same time, by automatically selecting test cases and automatically executing test cases, the execution efficiency of software regression testing is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of an embodiment of a software regression testing method in an embodiment of the present invention;

[0065] Figure 2 A schematic diagram of an embodiment of a software regression testing device in an embodiment of the present invention;

[0066] Figure 3 FIG. 1 is a schematic diagram of an embodiment of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0068] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , an embodiment of the software regression testing method in the embodiment of the present invention includes:

[0069] 101. Obtain a first software function description corresponding to a test case to be executed;

[0070] A test case is a combination of input, execution conditions, and expected results, used to verify specific functions or properties of software. Software regression testing is to retest the old code through test cases after modifying it to confirm that the modification has not introduced new errors or caused errors in other codes. Software functional description refers to the description of various functions of the software, such as the specific functions, specific roles, specific implementation processes, specific functional configurations, etc. implemented by the software.

[0071] In the prior art, testers usually rely on their understanding of the software system business to determine the possible impact of code modifications on the software system functions, and then manually select relevant test cases to perform software regression testing. The quality of regression test case selection often depends on the tester's experience and the depth of understanding of the business, so it is easy to miss and lead to insufficient test coverage. In addition, with the iteration of software system functions, the test case library continues to accumulate, and sometimes there are as many as thousands of test cases. It becomes very difficult and inefficient to manually select regression test cases from them, and it cannot meet the needs of rapid software iteration.

[0072] In this embodiment, the test cases are not directly selected manually. Instead, before performing software regression testing, the software function description corresponding to the test case to be tested is first obtained, and then the test case is obtained through the software function description.

[0073] In an optional embodiment, in order to obtain test cases through software function description, the following is further included before step 101:

[0074] S10, obtaining software function description document;

[0075] In this optional embodiment, the software function description document includes the function description of the entire software, which is a complete description of the software function. The complete software function description is used to facilitate the division of different software function points.

[0076] S20, dividing the software function description document into function points to obtain multiple software function points and function descriptions of each software function point;

[0077] In this optional embodiment, by splitting the software function description document into smaller and more specific function points, each function point has a corresponding description, which is convenient for associating with test cases, thereby ensuring the accuracy of the match between the selected test cases and the functions to be tested, thereby improving the quality of software regression testing.

[0078] In an optional embodiment, the function points are divided in the following manner, specifically including:

[0079] S201, dividing the software function description document into functional points based on the correlation between the text contents in the software function description document to obtain a plurality of document blocks;

[0080] In this optional embodiment, the software function description document is a functional description of the entire software, which is a text or document collection containing all functional descriptions of the software. Therefore, the software function description document needs to be divided into blocks in order to obtain multiple different functional points and functional descriptions of each functional point.

[0081] This optional embodiment analyzes the text content in the document through text mining, natural language processing (NLP) and other technologies to find the correlation between each paragraph of text, for example, using keyword matching, topic model (such as LDA) and other methods to identify the correlation between texts. Based on the correlation of the text content, the original software function description document is divided into multiple document blocks, each block represents a relatively independent functional area or function point.

[0082] S202, marking software function points of each document block respectively, to obtain multiple secondary software function points and function descriptions of each secondary software function point;

[0083] In this optional embodiment, after the software function description document is divided into multiple document blocks through the above step S201, each document block needs to be further analyzed and function-point marked. For example, a word frequency analysis tool (such as TF-IDF) can be used to screen the topic of each document block, and then the advanced summary generation algorithm can be used to refine the final topic. The TF-IDF model evaluates the importance of vocabulary in each document block by calculating word frequency (TF) and inverse document frequency (IDF).

[0084] After obtaining the software function point name corresponding to each document block, a secondary software function point tag is assigned to each independent function point name, so as to obtain the secondary software function point corresponding to each document block, and the content of each document block will serve as the functional description of each secondary software function point.

[0085] S203 , performing topic clustering on each secondary software function point to obtain a plurality of primary software function points, wherein each primary software function point contains one or more secondary software function points.

[0086] In this optional embodiment, considering that different software function points may be tests for different aspects of the same large function point, it is necessary to further use thematic clustering techniques (such as K-means, hierarchical clustering, etc.) to group these secondary function points. The purpose of clustering is to classify secondary function points with similar themes or functions into one category, thereby forming a higher-level function point (i.e., a first-level software function point). Specifically, thematic clustering can be performed based on keywords, themes, or semantic similarities in the function description, and then multiple first-level software function points, each of which contains one or more related second-level function points. The first-level function point represents a higher-level functional module or system component in the software.

[0087] In this optional embodiment, by dividing and marking the software function description document into function points, the complex software function description document is converted into a clearly structured and easy-to-understand function point structure, thereby realizing the automation of software regression testing and improving the testing quality and efficiency.

[0088] S30, associating each test case identifier with a functional description of a corresponding software function point;

[0089] In this optional embodiment, after the software function description document is divided into multiple software function points and the function description of each software function point, an identifier needs to be assigned to each test case and associated with the corresponding software function point, so that the test case corresponding to each function point can be quickly located when the test is executed later. In this embodiment, a unique identifier (such as ID) can be created for each test case, and then each function is bound to the test case representation, that is, the test case ID of each function is attached to the corresponding function description text, so as to achieve the association between the two.

[0090] In an optional embodiment, it is preferred to attach several associated test case identifiers in the function description of each secondary software function point, and use different priorities for multiple test cases. For example, function point A has 3 test cases, and the 3 test cases have 3 different priorities. When performing software testing, it is preferred to use the test case with the highest priority for software regression testing.

[0091] S40, performing vectorization processing on the function description of each software function point with the test case identifier to obtain a function description text vector corresponding to each software function point;

[0092] In this optional embodiment, the function description is converted from text form to vector form, so as to perform efficient search and matching in subsequent steps. Specifically, natural language processing (NLP) technology, such as word embedding or sentence embedding, can be used to convert the function description of each software function point into a vector to obtain a function description text vector corresponding to each software function point. The text vector is usually high-dimensional and can capture the semantic information in the software function description.

[0093] In an optional embodiment, the above step S40 further includes:

[0094] S401, converting the function description of each software function point with the test case identifier into block data of Json structure, wherein the block data includes a document path, a block title, a block content and a position of the block content in the document;

[0095] The functional description of each functional point is converted into a Json structure to obtain block data with a Json structure. For each block, a block data with a Json structure is generated, which includes the following data:

[0096] Document path: indicates the location of the software function description document;

[0097] Block title: refers to a secondary software function point corresponding to a block;

[0098] Block content: refers to the software function description content corresponding to a block;

[0099] The position of the block content in the document: indicates the specific position of the block content in the original document. For example, if the software function description document is divided into 20 blocks, the position of the block content in the document can refer to the specific block in the document, such as the 5th block.

[0100] S402: Input the block data into the vector model for vectorization processing and dense vector and sparse vector calculation to obtain a function description text vector corresponding to each software function point.

[0101] In this optional embodiment, a suitable vector model can be pre-selected to convert text data (block data) into vector representation. Common vector models include TF-IDF, Word2Vec, BERT, etc. The vector model can convert text data into numerical vectors, which is convenient for subsequent calculation and analysis.

[0102] The block data of the Json structure generated in step S401 is input into the vector model and vectorized, so that the text content of each block is converted into a vector representation. In the process of vectorization, different types of vector representations may be obtained according to the vector model used, such as dense vectors (usually with higher dimensions and dense values) and sparse vectors (usually with lower dimensions and sparse values). These vectors represent the characteristics of the text data and can be used for subsequent analysis and comparison. Through vectorization, the function description text vector corresponding to each software function point is finally obtained. The function description text vector can be used for tasks such as function point analysis, similarity calculation, and classification.

[0103] S50, storing the function description text vector corresponding to each software function point into a vector database.

[0104] In this optional embodiment, after obtaining the function description text vector corresponding to each software function point, the processed function description vector is further stored in a database for subsequent efficient retrieval and matching. For example, a database system that supports vector storage and retrieval (such as Elasticsearch vector database) is selected, and then the function description vector is imported into the vector database.

[0105] This optional embodiment introduces a vector model for slicing and vectorizing the software function description text, and stores the vectorization results in the vector database, and also stores the test case ID associated with the software function (one software function corresponds to one or more test case IDs). This optional embodiment binds each function description to the test case ID in advance, that is, the test case ID of each function is attached to the corresponding function description text, and when vectorizing, the test case ID is stored in a separate field in the vector database table, so as to facilitate subsequent processing.

[0106] In an optional embodiment, the specific fields stored in the vector database after vectorization processing include: the id field is an independent id; the context field is the original text; the file_path field is the document name; the instance_id field is the "nth block" of the document; the test_case_ids field is the test case ID list corresponding to the functional block; the doc_sparse_vector field is the sparse vector of the original text; and the doc_dense_vector field is the dense vector of the original text.

[0107] 102. Input the first software function description as a question into a preset vector model for retrieval, and output a number of retrieval results related to the question retrieved by the vector model from a preset vector database;

[0108] In this embodiment, the software function description obtained in step 101 is used as a question and input into a pre-trained vector model, which can convert the question used in the query into a vector form for efficient retrieval in a vector database. The vector database pre-stores a large number of software function descriptions and they are all converted into vector forms. The vector model compares the query question vector with the vectors in the vector database to find several retrieval results that are most similar to the query question.

[0109] In an optional embodiment, the above step 102 further includes:

[0110] 1021. Input the first software function description as a question into a preset vector model for vectorization processing to obtain an input vector;

[0111] In this optional embodiment, when a functional bug in the software under test is fixed, the tester needs to perform software regression testing on the function, so it is necessary to obtain the test case corresponding to the function and complete the regression test, specifically using the repair function as the prompt word input vector model for retrieval.

[0112] The vector model in this optional embodiment uses a trained model, which can convert the input text-based question into a vector form to obtain an input vector corresponding to the question.

[0113] 1022. Calculate, by using the vector model, a similarity score between the input vector and each function description text vector in the vector database;

[0114] The vector data used in this optional embodiment pre-stores a large number of text vectors describing software functions. Each function description text has been converted into a vector form, and then the similarity score between the input vector used for query and all the stored vectors in the vector database can be calculated, for example, by using cosine similarity, Euclidean distance, etc. to calculate the similarity score, and then return the query records that are greater than the preset score threshold.

[0115] 1023. Based on the similarity score, sort each function description text vector in the vector database using the vector model, and select several function description text vectors with the highest similarity scores as retrieval results related to the question.

[0116] In this optional embodiment, the function description text vectors in the vector database are sorted according to the similarity scores, and then several function description text vectors with the highest similarity scores are selected from the sorted results. These selected function description text vectors are the search results most relevant to the question (i.e., the first software function description).

[0117] 103. Input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output a second software function description that is most relevant to the question in the search result, wherein the second software function description carries a test case identifier;

[0118] In this optional embodiment, in order to further improve the accuracy of the retrieval results, after obtaining the retrieval results output by the vector model, the retrieval results related to the question (i.e., the function description text) and the original question output by the vector model are further provided to the large language model for semantic conversion.

[0119] The large language model understands the input text, including extracting keywords, understanding contextual relationships, and identifying semantic patterns. The large language model converts the input retrieval results and questions into an internal representation that captures the semantic connections and differences between the texts. The large language model conversion process may use deep learning algorithms (such as Transformer, BERT, etc.) to encode and decode the input to generate more accurate and relevant semantic representations. Based on the results of the semantic conversion, the large language model selects the second software function description that is most relevant to the question from the retrieval results. The selection process involves calculating indicators such as semantic similarity and evaluating relevance. The second software function description output by the large language model will retain the test case identifier carried in the software function description.

[0120] In an optional embodiment, the above step 103 further includes:

[0121] 1031. Combining the search result as a prompt word with the question to construct an input sequence;

[0122] The retrieval results in this optional embodiment are software function descriptions pre-stored in a vector database. The retrieval results are used as prompt words or context information and combined with the user's original question to form a more complete and specific input sequence, so that the subsequent large language model can more accurately understand the user's intentions and needs.

[0123] 1032. Input the input sequence into a preset large language model, and perform a preliminary analysis on the input sequence by using the large language model to identify words, phrases, and sentence structures in the input sequence;

[0124] The constructed input sequence is obtained through step 1031 and input into the preset large language model for processing. The large language model is a deep learning model that can process and understand natural language text. It performs preliminary analysis on the input sequence, identifies the vocabulary, phrases and sentence structures therein, and thus breaks the input sequence into smaller language units for subsequent deeper analysis and understanding.

[0125] 1033. Based on the vocabulary, phrases, and sentence structures in the input sequence, mining deep semantic information of the input sequence through the large language model;

[0126] After the initial analysis is completed, the large language model will further mine the deep semantic information of the input sequence, including understanding the association between words, the meaning of phrases, and the logical relationship between sentences. By mining the deep semantic information of the input sequence, it can further understand the user's intentions and needs, and obtain the association between the search results and the questions.

[0127] 1034. Based on the deep semantic information of the input sequence, output the second software function description most relevant to the question in the search results through the large language model.

[0128] After mining the deep semantic information of the input sequence, the large language model will filter out the second software function description that is most relevant to the question from the retrieval results based on the deep semantic information, and then further obtain a more accurate software function description from the multiple software function descriptions in the retrieval results to meet the needs of users.

[0129] In this optional embodiment, when a functional bug of the software needs to be regression tested, the tester only needs to input the functional description into the vector model. The vector model will search the vector database by similarity based on the input functional description, retrieve the functional points associated with it (specifically the functional descriptions corresponding to the functional points), and then hand them over to the large language model for further confirmation and screening of the returned functional points, thereby screening out the functional points that the large language model believes are truly associated. In addition, since the software functional description stored in the vector data carries a test case identifier, when the large language model returns the functional description corresponding to the functional point, it will also return the test case ID bound to the functional description, and the corresponding regression test case can be obtained, thereby improving the efficiency and effectiveness of selecting software regression test cases.

[0130] 104. Obtain the test script associated with the test case identifier and execute software regression testing.

[0131] In this embodiment, each test case ID is further associated with the test script, so that while automatically selecting the test case, the test script associated with the current regression test case can also be automatically selected and the software regression test can be executed, thereby achieving full automation of the entire regression testing process and improving the efficiency of software regression testing.

[0132] This embodiment pre-binds each function of the software with the corresponding test case ID, that is, the test case ID of each function is attached to the corresponding function description text. During vectorization, the test case ID is stored in a separate field in the vector database table. When searching the vector model, only the similarity of the function description is searched, and the function results that meet the similarity conditions are obtained and then filtered by the large language model, and finally the most relevant function descriptions and the test case IDs bound to these functions are returned.

[0133] In this embodiment, the test case corresponding to each software function is determined. By searching the vector model and filtering the large language model, the test case bound to the software function can be obtained. Compared with the prior art of directly searching the test case by text description, this method is more suitable for regression testing scenarios. In this embodiment, the software function description is vectorized, not the test case description.

[0134] In addition, this embodiment only searches for function descriptions, not test cases. Considering that test cases involve detailed test steps and expected results, the granularity is too fine and it is easy to be separated from the context, which is not conducive to finding use cases related to the function. If the entire use case is searched, the content is too general. Different function use cases may involve some of the same operation steps and expected results, which will also be retrieved. For example, it is obviously a test case for function B, but because there is a description in the test steps of the B function use case that is similar to the test steps of the A function use case, this use case is also returned, but this use case is useless for testing function A. That is, this embodiment avoids possible errors in the test case selection process. This embodiment mainly uses the vector model to complete the relevant function retrieval, and the large language model is only used to filter the retrieval results of the vector model again to improve the accuracy of the retrieval results. It is also possible even if the large language model is not used.

[0135] The above describes the software regression testing method in the embodiment of the present invention. The following describes the software regression testing device in the embodiment of the present invention. Figure 2 , an embodiment of the software regression testing device in the embodiment of the present invention includes:

[0136] An acquisition module 201 is used to acquire a first software function description corresponding to a test case to be executed;

[0137] A vector search module 202, configured to input the first software function description as a question into a preset vector model for search, and output a number of search results related to the question retrieved by the vector model from a preset vector database;

[0138] Semantic understanding module 203, used to input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output the second software function description most relevant to the question in the search result, wherein the second software function description carries a test case identifier;

[0139] The execution module 204 is used to obtain the test script associated with the test case identifier and execute software regression testing.

[0140] Optionally, in one embodiment, the software regression test further includes: a partitioning module 205, an association module 206, a vectorization module 207 and a storage module 208;

[0141] The acquisition module 201 is also used to: acquire a software function description document;

[0142] The division module 205 is used to: divide the software function description document into function points to obtain multiple software function points and function descriptions of each software function point;

[0143] The association module 206 is used to: associate each test case identifier with a functional description of a corresponding software function point;

[0144] The vectorization module 207 is used to: perform vectorization processing on the function description of each software function point with the test case identifier to obtain the function description text vector corresponding to each software function point;

[0145] The storage module 208 is used to store the function description text vector corresponding to each software function point into the vector database.

[0146] Optionally, in one embodiment, the division module 205 is specifically used for:

[0147] Based on the correlation between the text contents in the software function description document, the software function description document is divided into functional points to obtain a plurality of document blocks;

[0148] Marking the software function points of each document block respectively, and obtaining multiple secondary software function points and the function description of each secondary software function point;

[0149] The second-level software function points are subject-clustered to obtain a plurality of first-level software function points, wherein each first-level software function point contains one or more second-level software function points.

[0150] Optionally, in one embodiment, the functional description of each secondary software functional point is accompanied by several associated test case identifiers, and different test case identifiers have different priorities.

[0151] Optionally, in one embodiment, the vectorization module 207 is specifically used for:

[0152] Convert the functional description of each software function point with the test case identifier into block data of Json structure, wherein the block data includes the document path, block title, block content and the position of the block content in the document;

[0153] The block data is input into the vector model for vectorization processing and dense vector and sparse vector calculation to obtain the function description text vector corresponding to each software function point.

[0154] Optionally, in one embodiment, the vector retrieval module 202 is specifically used for:

[0155] Input the first software function description as a question into a preset vector model for vectorization processing to obtain an input vector;

[0156] Calculating the similarity score between the input vector and each function description text vector in the vector database respectively through the vector model;

[0157] Based on the similarity score, each function description text vector in the vector database is sorted by the vector model, and several function description text vectors with the highest similarity scores are selected as retrieval results related to the question.

[0158] Optionally, in one embodiment, the semantic understanding module 203 is specifically used for:

[0159] Combining the search results as prompt words with the question to construct an input sequence;

[0160] Inputting the input sequence into a preset large language model, and performing preliminary parsing of the input sequence by the large language model to identify words, phrases, and sentence structures in the input sequence;

[0161] Based on the vocabulary, phrases and sentence structures in the input sequence, mining the deep semantic information of the input sequence through the large language model;

[0162] Based on the deep semantic information of the input sequence, the second software function description most relevant to the question in the search results is outputted through the large language model.

[0163] The entire process of selecting software regression test cases in this embodiment is mainly completed by a vector model and a large language model. The vector model completes the vectorization of functional content and the retrieval of related functions. The large language model further screens the results retrieved by the vector model (recall results) based on semantic understanding. Finally, the test case IDs of associated functions are obtained, and the bound test scripts are executed according to the test case IDs, thereby realizing software regression testing. The combination of vector model retrieval and large language model screening ensures the effectiveness of the recall results, and the test cases are strongly related to the functions. The test case IDs of the recalled functions are directly returned, ensuring the integrity of the test cases. Both the accuracy and efficiency are far higher than manual selection.

[0164] Since the embodiments of the device part correspond to the embodiments of the above method, the description of the software regression testing device provided by the present invention can be referred to the above method embodiments, and the present invention will not be described in detail here. It has the same beneficial effects as the above software regression testing method.

[0165] Above Figure 2 The software regression testing device in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. Next, the computer device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0166] Figure 3 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present invention. The computer device 500 may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 for storing application programs 533 or data 532 (for example, one or more mass storage devices). Among them, the memory 520 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 500. Further, the processor 510 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the computer device 500.

[0167] The computer device 500 may further include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand, Figure 3The computer device structure shown does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0168] The present invention also provides a computer device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the software regression testing method in the above-mentioned various embodiments.

[0169] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the software regression testing method.

[0170] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0172] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A software regression testing method, characterized in that: The software regression testing method comprises: Obtaining a first software function description corresponding to the test case to be executed; Input the first software function description as a question into a preset vector model for retrieval, and output a number of retrieval results related to the question retrieved by the vector model from a preset vector database; Input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output a second software function description in the search result that is most relevant to the question, wherein the second software function description carries a test case identifier; Obtain the test script associated with the test case identifier and perform software regression testing.

2. The software regression testing method according to claim 1, characterized in that: Before obtaining the first software function description corresponding to the test case to be executed, the method further includes: Obtain software function description document; Dividing the software function description document into function points to obtain multiple software function points and function descriptions of each software function point; Associating each test case identifier with the functional description of the corresponding software function point; Vectorize the function description of each software function point with the test case identifier to obtain the function description text vector corresponding to each software function point; The function description text vector corresponding to each software function point is stored in the vector database.

3. The software regression testing method according to claim 2, characterized in that: The function point division of the software function description document to obtain a plurality of software function points and a function description of each software function point includes: Based on the correlation between the text contents in the software function description document, the software function description document is divided into functional points to obtain a plurality of document blocks; Marking the software function points of each document block respectively, and obtaining multiple secondary software function points and the function description of each secondary software function point; The second-level software function points are subject-clustered to obtain a plurality of first-level software function points, wherein each first-level software function point contains one or more second-level software function points.

4. The software regression testing method according to claim 3, characterized in that: The functional description of each secondary software functional point is accompanied by several associated test case identifiers, and different test case identifiers have different priorities.

5. The software regression testing method according to claim 3, characterized in that: The vectorization process of the function description of each software function point with the test case identifier to obtain the function description text vector corresponding to each software function point includes: Convert the functional description of each software function point with the test case identifier into block data of Json structure, wherein the block data includes the document path, block title, block content and the position of the block content in the document; The block data is input into the vector model for vectorization processing and dense vector and sparse vector calculation to obtain the function description text vector corresponding to each software function point.

6. The software regression testing method according to claim 2, characterized in that: The step of inputting the first software function description as a question into a preset vector model for retrieval, and outputting a number of retrieval results related to the question retrieved by the vector model from a preset vector database comprises: Input the first software function description as a question into a preset vector model for vectorization processing to obtain an input vector; Calculating the similarity score between the input vector and each function description text vector in the vector database respectively through the vector model; Based on the similarity score, each function description text vector in the vector database is sorted by the vector model, and several function description text vectors with the highest similarity scores are selected as retrieval results related to the question.

7. The software regression testing method according to any one of claims 1 to 6, characterized in that: The step of inputting the search result as a prompt word together with the question into a preset large language model for semantic conversion, and outputting the second software function description in the search result that is most relevant to the question comprises: Combining the search results as prompt words with the question to construct an input sequence; Inputting the input sequence into a preset large language model, and performing preliminary parsing of the input sequence by the large language model to identify words, phrases, and sentence structures in the input sequence; Based on the vocabulary, phrases and sentence structures in the input sequence, mining the deep semantic information of the input sequence through the large language model; Based on the deep semantic information of the input sequence, the second software function description most relevant to the question in the retrieval results is outputted through the large language model.

8. A software regression testing device, characterized in that: The software regression testing device comprises: An acquisition module, used to acquire a first software function description corresponding to a test case to be executed; A vector retrieval module, configured to input the first software function description as a question into a preset vector model for retrieval, and output a number of retrieval results related to the question retrieved by the vector model from a preset vector database; A semantic understanding module, used to input the search result as a prompt word together with the question into a preset large language model for semantic conversion, and output a second software function description in the search result that is most relevant to the question, wherein the second software function description carries a test case identifier; The execution module is used to obtain the test script associated with the test case identifier and execute software regression testing.

9. A computer device, characterized in that: The computer device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the computer device to execute the software regression testing method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the software regression testing method according to any one of claims 1 to 7 is implemented.