Software testing method and device, computer equipment, readable storage medium and program product

By employing a multi-stage training strategy and a multi-task model, the mapping test between source code files and requirements design documents, and between requirements design documents and test documents, is enhanced in software testing. This solves the problems of incomplete and inconsistent mapping in traditional software testing, and improves the quality of software testing as well as the efficiency and accuracy of test document generation.

CN121029629AActive Publication Date: 2025-11-28CHINA ELECTRONICS RELIABILITY AND ENVIRONMENTAL TESTING INSTITUTE ((THE FIFTH INSTITUTE OF ELECTRONICS MINISTRY OF INDUSTRY AND INFORMATION TECHNOLOGY) (CHINA SAIBAO LABORATORY)

Patent Information

Application Number
CN202511574929.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-11-28
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

In traditional software testing, the incomplete mapping between source code files and requirements design documents leads to missing or incorrect software functions; the inconsistency between requirements design documents and test documents reduces software reliability and controllability; and traditional methods for generating test documents are inefficient, inaccurate, and lack real-time performance.

Method used

A multi-stage training strategy is adopted to obtain a deep learning model. Through a multi-task model and a relation extraction model, the mapping test between source code files and requirement design documents, and between requirement design documents and test documents is enhanced. Test documents are optimized to improve semantic association, including pre-training, intermediate training and model tuning. Semantic association between entities is optimized using a multi-neural network collaboration model and a relation extraction model.

Benefits of technology

It improved the quality and accuracy of software testing, ensured the integrity and reliability of software functions, increased the efficiency and accuracy of test documentation generation, and solved the problems of incomplete and inconsistent mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029629A_ABST
    Figure CN121029629A_ABST
Patent Text Reader

Abstract

The invention relates to a software testing method and device, computer equipment, a readable storage medium and a program product. The method comprises the steps of obtaining a deep learning model based on a multi-stage training strategy, wherein the deep learning model is used for performing mapping test on a source code file and a demand design document; obtaining a multi-task model based on the deep learning model, wherein the multi-task model is used for performing a mapping test on the demand design document and the test document; based on a relation extraction model, optimizing the test document; the optimization at least comprises enhancement of semantic association between entities. By adopting the method, the software testing quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of software evaluation, in particular to a software testing method and device, computer equipment, readable storage medium and program product. BACKGROUND

[0002] In recent years, with the increase of the complexity of basic software functions, the traceability of software becomes more and more important, and the relationship between software source code, design documents and test documents plays a key role in ensuring software quality and functional integrity.

[0003] However, the traditional basic software evaluation process has the problem of low software test quality. SUMMARY

[0004] Therefore, it is necessary to provide a software testing method, device, computer equipment, readable storage medium and program product capable of improving software test quality.

[0005] In a first aspect, the present application provides a software testing method, comprising:

[0006] obtaining a deep learning model based on a multi-stage training strategy, the deep learning model being used for mapping test of source code files and requirement design documents; the multi-stage training strategy at least including pre-training, intermediate training and model adjustment;

[0007] obtaining a multi-task model based on the deep learning model, the multi-task model being used for mapping test of the requirement design documents and test documents;

[0008] optimizing the test documents based on a relationship extraction model; the optimization at least including enhancing semantic association between entities; the relationship extraction model at least including an encoding layer, an enhanced entity layer and a classification layer.

[0009] In one embodiment, the deep learning model is obtained based on a multi-stage training strategy, comprising:

[0010] pre-training the model based on unsupervised text data to obtain the deep learning model, the deep learning model including a general language representation;

[0011] adjusting the deep learning model to migrate the knowledge obtained by the deep learning model from the code search task to the mapping link task between the requirement design documents and the source code files.

[0012] In one of the embodiments, the adjusting of the deep learning model, the migration of the knowledge obtained by the deep learning model from the code search task to the mapping link task between the requirement design document and the source code file, comprises:

[0013] Intermediate training of the deep learning model for the code search problem for the general principle of the deep learning model to identify the correlation between natural language and programming language; the code search problem includes obtaining the corresponding code segment based on the natural language description;

[0014] Based on the sample data of the associated source code file and the requirement design document, the knowledge obtained by the deep learning model from the intermediate training is migrated to the adjustment stage of the deep learning model, and the deep learning model is adjusted.

[0015] In one of the embodiments, the mapping test of the source code file and the requirement design document comprises:

[0016] The source code file and the requirement design document are decomposed and input to the deep learning model in turn;

[0017] The deep learning model determines whether there is an effective mapping link between the source code file and the requirement design document based on the feature conversion technology, and outputs the evaluation results.

[0018] In one of the embodiments, the deep learning model obtains a multi-task model, comprising:

[0019] The text sequence data of the cleaned and pretreated requirement design document and test document is input into the deep learning model to obtain the hidden representation of the text;

[0020] The hidden representation is used as a shared representation, and an independent output layer is designed for each task to obtain the multi-task model; each output layer includes a fully connected layer and an activation function, and each output layer is connected with the shared representation;

[0021] The multi-task model is adjusted based on the associated data set of the requirement design document and the test document.

[0022] In one of the embodiments, the optimization of the test document based on the relationship extraction model comprises:

[0023] The requirement design document is vectorized based on the encoding layer of the relationship extraction model;

[0024] constructing a context representation of the entity based on an enhanced entity layer of the relation extraction model, and creating a final representation of the entity node;

[0025] performing entity type prediction based on a classification layer of the relation extraction model.

[0026] In a second aspect, the present application further provides a software testing device, comprising:

[0027] a multi-stage training module configured to obtain a deep learning model based on a multi-stage training strategy, the deep learning model being configured to perform mapping testing on a source code file and a requirement design document;

[0028] a multi-task model obtaining module configured to obtain a multi-task model based on the deep learning model, the multi-task model being configured to perform mapping testing on the requirement design document and a test document;

[0029] a test document optimization module configured to optimize the test document based on a relation extraction model, the optimization at least including enhancing semantic association between entities.

[0030] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, the memory storing a computer program, and the processor implementing steps of the method of the first aspect when executing the computer program.

[0031] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements steps of the method of the first aspect when executed by a processor.

[0032] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program implements steps of the method of the first aspect when executed by a processor.

[0033] The above software testing method, device, computer device, readable storage medium and program product obtain a deep learning model based on a multi-stage training strategy, wherein the deep learning model is configured to perform mapping testing on a source code file and a requirement design document; obtain a multi-task model based on the deep learning model, wherein the multi-task model is configured to perform mapping testing on the requirement design document and a test document; and optimize the test document based on a relation extraction model, wherein the optimization at least includes enhancing semantic association between entities, thereby improving software testing quality. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained from these drawings without creative effort.

[0035] Figure 1 An application environment diagram of the software testing method in one embodiment;

[0036] Figure 2 A flowchart of the software testing method in one embodiment;

[0037] Figure 3 A detailed flowchart of the software testing method in one embodiment;

[0038] Figure 4 A flowchart of the deep learning model based on the multi-stage training strategy in one embodiment;

[0039] Figure 5 A diagram of the multi-neural network cooperation model in one embodiment;

[0040] Figure 6 A flowchart of the test document optimization based on the relation extraction model in one embodiment;

[0041] Figure 7 A structural block diagram of the software testing device in one embodiment;

[0042] Figure 8 An internal structure diagram of the computer device in one embodiment. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0044] The terms "include" and "have" and any variations thereof used in the present application are intended to cover the non-exclusive inclusion. The term "multiple" used in the present application refers to two and more than two. The term "and / or" used in the present application refers to one of the schemes, or any combination of multiple schemes.

[0045] Regarding research on software testing techniques based on large models, in recent years, with the increasing complexity of basic software functions, software traceability has become increasingly important. Clearly defining the relationship between software source code, design documents, and test documents plays a crucial role in ensuring software quality and functional integrity. However, traditional software testing suffers from the following technical problems: incomplete mapping between source code files and requirements design documents, leading to missing or incorrect software functions; inconsistency between requirements design documents and test documents, resulting in reduced software reliability and controllability; and the low efficiency, low accuracy, and poor real-time performance of traditional methods for generating test documents.

[0046] Based on the aforementioned traditional technologies, the software testing method proposed in this application addresses the consistency requirements and test case generation requirements among basic software source code, requirement design documents, and test documents. By combining large-model technology, it forms a software testing technology based on a large model, as shown in the technology roadmap. Figure 3 As shown, a test-oriented intelligent software document review technology is proposed, which breaks through the mapping enhancement technology between source code files and requirements design documents based on multi-stage training strategy, the mapping enhancement technology between requirements design documents and test documents based on multi-neural network collaboration model, and the test document optimization technology based on relation extraction model. It effectively enhances the semantic association between entities and deeply mines the inherent requirements of test documents.

[0047] Specifically, embodiments of this application provide a software testing method that, based on the correlation between basic software source code, requirements design documents, and test documents, and combined with large-scale modeling technology, forms an intelligent software documentation review technology oriented towards testing, such as... Figure 3 As shown. First, this application provides a mapping enhancement technique for source code files and requirement design documents based on a multi-stage training strategy. It introduces a pre-trained large model and reinforcement learning mechanism to improve the accuracy and completeness of model mapping, reduce mapping bias, and ensure the accuracy and completeness of software functions. Second, this application provides a mapping enhancement technique for requirement design documents and test documents based on a multi-neural network collaborative model. It employs a multi-task learning model to adapt to multi-level features, alleviating the problem of insufficient domain annotation corpus when extracting entity relationships. It can effectively extract relationships between different entities even with limited data. Finally, this application provides a test document optimization technique based on a relationship extraction model. By establishing a document graph structure and utilizing self-attention mechanisms and neighbor information aggregation technology, it establishes relationships between entities, effectively enhancing the semantic association between entities and deeply mining the inherent requirements of the test document.

[0048] It should be noted that the beneficial effects or technical problems solved by the embodiments of this application are not limited to this one, but may also be other implicit or related problems. For details, please refer to the description of the embodiments below.

[0049] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0050] The software testing method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0051] In one exemplary embodiment, such as Figure 2 As shown, a software testing method is provided, which is applied to... Figure 1 Taking terminal 102 or server 104 as an example, the explanation includes the following steps 202 to 206. Wherein:

[0052] Step 202: Obtain a deep learning model based on a multi-stage training strategy. The deep learning model is used to perform mapping tests on source code files and requirement design documents. The multi-stage training strategy includes at least pre-training, intermediate training, and model tuning.

[0053] Specifically, a deep learning model is obtained based on a multi-stage training strategy, whereby the deep learning model is used to perform mapping tests between source code files and requirement design documents.

[0054] For example, such as Figure 3As shown in the embodiments of this application, a semantic association algorithm based on a multi-stage deep learning training strategy is proposed. First, rich semantic features are obtained on a large-scale pre-trained deep learning model. Then, adjustments are made to the file mapping task to adapt to the specific semantic structure of the source code and the requirement design document, thereby solving the problem of incomplete mapping between source code files and requirement design documents, which leads to missing or incorrect software functions.

[0055] Step 204: Obtain a multi-task model based on a deep learning model. The multi-task model is used to map and test the requirements design document and the test document.

[0056] Among them, the multi-task model can refer to the multi-neural network collaborative model, or it can also be called the multi-task learning model.

[0057] Specifically, a multi-task model is obtained based on a deep learning model, which is used to perform mapping tests between requirement design documents and test documents.

[0058] For example, such as Figure 3 As shown in the embodiments of this application, an entity information extraction algorithm based on a multi-neural network collaborative model is proposed. By using a pre-trained deep learning model to learn the semantic information implied in the samples, the model's relation extraction capability is improved, thereby solving the problem of inconsistent mapping between requirement design documents and test documents, which leads to reduced software reliability and controllability.

[0059] Step 206: Optimize the test document based on the relation extraction model; the optimization includes at least enhancing the semantic association between entities; the relation extraction model includes at least an encoding layer, an enhanced entity layer, and a classification layer.

[0060] Specifically, the test document is optimized based on the relation extraction model, and the optimization of the test document includes at least enhancing the semantic association between entities in the test document.

[0061] For example, such as Figure 3 As shown in the embodiments, this application proposes a document-level relation extraction model based on enhanced entity representation. It utilizes a deep learning model to design the document graph structure based on fundamental requirements and aggregates information from adjacent nodes through a graph neural network propagation mechanism. Sentence context and topic information relevant to entity relation prediction are integrated into the entity node representation of the basic document graph, thereby obtaining enhanced entity node representations and improving the quality of test documents. This solves the problems of low efficiency, low accuracy, and poor real-time performance in traditional methods for generating test documents.

[0062] In the aforementioned software testing methods, a deep learning model based on a multi-stage training strategy is used for mapping testing between source code files and requirement design documents. This involves introducing a pre-trained large model with a multi-stage training strategy and a reinforcement learning mechanism to improve the mapping accuracy and completeness of the deep learning model, reduce mapping bias, and ensure the accuracy and completeness of software functions. Secondly, a mapping test between requirement design documents and test documents based on a multi-neural network collaborative model is used. A multi-task learning model adapts to multi-level features, alleviating the problem of insufficient domain annotation corpus when extracting entity relationships, and effectively extracting relationships between different entities even with limited data. Finally, a test document optimization technique based on a relationship extraction model is used. By establishing a document graph structure and utilizing self-attention mechanisms and nearest-neighbor information aggregation techniques, relationships between entities are established, effectively enhancing the semantic associations between entities and deeply mining the inherent requirements of the test documents.

[0063] In one exemplary embodiment, obtaining a deep learning model based on a multi-stage training strategy includes:

[0064] Model pre-training is performed on unsupervised text data to obtain a deep learning model, which includes a general language representation.

[0065] The deep learning model is adjusted to transfer the knowledge it gains from code search tasks to the task of mapping and linking requirement design documents and source code files.

[0066] Unsupervised text data may refer to codebase data, but this application embodiment is not limited thereto; deep learning models have powerful semantic recognition capabilities and flexible transfer learning capabilities, which can effectively bridge the semantic differences between source code and requirement design documents; deep learning models may be simply referred to as models.

[0067] Specifically, a deep learning model is obtained by pre-training the model based on unsupervised text data. The deep learning model includes a general language representation. The deep learning model is then adjusted to transfer the knowledge obtained by the deep learning model from the code search task to the mapping and linking task between the deep learning model and the requirement design document and the source code file.

[0068] For example, embodiments of this application employ a method for enhancing the mapping between source code files and requirements design documents based on a multi-stage training strategy, such as... Figure 4As shown, deep learning models possess powerful semantic recognition capabilities and flexible transfer learning abilities, effectively bridging the semantic differences between source code and requirement design documents. This project proposes a three-stage training strategy based on a deep learning model, sequentially using pre-training, intermediate training, and tuning to enhance the mapping relationship layer by layer. First, the model is pre-trained on a large amount of unsupervised text data to obtain a general language representation. Then, through tuning, knowledge gained from data-rich code search tasks is transferred to the mapping linking task between requirement design documents and source code files. Unsupervised learning allows for the effective deployment of deep learning models in environments with sparse multi-source document data in basic software, effectively alleviating the problem of sparse mapping file data.

[0069] Optionally, model pre-training is performed on large amounts of codebase data (such as the CodeSearchNet dataset), with the primary goal of enabling the model to learn the lexical distribution and semantic relationships between natural language documents and programming language documents. By utilizing Masked Language Modeling (MLM), certain words in the input text are randomly masked, and the model is then trained to predict these masked words, thereby capturing deep linguistic features and patterns. Figure 4 As shown, the pre-training process includes obtaining a model from training data (functions).

[0070] In one exemplary embodiment, the deep learning model is adjusted to transfer the knowledge gained by the deep learning model from the code search task to the task of mapping and linking requirements design documents and source code files, including:

[0071] Intermediate training of deep learning models for the code search problem aims to establish general principles for the deep learning models to identify the correlation between natural language and programming languages; the code search problem includes obtaining corresponding code snippets based on natural language descriptions.

[0072] Based on sample data from associated source code files and requirements design documents, the knowledge acquired from the intermediate training of the deep learning model is transferred to the adjustment stage of the deep learning model, and the deep learning model is adjusted accordingly.

[0073] Specifically, during intermediate training, a deep learning model is trained on a code search problem, which involves retrieving corresponding code snippets (such as function definitions) using natural language descriptions (e.g., docstrings of functions). By utilizing a labeled dataset, the model is trained to learn general rules for discerning correlations between natural and programming languages. This stage is set as a binary classification problem, where the model needs to determine whether a given docstring accurately describes the corresponding function. A balanced training set containing equal amounts of positive and negative samples is constructed, and a dynamic random negative sampling strategy is employed. At the beginning of each training cycle, the training set is updated to include all function and docstring pairs to ensure the model learns from previously unseen negative samples. Simultaneously, a reinforcement learning mechanism is used to continuously improve the performance of mapping predictions through feedback, thereby increasing the accuracy of the mapping.

[0074] Optionally, such as Figure 4 As shown, during intermediate training, a connection classifier is formed based on the pre-trained model, and the model is trained on the code search problem based on the pre-trained training data (function) and the corresponding documentation. Subsequently, knowledge transfer forms a new connection classifier.

[0075] When tuning deep learning models, such as Figure 4 As shown, sample data from the associated source code and requirements design document (i.e.) Figure 4 The process (in terms of functions and requirements) requires only a small amount of manually labeled data, effectively transferring the knowledge gained in the intermediate training phase to the adjustment phase. This allows the model to achieve effective classification performance in downstream tasks with only a minimal amount of labeled data. The training in the adjustment phase also employs a similar binary classification method to evaluate the model's ability to distinguish between relevant and irrelevant mappings. Furthermore, by applying a dynamic random negative sampling strategy, the model's ability to identify difficult negative samples is enhanced, which helps improve the model's generalization ability and avoids premature overfitting.

[0076] In one exemplary embodiment, mapping tests between source code files and requirements design documents are performed, including:

[0077] The source code files and requirements design documents are broken down and then sequentially fed into the deep learning model;

[0078] The deep learning model uses feature transformation technology to determine whether there is a valid mapping link between the source code file and the requirement design document, and then summarizes and outputs the evaluation results.

[0079] Mapping tests can also refer to model inference in deep learning models.

[0080] Specifically, during model inference, the source code is segmented by function, and the requirements design document is decomposed into individual requirement units. Each source code segment and its corresponding requirements design document segment are sequentially input into the model. By comparing the similarity between each requirement in the requirements design document and each part of the source code, valid mapping links are determined, and the evaluation results are summarized and output. Furthermore, to improve inference accuracy, the model uses feature transformation techniques learned during training, such as text embedding and contextual analysis, when processing input data to ensure a deep understanding of the semantic relationships between requirements and code.

[0081] In one exemplary embodiment, obtaining a multi-task model based on a deep learning model includes:

[0082] The cleaned and preprocessed text sequence data of the requirements design documents and test documents are input into a deep learning model to obtain the hidden representation of the text;

[0083] Hidden representations are used as shared representations, and independent output layers are designed for each task to obtain a multi-task model; each output layer includes a fully connected layer and an activation function, and each output layer is connected to the shared representation.

[0084] The multi-task model was adjusted based on the associated dataset of requirements design documents and test documents.

[0085] Requirements design documents and test documents play a crucial role in ensuring software quality and functional integrity. Requirements design documents typically describe details of the software system's architecture, module design, data flow, etc., serving as a guide for developers when implementing software functions. Test documents, on the other hand, contain software testing plans, methodologies, and test cases, used to verify the correctness and stability of the software under different conditions. There is a close relationship between design and test documents; they complement each other. The requirements descriptions and design solutions in the requirements design document directly influence the design and execution of test cases in the test document; therefore, the mapping relationship between the requirements design document and the test document is essential. Intelligent review of the mapping relationship between design and test documents can help the development team better understand the connection between requirements, design, and testing, thereby improving the efficiency and quality of software development.

[0086] Specifically, this application employs a method for enhancing the mapping between requirement design documents and test documents based on a multi-neural network collaborative model. This method enables intelligent review of the mapping relationship between design documents and test documents. One of the key technologies is information extraction. The goal of information extraction is to identify entities and events in the text and their related relationships, extracting specific information from structured or semi-structured natural language text and organizing it into a structured form. However, since entities are scattered throughout the text, it is impossible to determine the possible relationships between entities or to summarize them into structured information. Therefore, entity relationship extraction technology is needed to determine the relationships between entities in the text and to structure the extracted data. The mapping enhancement method in this application utilizes deep learning models, semantic role embedding, entity attention, and multi-task learning to effectively extract unit features from target entity pairs and merge them into fusion features, thereby capturing abstract semantics and sentence structure.

[0087] Multi-task learning combines multiple deep learning models into a single multi-task model, for example... Figure 5 As shown, it has multiple input layers connected to a parameter sharing layer. Figure 5 Layers A, B, and C, along with their corresponding output layers, simultaneously execute tasks A, B, and C, respectively. Each task has an independent branch, used for capturing local patterns, extracting global context, and identifying word relationships, while lower layers share representations. By integrating the captured local and global context, the model learns more supervised information, thereby improving the effectiveness of entity relationship extraction.

[0088] For example, the specific steps of this application embodiment can be as follows:

[0089] First, a high-dimensional vector representation of the text is created. The cleaned and preprocessed text sequence data of the requirements design documents and test documents are input into a deep learning model. Pre-trained weights are loaded into the model to encode the text sequences of the requirements design documents and test documents. This step generates a hidden representation of the text, mapping the text sequences to high-dimensional vector representations.

[0090] Secondly, multi-task learning. The hidden representations generated by the deep learning model are used as shared representations, i.e. Figure 5 The model employs a parameter-sharing layer for multi-task learning, ensuring that the learned text representations are consistent across all tasks. An independent output layer is designed for each task, comprising a fully connected layer and an activation function. Each task's output layer is connected to the shared representations of the deep learning model, enabling simultaneous learning and optimization across multiple tasks. A custom comprehensive loss function is also defined, weighted for different tasks, allowing the model to balance and correlate different tasks during training.

[0091] Finally, model tuning, or model fine-tuning, involves using a dataset that associates small requirement design documents with test documents to adjust certain layers of the model to adapt to the needs of a specific task, thereby improving the model's performance and generalization ability.

[0092] In one exemplary embodiment, the test document is optimized based on a relation extraction model, including:

[0093] The encoding layer based on the relation extraction model is used to vectorize the requirements design document;

[0094] An enhanced entity layer based on a relation extraction model is used to construct a contextual representation of an entity and create the final representation of the entity node.

[0095] Entity type prediction is performed using a classification layer based on a relation extraction model.

[0096] Test documentation plays a crucial role in software development, recording information such as functional requirements, test plans, test cases, and test results. This helps ensure software quality, improve development efficiency, and reduce maintenance costs. In development environments with frequent changes in software requirements and rapid iteration, test documentation has high real-time requirements and needs continuous updates. If test documentation is not updated in a timely manner, using outdated documentation will negatively impact the testing results. With the development of artificial intelligence, generative techniques have shown great potential in test documentation enhancement and optimization. Generative techniques can achieve dynamic updates and real-time generation of documents, maintaining consistency between the documentation and the actual software implementation. When software requirements or functions change, generative techniques can quickly update the documentation, ensuring the team always uses the latest documentation for testing, improving the effectiveness and reliability of testing. However, using generative techniques to optimize test documentation often faces the problem of long-distance dependencies between entity nodes. Long-distance dependencies make it difficult for models to effectively capture distant entity relationships, increasing training complexity and computational costs, and resulting in a lack of logic and coherence in the generated text. These problems limit the application of generative models in test documentation optimization.

[0097] Specifically, embodiments of this application employ a test document optimization method based on a document-level relation extraction model with enhanced entity representation, such as... Figure 6As shown, this design aims to further optimize the quality of the test documents. The entire model consists of three parts: an encoding layer, an augmented entity representation layer, and a classification layer. In the encoding layer, a pre-trained model is used to encode the input document. The augmented entity representation layer is the core of the model. To improve the ability to extract intra-sentence relations, sentence context representations are added to the entity representations. To enhance inter-sentence reasoning capabilities, document topic information is added to the constructed entity representations. Thus, the final entity representation consists of the entity representation itself, the sentence context representation, and the document topic information representation. Then, the representational capability of entity nodes is further enhanced by aggregating information from neighboring nodes. In the classification layer, a bilinear function is used to classify entity pairs.

[0098] Optionally, such as Figure 6 As shown, the specific steps of this application embodiment can be as follows:

[0099] First, data vectorization is performed. Each word in the requirements design document is mapped to a vector, and the word's embedding representation is concatenated with its corresponding entity type embedding representation. Then, the embedding representation of each word is fed into the model encoder to obtain the final vectorized representation of each word. Figure 6 The word embeddings and entity type embeddings shown are illustrated.

[0100] Subsequently, an entity context representation is constructed. A document graph is constructed, containing entity nodes and edges between entities. A self-attention mechanism is employed, assigning higher weights to information that has a greater impact on an entity and lower weights to information that has less impact. This is achieved by aggregating different words in a sentence to form the sentence's context representation. Figure 6 The diagram shows the construction of sentence context representations and document topic representations based on the model and sentences 1, 2, 3, etc.

[0101] Create the final representation of the entity node. Connect the entity node representation with the representation of its containing sentence and the document body information representation, so that the entity representation possesses information from both, enabling both intra-sentence and inter-sentence reasoning. Use a convolutional network to aggregate neighbor node information, and then connect them to obtain the final representation of the entity node. Figure 6 The entity itself is represented in the text.

[0102] Finally, entity type prediction is performed. After multiple network aggregations, representations of all entity nodes are obtained. Then, the sigmoid function is used to calculate the probability of each relation type for the entity pair, i.e. Figure 6 The system uses convolutional networks to classify entities and calculate relationship types.

[0103] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0104] Based on the same inventive concept, this application also provides a software testing apparatus for implementing the software testing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more software testing apparatus embodiments provided below can be found in the limitations of the software testing method described above, and will not be repeated here.

[0105] In one exemplary embodiment, such as Figure 7 As shown, a software testing device 900 is provided, including: a multi-stage training module 901, a multi-task model acquisition module 902, and a test document optimization module 903, wherein:

[0106] The multi-stage training module 901 is used to obtain a deep learning model based on a multi-stage training strategy. The deep learning model is used to perform mapping tests between source code files and requirement design documents.

[0107] The multi-task model acquisition module 902 is used to acquire a multi-task model based on a deep learning model. The multi-task model is used to map the requirement design document and the test document for testing.

[0108] Test document optimization module 903 is used to optimize test documents based on the relation extraction model; optimization includes at least enhancing the semantic relationship between entities.

[0109] In one embodiment, the multi-stage training module 901 is further configured to perform model pre-training based on unsupervised text data to obtain a deep learning model, the deep learning model including a general language representation; and to adjust the deep learning model by transferring the knowledge obtained by the deep learning model from the code search task to the mapping and linking task between the requirement design document and the source code file.

[0110] In one embodiment, the multi-stage training module 901 is further used to perform intermediate training on the deep learning model for the code search problem, which enables the deep learning model to identify the general principles of correlation between natural language and programming language; the code search problem includes obtaining corresponding code snippets based on natural language descriptions; based on sample data from associated source code files and requirement design documents, the knowledge obtained by the deep learning model from intermediate training is transferred to the adjustment of the deep learning model, i.e., the fine-tuning stage, and the deep learning model is fine-tuned.

[0111] In one embodiment, the multi-stage training module 901 is further used to decompose the source code file and the requirement design document, and input them into the deep learning model in sequence; the deep learning model uses feature transformation technology to determine whether there is a valid mapping link between the source code file and the requirement design document, and summarizes and outputs the evaluation results.

[0112] In one embodiment, the multi-task model acquisition module 902 is further configured to input the cleaned and preprocessed text sequence data of the requirement design document and the test document into the deep learning model to obtain the hidden representation of the text; use the hidden representation as a shared representation to design independent output layers for each task to obtain the multi-task model; wherein each output layer includes a fully connected layer and an activation function, and each output layer is connected to the shared representation; and adjust the multi-task model based on the associated dataset of the requirement design document and the test document.

[0113] In one embodiment, the test document optimization module 903 is further configured to: use an encoding layer based on the relation extraction model to vectorize the requirement design document; use an enhanced entity layer based on the relation extraction model to construct the context representation of the entity and create the final representation of the entity node; and use a classification layer based on the relation extraction model to predict the entity type.

[0114] Each module in the aforementioned software testing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0115] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores software test data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a software testing method.

[0116] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0117] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps shown in the method described above.

[0118] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0119] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps shown in the method described above.

[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0123] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A software testing method, characterized in that, The method includes: A deep learning model is obtained based on a multi-stage training strategy. The deep learning model is used to perform mapping tests between source code files and requirement design documents. The multi-stage training strategy includes at least pre-training, intermediate training, and model tuning. A multi-task model is obtained based on the deep learning model, and the multi-task model is used to perform mapping tests on the requirement design document and the test document. The test document is optimized based on the relation extraction model; the optimization includes at least enhancing the semantic association between entities; the relation extraction model includes at least an encoding layer, an enhanced entity layer, and a classification layer.

2. The method according to claim 1, characterized in that, The method for obtaining a deep learning model based on a multi-stage training strategy includes: The deep learning model is obtained by pre-training the model based on unsupervised text data, and the deep learning model includes a general language representation. The deep learning model is adjusted to transfer the knowledge obtained by the deep learning model from the code search task to the mapping and linking task between the requirement design document and the source code file.

3. The method according to claim 2, characterized in that, The adjustment of the deep learning model, transferring the knowledge obtained by the deep learning model from the code search task to the mapping and linking task between the requirements design document and the source code file, includes: Intermediate training is performed on the deep learning model for the code search problem to enable the deep learning model to identify the general principles of correlation between natural language and programming language; the code search problem includes obtaining corresponding code snippets based on natural language descriptions; Based on the associated source code files and sample data from the requirements design document, the knowledge acquired by the deep learning model from the intermediate training is transferred to the adjustment stage of the deep learning model, and the deep learning model is adjusted.

4. The method according to claim 1, characterized in that, The mapping test between the source code file and the requirement design document includes: The source code file and the requirements design document are decomposed and then sequentially input into the deep learning model; The deep learning model, based on feature transformation technology, determines whether there is a valid mapping link between the source code file and the requirement design document, and summarizes and outputs the evaluation results.

5. The method according to claim 1, characterized in that, The process of obtaining a multi-task model based on the deep learning model includes: The cleaned and preprocessed text sequence data of the requirement design document and the test document are input into the deep learning model to obtain the hidden representation of the text; The hidden representation is used as a shared representation, and independent output layers are designed for each task to obtain the multi-task model; wherein each output layer includes a fully connected layer and an activation function, and each output layer is connected to the shared representation; The multi-task model is adjusted based on the associated dataset of the requirements design document and the test document.

6. The method according to claim 1, characterized in that, The optimization of the test document based on the relation extraction model includes: Based on the encoding layer of the relation extraction model, the requirement design document is vectorized. Based on the enhanced entity layer of the relation extraction model, a contextual representation of the entity is constructed, and the final representation of the entity node is created. Based on the classification layer of the aforementioned relation extraction model, entity type prediction is performed.

7. A software testing device, characterized in that, The device includes: A multi-stage training module is used to obtain a deep learning model based on a multi-stage training strategy. The deep learning model is used to perform mapping tests between source code files and requirement design documents. A multi-task model acquisition module is used to acquire a multi-task model based on the deep learning model, and the multi-task model is used to perform mapping tests on the requirement design document and the test document. The test document optimization module is used to optimize the test document based on the relation extraction model; the optimization includes at least enhancing the semantic association between entities.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for enhancing classification-based software demand tracking link recovery through knowledge learning and electronic device

    CN113011461A

  • Rail transit signal system function demand tracing method and device and electronic equipment

    CN115758135A

  • Test case generation method and system based on large model and retrieval enhancement generation

    CN119537222A

  • Document-level relation extraction method based on multi-task learning and knowledge distillation

    CN119761495A

  • Establishing and maintaining a relationship between a three-dimensional model and related data

    US20040250236A1

Cited By

  • Software tracing method and system based on multi-agent collaborative decision

    CN121680790A