Large model construction method and device for basic software testing
By constructing a large model and combining it with multi-dimensional scoring to select the target model, the problem of reliance on manual operation in basic software testing has been solved, achieving an efficient and accurate testing process, reducing labor costs and improving testing quality and efficiency.
Patent Information
- Application Number
- CN202511567857.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2025-11-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing basic software testing process relies too heavily on manual operation, which makes testing time-consuming, labor-intensive, and difficult to guarantee quality and efficiency.
By constructing a large model, obtaining an evaluation dataset, and calling different preset models to process it, the target model is selected for testing based on multi-dimensional scoring. The dataset includes real data samples and synthetic data samples generated by adversarial networks. The target model is determined by combining multi-dimensional scoring.
It reduced labor costs, improved the quality and efficiency of testing, enhanced the generalization ability of the target model, and ensured comprehensive coverage of basic software testing and accurate identification of potential defects.
Smart Images

Figure CN121029628A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software testing technology, and in particular to a method and apparatus for building large models for basic software testing. Background Technology
[0002] With the rapid development of information technology, the functions of basic software are becoming increasingly rich and complex, which places higher demands on the testing of basic software. Currently, the testing process of basic software generally relies excessively on manual operation, which not only makes the testing process time-consuming and labor-intensive, but also makes it difficult to guarantee the quality and efficiency of the testing. Summary of the Invention
[0003] Therefore, it is necessary to provide a method and apparatus for building large models for basic software testing that can reduce the manpower cost of basic software testing while ensuring the quality and efficiency of testing, in order to address the above-mentioned technical problems.
[0004] Firstly, this application provides a method for constructing a large model for basic software testing, the method comprising:
[0005] Obtain the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network;
[0006] Different preset models were called to process the evaluation dataset, resulting in multiple processing results. The preset models included ChatGLM model, LLaMA model, Baichuan model and Qwen model.
[0007] Based on the processing results, the corresponding preset models are scored in multiple dimensions to obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring.
[0008] The target model is determined from multiple preset models based on the multi-dimensional scoring results. The target model is used to test the basic software.
[0009] In one embodiment, obtaining the assessment dataset includes:
[0010] Construct an adversarial network, which includes a generator and a discriminator; use a random noise vector as input to the generator and output test data samples; use real data samples and test data samples as input to the discriminator.
[0011] Construct an adversarial loss function and train an adversarial network based on the adversarial loss function. Input a random noise vector into the trained adversarial network to obtain synthetic data samples. Construct an evaluation dataset based on the synthetic data samples and real data samples.
[0012] In one embodiment, the adversarial loss function is as follows:
[0013] ;
[0014] Where G is the generator; D is the discriminator; E is the expected value; x is the real data sample; z is the random noise vector; X r To extract a predetermined number of real data samples from the real data sample; P r (.) represents the probability distribution of the real data sample; P g (.) represents the probability distribution of the synthetic data sample; G(z) represents the synthetic data sample; D(x) represents the probability estimate that the discriminator confirms the real data sample is real; D(G(z)) represents the probability estimate that the discriminator confirms the synthetic data sample is real.
[0015] In one embodiment, training an adversarial network based on an adversarial loss function includes:
[0016] The adversarial network is iteratively trained based on the adversarial loss function until the preset number of iterations is reached.
[0017] In each iteration of training, the parameters of the adversarial network are adjusted based on the loss value calculated by the adversarial loss function to train the adversarial network model until convergence. The real data samples and test data samples used in this iteration of training are used as historical data, and a preset proportion of historical data is added to the test data samples used in the next iteration of training.
[0018] In one embodiment, the method further includes:
[0019] The performance of the corresponding preset models is scored based on performance evaluation metrics and processing results; performance evaluation metrics include precision, recall, and F1 score.
[0020] The corresponding preset models are scored in a fine-grained dimension based on the fine-grained evaluation indicators and processing results. The fine-grained evaluation indicators include correctness, completeness and semantic consistency.
[0021] Calculate the consistency index between the human scoring results and the fine-grained dimension scores, and then score the subjective dimension based on the consistency index. The consistency index is either Cohen's Kappa coefficient or Spearman correlation coefficient.
[0022] In one embodiment, the target model is also used to generate unit test cases and system test cases, both of which are used to test the underlying software.
[0023] The target model is also used for fuzz testing, error analysis and debugging of underlying software, as well as program repair.
[0024] Secondly, this application also provides a large model building apparatus for basic software testing, the apparatus comprising:
[0025] The acquisition module is used to acquire the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network.
[0026] The model invocation module is used to invoke different preset models to process the evaluation dataset and obtain multiple processing results. The preset models include ChatGLM model, LLaMA model, Baichuan model and Qwen model.
[0027] The scoring module is used to score the corresponding preset models in multiple dimensions based on each processing result, and obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring;
[0028] The determination module is used to identify the target model from multiple preset models based on the multi-dimensional scoring results. The target model is used to conduct tests on the basic software.
[0029] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0030] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0031] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0032] The aforementioned method and apparatus for constructing large models for basic software testing acquires an evaluation dataset and processes it using different preset models to obtain multiple processing results. Based on each processing result, the corresponding preset models are scored in multiple dimensions to obtain multi-dimensional scoring results. The target model for testing the basic software is then determined from among the multiple preset models based on the multi-dimensional scoring results. In this application, the evaluation dataset includes real data samples and synthetic data samples generated by adversarial networks, thereby improving the generalization ability of the subsequent target model. Furthermore, the use of multi-dimensional scoring to determine the target model ensures the quality and efficiency of subsequent testing of the basic software model by the target model, and the use of the target model for testing the basic software reduces labor costs. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart illustrating a method for building a large model for basic software testing in one embodiment.
[0035] Figure 2 This is a flowchart illustrating a large model building method for basic software testing in another embodiment;
[0036] Figure 3 This is a schematic diagram of the functional coverage of the target model in one embodiment;
[0037] Figure 4 This is a structural block diagram of a large model building apparatus for basic software testing in one embodiment;
[0038] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0040] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0041] Software testing plays a crucial role in the lifecycle of foundational software development, ensuring that the quality and security of the software meet stringent industry standards. With advancements in industrial technology, the complexity and functionality of foundational software systems are constantly increasing, posing greater challenges to testing. Traditional testing methods are proving inadequate in handling these complex systems, especially in the face of rapidly changing environments and high security requirements. Traditional foundational software testing processes often rely excessively on manual intervention, making the process time-consuming and labor-intensive, and failing to guarantee both quality and efficiency.
[0042] The large model construction method for basic software testing provided in this application involves acquiring an evaluation dataset and processing it using different preset models to obtain multiple processing results. Based on each processing result, the corresponding preset models are scored in multiple dimensions to obtain multi-dimensional scoring results. A target model for testing basic software is determined from the multiple preset models based on the multi-dimensional scoring results. That is, a suitable large model base (target model) is selected to build a dedicated large model for basic software testing. This aims to help testers accurately identify potential defects and risks through intelligent analysis and prediction, reducing labor costs. Furthermore, the evaluation dataset includes real data samples and synthetic data samples generated by adversarial networks, thereby improving the generalization ability of the subsequent target model. Simultaneously, the use of multi-dimensional scoring to determine the target model ensures the quality and efficiency of the subsequent target model in testing basic software models.
[0043] It should be noted that the functions of the basic software involved in the embodiments of this application can be set according to the actual situation, and are not limited in the embodiments of this application.
[0044] In one exemplary embodiment, such as Figure 1 As shown, a method for building large models for basic software testing is provided, the method including:
[0045] S102, Obtain the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network.
[0046] Specifically, a knowledge base (basic software testing resource library) covering the field of quality and reliability can be built within the computer equipment. This resource includes historical test cases, defect reports, and standards and specifications. By utilizing natural language processing techniques such as named entity recognition, relation extraction, and knowledge graph construction, valuable knowledge units are extracted from this unstructured data. Through entity recognition and relation extraction, textual information is transformed into a structured knowledge graph, forming a "entity-relationship-entity" triple structure. To ensure the accuracy and consistency of knowledge, knowledge fusion and deduplication are also required to eliminate redundant and conflicting information, thereby obtaining realistic data samples. To enhance the diversity and coverage of the data, synthetic data samples generated by adversarial networks are added to the testing dataset. These synthetic data samples can simulate the distribution and characteristics of real data samples, improving the model's generalization ability.
[0047] S104 calls different preset models to process the evaluation dataset and obtain multiple processing results. The preset models include ChatGLM model, LLaMA model, Baichuan model and Qwen model.
[0048] The preset model may also include other large models, and is not limited to the examples mentioned in the embodiments of this application.
[0049] Specifically, large-scale pre-trained models have become a core driving force for technological progress. ChatGLM, Baichuan, and Qwen models, as leading large-scale models in the industry, have demonstrated outstanding performance in fields such as Natural Language Processing (NLP), knowledge understanding, and human-computer interaction. These models are not only the results of the latest artificial intelligence research but also indispensable tools in practical applications. With the continuous advancement of technology, selecting appropriate large-scale models has become particularly important for building dedicated large-scale models for basic software testing.
[0050] Table 1 below provides an exemplary overview of common foundational models used for basic software testing:
[0051] Table 1
[0052]
[0053] The computer equipment processes the evaluation dataset by calling different preset models to obtain multiple processing results, so that the final large model base used for basic software testing can be determined based on the processing results.
[0054] S106. Based on each processing result, the corresponding preset model is scored in multiple dimensions to obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring.
[0055] Specifically, computer equipment can perform multi-dimensional scoring on the corresponding preset models based on each processing result, namely, performance dimension scoring, fine-grained dimension scoring, and subjective dimension scoring. By combining the aforementioned scoring results, the final score of the model (multi-dimensional scoring result) is calculated according to preset weights, clarifying the comprehensive performance of different preset models in each evaluation dimension, determining their applicable scenarios and limitations, and, for example, generating a complete model evaluation report to provide a basis for subsequent model optimization and application.
[0056] S108 determines the target model from multiple preset models based on multi-dimensional scoring results. The target model is used to test the basic software.
[0057] Specifically, based on the final score (multi-dimensional scoring results) and evaluation analysis, the basic large model that performs best in terms of quality, reliability, and knowledge representation is selected as the target model for testing the basic software.
[0058] It should be noted that the basis for model selection, analysis of advantages and disadvantages, and potential optimization space can also be recorded, ultimately forming a complete basic large model selection report and application suggestions to guide the deployment and application of the model.
[0059] In the aforementioned method for constructing a large model for basic software testing, an evaluation dataset is acquired, and different preset models are called to process the dataset, resulting in multiple processing results. Based on each processing result, the corresponding preset model is scored in multiple dimensions to obtain multi-dimensional scoring results. The target model for testing basic software is determined from these preset models based on the multi-dimensional scoring results. In other words, a suitable large model base (target model) is selected to build a dedicated large model for basic software testing. This aims to help testers accurately identify potential defects and risks through intelligent analysis and prediction, reducing labor costs. Furthermore, the evaluation dataset includes real data samples and synthetic data samples generated by adversarial networks, thereby improving the generalization ability of the subsequent target model. Simultaneously, the use of multi-dimensional scoring to determine the target model ensures the quality and efficiency of the subsequent target model in testing the basic software model.
[0060] In one embodiment, obtaining the assessment dataset includes:
[0061] Construct an adversarial network, which includes a generator and a discriminator; use a random noise vector as input to the generator and output test data samples; use real data samples and test data samples as input to the discriminator.
[0062] Construct an adversarial loss function and train an adversarial network based on the adversarial loss function. Input a random noise vector into the trained adversarial network to obtain synthetic data samples. Construct an evaluation dataset based on the synthetic data samples and real data samples.
[0063] Specifically, to address the scarcity of test data for large-scale basic software evaluation models, during the data collection phase, basic software evaluation knowledge is retrieved from the basic software evaluation resource library. A basic software evaluation corpus is designed using true / false, multiple-choice, and subjective questions. Real data samples generated in the previous phase are used as input, and a large number of synthetic data samples with similar statistical characteristics are generated using a generative adversarial network (GAN). The adversarial network consists of a generator and a discriminator, and is trained using an adversarial loss function. During training, the generator and discriminator are alternately optimized to achieve the purpose of adversarial training.
[0064] When the adversarial network completes training, it can be understood that the test data samples generated by the generator at this time follow the same distribution as the real data samples, thus realizing the effective simulation of the real data samples. In the dataset construction and training stage, the real data samples and the above-mentioned synthetic data samples are mixed in proportion (set according to the actual situation, which is not limited in this application embodiment) to construct a dynamically enhanced basic software evaluation knowledge evaluation dataset.
[0065] In this embodiment, an adversarial network and an adversarial loss function are constructed respectively. The adversarial network is trained according to the adversarial loss function. Random noise vectors are input into the trained adversarial network to obtain synthetic data samples. An evaluation dataset is constructed based on the synthetic data samples and real data samples. This solves the problem of scarce test data for large models in basic software evaluation and improves the generalization ability of subsequent target models.
[0066] In one embodiment, the adversarial loss function is as follows:
[0067] ;
[0068] Where G is the generator; D is the discriminator; E is the expected value; x is the real data sample; z is the random noise vector; X r To extract a predetermined number of real data samples from the real data sample; P r (.) represents the probability distribution of the real data sample; P g (.) represents the probability distribution of the synthetic data sample; G(z) represents the synthetic data sample; D(x) represents the probability estimate that the discriminator confirms the real data sample is real; D(G(z)) represents the probability estimate that the discriminator confirms the synthetic data sample is real.
[0069] In one embodiment, training an adversarial network based on an adversarial loss function includes:
[0070] The adversarial network is iteratively trained based on the adversarial loss function until the preset number of iterations is reached.
[0071] In each iteration of training, the parameters of the adversarial network are adjusted based on the loss value calculated by the adversarial loss function to train the adversarial network model until convergence. The real data samples and test data samples used in this iteration of training are used as historical data, and a preset proportion of historical data is added to the test data samples used in the next iteration of training.
[0072] The preset ratio and preset number of times can be set according to the actual situation, and are not limited in this embodiment.
[0073] Specifically, the iterative training process of the adversarial network based on the adversarial loss function can be as follows: (1) Initial stage: The generator accepts random noise vectors as input and generates synthetic data samples through a multi-layer neural network structure. The discriminator is input with real data samples from the real dataset and synthetic data samples generated by the generator; (2) Discriminator optimization stage: In this stage, the discriminator aims to distinguish between real data and generated data. By labeling real data as the real category and the data generated by the generator as the fake category, the discriminator minimizes the classification error through the gradient descent algorithm to improve its ability to distinguish between real and fake data; (3) Generator optimization stage: Subsequently, the generator aims to deceive the discriminator so that it misclassifies the generated data samples as real data. The generator updates the model parameters through the gradient descent algorithm to maximize the probability that the discriminator identifies the synthetic data as real data, thereby continuously optimizing the quality of the synthetic data; (4) Alternating iterative optimization: By repeatedly alternating between stages (2) and (3), the generator and discriminator continuously optimize their performance. The generator gradually generates data that is closer to the real data distribution, while the discriminator’s ability to distinguish data gradually improves with training; (5) Convergence state: Finally, when the data distribution generated by the generator is difficult to distinguish from the real data distribution by the discriminator, that is, when the discriminator’s accuracy is close to the level of random guessing, the training reaches the convergence state.
[0074] It's important to note that an incremental learning strategy with stepwise fine-tuning is employed to train the adversarial network (ANN). Initially, the AAN is trained on the original dataset until convergence. After each training iteration, a portion of historical data is retained to maintain the AAN's memory and stability of existing knowledge. Subsequently, newly collected data is introduced and combined with the retained historical data to form a new training set (the test data samples used in the next iteration). This updated training set is then used to fine-tune the AAN. In this way, each training round effectively absorbs new knowledge while preventing the AAN from forgetting existing knowledge, thus continuously optimizing the AAN. After each stage converges, a portion of historical data is retained for the next training round, while newly added data is used for fine-tuning the AAN. Finally, the incrementally trained AAN is evaluated, and feedback data streams are collected. This feedback data is preprocessed and incorporated into the generation of synthetic data samples and fine-tuning of the AAN in the next round, forming a closed-loop mechanism; this ensures the continuous evolution of the AAN's capabilities and the continuous enrichment of the dataset.
[0075] In this embodiment, the combination of synthetic data technology and incremental learning solves the problem of scarce evaluation data for large-scale models used in basic software testing. The dynamically augmented dataset ensures the continuity and diversity of model inputs, and the incremental fine-tuning strategy enables the model to continuously absorb new knowledge and dynamically adapt to changes through iterative cycles.
[0076] In one embodiment, the method further includes:
[0077] The performance of the corresponding preset models is scored based on performance evaluation metrics and processing results; performance evaluation metrics include precision, recall, and F1 score.
[0078] The corresponding preset models are scored in a fine-grained dimension based on the fine-grained evaluation indicators and processing results. The fine-grained evaluation indicators include correctness, completeness and semantic consistency.
[0079] Calculate the consistency index between the human scoring results and the fine-grained dimension scores, and then score the subjective dimension based on the consistency index. The consistency index is either Cohen's Kappa coefficient or Spearman correlation coefficient.
[0080] Specifically, performance evaluation indicators and fine-grained evaluation indicators are formulated to complete the performance dimension scoring and fine-grained dimension scoring of the preset model. The performance dimension score can be used to characterize the performance of the preset model in terms of quality reliability knowledge expression, and the fine-grained dimension score can be used to characterize the performance of the preset model in each fine-grained dimension.
[0081] The manual scoring results are determined by manual annotation or expert scoring, which will not be elaborated in this embodiment. By calculating the consistency index between the manual scoring results and the fine-grained dimension scoring, subjective dimension scoring is performed based on the consistency index. The subjective dimension scoring can be used to characterize the reliability and credibility of the preset model to ensure that the output of the preset model meets human judgment standards.
[0082] In this embodiment, multi-dimensional scoring is used to determine the target model, ensuring the quality and efficiency of subsequent testing of the target model against the basic software model.
[0083] To facilitate understanding by those skilled in the art, the following example illustrates the method for constructing large models for basic software testing. Figure 2 As shown:
[0084] Step 1: Extract knowledge from the knowledge source repository of quality and reliability.
[0085] A knowledge base covering the field of quality and reliability can be constructed, including historical test cases, defect reports, and standards and specifications. Using natural language processing techniques, such as named entity recognition, relation extraction, and knowledge graph construction methods, valuable knowledge units can be extracted from this unstructured data. Through entity recognition and relation extraction, textual information is transformed into a structured knowledge graph, forming a "entity-relationship-entity" triple structure. To ensure the accuracy and consistency of knowledge, knowledge fusion and deduplication are also required to eliminate redundant and conflicting information, thereby obtaining realistic data samples.
[0086] Step 2: Construct an assessment dataset based on a self-built question bank and data synthesis enhancement technology.
[0087] When constructing the assessment dataset, the first step is to design corresponding test questions based on the knowledge content extracted in step one (real data samples), forming a preliminary question bank. To enhance the diversity and coverage of the data, data augmentation techniques, including generative adversarial networks (GANs) and variational autoencoders (VAEs), are used to expand the original data and generate new data samples. These synthetic data samples can simulate the distribution and characteristics of real data, improving the model's generalization ability. Finally, the original question bank data and the augmented data are integrated to form a complete dataset for model training and evaluation. That is, synthetic data samples generated by adversarial networks are added to the assessment dataset, which can simulate the distribution and characteristics of real data samples and improve the model's generalization ability.
[0088] Step 3: Evaluate the model's quality, reliability, and knowledge representation capabilities based on evaluation criteria.
[0089] During the model evaluation phase, clear evaluation criteria and metrics need to be established, such as accuracy, recall, and F1 score. Specific evaluation tasks are designed, the constructed dataset is input, and the model's output results are recorded. Then, based on the predefined evaluation criteria, the model's performance in terms of quality, reliability, and knowledge representation is objectively analyzed to identify its strengths and weaknesses.
[0090] Step 4: Perform fine-grained scoring on the model based on the scoring rules.
[0091] To more deeply evaluate model performance, fine-grained scoring rules are developed, specifying scoring dimensions such as correctness, completeness, and semantic consistency. Experts or automated scoring tools are then used to score each item of the model's output. Statistical analysis of the scoring results reveals the model's performance in each fine-grained dimension, identifying specific areas for improvement.
[0092] Step 5: Conduct an assessment and human consistency analysis.
[0093] At this stage, the model's scoring results are compared with those of human annotation or expert scoring, and consistency indices, such as Cohen's Kappa coefficient or Spearman correlation coefficient, are calculated. By analyzing the consistency between model evaluation and human evaluation, the reliability and credibility of the model are assessed, ensuring that the model's output conforms to human judgment standards.
[0094] Step 6: Determine the final score of the model.
[0095] Based on the aforementioned evaluation results, the final score of the model (multi-dimensional scoring result) is calculated according to the preset weights. The overall performance of the model across each evaluation dimension is clarified, and its applicable scenarios and limitations are determined. Finally, a complete model evaluation report is generated to provide a basis for subsequent model optimization and application.
[0096] Step 7: Complete the selection of the base model.
[0097] Based on the final scoring and evaluation analysis, the optimal basic large-scale model (target model) in terms of quality reliability knowledge representation is selected. The basis for model selection, advantages and disadvantages analysis, and potential optimization space are recorded. Finally, a complete basic large-scale model selection report and application recommendations are generated to guide the deployment and application of the model.
[0098] Meanwhile, with the continuous emergence of new open-source large-scale pre-trained models, the evaluation dataset generation method provided in this application can continuously enrich the basic software evaluation knowledge resource library, improve the evaluation system and benchmark test dataset, and use the reliable large model of project development quality as a tool to automatically evaluate the open-source large model, and determine whether to update the base large model based on the evaluation results.
[0099] In one embodiment, the target model is also used to generate unit test cases and system test cases, both of which are used to test the underlying software.
[0100] The target model is also used for fuzz testing, error analysis and debugging of underlying software, as well as program repair.
[0101] It is important to note that software testing, especially basic software testing, is an indispensable part of the basic software development cycle and bears the significant responsibility of ensuring the quality, security, and reliability of basic software products. With the rapid development of information technology, the functions of basic software systems are becoming increasingly complex, and security requirements are extremely high. This undoubtedly places more stringent demands on basic software testing. Traditional software testing methods, such as manual testing and automated script testing, show increasingly obvious limitations when faced with large-scale and highly integrated basic software and ever-changing requirements. To address this challenge, innovative technologies such as large-scale models are introduced, aiming to help testers accurately identify potential defects and risks through intelligent analysis and prediction, achieving comprehensive coverage of basic software testing.
[0102] Specifically, the functional coverage of the large-scale model (target model, hereinafter referred to as the large model) for basic software testing is as follows: Figure 3 As shown, the target model is also used to generate unit test cases. In basic software testing, the accurate generation of unit test cases is the cornerstone of ensuring software reliability. The large model can deeply analyze the source code and detailed requirements documents of the basic software, automatically generating highly targeted unit test cases. These test cases not only cover normal operating scenarios but also encompass possible extreme conditions and boundary situations, ensuring that each software unit can still operate stably under extreme environments. This method of automatically generating test cases greatly improves the comprehensiveness and efficiency of testing, while reducing the burden on testers in high-intensity basic software testing.
[0103] The target model is also used to generate system test cases. During the software system testing phase, the large model can simulate complex and special environments, generating diverse system-level test inputs. These inputs simulate various unexpected situations and abnormal behaviors in special environments, such as electromagnetic interference and network communication interruptions, to verify the performance and stability of the underlying software under extreme conditions. By deeply understanding and analyzing the functional characteristics of the underlying software and special environmental usage scenarios, the large model can generate highly realistic test inputs, ensuring that the software can perform its intended function in the production environment.
[0104] Target models are also used for fuzz testing, a technique that evaluates the behavior of a software system by injecting invalid, unexpected, or random data. The application of models in fuzz testing primarily lies in generating diverse input data. Fuzz testing is particularly important in basic software testing because it can assess the software's behavior when handling invalid or unexpected input. Large models are mainly used in fuzz testing to generate highly randomized input data that simulates various anomalies and unforeseen situations that may occur in specific environments. Through the generation capabilities of large models, a large amount of data conforming to the specific format of the basic software but with randomized content can be generated to reveal potential vulnerabilities in the software's handling of anomalous data. Furthermore, large models can iterate and optimize based on discovered errors, improving the effectiveness of fuzz testing and thus enhancing the robustness of the basic software.
[0105] The target model is also used for error analysis and debugging, where speed and accuracy are crucial in fundamental software testing. The large model can parse complex error logs and stack traces, generating easy-to-understand explanations and suggestions to help developers quickly locate and understand the root causes of errors. For specific errors in typical application scenarios, the large model can also generate repair suggestions or code patches based on the error description, assisting developers in efficient error fixing. This capability is especially important in emergency situations, enabling rapid restoration of software functionality and ensuring the smooth progress of tasks.
[0106] The target model is also used for program repair. For errors discovered during basic software testing, the large model can automatically generate repair suggestions or code patches based on error descriptions, source code context, and other information. This automatic or semi-automatic repair method is particularly important in special environments because it can quickly respond to and fix defects in software, ensuring the stability and reliability of the underlying software at critical moments. By combining natural language understanding and code generation capabilities, the large model can simulate human repair strategies, successfully fixing errors in multiple underlying software programs and providing solid support for immediate task operations.
[0107] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0108] Based on the same inventive concept, this application also provides a large model building apparatus for implementing the large model building method for basic software testing described above. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the large model building apparatus for basic software testing provided below can be found in the limitations of the large model building method for basic software testing described above, and will not be repeated here.
[0109] In one exemplary embodiment, such as Figure 4 As shown, a large model building apparatus 400 for basic software testing is provided. The apparatus 400 includes:
[0110] The acquisition module 401 is used to acquire the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network.
[0111] The model invocation module 402 is used to invoke different preset models to process the evaluation dataset and obtain multiple processing results. The preset models include ChatGLM model, LLaMA model, Baichuan model and Qwen model.
[0112] The scoring module 403 is used to score the corresponding preset model in multiple dimensions based on each processing result, and obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring.
[0113] The determination module 404 is used to determine the target model from multiple preset models based on the multi-dimensional scoring results. The target model is used to conduct tests on the basic software.
[0114] In one embodiment, the acquisition module 401 is further configured to construct an adversarial network, which includes a generator and a discriminator; use a random noise vector as input to the generator and output test data samples; use real data samples and test data samples as input to the discriminator;
[0115] Construct an adversarial loss function and train an adversarial network based on the adversarial loss function. Input a random noise vector into the trained adversarial network to obtain synthetic data samples. Construct an evaluation dataset based on the synthetic data samples and real data samples.
[0116] In one embodiment, the adversarial loss function is as follows:
[0117] ;
[0118] Where G is the generator; D is the discriminator; E is the expected value; x is the real data sample; z is the random noise vector; X r To extract a predetermined number of real data samples from the real data sample; P r (.) represents the probability distribution of the real data sample; P g (.) represents the probability distribution of the synthetic data sample; G(z) represents the synthetic data sample; D(x) represents the probability estimate that the discriminator confirms the real data sample is real; D(G(z)) represents the probability estimate that the discriminator confirms the synthetic data sample is real.
[0119] In one embodiment, the acquisition module 401 is further configured to perform iterative training on the adversarial network according to the adversarial loss function until the number of iterative training times reaches a preset number;
[0120] In each iteration of training, the parameters of the adversarial network are adjusted based on the loss value calculated by the adversarial loss function to train the adversarial network model until convergence. The real data samples and test data samples used in this iteration of training are used as historical data, and a preset proportion of historical data is added to the test data samples used in the next iteration of training.
[0121] In one embodiment, the scoring module 403 is further configured to score the performance dimensions of the corresponding preset model based on the performance evaluation metrics and processing results; the performance evaluation metrics include precision, recall and F1 score.
[0122] The corresponding preset models are scored in a fine-grained dimension based on the fine-grained evaluation indicators and processing results. The fine-grained evaluation indicators include correctness, completeness and semantic consistency.
[0123] Calculate the consistency index between the human scoring results and the fine-grained dimension scores, and then score the subjective dimension based on the consistency index. The consistency index is either Cohen's Kappa coefficient or Spearman correlation coefficient.
[0124] In one embodiment, the target model is also used to generate unit test cases and system test cases, both of which are used to test the underlying software.
[0125] The target model is also used for fuzz testing, error analysis and debugging of underlying software, as well as program repair.
[0126] The modules in the large-scale model building device for basic software testing described above can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0127] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores evaluation datasets and multi-dimensional scoring results. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a large-scale model building method for basic software testing.
[0128] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0129] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described large model building method for basic software testing.
[0130] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method for building large models for basic software testing.
[0131] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method for building large models for basic software testing.
[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0133] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0134] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0135] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing a large model for basic software testing, characterized in that, The method includes: Obtain the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network; The evaluation dataset is processed by calling different preset models to obtain multiple processing results. The preset models include ChatGLM model, LLaMA model, Baichuan model and Qwen model. Based on the processing results, the corresponding preset models are scored in multiple dimensions to obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring. A target model is determined from multiple preset models based on the multi-dimensional scoring results. The target model is used to test the basic software.
2. The method according to claim 1, characterized in that, The acquisition of the assessment dataset includes: The adversarial network is constructed, which includes a generator and a discriminator; a random noise vector is used as the input to the generator, and a test data sample is output; the real data sample and the test data sample are used as the input to the discriminator. Construct an adversarial loss function and train the adversarial network based on the adversarial loss function. Input the random noise vector into the trained adversarial network to obtain the synthetic data sample. Construct the evaluation dataset based on the synthetic data sample and the real data sample.
3. The method according to claim 2, characterized in that, The adversarial loss function is as follows: ; Wherein, G is the generator; D is the discriminator; E is the expected value; x is the real data sample; z is the random noise vector; and X... r To extract a predetermined number of real data samples from the real data samples; the P r (.) represents the probability distribution of the real data sample; the P g (.) represents the probability distribution of the synthetic data sample; G(z) represents the synthetic data sample; D(x) represents the probability estimate by which the discriminator confirms that the real data sample is real; D(G(z)) represents the probability estimate by which the discriminator confirms that the synthetic data sample is real.
4. The method according to claim 2, characterized in that, Training the adversarial network based on the adversarial loss function includes: The adversarial network is iteratively trained according to the adversarial loss function until the preset number of iterations is reached. In each iteration of training, the parameters of the adversarial network are adjusted according to the loss value calculated by the adversarial loss function to train the adversarial network model until convergence. The real data samples and test data samples used in this iteration of training are used as historical data, and a preset proportion of the historical data is added to the test data samples used in the next iteration of training.
5. The method according to claim 1, characterized in that, The method further includes: The corresponding preset model is scored according to the performance evaluation metrics and the processing results; the performance evaluation metrics include precision, recall and F1 score. The corresponding preset model is scored in fine-grained dimensions based on the fine-grained evaluation metrics and the processing results. The fine-grained evaluation metrics include correctness, completeness, and semantic consistency. Calculate the consistency index between the manual rating results and the fine-grained dimension ratings, and score the subjective dimension based on the consistency index, wherein the consistency index is the Cohen's Kappa coefficient or the Spearman correlation coefficient.
6. The method according to claim 1, characterized in that, The target model is also used to generate unit test cases and system test cases, both of which are used to test the basic software. The target model is also used for fuzz testing, error analysis and debugging of the underlying software, as well as program repair.
7. A large model building device for basic software testing, characterized in that, The device includes: The acquisition module is used to acquire the evaluation dataset, which includes real data samples and synthetic data samples generated by the adversarial network. The model invocation module is used to invoke different preset models to process the evaluation dataset and obtain multiple processing results. The preset models include ChatGLM model, LLaMA model, Baichuan model and Qwen model. The scoring module is used to perform multi-dimensional scoring on the corresponding preset models based on each of the processing results, and obtain multi-dimensional scoring results; the multi-dimensional scoring includes performance dimension scoring, fine-grained dimension scoring and subjective dimension scoring; The determination module is used to determine a target model from multiple preset models based on the multi-dimensional scoring results. The target model is used for testing the basic software.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Generative data enhancement method, system and equipment for vehicle body part quality model
CN117495712A
Method for determining target model and method and system for determining image category
CN118038104A
Software test data generation method and system
CN119166503A
Interface test method and device, storage medium and program product
CN120336182A