Dynamic test case sorting method and system based on hot area identification and defect prediction

Through a dynamic test case sorting method based on hot spot identification and defect prediction, multimodal data is collected in real time and the LSTM-GRU network is used to predict defect probability, which solves the problems of dynamic and accurate test case sorting in the software continuous integration environment and improves the efficiency and adaptability of defect detection.

CN120780609AActive Publication Date: 2025-10-14四川互慧软件有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511254442.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-14
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve dynamic and accurate test case sorting in a software continuous integration environment, resulting in low defect detection efficiency, serious waste of resources, and inability to meet quality requirements in high-reliability fields.

Method used

A dynamic test case sorting method based on hot spot identification and defect prediction is adopted. By collecting multimodal data in real time, a DefectBERT model is constructed, four-dimensional features are calculated, and the defect probability is predicted using the LSTM-GRU joint network. The weight parameters are dynamically adjusted to generate a test case priority queue.

Benefits of technology

It improves testing efficiency, increases defect detection rate, reduces false alarm rate, enhances dynamic adaptability, adapts to high-frequency code change scenarios, and meets the real-time requirements of continuous integration environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780609A_ABST
    Figure CN120780609A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of software testing, and relates to a test case dynamic sorting method and system based on hot area identification and defect prediction. The method comprises the steps of collecting multi-modal data in real time; a time decay sliding window algorithm is adopted to calculate the code change density, and hot area self-adaptive classification is carried out; a defect feature vector is output through a DefectBERT model, and the correlation degree between the test case and the defect is calculated; dynamically distributing the weight of each feature through an adaptive weight adjustment mechanism, and carrying out adaptive feature fusion; predicting a defect probability by adopting an LSTM-GRU combined network, and generating a test case priority queue; the weight parameters of the four-dimensional features are adjusted in real time through a PID weight controller. According to the method, the test efficiency is improved, the defect detection rate is increased, the false alarm rate is reduced, the dynamic adaptability is enhanced, a high-frequency code change scene can be adapted through hot area dynamic division and weight self-adaptive adjustment, and the real-time requirement of a continuous integration environment is met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of software testing, and specifically relates to a test case dynamic sequencing method and system based on hot area identification and defect prediction. BACKGROUND

[0002] In the software development continuous integration environment, the efficiency optimization of large-scale regression testing has become a key bottleneck to improve the development iteration speed, especially in the fields such as finance and automatic driving which have strict requirements on software quality. The execution order of test cases directly affects the defect detection efficiency and resource investment cost.

[0003] The current mainstream test case sequencing method has significant technical limitations, which is difficult to meet the dynamic and accurate needs in the continuous integration scene. The specific performance is as follows: (1) Static rule adaptability is insufficient: existing technologies mostly rely on fixed weight strategy or single dimension index for sequencing, which cannot adapt to the dynamic scene of high-frequency code changes in agile development. For example, the patent file "Data processing method, interface conversion structure and equipment" adopts a fixed weight allocation method. When the code change frequency and range change dramatically, the sequencing logic cannot be automatically adjusted, resulting in an increased risk of missing critical defects. Patent US20180075109A1 only uses code coverage as the basis for sequencing, ignoring key factors such as defect history and change impact range, limiting the sequencing accuracy; (2) Multi-source data utilization is insufficient: existing systems fail to establish a correlation analysis mechanism for code changes, defect history and call topology, resulting in insufficient information value mining. For example, IBM system US20200349072A1 only processes code submission data in isolation, without comprehensive analysis of text features in defect reports and function call link topology, making it difficult to accurately locate high-risk code areas; (3) Weak model capturing ability for complex features: traditional methods mostly use single models (such as LSTM model used in patent EP3564886), which cannot simultaneously consider the time sequence characteristics and structural features of code changes. Code changes contain both the dynamic rules of submission time sequence and the topology structure of function call relationship, and single models are difficult to fully capture these multi-modal features, resulting in low defect prediction accuracy.

[0004] In continuous integration scenarios, these flaws further exacerbate the core contradiction between limited testing resources and comprehensive defect detection. Traditional methods, executing all test cases, take excessively long times, significantly slowing down development iterations. Using random sampling to reduce test coverage significantly increases the rate of missed critical defects, failing to meet the quality requirements of high-reliability applications. Therefore, an intelligent ranking system integrating dynamic code feature recognition, multi-dimensional data association, and adaptive models is urgently needed to overcome existing technical bottlenecks and achieve efficient defect detection within limited resources. Summary of the Invention

[0005] In order to solve the problems of low test case sorting efficiency and insufficient defect detection rate in large-scale regression testing under a continuous integration environment, the present invention provides a test case dynamic sorting method and system based on hot spot identification and defect prediction.

[0006] In a first aspect, the present invention provides a method for dynamically sorting test cases based on hotspot identification and defect prediction, comprising: Real-time collection of multimodal data, including change logs from code repositories and defect reports from defect management systems; A time-decay sliding window algorithm is used to calculate code change density and perform adaptive classification of hot spots. Build and train the DefectBERT model. Use the DefectBERT model to output defect feature vectors for multimodal data after adaptive hotspot classification. Calculate the correlation between test cases and defects based on the defect feature vectors. Calculate four-dimensional features, dynamically assign weights to each feature through an adaptive weight adjustment mechanism, and perform adaptive feature fusion; the four-dimensional features include hot zone coverage, defect correlation, call path weight, and time decay coefficient; The LSTM-GRU joint network is used to predict defect probability, and the simulated annealing algorithm is used to generate the test case priority queue; According to the test case execution feedback, the weight parameters of the four-dimensional features are adjusted in real time through the PID weight controller.

[0007] In a second aspect, the present invention provides a test case dynamic sorting system based on hotspot identification and defect prediction, comprising a code repository, a defect management system, a real-time acquisition unit, a dynamic hotspot identification unit, a defect feature extraction unit, a feature fusion unit, a hybrid prediction unit, a test case priority queue generation unit, and a PID weight control unit; Code repository, used to generate change logs; Defect management system, used to generate defect reports for the defect management system; Real-time collection unit, used to collect multimodal data in real time, including change logs of the code repository and defect reports from the defect management system; The dynamic hotspot identification unit is used to calculate the code change density based on the code repository change log using a time-decay sliding window algorithm and perform adaptive hotspot classification. The defect feature extraction unit is used to build and train the DefectBERT model. The DefectBERT model outputs defect feature vectors based on the multimodal data after adaptive classification of hot spots. The correlation between the test case and the defect is calculated based on the defect feature vectors. The feature fusion unit is used to calculate four-dimensional features and dynamically assign weights to each feature through an adaptive weight adjustment mechanism to perform adaptive feature fusion. The four-dimensional features include hot zone coverage, defect correlation, call path weight, and time decay coefficient. Hybrid prediction unit, used to predict defect probability using a LSTM-GRU joint network; A test case priority queue generation unit, used for generating a test case priority queue using a simulated annealing algorithm; The PID weight control unit is used to adjust the weight parameters of the four-dimensional features in real time through the PID weight controller according to the test case execution feedback.

[0008] On the basis of the above technical solution, the present invention can also be improved as follows.

[0009] Furthermore, the change log of the code repository includes the modified file path, the number of changed code lines and the timestamp; the defect report includes the text content, defect status and the associated code version.

[0010] Furthermore, real-time collection of multimodal data also includes: cleaning the change log of the code repository, extracting valid fields, and performing text standardization on defect reports; valid fields include commit hash, timestamp, and number of changed lines; text standardization includes removing redundant symbols and unifying the format.

[0011] Furthermore, a time-decay sliding window algorithm is used to calculate the code change density and perform adaptive classification of hot zones, including: setting the sliding window size to the size of the code submitted the most recently set number of times; applying an exponential decay function to the number of changed lines submitted each time; calculating the weighted change density within the window; based on the mean and standard deviation of the change density, dynamically dividing the hot zones into three levels through thresholds, and adaptively classifying the hot zones.

[0012] Furthermore, feature weights are dynamically allocated through an adaptive weight adjustment mechanism to perform adaptive feature fusion, including: allocating initial weights, dynamically adjusting feature weights through PID controller, gradient descent and simulated annealing algorithms to perform adaptive feature fusion.

[0013] Furthermore, the DefectBERT model is used to adaptively classify the multimodal data of hot spots to output defect feature vectors, including: input defect report text, code change context, and call stack topology diagram; the domain-adaptive pre-trained DefectBERT model with a multi-head attention mechanism is used to output the CLS vector as the defect feature vector.

[0014] Furthermore, the cosine similarity algorithm is used to calculate the correlation between test cases and defects.

[0015] Furthermore, we build the DefectBERT model and train it, including: Hyperparameters are determined through iterative search using Bayesian optimization methods; hyperparameters include learning rate, batch size, simulated annealing parameters, and cooling rate; Calculate change density based on historical code change data, divide hot zones, and label samples; Use the DefectBERT model to pre-train multimodal defect data and output defect feature vectors; The four-dimensional features are integrated, the defect prediction model is trained through the LSTM-GRU joint network, and the cross entropy loss function is used to optimize the parameters.

[0016] Furthermore, the LSTM-GRU joint network includes an input layer, a bidirectional LSTM layer, a GRU layer, a fully connected layer and an output layer; the LSTM-GRU joint network is used to predict defect probabilities, and the simulated annealing algorithm is used to generate a test case priority queue, including: the bidirectional LSTM layer captures the timing features of code changes; the GRU layer learns to call topological structure features; the fully connected layer integrates the timing features of code changes with the topological structure features; the fully connected layer integrates the timing features of code changes with the topological structure features; and the output layer generates defect prediction probabilities through a sigmoid function.

[0017] The beneficial effects of the present invention are: the test efficiency of the present invention is improved, the defect detection rate is increased, the false alarm rate is reduced, and the dynamic adaptability is enhanced. Through dynamic division of hot zones and adaptive adjustment of weights, it can adapt to high-frequency code change scenarios and meet the real-time requirements of the continuous integration environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the test case dynamic sorting method based on hot spot identification and defect prediction provided by the present invention; Figure 2 The figure is a flowchart of a dynamic test case ranking method based on hotspot identification and defect prediction; Figure 3 Schematic diagram of the adaptive classification of hot zones; Figure 4 This is the structural diagram of the LSTM-GRU joint network; Figure 5 The principle block diagram of the test case dynamic sequencing system based on hot area identification and defect prediction provided for Embodiment 2 of the present application is shown in the figure. DETAILED DESCRIPTION

[0019] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0020] Embodiment 1 As an embodiment, as shown in the figure, to solve the above technical problems, the present embodiment provides a test case dynamic sequencing method based on hot area identification and defect prediction, comprising: Figure 1 Real-time collection of multi-modal data, including change logs of code repositories and defect reports of defect management systems; Adopting a time-decay sliding window algorithm to calculate the code change density and perform hot area adaptive classification; Building a DefectBERT model and training it, outputting a defect feature vector from the DefectBERT model for the multi-modal data after hot area adaptive classification, and calculating the correlation between test cases and defects according to the defect feature vector; Calculating four-dimensional features, dynamically allocating the weights of each feature through an adaptive weight adjustment mechanism, and performing adaptive feature fusion; the four-dimensional features include hot area coverage, defect correlation, call path weight and time decay coefficient; Adopting an LSTM-GRU joint network to predict defect probability, and using a simulated annealing algorithm to generate a test case priority queue; According to the test case execution feedback, a PID (Proportional-Integral-Derivative) weight controller is used to adjust the weight parameters of the four-dimensional features in real time.

[0021] As shown in the figure, the flowchart of the test case dynamic sequencing method based on hot area identification and defect prediction. Figure 2

[0022] The method adopts an exponential weighted moving average for three-level hot area classification, specifically: real-time monitoring of code repository change records, extracting change logs for each submission; dynamically calculating the change density using an exponential decay function; dynamically window classification to divide the hot area into three levels.

[0023] ​​Optionally, the code repository change log includes the modified file path, the number of changed code lines, and the timestamp; the defect report includes the text content, defect status, and the associated code version.

[0024] Specifically, the modified file path refers to the specific file location involved; the number of changed code lines refers to the number of code lines added, deleted, or modified in this submission; the timestamp refers to the specific time when the submission log occurs.

[0025] An exponential decay function is used to dynamically calculate the change density. Let the change density be decay, the current time be current_time, the time when the log is submitted be commit.time, exp be an exponential function, and the formula of the exponential decay function is decay=exp(-λ*(current_time-commit.time)), where λ=0.05 is a parameter determined after grid search optimization.

[0026] Optionally, real-time multimodal data collection also includes: cleaning the code repository change log, extracting valid fields, and performing text standardization on defect reports; valid fields include commit hash, timestamp, and number of changed lines; text standardization includes removing redundant symbols and unifying the format.

[0027] Optional, as attached Figure 3 As shown, a time-decay sliding window algorithm is used to calculate the code change density and perform hot zone adaptive classification, including: setting the sliding window size to the code size of the most recently set number of submissions; applying an exponential decay function to the number of changed lines for each submission; calculating the weighted change density within the window; based on the mean and standard deviation of the change density, dynamically dividing the hot zones into three levels through thresholds, and adaptively classifying the hot zones.

[0028] Adaptive window size is used. Generally, the default window contains 50 submissions, which can be adjusted according to actual conditions to ensure the timeliness and accuracy of classification.

[0029] Classification is based on the statistical characteristics of the data within the window, as follows: Calculate the mean and standard deviation of the weighted change values ​​within the window; let the mean be μ and the standard deviation be σ; Determine the thresholds for high, medium, and low heat zones; let the average value of the weighted change value in the window be mean, the threshold for the high heat zone be threshold_high, and the threshold for the low heat zone be threshold_low, where threshold_high = mean + 1.5*σ and threshold_low = mean - 1.5*σ; When the weighted change value is greater than or equal to threshold_high, it is classified into a high heat region; when the weighted change value is less than threshold_high and greater than or equal to threshold_low, it is classified into a medium heat region; and when the weighted change value is less than threshold_low, it is classified into a low heat region.

[0030] Optionally, the adaptive feature fusion is performed through a dynamic allocation of feature weights by an adaptive weight adjustment mechanism, including: allocating initial weights, dynamically adjusting the feature weights by a PID controller, gradient descent and simulated annealing algorithm, and performing adaptive feature fusion.

[0031] Optionally, the defect feature vector is output by the DefectBERT model for the multi-modal data classified by the hot area in an adaptive manner, including: inputting the defect report text, the code change context and the call stack topology graph; and outputting the CLS vector as the defect feature vector by the DefectBERT model with a multi-head attention mechanism pre-trained in an adaptive manner.

[0032] In the forward propagation method, the input data is transmitted into the BERT module for processing to obtain an output result pre-trained in an adaptive manner, so that the model can better adapt to the data features of a specific defect field; the output result is transmitted into the attention module in the model as a query, a key and a value for attention calculation to obtain an output result processed by the attention mechanism; through the attention mechanism, information more important to the defect features can be highlighted; the 0th vector 0 is extracted from the obtained output result, and this CLS vector is used as the final defect pattern representation result for subsequent correlation calculation.

[0033] In the multi-modal data fusion stage, three types of multi-modal data are collected and integrated, i.e., defect reports (text type), code change context (diff format) and call stack topology graph, to provide a comprehensive data basis for subsequent defect analysis.

[0034] Optionally, the cosine similarity algorithm is used to calculate the correlation between the test case and the defect.

[0035] The cosine similarity algorithm is used to calculate the matching degree between the test case and the historical defect to obtain the defect correlation, which can be used to measure the similarity between the two, and provide a basis for defect analysis and positioning. Compared with the traditional TF-IDF (Term Frequency-Inverse Document Frequency) method, the accuracy of the defect correlation calculation is significantly improved.

[0036] The weights of each feature are dynamically allocated through the adaptive weight adjustment mechanism to perform adaptive feature fusion. The weights of the hot zone coverage, defect correlation, call path weight, and time decay coefficient are set to α, β, γ, and δ, respectively, α+β+γ+δ=1. Optionally, the initial weights are set to α=0.4, β=0.3, γ=0.2, and δ=0.1, and are updated in real time based on test feedback.

[0037] Assume that the coverage of the hot zone is , is the total number of changed rows, is the index variable, Represents the total number of calculated hotspots, is the file weight, is the number of file change lines, the weight α is dynamically adjusted by the PID controller, and the hot zone coverage is calculated as follows: .

[0038] Assume the defect correlation is , is the test case vector, is the defect feature vector, and the cosine similarity is used to calculate the matching degree between the test case and the defect feature vector to obtain the defect correlation degree: .

[0039] Assume the call path weight is , is the node importance evaluation algorithm, is the call path depth, and the call path weight is calculated as: .

[0040] Assume that the time decay coefficient is , is the code change time interval, the time decay coefficient is calculated as: .

[0041] Optionally, build and train a DefectBERT model, including: Hyperparameters are determined through iterative search using Bayesian optimization methods; hyperparameters include learning rate, batch size, simulated annealing parameters, and cooling rate; Calculate change density based on historical code change data, divide hot zones, and label samples; Use the DefectBERT model to pre-train multimodal defect data and output defect feature vectors; The four-dimensional features are integrated, the defect prediction model is trained through the LSTM-GRU joint network, and the cross entropy loss function is used to optimize the parameters.

[0042] Optional, as attachedFigure 4 As shown in the figure, the LSTM-GRU joint network consists of an input layer, a bidirectional LSTM (Long Short-Term Memory) layer, a GRU (Gated Recurrent Unit) layer, a fully connected layer, and an output layer. This network predicts defect probabilities and generates a test case priority queue using a simulated annealing algorithm. The network includes: a bidirectional LSTM layer captures the temporal features of code changes; a GRU layer learns call topology features; a fully connected layer integrates the temporal and topological features of code changes; and an output layer generates defect prediction probabilities using a sigmoid function. The temporal features output by the LSTM and the structural features output by the GRU are processed uniformly to generate defect probabilities. The fully connected layer implements nonlinear combinations through a weight matrix, enhancing the model's expressiveness.

[0043] During the model initialization phase, a hybrid neural network is created, the bidirectional long short-term memory layer is initialized, and the input and hidden layer dimensions are set to capture temporal features. The gated recurrent unit layer is initialized, and the input and hidden layer dimensions are set to extract structural association features. During the model forward propagation phase, the input data is passed to the bidirectional long short-term memory layer for processing, resulting in an output that captures temporal features in the data. The gated recurrent unit processes the output of the bidirectional long short-term memory layer and extracts structural association features in the data. The prediction result is output by processing the features of the last time step of the gated recurrent unit input with a fully connected layer and then activating it with a sigmoid function to obtain the final prediction result, which represents the probability value of the predicted target. In other words, the LSTM layer captures the temporal pattern of code changes, and the GRU layer learns and calls topological structural features.

[0044] Practical application examples show that the test efficiency of the present invention is improved, and the test time is shortened from 6.8 hours of the traditional method to 2.1 hours, an increase of 69.1%; the defect detection rate is improved: the defect detection rate is increased from 82% to 97%, an increase of 15 percentage points; the false alarm rate is reduced: the false alarm rate is reduced from 23% to 8%, a decrease of 65.2%; dynamic adaptability is enhanced: through dynamic division of hot zones and adaptive adjustment of weights, it can adapt to high-frequency code change scenarios and meet the real-time requirements of the continuous integration environment.

[0045] Example 2 Based on the same principle as the method shown in Example 1 of the present invention, as shown in the attached Figure 5As shown, an embodiment of the present invention further provides a test case dynamic sorting system based on hot zone identification and defect prediction, which includes a code repository, a defect management system, a real-time acquisition unit, a dynamic hot zone identification unit, a defect feature extraction unit, a feature fusion unit, a hybrid prediction unit, a test case priority queue generation unit, and a PID weight control unit; Code repository, used to generate change logs; Defect management system, used to generate defect reports for the defect management system; Real-time collection unit, used to collect multimodal data in real time, including change logs of the code repository and defect reports from the defect management system; The dynamic hotspot identification unit is used to calculate the code change density based on the code repository change log using a time-decay sliding window algorithm and perform adaptive hotspot classification. The defect feature extraction unit is used to build and train the DefectBERT model. The DefectBERT model outputs defect feature vectors based on the multimodal data after adaptive classification of hot spots. The correlation between the test case and the defect is calculated based on the defect feature vectors. The feature fusion unit is used to calculate four-dimensional features and dynamically assign weights to each feature through an adaptive weight adjustment mechanism to perform adaptive feature fusion. The four-dimensional features include hot zone coverage, defect correlation, call path weight, and time decay coefficient. Hybrid prediction unit, used to predict defect probability using LSTM-GRU joint network; A test case priority queue generation unit, used for generating a test case priority queue using a simulated annealing algorithm; The PID weight control unit is used to adjust the weight parameters of the four-dimensional features in real time through the PID weight controller according to the test case execution feedback.

[0046] Optionally, a time-decay sliding window algorithm is used to calculate the code change density and perform adaptive classification of hot zones, including: setting the sliding window size to the size of the code submitted the most recently set number of times; applying an exponential decay function to the number of changed lines submitted each time; calculating the weighted change density within the window; based on the mean and standard deviation of the change density, dynamically dividing the hot zones into three levels through thresholds, and adaptively classifying the hot zones.

[0047] Optionally, feature weights are dynamically assigned through an adaptive weight adjustment mechanism to perform adaptive feature fusion, including: assigning initial weights, dynamically adjusting feature weights through a PID controller, gradient descent, and simulated annealing algorithms to perform adaptive feature fusion.

[0048] Optionally, the DefectBERT model is used to adaptively classify the multimodal data of hotspots to output a defect feature vector, including: input defect report text, code change context, and call stack topology; the domain-adaptive pre-trained DefectBERT model with a multi-head attention mechanism is used to output a CLS vector as the defect feature vector.

[0049] Optionally, a cosine similarity algorithm is used to calculate the correlation between the test cases and the defects.

[0050] Optionally, build and train a DefectBERT model, including: Hyperparameters are determined through iterative search using Bayesian optimization methods; hyperparameters include learning rate, batch size, simulated annealing parameters, and cooling rate; Calculate change density based on historical code change data, divide hot zones, and label samples; Use the DefectBERT model to pre-train multimodal defect data and output defect feature vectors; The four-dimensional features are integrated, the defect prediction model is trained through the LSTM-GRU joint network, and the cross entropy loss function is used to optimize the parameters.

[0051] Optionally, an LSTM-GRU joint network includes an input layer, a bidirectional LSTM layer, a GRU layer, a fully connected layer, and an output layer. The LSTM-GRU joint network is used to predict defect probabilities, and a simulated annealing algorithm is used to generate a test case priority queue, including: a bidirectional LSTM layer to capture code change timing features; a GRU layer to learn and call topological structure features; a fully connected layer to integrate code change timing features and topological structure features; and an output layer to generate defect prediction probabilities through a sigmoid function.

[0052] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A dynamic test case sorting method based on hotspot identification and defect prediction, characterized in that: include: Real-time collection of multimodal data, including change logs from code repositories and defect reports from defect management systems; A time-decay sliding window algorithm is used to calculate code change density and perform adaptive classification of hot spots. Build and train the DefectBERT model. Use the DefectBERT model to output defect feature vectors for multimodal data after adaptive hotspot classification. Calculate the correlation between test cases and defects based on the defect feature vectors. Calculate four-dimensional features, dynamically assign weights to each feature through an adaptive weight adjustment mechanism, and perform adaptive feature fusion; The four-dimensional features include hot zone coverage, defect correlation, call path weight and time decay coefficient; The LSTM-GRU joint network is used to predict defect probability, and the simulated annealing algorithm is used to generate the test case priority queue; According to the test case execution feedback, the weight parameters of the four-dimensional features are adjusted in real time through the PID weight controller.

2. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: The change log of the code repository includes the modified file path, the number of changed code lines and the timestamp; the defect report includes the text content, defect status and the associated code version.

3. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: Real-time multimodal data collection also includes: cleaning the change log of the code repository, extracting valid fields, and performing text standardization on defect reports; valid fields include commit hash, timestamp, and number of changed lines; text standardization includes removing redundant symbols and unifying the format.

4. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: A time-decay sliding window algorithm is used to calculate the code change density and perform adaptive classification of hot zones, including: setting the sliding window size to the size of the code submitted the most recently set number of times; applying an exponential decay function to the number of changed lines submitted each time; calculating the weighted change density within the window; based on the mean and standard deviation of the change density, dynamically dividing the hot zones into three levels through thresholds and adaptively classifying the hot zones.

5. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: The adaptive weight adjustment mechanism is used to dynamically allocate feature weights and perform adaptive feature fusion, including: allocating initial weights, dynamically adjusting feature weights through PID controller, gradient descent and simulated annealing algorithm, and performing adaptive feature fusion.

6. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: Build the DefectBERT model and train it, including: Hyperparameters are determined through iterative search using Bayesian optimization methods; hyperparameters include learning rate, batch size, simulated annealing parameters, and cooling rate; Calculate change density based on historical code change data, divide hot zones, and label samples; Use the DefectBERT model to pre-train multimodal defect data and output defect feature vectors; The four-dimensional features are integrated, the defect prediction model is trained through the LSTM-GRU joint network, and the cross entropy loss function is used to optimize the parameters.

7. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: The DefectBERT model outputs defect feature vectors from multimodal data after adaptive hotspot classification, including: input defect report text, code change context, and call stack topology diagram; the domain-adaptive pre-trained DefectBERT model with a multi-head attention mechanism outputs a CLS vector as the defect feature vector.

8. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1 is characterized in that: The cosine similarity algorithm is used to calculate the correlation between test cases and defects.

9. The method for dynamic sorting of test cases based on hotspot identification and defect prediction according to claim 1, characterized in that: The LSTM-GRU joint network includes an input layer, a bidirectional LSTM layer, a GRU layer, a fully connected layer, and an output layer. The LSTM-GRU joint network is used to predict defect probabilities and a simulated annealing algorithm is used to generate a test case priority queue. The network includes: a bidirectional LSTM layer that captures the timing features of code changes; a GRU layer that learns to call topological structure features; a fully connected layer that integrates the timing features of code changes with the topological structure features; and an output layer that generates defect prediction probabilities using a sigmoid function.

10. A dynamic test case ranking system based on hotspot identification and defect prediction, characterized by: It includes code repository, defect management system, real-time acquisition unit, dynamic hotspot identification unit, defect feature extraction unit, feature fusion unit, hybrid prediction unit, test case priority queue generation unit and PID weight control unit; Code repository, used to generate change logs; Defect management system, used to generate defect reports for the defect management system; Real-time collection unit, used to collect multimodal data in real time, including change logs of the code repository and defect reports from the defect management system; The dynamic hotspot identification unit is used to calculate the code change density based on the code repository change log using a time-decay sliding window algorithm and perform adaptive hotspot classification. The defect feature extraction unit is used to build and train the DefectBERT model. The DefectBERT model outputs defect feature vectors based on the multimodal data after adaptive classification of hot spots. The correlation between the test case and the defect is calculated based on the defect feature vectors. The feature fusion unit is used to calculate the four-dimensional features and dynamically assign the weights of each feature through the adaptive weight adjustment mechanism to perform adaptive feature fusion; The four-dimensional features include hot zone coverage, defect correlation, call path weight and time decay coefficient; Hybrid prediction unit, used to predict defect probability using LSTM-GRU joint network; A test case priority queue generation unit, used for generating a test case priority queue using a simulated annealing algorithm; The PID weight control unit is used to adjust the weight parameters of the four-dimensional features in real time through the PID weight controller according to the test case execution feedback.

Citation Information

Patent Citations

  • Spatial change detector and check and set operation

    US20180075109A1

  • Handling metadata corruption to avoid data unavailability

    US20200349072A1

  • Protocol conformance test case priority ordering method based on risk analysis

    CN108446231A

  • Regression test case determination method and device, computer equipment and storage medium

    CN110109820A

  • Priority ranking method and device for large-scale continuously-integrated test cases and medium

    CN115470133A