Test risk identification method and system based on artificial intelligence

By introducing an adaptive learning rate and feedback mechanism into the test risk identification method, the bias of machine learning models to successful use cases is solved, the sensitivity to failed use cases is improved, and more efficient and accurate test risk identification and defect discovery is achieved.

CN119988240AActive Publication Date: 2025-05-13SHANDONG YUNZE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510473834.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing artificial intelligence-based testing risk identification methods have caused machine learning models to predict successful use cases due to data imbalance, ignore failed use cases, and thus may miss key defects and lead to wasting test resources and time.

Method used

By collecting test data, extracting the dependency outliers and network delay high-frequency values ​​between code modules, building a data prediction model, and determining whether the prediction results of the machine learning model are biased. When bias occurs, an adaptive learning rate and feedback mechanism are introduced to gradually reduce the model's bias towards successful use cases and improve the sensitivity to failed use cases.

Benefits of technology

Accurately evaluate the deviation of predicted results of machine learning models, dynamically adjust model parameters, reduce the possibility of high-risk areas being ignored, improve testing efficiency, reduce testing costs, and ensure timely discovery and handling of key defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988240A_ABST
    Figure CN119988240A_ABST
Patent Text Reader

Abstract

The invention discloses a test risk identification method and system based on artificial intelligence, and particularly relates to the technical field of data processing. By introducing an adaptive learning rate and a feedback mechanism, the improved test risk identification method is provided for the problem of data imbalance existing in the test process of an existing machine learning model, and the test risk identification method is provided by comprehensively collecting test data including test case execution results and test environment states. Extracting a dependency relationship abnormal value and a network delay high-frequency value between code modules, constructing a data prediction model, and comprehensively calculating a deviation of a prediction result based on the data; by dynamically adjusting the deviation of the model to the successful use case, the identification capability of the failed use case and the high-risk area is enhanced, and the high-risk area is effectively prevented from being underestimated or the normal use case is misjudged to be high in risk, so that the prediction precision is improved, the configuration of test resources is optimized, the invalid test reexamination is reduced, the test efficiency is remarkably improved, and the cost and the period are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a test risk identification method and system based on artificial intelligence. Background Art

[0002] AI-based test risk identification refers to the use of artificial intelligence technology, especially machine learning and data analysis algorithms, to automatically identify possible risks in the software testing process. These risks may include incomplete test coverage, unstable test environment, defects that may be overlooked, and other issues. Specifically, AI-based test risk identification systems usually rely on big data analysis, pattern recognition, and predictive models to monitor and evaluate the test process in real time. The AI ​​model can automatically identify which test cases may have a high risk of execution failure and which areas may be ignored due to insufficient resources or limited testing time, thereby providing decision support for the test team. This can not only effectively save time and resources, but also prevent software defects and quality issues at an earlier stage.

[0003] The prior art has the following deficiencies: In the prior art, data is trained through machine learning algorithms (such as decision trees, support vector machines, deep learning, etc.) to identify common failure modes and potential high-risk areas in the testing process. However, in the test data, the number of test cases that are executed normally is much larger than the number of failed test cases. This data imbalance will cause the machine learning model to tend to predict successful test cases and ignore the few failed test cases. This bias may cause high-risk test areas to be underestimated, thereby missing critical defects. In addition, because the machine learning model is too biased towards predicting successful cases, it may also make incorrect failure predictions for some successful test cases. As a result, originally normal test cases are misjudged as high-risk, thereby wasting test resources and time on unnecessary test reviews. This misjudgment may lead to inefficiency in the entire testing process and increase testing costs and cycles. Summary of the invention

[0004] The purpose of the present invention is to provide a test risk identification method and system based on artificial intelligence to solve the shortcomings of the background technology.

[0005] In order to achieve the above object, the present invention provides the following technical solution: a test risk identification method based on artificial intelligence, comprising: Collect comprehensive test data, including test case execution results and test environment status; Extract dependency anomalies between code modules from test case execution results, and extract high-frequency network latency values ​​from test environment status; Build a data prediction model, input the dependency outliers and high-frequency network latency values into the data prediction model for comprehensive calculation, and judge whether there is a deviation in the prediction result of the machine learning model according to the calculation result; When a prediction deviation occurs, for failed test cases and high-risk test areas, by introducing an adaptive learning rate and a feedback mechanism, gradually reduce the bias of the machine learning model towards successful test cases, improve the sensitivity to failed test cases, and feedback the adjusted prediction result to the test team.

[0006] Preferably, extract the dependency outliers between code modules from the execution results of test cases, and extract the high-frequency network latency values from the test environment status, including: Use a directed graph to represent code modules and their dependencies. Nodes represent code modules, and edges represent dependencies between modules. Calculate the current shortest distance from the source node Vi to node u and mark it as dist[u]. Let S represent the set of nodes that have been processed, indicating that the shortest paths from the source node Vi to these nodes have been found. Let Q represent the set of unprocessed nodes, indicating the set of nodes closest to the source node. Set the weight of edge eij as w(eij), representing the dependency strength from Vi to Vj. Set the shortest distance of each node dist[Vi]=0, indicating that the distance of the source node is 0. For node Vj, initialize dist[Vj]=∞. Initialize S and Q. Select a node u with the shortest distance in Q, move node u to the processed set S, and remove u from Q. For all adjacent nodes v of node u, check if there is a shorter path: If dist[u]+w(euv)<dist[v], then update dist[v]=dist[u]+w(euv), where w(euv) is the weight of edge euv, representing the dependency strength between code modules Vu and Vv. Set the dependency strength calculation formula: ; where: represents the number of function calls between modules, represents the code coupling degree, reflects the data interaction frequency. c, β, and γ are weight coefficients. Update the shortest path of v until there are no nodes in Q. After calculating the shortest path of each module through the Dijkstra algorithm, calculate the dependency outlier, and the expression is: ; In the formula, AWD is the dependency outlier, which is the shortest path from Vi to Vj. mean(dist) is the average of the shortest paths between all modules, and std(dist) is the standard deviation of the shortest paths between all modules.

[0007] Preferably, the high-frequency value of network delay is extracted from the test environment state. The method for obtaining the high-frequency value of network delay is: collecting network delay data within a period of time, converting the time domain signal of network delay to the frequency domain through fast Fourier transform, and obtaining the amplitude spectrum by calculating the complex amplitude of Fourier transform: ; Where: A(k) is the amplitude of the kth frequency component, Re(F(k)) and Im(F(k)) represent the real and imaginary parts of the Fourier transform result respectively, set a frequency threshold, and regard the components with frequencies greater than the threshold as high-frequency components. The calculation expression for extracting the amplitude of all high-frequency parts is: ; It is the sum of the amplitudes of all high-frequency components. The calculated sum of the amplitudes of all high-frequency components is used as the high-frequency value of network delay.

[0008] Preferably, a data prediction model is constructed, and the dependency outliers and high-frequency values ​​of network delay are input into the data prediction model for comprehensive calculation, specifically: the dependency outliers and high-frequency values ​​of network delay are normalized so that they are both between [0,1], and the prediction result deviation index of the machine learning model is calculated based on the normalized dependency outliers and high-frequency values ​​of network delay.

[0009] Preferably, the prediction result deviation index of the obtained machine learning model is compared with a pre-set standard threshold. If the prediction result deviation index of the machine learning model is greater than or equal to the pre-set standard threshold, it means that the prediction result of the machine learning model is deviated, and a warning signal is generated at this time; if the prediction result deviation index of the machine learning model is less than the pre-set standard threshold, it means that the prediction result of the machine learning model is not deviated, and no warning signal is generated at this time.

[0010] Preferably, when there is a deviation in the prediction result, the weight of the machine learning model is adjusted by introducing an adaptive learning rate and a feedback mechanism, specifically: Set the learning rate of the current model to η, the strength of the feedback signal to F, the prediction result deviation index to PDI, and update the learning rate. The expression is: ;in: is the updated learning rate, is the current learning rate, γ is the adjustment coefficient of the feedback mechanism, which controls the amplitude of the learning rate change, and F is the strength of the feedback signal; According to the feedback signal, the weights in the training data are dynamically adjusted to reduce the weight of successful cases. The prediction weight of the current model for each case is set to , where represents the weight of case i, and the feedback signal of the deviation is F. The weight of each case is adjusted: ;in: is the updated weight of use case i, and v is the adjustment coefficient, which determines the influence of deviation feedback on weight adjustment.

[0011] The present invention also provides a test risk identification system based on artificial intelligence, including a data acquisition module, a data extraction module, a data prediction module and a test adjustment module; Data acquisition module: collects comprehensive test data, including test case execution results and test environment status; Data extraction module: extracts abnormal values ​​of dependencies between code modules from the test case execution results, and extracts high-frequency values ​​of network delay from the test environment status; Data prediction module: Build a data prediction model, input dependency anomalies and high-frequency values ​​of network delay into the data prediction model for comprehensive calculation, and judge whether the prediction results of the machine learning model are biased based on the calculation results; Test adjustment module: When prediction deviation occurs, for failed cases and high-risk test areas, by introducing adaptive learning rate and feedback mechanism, the machine learning model's bias towards successful cases is gradually reduced, the sensitivity to failed cases is increased, and the adjusted prediction results are fed back to the test team.

[0012] In the above technical solution, the technical effects and advantages provided by the present invention are: 1. The present invention collects comprehensive test data, extracts dependency anomalies between code modules and high-frequency values ​​of network delay, and inputs them into the data prediction model for comprehensive calculation. The present invention can accurately evaluate whether there is a deviation in the prediction results of the machine learning model. When a deviation is found, the adaptive learning rate and feedback mechanism are used to adjust the model's bias towards successful cases and enhance the model's sensitivity to failed cases, thereby reducing the possibility of high-risk areas being ignored, improving test efficiency and reducing test costs.

[0013] 2. The present invention avoids the phenomenon of normal use cases being misjudged as high-risk during the test process by dynamically adjusting the learning process of the model and optimizing the allocation of test resources, reducing invalid reviews and waste of resources. At the same time, the optimized prediction results can provide more accurate decision support for the test team, making the test work more efficient and accurate, thereby significantly improving the quality of software testing, shortening the test cycle, and reducing the overall test cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0015] Figure 1 The figure is a flow chart of the method of the present invention.

[0016] Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] Example 1, please refer to Figure 1 As shown, the test risk identification method based on artificial intelligence described in this embodiment includes: Collect comprehensive test data, including test case execution results and test environment status; Extract dependency anomalies between code modules from test case execution results, and extract high-frequency network latency values ​​from test environment status; Build a data prediction model, input dependency anomalies and high-frequency network delay values ​​into the data prediction model for comprehensive calculation, and determine whether the prediction results of the machine learning model are biased based on the calculation results; When prediction deviation occurs, for failed cases and high-risk test areas, by introducing adaptive learning rate and feedback mechanism, the bias of machine learning model towards successful cases is gradually reduced, the sensitivity to failed cases is increased, and the adjusted prediction results are fed back to the test team.

[0019] Collecting comprehensive test data is a key step in achieving AI-based test risk identification. This process involves collecting and organizing all test-related data to ensure the comprehensiveness and multi-dimensionality of the data, thereby providing sufficient information basis for subsequent analysis and prediction. Specifically, the collected data can be divided into the following two categories: Test case execution result data: Test case status: Record the execution result of each test case, including whether it passed, failed, and the reason for failure. For example, for functional test cases, it may include "normal function" or "abnormal function"; for performance testing, it may include "passed performance benchmark" or "failed to meet the standard". Execution time and duration: The execution timestamp and duration of each test case. By analyzing this data, it can be determined whether some test cases failed to complete as expected due to excessive resource consumption or long execution time. Error type and error message: Detailed record of error information, exception stack, error code segment or related logs when the test case fails. The error message provides potential clues to failure modes for subsequent analysis. Coverage data: Use coverage tools to record which code paths or functional modules are covered by test cases and which are not. This helps identify potential test blind spots, especially high-risk areas. Input data and expected output: Record the input conditions and expected output of the test case. When the actual output is inconsistent with the expected output, it can help analyze the cause of the problem. Retry times and repair history: If a test case fails and is retried, record the number of failed retries, whether a repair was performed after the failure, and the results of the repair. These data can be used to measure the stability of certain test cases.

[0020] Test environment status data: Hardware resource usage: Record the usage of hardware resources during the test, such as the CPU, memory, disk space and other resource usage. These data can help analyze whether resource bottlenecks affect the execution results of test cases. Network delay and bandwidth: For tests that rely on the network (such as distributed systems, cloud services, etc.), record network indicators such as network delay, bandwidth, packet loss rate, etc. during the test to evaluate the impact of network performance on test results. Operating system and dependent environment: Record environmental data such as the operating system version, dependent library version, and framework version during test execution. Different versions of operating systems or dependent libraries may lead to inconsistent test results. Environmental stability data: Records of events such as system crashes, restarts, and failures, which may cause test cases to fail to execute smoothly or inaccurate results. Concurrency and load: When performing performance or stress testing, record data such as the number of concurrent users, request rate, and response time to evaluate the performance of the system under different loads. External system and interface status: If the test relies on external systems or interfaces (such as databases, third-party APIs), recording the status of external systems (such as response time, whether the service is available, etc.) will help analyze whether the test failure is related to external factors.

[0021] Data collection methods and tools: Automated testing tools: Use automated testing tools (such as Selenium, JUnit, TestNG) to record the test case execution status, logs, and execution time. The tool can automatically collect detailed test data and reduce human intervention. Monitoring tools: Use performance monitoring tools (such as New Relic, Prometheus, Nagios) to collect real-time resource usage, network performance, operating system status, etc. of the test environment. Log management system: Use centralized log management tools (such as ELK Stack, Splunk) to collect detailed test environment and application logs, and extract key information from them during the test execution process. Version control system: By combining with the version control system (such as Git), records of each code submission, change, repair, and rollback are recorded to analyze the impact of code changes on test results.

[0022] Integrate data from different sources into a unified database or data warehouse. Ensure that data formats are unified, timestamps are consistent, and data is complete. For multi-dimensional data, appropriate association and classification are required. For example, the test case execution results and the corresponding test environment status data need to be associated by time or test ID to ensure that the relationship between the environment status and the execution results can be compared during analysis. Deduplication and cleanup: Remove redundant or duplicate test records to ensure data uniqueness and accuracy. Handle missing data through interpolation, prediction, or other supplementary methods, and use anomaly detection algorithms to identify and handle outliers to ensure data reliability.

[0023] Extract dependency anomalies between code modules from the test case execution results, and extract high-frequency network latency values ​​from the test environment status, including: Model code modules and their dependencies as a directed graph. In this graph: nodes represent code modules. Edges represent dependencies between modules. The weight of the edge represents the strength of the dependency or some measure (for example: execution time, call frequency, resource consumption, coupling between modules, etc.). In the dependency graph, calculate the shortest path from a module (node) to other modules (nodes). The shortest path may mean stronger dependencies or resource consumption.

[0024] Calculate the current shortest distance from the source node Vi to the node u, denoted as dist[u]. Let S represent the set of nodes that have been processed, indicating that the shortest paths from the source node Vi to these nodes have been found. Let Q represent the set of unprocessed nodes, which is the set of nodes closest to the source node. Set the weight of the edge eij as w(eij), representing the strength of the dependency from Vi to Vj. Set the shortest distance of each node dist[Vi]=0, indicating that the distance of the source node is 0. For other nodes Vj, initialize dist[Vj]=∞, indicating that the distance is unknown; initialize S and Q. Select a node u with the shortest distance in Q, move the node u to the processed set S, and remove u from Q. For all adjacent nodes v of the node u, check if there is a shorter path: If dist[u]+w(euv)<dist[v], then update dist[v]=dist[u]+w(euv), where w(euv) is the weight of the edge euv, and w(euv) represents the dependency strength between the code modules Vu and Vv. The dependency strength can be measured based on the number of function calls, code coupling degree, and data interaction frequency. Set the formula for calculating the dependency strength: ; where: represents the number of function calls between modules, (such as using static analysis to calculate dependencies) represents the code coupling degree, reflects the data interaction frequency, and c, β, γ are weight coefficients that can be adjusted according to the actual situation. This means that if the shortest path from the source node Vi to u plus the dependency weight from u to v is less than the current shortest path of v, then update the shortest path of v until there are no nodes in Q and the shortest paths of all nodes have been calculated. After calculating the shortest paths of each module through the Dijkstra algorithm, calculate the dependency relationship outlier, and the expression is: ; In the formula, AWD is the dependency relationship outlier, which is the shortest path from Vi to Vj, mean(dist) is the average value of the shortest paths between all modules, and std(dist) is the standard deviation of the shortest paths between all modules. AWD is used to measure the abnormal degree of the path length of a certain module pair compared to the overall distribution, similar to the standardized Z-score. If AWD is greater than 0, it indicates that the dependency relationship between the modules is relatively far, and there may be abnormal dependencies. If AWD is less than 0, it indicates that the dependency between the modules is too tight, which may lead to high coupling risks.

[0025] The core of the Dijkstra algorithm is based on the greedy selection of the shortest path, that is: if the length of the shortest path from the source node Vi to the current node Vu is dist[u], now consider passing through the edge euv from Vu to reach the adjacent node Vv. If dist[u]+w(euv)<dist[v], it means that reaching Vv through this path of Vu is shorter than the previously recorded path. Therefore, dist[v] is updated.

[0026] If the outlier value of a certain dependency is relatively high, it indicates that the strength or complexity of this dependency is too large, which may be a potential risk area. When the calculated outlier value exceeds a certain threshold (for example, more than 2 standard deviations), this dependency can be marked as "abnormal". This may indicate that the coupling between modules is too complex, or there are potential errors or bottlenecks.

[0027] Extract the high-frequency value of network latency from the test environment status. The method for obtaining the high-frequency value of network latency is as follows: Collect network latency data over a period of time. Network latency data is usually continuous and can be obtained by periodically measuring the network request response time.

[0028] Convert the time-domain signal of network latency to the frequency domain through the Fast Fourier Transform (FFT). The calculation formula: ; where: F(k) is the kth frequency component in the frequency-domain data, is the nth data point in the time-domain data, and N is the length of the data (i.e., the number of network latency data points). k is the frequency index, representing different frequency components. is the kernel function of the Fast Fourier Transform; by calculating the complex amplitude of the Fourier transform, the amplitude spectrum is obtained: ; where: A(k) is the amplitude of the kth frequency component. Re(F(k)) and Im(F(k)) respectively represent the real part and the imaginary part of the Fourier transform result. The amplitude spectrum A(k) describes the intensity of each frequency component. The larger the amplitude of the frequency component, the greater the proportion of this frequency component in the signal.

[0029] High-frequency components are usually related to the rapid fluctuations or instabilities of network latency. In the frequency-domain signal, high-frequency components correspond to higher frequency values. In order to extract high-frequency components, a frequency threshold needs to be set , and the components with frequencies greater than this threshold are regarded as high-frequency components. The calculation expression for extracting the amplitudes of all high-frequency parts is: ; is the sum of the amplitudes of all high-frequency components, representing the total intensity of high-frequency fluctuations in the network latency signal. The sum of the amplitudes of all calculated high-frequency components is used as the high-frequency value of network latency.

[0030] If the amplitude of the high-frequency component is large, it may indicate that the network is experiencing rapid fluctuations or instability, which may be a manifestation of network congestion, packet loss, bandwidth fluctuations, etc. High-frequency fluctuations may be related to certain system loads or external factors. Analyzing these high-frequency components can help locate potential performance bottlenecks or problems.

[0031] Build a data prediction model, input dependency anomalies and high-frequency network delay values ​​into the data prediction model for comprehensive calculation, and determine whether the prediction results of the machine learning model are biased based on the calculation results. Specifically: The dependency outliers and network delay high-frequency values ​​are normalized so that they are both between [0, 1]. The prediction result deviation index of the machine learning model is calculated based on the normalized dependency outliers and network delay high-frequency values.

[0032] For example, the present invention can use the following formula to calculate the prediction result deviation index of the machine learning model, and the calculation expression is: ; Wherein, is the prediction result deviation index of the machine learning model, AWD is the dependency anomaly value, is the network delay high-frequency value, and is the weight coefficient of the dependency anomaly value and the network delay high-frequency value (which can be optimized based on experimental experience or machine learning), and all are greater than 0.

[0033] The obtained prediction result deviation index of the machine learning model is compared with the pre-set standard threshold. If the prediction result deviation index of the machine learning model is greater than or equal to the pre-set standard threshold, it means that the prediction result of the machine learning model is deviated, and a warning signal is generated at this time; if the prediction result deviation index of the machine learning model is less than the pre-set standard threshold, it means that the prediction result of the machine learning model is not deviated, and no warning signal is generated at this time.

[0034] When there is a bias in the prediction results, the weights of the machine learning model can be adjusted by introducing an adaptive learning rate and feedback mechanism to reduce the bias towards successful cases, thereby increasing the sensitivity to failed cases and high-risk test areas. This mechanism can gradually adjust the parameters of the model so that it can better identify potential high-risk test cases.

[0035] Adaptive learning rate is a strategy that dynamically adjusts the learning rate based on the current prediction deviation. The learning rate determines the adjustment range of the model parameters at each update. Too large a learning rate may cause the model to be unstable, while too small a learning rate will lead to slow convergence. The adaptive learning rate can dynamically adjust the model update rate according to the strength of the feedback signal, thereby improving the model's responsiveness to deviations.

[0036] Adaptive learning rate update formula: Set the learning rate of the current model as η, the intensity of the feedback signal as F, and the prediction result deviation index as PDI. Update the learning rate, and the expression is: ; where: is the updated learning rate, is the current learning rate. γ is the adjustment coefficient of the feedback mechanism, which controls the amplitude of the change in the learning rate. Usually, 0 < γ < 1. F is the intensity of the feedback signal, which reflects the degree of prediction deviation. If the deviation is large, the F value is high, and the adjustment amplitude of the learning rate increases; if the deviation is small, the F value is low, and the adjustment amplitude of the learning rate decreases.

[0037] The bias of the machine learning model may lead to over-prediction of successful use cases while ignoring failed use cases or high-risk areas. When there is a prediction deviation, it is necessary to gradually reduce the bias of the model towards successful use cases and make it pay more attention to failed use cases and high-risk areas. To achieve this goal, the weights in the training data can be dynamically adjusted according to the feedback signal to reduce the weight of successful use cases. Set the prediction weight of the current model for each use case as, where represents the weight of use case i, and the feedback signal of the deviation is F. Adjust the weight of each use case: ; where: is the updated weight of use case i, v is the adjustment coefficient, which determines the influence degree of the deviation feedback on the weight adjustment. Usually, 0 < v < 1. F is the feedback signal of the deviation (such as the deviation index or other calculation results), which represents the accuracy of the prediction. If the deviation is large, F is large, and the weight of the model for successful use cases will be reduced, thereby reducing the bias towards successful use cases.

[0038] Through the adaptive learning rate and the feedback mechanism, in addition to adjusting the learning rate and the weights of use cases, other parameters of the model can also be adjusted (for example, the depth of the decision tree, the penalty parameter of the support vector machine, etc.). Set that certain parameters of the current model are obtained in a certain way, and these parameters will be affected by the deviation feedback. The adjustment expression is: ; where: is the updated model parameter, is the parameter of the current model, L is the loss function of the model, which represents the prediction error, is the gradient of the loss function with respect to the model parameter, which represents the contribution of the current model parameter to the prediction error.

[0039] During the whole process, the feedback mechanism will determine the adjustment amplitude of the learning rate and the weights of use cases according to the size of the deviation index PDI. When the deviation of the prediction result is large (i.e., PDI is high), the model will more strongly adjust the learning rate and the weights of use cases to reduce the bias towards successful use cases, so as to better identify failed use cases and high-risk areas.

[0040] Ultimately, by continuously adjusting the model's parameters and weights, the model can gradually reduce bias and increase sensitivity to high-risk areas during repeated learning.

[0041] The adjusted prediction results are promptly delivered to the test team through the feedback mechanism so that they can re-evaluate the test strategy based on the latest prediction information. Specifically, when the machine learning model is adjusted according to the adaptive learning rate and feedback mechanism, the generated prediction results will indicate which test cases or test areas have higher risks or potential defects. Based on these feedback signals, the test team can prioritize the review of high-risk areas and optimize resource allocation to avoid wasting time and energy in low-risk areas. This process can improve testing efficiency, ensure that key defects are discovered and handled in a timely manner, and thus improve the quality and accuracy of the entire software testing.

[0042] Example 2, please refer to Figure 2 As shown, the test risk identification system based on artificial intelligence described in this embodiment includes a data acquisition module, a data extraction module, a data prediction module and a test adjustment module; Data acquisition module: collects comprehensive test data, including test case execution results and test environment status; Data extraction module: extracts abnormal values ​​of dependencies between code modules from the test case execution results, and extracts high-frequency values ​​of network delay from the test environment status; Data prediction module: Build a data prediction model, input dependency anomalies and high-frequency values ​​of network delay into the data prediction model for comprehensive calculation, and judge whether the prediction results of the machine learning model are biased based on the calculation results; Test adjustment module: When prediction deviation occurs, for failed cases and high-risk test areas, by introducing adaptive learning rate and feedback mechanism, the machine learning model's bias towards successful cases is gradually reduced, the sensitivity to failed cases is increased, and the adjusted prediction results are fed back to the test team.

[0043] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0044] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0045] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0046] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A test risk identification method based on artificial intelligence, characterized by: include: Collect comprehensive test data, including test case execution results and test environment status; Extract dependency anomalies between code modules from test case execution results, and extract high-frequency network latency values ​​from test environment status; Build a data prediction model, input dependency anomalies and high-frequency network delay values ​​into the data prediction model for comprehensive calculation, and determine whether the prediction results of the machine learning model are biased based on the calculation results; When prediction deviation occurs, for failed cases and high-risk test areas, by introducing adaptive learning rate and feedback mechanism, the bias of machine learning model towards successful cases is gradually reduced, the sensitivity to failed cases is increased, and the adjusted prediction results are fed back to the test team.

2. The test risk identification method based on artificial intelligence according to claim 1, characterized in that: Extract dependency anomalies between code modules from the test case execution results, and extract high-frequency network latency values ​​from the test environment status, including: Use a directed graph to represent code modules and their dependencies. Nodes represent code modules, and edges represent dependencies between modules. Calculate the current shortest distance from the source node Vi to node u, denoted as dist[u]. Let S represent the set of nodes that have been processed, indicating that the shortest paths from the source node Vi to these nodes have been found. Let Q represent the set of unprocessed nodes, indicating the set of nodes closest to the source node. Set the weight of edge eij as w(eij), representing the strength of the dependency from Vi to Vj. Set the shortest distance of each node dist[Vi]=0, indicating that the distance of the source node is 0. For node Vj, initialize dist[Vj]=∞. Initialize S and Q. Select a node u with the shortest distance in Q, move node u to the processed set S, and remove u from Q. For all adjacent nodes v of node u, check if there is a shorter path: If dist[u]+w(euv)<dist[v], then update dist[v] =dist[u]+w(euv), where w(euv) is the weight of edge euv, representing the dependency strength between code modules Vu and Vv. Set the dependency strength calculation formula: ; where: represents the number of function calls between modules, represents the code coupling degree, reflects the data interaction frequency. c, β, and γ are weight coefficients. Update the shortest path of v until there are no nodes in Q. After calculating the shortest paths of each module through the Dijkstra algorithm, calculate the dependency relationship outlier, and the expression is: ; In the formula, AWD is the dependency relationship outlier, which is the shortest path from Vi to Vj. mean(dist) is the average value of the shortest paths between all modules, and std(dist) is the standard deviation of the shortest paths between all modules.

3. The test risk identification method based on artificial intelligence according to claim 2 is characterized in that: Extract the high-frequency value of network delay from the test environment state. The method for obtaining the high-frequency value of network delay is as follows: collect network delay data over a period of time, convert the time domain signal of network delay to the frequency domain through fast Fourier transform, and obtain the amplitude spectrum by calculating the complex amplitude of Fourier transform: ; Where: A(k) is the amplitude of the kth frequency component, Re(F(k)) and Im(F(k)) represent the real and imaginary parts of the Fourier transform result respectively, set a frequency threshold, and regard the components with frequencies greater than the threshold as high-frequency components. The calculation expression for extracting the amplitude of all high-frequency parts is: ; It is the sum of the amplitudes of all high-frequency components. The calculated sum of the amplitudes of all high-frequency components is used as the high-frequency value of network delay.

4. The test risk identification method based on artificial intelligence according to claim 3 is characterized in that: A data prediction model is constructed, and the dependency anomaly values ​​and high-frequency values ​​of network delay are input into the data prediction model for comprehensive calculation. Specifically, the dependency anomaly values ​​and high-frequency values ​​of network delay are normalized so that they are both between [0,1], and the prediction result deviation index of the machine learning model is calculated based on the normalized dependency anomaly values ​​and high-frequency values ​​of network delay.

5. The test risk identification method based on artificial intelligence according to claim 4 is characterized in that: The obtained prediction result deviation index of the machine learning model is compared with the pre-set standard threshold. If the prediction result deviation index of the machine learning model is greater than or equal to the pre-set standard threshold, it means that the prediction result of the machine learning model is deviated, and a warning signal is generated at this time; If the prediction result deviation index of the machine learning model is less than the pre-set standard threshold, it means that the prediction result of the machine learning model has not deviated, and no warning signal is generated.

6. The test risk identification method based on artificial intelligence according to claim 5 is characterized in that: When there is a deviation in the prediction results, the weight of the machine learning model is adjusted by introducing an adaptive learning rate and feedback mechanism, specifically: Set the learning rate of the current model to η, the strength of the feedback signal to F, the prediction result deviation index to PDI, and update the learning rate. The expression is: ;in: is the updated learning rate, is the current learning rate, γ is the adjustment coefficient of the feedback mechanism, which controls the amplitude of the learning rate change, and F is the strength of the feedback signal; According to the feedback signal, the weights in the training data are dynamically adjusted to reduce the weight of successful cases. The prediction weight of the current model for each case is set to , where represents the weight of case i, and the feedback signal of the deviation is F. The weight of each case is adjusted: ;in: is the updated weight of use case i, and v is the adjustment coefficient, which determines the influence of deviation feedback on weight adjustment.

7. A test risk identification system based on artificial intelligence, used to implement the test risk identification method based on artificial intelligence according to any one of claims 1 to 6, characterized in that: It includes data acquisition module, data extraction module, data prediction module and test adjustment module; Data acquisition module: collects comprehensive test data, including test case execution results and test environment status; Data extraction module: extracts abnormal values ​​of dependencies between code modules from the test case execution results, and extracts high-frequency values ​​of network delay from the test environment status; Data prediction module: Build a data prediction model, input dependency anomalies and high-frequency values ​​of network delay into the data prediction model for comprehensive calculation, and judge whether the prediction results of the machine learning model are biased based on the calculation results; Test adjustment module: When prediction deviation occurs, for failed cases and high-risk test areas, by introducing adaptive learning rate and feedback mechanism, the machine learning model's bias towards successful cases is gradually reduced, the sensitivity to failed cases is increased, and the adjusted prediction results are fed back to the test team.

Citation Information

Patent Citations

  • Modbus TCP (Transmission Control Protocol) fuzzy test method based on QRNN (Quantitative Recurrent Neural Network)

    CN116094972A

  • Detection control method and system based on intelligent detection robot

    CN117633722A

  • Automatic test platform and method adaptive to operating system environment

    CN118170685A

  • Automobile MCU automatic test method and system

    CN118760143A

  • Test method and device, storage medium and program product

    CN119226174A

Cited By

  • Game function development system and method based on natural language instruction

    CN120560648A

  • Game function development system and method based on natural language instructions

    CN120560648B

  • Mobile phone application performance automatic test system and method based on artificial intelligence

    CN121070747A