High-speed interface fault-tolerant test process self-adaptive generation method based on reinforcement learning
Through the adaptive generation method of high-speed interface fault-tolerant testing process based on reinforcement learning, combined with the fault-tolerant testing impact prediction module and test execution and monitoring platform, the problem of the test strategy in the existing technology cannot be adaptively adjusted, and the efficiency and reliability of high-speed interface fault-tolerant testing are achieved.
Patent Information
- Application Number
- CN202510526119.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing high-speed interface testing and monitoring methods cannot adaptively adjust the testing strategy based on real-time test feedback, resulting in wasted testing resources in low-risk areas, making it difficult to dig deep into the corners and failure boundaries that may have problems, affecting the efficiency and thoroughness of error detection.
Adaptive generation method of high-speed interface fault-tolerant testing process based on reinforcement learning is adopted. By obtaining high-speed interface data, environment data and historical fault-tolerant testing data, the fault-tolerant testing impact prediction module is trained, the status, actions, strategies and reward elements of the reinforcement learning agent are defined, and high-speed interface adaptive testing is carried out in combination with reinforcement learning agents, pre-training modules, and test execution and monitoring platforms.
It improves the effectiveness and reliability of the high-speed interface fault-tolerant testing process, reduces manual intervention by automatically executing the test process, shortens the test cycle, and assists reinforcement learning agents to generate more effective testing strategies.
Smart Images

Figure CN120045399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-speed interface fault tolerance testing, and specifically to a method for adaptively generating a high-speed interface fault tolerance testing process based on reinforcement learning. Background Art
[0002] The overall performance and reliability of high-computing power chips highly depend on the stable operation of their integrated high-speed interfaces. These interfaces not only have extremely high data transfer rates but also adopt complex communication protocols with a vast configurable parameter space. Moreover, their normal operation is extremely sensitive to physical layer conditions and environmental factors.
[0003] In the design verification and production testing phases of chips, one of the core challenges lies in effectively and efficiently detecting potential design defects, manufacturing flaws, or runtime functional errors in these high-speed interfaces. These errors can manifest in various forms, including but not limited to: bit errors in data transfer, cyclic redundancy check failures, protocol state machine errors, link training failures, packet loss or corruption, timing violations, and intermittent failures induced by signal integrity issues or environmental stress.
[0004] Existing high-speed interface testing and monitoring methods, such as specification-based or manually designed test cases, random or pseudo-random testing, and traditional coverage-driven verification, even with built-in error detection mechanisms, cannot adaptively adjust the test strategy based on real-time test feedback (such as the detected error types, frequencies, or interface state changes). At the same time, the lack of an effective adaptive error detection mechanism results in a large amount of test resources being consumed in low-risk areas, making it difficult to concentrate efforts on deeply exploring the corner cases and failure boundaries where problems are most likely to exist, which directly affects the efficiency and thoroughness of error detection.
[0005] Therefore, a method for adaptively generating a high-speed interface fault tolerance testing process based on reinforcement learning is proposed. Summary of the Invention
[0006] The object of the present invention is to provide a method for adaptively generating a high-speed interface fault tolerance test process based on reinforcement learning. First, obtain high-speed interface data, environment data, and historical fault tolerance test data; use the historical fault tolerance test data to train a fault tolerance test impact prediction module to obtain a pre-trained module and historical fault tolerance test impact weights; define the state, action, policy, and reward elements of a reinforcement learning agent according to the high-speed interface data and environment data; and perform high-speed interface adaptive testing in combination with the reinforcement learning agent, the pre-trained module, and a test execution and monitoring platform. During the testing process, use the pre-trained module to update the historical fault tolerance test impact weights and the current state; obtain an adaptive reward according to the test results and state changes obtained by the test execution and monitoring platform by monitoring the interface test state, error information, and performance metrics; the present invention improves the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0007] To achieve the above object, the present invention provides the following technical solutions: A method for adaptively generating a high-speed interface fault tolerance test process based on reinforcement learning, comprising: Obtain high-speed interface data, environment data, and historical fault tolerance test data; Use the historical fault tolerance test data to train a fault tolerance test impact prediction module to obtain a pre-trained fault tolerance test impact prediction module and historical fault tolerance test impact weights; Define the state, action, policy, and reward elements of a reinforcement learning agent according to the high-speed interface data and the environment data; Perform high-speed interface adaptive testing in combination with the reinforcement learning agent, the pre-trained fault tolerance test impact prediction module, and a test execution and monitoring platform. The testing process is as follows: set initial interface configuration and environmental conditions, select an initial action based on interface prior knowledge; obtain the current state, and use the pre-trained fault tolerance test impact prediction module to update the historical fault tolerance test impact weights and the current state to obtain an updated state; select an action according to the current policy and the updated state, and execute the action instruction by the test execution and monitoring platform; obtain the test result and the next state after executing the action according to the platform's monitoring of the interface test state, test error information, and test performance metrics; obtain an adaptive reward according to the test result and state change; update the agent policy using an experience tuple; and perform cyclic iteration by continuously interacting with the environment and learning until the test reaches a preset time, number of error discoveries, coverage target, or the agent policy converges and then stops.
[0008] Further, the high-speed interface data includes interface configuration parameter data and interface operation state data; the environment data includes environment parameter data and interference data; the historical fault tolerance test data includes: historical test configuration data, historical test process data, and historical test result data.
[0009] Further, the process of training the fault tolerance test impact prediction module using the historical fault tolerance test data to obtain the pre-trained fault tolerance test impact prediction module and the historical fault tolerance test impact weight includes: Construct the fault tolerance test impact prediction module; Select historical test process data and historical test result data that match the configuration parameters of the interface to be tested from the historical fault tolerance test data; wherein, the historical test process data includes historical interface operation status data and historical environment data; Input the historical test process data and the historical test result data into the fault tolerance test impact prediction module for processing to obtain the pre-trained fault tolerance test impact prediction module and the historical fault tolerance test impact weight under the configuration of the interface to be tested.
[0010] Further, the process of updating the historical fault tolerance test impact weight and the current state using the pre-trained fault tolerance test impact prediction module to obtain the updated state includes: Obtain the current state; Wherein, the current state includes the high-speed interface data and the environment data; Input the interface operation status data, environment parameter data, and interference data in the current state into the pre-trained fault tolerance test impact prediction module for processing, and update the historical fault tolerance test impact weight to generate an impact prediction vector; Combine the impact prediction vector with the current state to obtain the updated state.
[0011] Further, the process of obtaining the test result and the next state after the execution action according to the monitoring of the interface test status, test error information, and test performance metrics by the platform, and obtaining the adaptive reward according to the test result and state change is: The test execution and monitoring platform monitors the status, errors, and performance of the high-speed interface after receiving the action instruction; Compare the monitored interface test status with the next state to obtain a state distance reward; Evaluate the severity and novelty of the monitored test error information to obtain an error reward; Compare the monitored test performance metrics with the performance benchmark values to obtain a performance reward; Perform a weighted sum of the state distance reward, the error reward, and the performance reward to obtain the adaptive reward.
[0012] Further, the experience tuple includes: the updated state, the action, the adaptive reward, and the next state.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention proposes an adaptive generation method for high-speed interface fault-tolerant test processes based on deep learning to achieve adaptive interface fault-tolerant testing. This method combines reinforcement learning, a fault-tolerant test impact prediction module, and a test execution and monitoring platform. By using reinforcement learning and the test execution and monitoring platform, the test process can be automatically executed, reducing manual intervention and shortening the test cycle. The fault-tolerant test impact prediction module can assist the reinforcement learning agent in generating more effective test strategies.
[0014] 2. The present invention proposes a fault-tolerant test impact prediction module for generating impact prediction vectors. This module is trained using historical fault-tolerant test data to learn the correlation between the interface operating state, environmental parameters, and the occurrence of fault-tolerant test errors. The generated impact prediction vectors are used to update the current state in the high-speed interface adaptive test, thereby helping the reinforcement learning agent generate effective test strategies.
[0015] 3. The present invention proposes an adaptive reward for guiding the agent to learn better test strategies. This reward function combines the monitoring information of the test execution and monitoring platform and the agent state, comprehensively considering the state, errors, and performance metrics of the high-speed interface. This reward function is obtained by weighting the state distance reward, error reward, and performance reward, and can evaluate the state exploration value, the severity and novelty of errors, and the actual deviation of performance, which is beneficial to assisting the agent in learning better test strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flow chart of the adaptive generation method for high-speed interface fault-tolerant test processes based on reinforcement learning of the present invention; Figure 2 is a schematic structural diagram of the fault-tolerant test impact prediction module of the present invention; Figure 3 is a schematic flow chart of the high-speed interface adaptive test process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figures 1 to 3 , the present invention provides an adaptive generation method for high-speed interface fault-tolerant test processes based on reinforcement learning, and the technical solutions are as follows: Embodiment 1: In order to adaptively generate a high-speed interface fault tolerance test process, a certain company uses the method for adaptively generating a high-speed interface fault tolerance test process based on reinforcement learning proposed by the present invention. The process schematic of this method is as shown in Figure 1 shown, and specifically includes: Obtain high-speed interface data, environmental data, and historical fault tolerance test data; Furthermore, the high-speed interface data, environmental data, and historical fault tolerance test data can be obtained by calling an interface test management platform; Furthermore, the high-speed interface data includes interface configuration parameter data and interface operating status data; the environmental data includes environmental parameter data and interference data; the historical fault tolerance test data includes: historical test configuration data, historical test process data, and historical test result data; Furthermore, the interface configuration parameter data includes: rate, bit width, pre-emphasis / equalization setting, timing parameters, etc.; the interface operating status data includes: link training status, bit error rate (BER), cyclic redundancy check (CRC) error count, protocol error flag, throughput, etc.; the environmental parameter data includes: core voltage, IO voltage, chip temperature, environmental temperature, etc.; the interference data includes: jitter, noise, voltage drop, etc.; Furthermore, the historical test configuration data is historical interface configuration parameter data; the historical test process data includes historical interface operating status data and historical environmental data; the historical test result data includes historical test normal result data and historical test abnormal result data; Furthermore, the abnormal error types in the historical test abnormal result data include: CRC error, link retraining, link loss, BER exceeding the standard, etc.
[0019] By introducing high-speed interface data, environmental data, and historical fault tolerance test data, data support is provided for subsequent fault tolerance test impact prediction and high-speed interface adaptive test, thereby improving the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0020] Train the fault tolerance test impact prediction module using the historical fault tolerance test data to obtain a pre-trained fault tolerance test impact prediction module and historical fault tolerance test impact weights; Furthermore, the process of training the fault tolerance test impact prediction module using the historical fault tolerance test data to obtain a pre-trained fault tolerance test impact prediction module and historical fault tolerance test impact weights includes: Construct a fault tolerance test impact prediction module; Select historical test process data and historical test result data that match the configuration parameters of the interface to be tested from the historical fault tolerance test data; among them, the historical test process data includes historical interface operating status data and historical environmental data; Input the historical test process data and historical test result data into the fault-tolerant test impact prediction module for processing, and obtain the pre-trained fault-tolerant test impact prediction module and historical fault-tolerant test impact weights under the interface configuration to be tested.
[0021] Further, the fault-tolerant test impact prediction module is a temporal-environment interaction attention prediction network, and its structure is as Figure 2 shown, including: an input layer, a parallel temporal feature encoding layer, a cross-modal interaction attention layer, a feature fusion and compression layer, an impact prediction head, and an impact weight generation layer; Further, the input layer is used to receive historical interface operation status data, historical environment data, and historical test result data; among them, the historical test result data is used as the label during module training; Further, the parallel temporal feature encoding layer is used to extract their respective time-dependent features from the interface status data sequence and the environment status sequence; the parallel temporal feature encoding layer includes an interface status encoder and an environment status encoder; the interface status encoder and the environment status encoder are processed using a long short-term memory network and a temporal convolutional network respectively to obtain the encoded interface status sequence features and environment status sequence features; Further, the cross-modal interaction attention layer is used to model the dynamic interaction impact between the interface status sequence and the environment status sequence; the cross-modal interaction attention layer adopts a cross-modal attention mechanism, and by calculating the attention weights of the environment on the interface status and the interface status on the environment, the weighted interface status features and environment status features are obtained respectively; In this embodiment, the attention weight of the environment on the interface status is calculated by using the encoded interface status sequence features as the query, and the encoded environment status sequence features as the key and value.
[0022] Further, the feature fusion and compression layer is used to fuse the original encoded features and the interaction features weighted by attention into a fixed-length fusion feature vector; the processing process of the feature fusion and compression layer can be expressed as: ; where, is the fusion feature vector; is a multi-layer perceptron; is the scratch weight value; is the encoded interface status sequence features; is the encoded environment status sequence features; is the weighted interface status features; is the weighted environment status features; is a pooling operation in the time dimension to extract global context information; Furthermore, the impact prediction head is used to output the final fault-tolerant test impact prediction value; this structure mainly processes the fused feature vector using a multi-layer perceptron and an activation function to obtain the fault-tolerant test impact prediction value; Furthermore, the impact weight generation layer is used to output the historical fault-tolerant test impact weight; the impact weight generation layer aggregates the attention weights calculated in the cross-modal interaction attention layer along the time dimension to obtain the historical fault-tolerant test impact weight; Furthermore, the fault-tolerant test impact prediction value and the fault-tolerant test impact weight constitute the impact prediction vector.
[0023] By using the fault-tolerant test impact prediction module, the correlation between the interface operation state, environmental parameters, and the generation of fault-tolerant test errors can be learned; and the current state in the high-speed interface adaptive test is updated using the generated impact prediction vector to help the reinforcement learning agent generate effective test strategies, thereby improving the effectiveness and reliability of the high-speed interface fault-tolerant test process.
[0024] Define the state, action, policy, and reward elements of the reinforcement learning agent based on the high-speed interface data and environmental data; Furthermore, the state includes: interface configuration parameters, interface operation state, environmental parameters, interference data, and historical test information; among them, the historical test information includes the types of discovered errors, covered configuration / state regions, etc.; Furthermore, the actions include: adjusting interference and environmental parameters, selecting test patterns or data traffic modes, adjusting system load or working state, etc.; Furthermore, the policy of the reinforcement learning agent is based on Q-values, that is, DQN is used as the reinforcement learning algorithm; Furthermore, the rewards include: state distance reward, error reward, and performance reward.
[0025] Perform high-speed interface adaptive testing by combining the reinforcement learning agent, the pre-trained fault-tolerant test impact prediction module, and the test execution and monitoring platform.
[0026] Furthermore, the process flow diagram of the high-speed interface adaptive test process is as Figure 3As shown below, specifically: Set the initial interface configuration and environmental conditions, and select the initial action based on the interface prior knowledge; Obtain the current state, and use the pre-trained fault tolerance test impact prediction module to update the historical fault tolerance test impact weight and the current state to obtain the updated state; Select an action according to the current policy and the updated state, and the test execution and monitoring platform executes the action instruction; According to the platform's monitoring of the interface test status, test error information, and test performance indicators, obtain the test result and the next state after executing the action; According to the test result and the state change, obtain the adaptive reward; Update the agent policy using the experience tuple; Continuously perform loop iterations through interaction and learning with the environment until the test reaches the preset time, the number of errors found, the coverage target, or the agent policy converges and then stop; Further, the interface prior knowledge includes: historical experience data, interface protocol specifications, and physical and environmental characteristic knowledge; This prior knowledge can be obtained through interface technical documents and historical test reports; By combining reinforcement learning, the fault tolerance test impact prediction module, and the test execution and monitoring platform, the reinforcement learning and the test execution and monitoring platform can automatically execute the test process, reduce manual intervention, and shorten the test cycle; Using the fault tolerance test impact prediction module can assist the reinforcement learning agent to generate more effective test strategies, thereby improving the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0027] Further, the process of using the pre-trained fault tolerance test impact prediction module to update the historical fault tolerance test impact weight and the current state to obtain the updated state includes: Obtain the current state; Among them, the current state includes high-speed interface data and environmental data; Input the interface running state data, environmental parameter data, and interference data in the current state into the pre-trained fault tolerance test impact prediction module for processing, and update the historical fault tolerance test impact weight to generate an impact prediction vector; Combine the impact prediction vector with the current state to obtain the updated state.
[0028] By introducing the impact prediction vector obtained by the fault tolerance test impact prediction module into the current state, it can endow the state information with predictability, interpretability, and context awareness based on historical experience, highlight the key influencing factors in the current state, improve the intelligence level of subsequent decisions, significantly improve the test efficiency and defect discovery ability, and thus ensure the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0029] Further, the process of obtaining the test result and the next state after executing the action according to the platform's monitoring of the interface test status, test error information, and test performance indicators, and obtaining the adaptive reward according to the test result and the state change is: After receiving the action instruction, the test execution and monitoring platform monitors the status, errors, and performance of the high-speed interface; Compare the monitored interface test status with the next status to obtain the status distance reward; Evaluate the severity and novelty of the monitored test error information to obtain the error reward; Compare the monitored test performance indicators with the performance benchmark values to obtain the performance reward; Perform a weighted sum of the status distance reward, error reward, and performance reward to obtain the adaptive reward; Furthermore, the adaptive reward can be expressed as: ; where, is the adaptive reward at time step ; is the status distance reward; is the error reward; is the performance reward; Furthermore, the status distance reward can be expressed as: ; where, is the status distance reward; is the status reward weight; is the number of parameter types in the state space; is the Euclidean norm; is the current updated state at time step ; is the next state at time step ; Furthermore, the error reward can be expressed as: ; ; where, is the error reward; is the severity evaluation weight; is the severity score; is the novelty evaluation weight; is the novelty score; Furthermore, the severity evaluation weight and novelty evaluation weight are respectively set to 0.5; the setting of the weights can be adjusted according to actual needs; Further, the severity score is a normalized value assigned according to the error type. For example, it is set to 0.8 when there is a link loss, 0.6 when there are a large number of CRC errors, and 0.4 when there are a small number of correctable errors. The novelty score is set according to whether the state-action combination that causes the error to occur is discovered for the first time. If so, it is set to 1; otherwise, it is 0. Further, the performance reward can be expressed as: ; Where is the performance reward; is the performance reward weight; is the number of performance parameters; is the expected value of the th performance parameter; is the measured value of the
[0030] By comprehensively considering the status, errors, and performance metrics of the high-speed interface, the adaptive reward can evaluate the value of state exploration, the severity and novelty of errors, and the actual deviation of performance, which is beneficial to assisting the agent in learning better test strategies, thereby ensuring the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0031] Further, the experience tuple includes: updated state, action, adaptive reward, and next state.
[0032] In this embodiment, the experience tuple contains rich state prediction information and refined reward signals, enabling the reinforcement learning agent to be fully applied to complex fault tolerance test scenarios, making the adaptive test process learn faster, explore more intelligently, and produce more effective results, thereby ensuring the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0033] This embodiment provides an adaptive generation method for the high-speed interface fault tolerance test process based on reinforcement learning. First, obtain high-speed interface data, environmental data, and historical fault tolerance test data; use the historical fault tolerance test data to train the fault tolerance test impact prediction module to obtain the pre-trained module and the historical fault tolerance test impact weight; define the state, action, policy, and reward elements of the reinforcement learning agent according to the high-speed interface data and environmental data; combine the reinforcement learning agent, the pre-trained module, and the test execution and monitoring platform to perform high-speed interface adaptive testing. During the testing process, use the pre-trained module to update the historical fault tolerance test impact weight and the current state; obtain the adaptive reward according to the test results and state changes obtained by the test execution and monitoring platform through monitoring the interface test status, error information, and performance metrics; the present invention improves the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0034] Embodiment Two: The present invention proposes a method for adaptively generating a high-speed interface fault tolerance test process based on reinforcement learning. To further verify the effectiveness of the fault tolerance test impact prediction module and the high-speed interface adaptive test process proposed by the present invention, the present invention conducts effectiveness tests on the modules and processes respectively by selecting different modules and different test processes. The present invention selects two enterprises, A and B, to conduct the above two groups of tests.
[0035] The present invention selects the historical fault tolerance test data of enterprise A in the past 3 years as the data set of the module, where the data in the first and second years is used as the training set of the model, and the data in the third year is used as the validation set; the sampling of the data set refers to the following rules: the data comes from the same interface configuration environment, and 4 days of data per week are randomly selected on a weekly basis; among them, 3 groups of data are extracted in the morning and afternoon every day.
[0036] The present invention inputs the training sets collected from enterprise A into different fault tolerance test impact prediction modules for training respectively to obtain their respective pre-trained modules; then, the validation sets are input into each pre-trained module to obtain the impact prediction vectors of each module; then, the impact prediction vectors are compared with the actual situation through manual verification to obtain the proportion of the impact prediction vectors of each module within a reasonable range.
[0037] The modules are respectively: the fault tolerance test impact prediction module proposed by the present invention, denoted as module one; the cross-modal interaction attention layer in the fault tolerance test impact prediction module is removed, denoted as model two; the pooling operation in the feature fusion and compression layer is removed, and only the encoded features and the weighted features are fused, denoted as module three.
[0038] The results of the module effectiveness test are shown in Table 1.
[0039] Table 1 Results of Module Effectiveness Test Test module Proportion of influence prediction vectors within a reasonable range Model 1 90.59% Model 2 86.82% Model 3 88.78% It can be seen from the results in Table 1 that the fault tolerance test impact prediction module proposed by the present invention is better than other modules in the effectiveness test results; thus, it can be explained that the module proposed by the present invention can output accurate impact prediction vectors by combining cross-modal interaction attention and multi-parameter feature fusion, further improving the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0040] To further verify the effectiveness of the high-speed interface adaptive test process, this embodiment collected the historical data of Company B in the past year, sampled it according to the same data sampling rules, and then processed the sampled data sets according to Process 1, Process 2, and Process 3 respectively to obtain their respective test results. Among them, Process 1 is the high-speed interface adaptive test process proposed by the present invention; Process 2 is to remove the operation of updating the state of the prediction module affected by the pre-training fault tolerance test; Process 3 is to calculate the reward only by using the state change and not to use the platform to monitor the interface. Finally, the effectiveness of each test process was verified through manual verification. The test results of process effectiveness are shown in Table 2. Table 2 Test Results of Process Effectiveness Test process Proportion of test results within a reasonable range Process 1 90.43% Process 2 88.57% Process 3 87.92% As can be seen from the results in Table 2, the effectiveness test results obtained by using Process 1, that is, the high-speed interface adaptive test process proposed by the present invention, are better than those obtained by using other processes. This shows the necessity of combining the impact prediction vector with the platform test results for the high-speed interface adaptive test, which can improve the effectiveness and reliability of the high-speed interface fault tolerance test process.
[0041] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for adaptively generating high-speed interface fault-tolerant test flow based on reinforcement learning, characterized in that: include: Obtain high-speed interface data, environmental data, and historical fault-tolerance test data; The fault-tolerant test impact prediction module is trained using historical fault-tolerant test data to obtain a pre-trained fault-tolerant test impact prediction module and historical fault-tolerant test impact weights; Define the state, action, strategy, and reward elements of the reinforcement learning agent based on high-speed interface data and environmental data; Combined with reinforcement learning agent, pre-trained fault-tolerant test impact prediction module and test execution and monitoring platform, high-speed interface adaptive testing is carried out. The testing process is as follows: set the initial interface configuration and environmental conditions, select the initial action based on the interface prior knowledge; obtain the current state, and use the pre-trained fault-tolerant test impact prediction module to update the historical fault-tolerant test impact weight and the current state to obtain the updated state; select the action according to the current strategy and the updated state, and the test execution and monitoring platform executes the action instruction; Based on the platform's monitoring of interface test status, test error information, and test performance indicators, the test results and the next status after the action is executed are obtained; Get adaptive rewards based on test results and state changes; use experience tuples to update agent strategies; The cycle iterates by continuously interacting with the environment and learning until the test reaches the preset condition.
2. The method for adaptively generating a high-speed interface fault-tolerant test process based on reinforcement learning according to claim 1, characterized in that: The high-speed interface data includes interface configuration parameter data and interface operation status data; the environmental data includes environmental parameter data and interference data; the historical fault-tolerant test data includes: historical test configuration data, historical test process data and historical test result data.
3. The method for adaptively generating a high-speed interface fault-tolerant test process based on reinforcement learning according to claim 1, characterized in that: The process of training the fault-tolerant test impact prediction module using the historical fault-tolerant test data to obtain the pre-trained fault-tolerant test impact prediction module and the historical fault-tolerant test impact weights includes: Constructing the fault-tolerant test impact prediction module; Selecting historical test process data and historical test result data that match the configuration parameters of the interface to be tested from the historical fault-tolerant test data; wherein the historical test process data includes historical interface operation status data and historical environment data; The historical test process data and the historical test result data are input into the fault-tolerant test impact prediction module for processing to obtain the pre-trained fault-tolerant test impact prediction module and the historical fault-tolerant test impact weight under the interface configuration to be tested.
4. The method for adaptively generating a high-speed interface fault-tolerant test process based on reinforcement learning according to claim 1, characterized in that: The process of updating the historical fault-tolerant test impact weight and the current state by using the pre-trained fault-tolerant test impact prediction module to obtain the updated state includes: Get the current state; Wherein, the current state includes the high-speed interface data and the environmental data; Inputting the interface operation status data, environmental parameter data and interference data in the current state into the pre-trained fault-tolerant test impact prediction module for processing, and updating the historical fault-tolerant test impact weight to generate an impact prediction vector; The impact prediction vector is combined with the current state to obtain the updated state.
5. The method for adaptively generating a high-speed interface fault-tolerant test process based on reinforcement learning according to claim 1, characterized in that: According to the platform's monitoring of the interface test status, test error information, and test performance indicators, the test results and the next state after the action is executed are obtained, and the process of obtaining adaptive rewards based on the test results and state changes is as follows: The test execution and monitoring platform monitors the status, errors and performance of the high-speed interface after receiving the action instruction; Compare the interface test state obtained by monitoring with the next state to obtain a state distance reward; Evaluate the severity and novelty of the monitored test error information and obtain error rewards; Compare each test performance indicator obtained through monitoring with the performance benchmark value to obtain performance rewards; The state distance reward, the error reward, and the performance reward are weightedly summed to obtain the adaptive reward.
6. The method for adaptively generating a high-speed interface fault-tolerant test process based on reinforcement learning according to claim 1, characterized in that: The experience tuple includes: the updated state, the action, the adaptive reward, and the next state.
Citation Information
Patent Citations
Interface test case generation method, device and equipment based on reinforcement learning
CN114911704A
Reinforcement learning sequence decision-making method, system, equipment and medium
CN117972588A
Reinforcement learning agent training environment construction method and device and electronic equipment
CN117993471A
Systems and methods for managing network performance based on defining rewards for a reinforcement learning model
US20210152439A1