Software test data management system and method based on cloud platform
Through a cloud-based software test data management system, using methods such as data normalization and information entropy calculation, we quantify environmental differences and automatically adjust the test data, solving the data deviation problem between the test environment and the production environment, and improving testing efficiency and accuracy.
Patent Information
- Application Number
- CN202510756712.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, there are systematic differences between the test environment and the production environment during software testing in terms of hardware configuration, network delay, concurrent pressure, etc., resulting in deviations from the behavioral characteristics of the test data from the real scenario. Traditional methods rely on manual experience to adjust the test data parameters, which is inefficient, has large errors, and is difficult to quantify.
Through a cloud-based software test data management system, mathematical tools such as data normalization, information entropy calculation, probability distribution analysis, etc. are used to quantify environmental differences and automatically adjust the test data to achieve accurate compensation.
Quantitative comparison and automatic compensation of test environment and production environment data is realized, manual intervention is reduced, testing efficiency and accuracy is improved, and misjudgment and failure caused by environmental differences is avoided.
Smart Images

Figure CN120276997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a software test data management system and method based on a cloud platform. Background Art
[0002] With the rapid development of computer technology, computer automatic test systems have been widely used, and a large number of automatic test systems generate a huge amount of test data, covering the original data, judgment result data, test environment data, and index judgment data of each test stage of the product under test. These data are stored dispersedly in each test system, showing an island phenomenon, and it is difficult for manual management to comprehensively analyze the data; with the popularization and development of cloud computing, it provides a new solution for software test data management. The cloud platform has advantages such as elastic scaling and rapid construction of test environments, can provide computing resources on demand, reduce test costs, and at the same time can store and manage a huge amount of data, and support data sharing and collaboration; However, when testing software today, there are systematic differences in the hardware configuration, network latency, concurrent pressure, etc. between the test environment and the production environment, resulting in deviations in the behavioral characteristics of test data from the real scenario (such as distorted response time and inflated transaction success rate). Traditional methods rely on manual experience to adjust test data parameters, which have problems such as low efficiency, large errors, and difficulty in quantification. Summary of the Invention
[0003] The purpose of the present invention is to provide a software test data management system and method based on a cloud platform to solve the problems raised in the prior art.
[0004] To achieve the above purpose, the present invention provides the following technical solutions: A software test data management method based on a cloud platform, the method comprising the following steps: S100. Collect data indicators in the software production environment and data indicators in the test environment, normalize the collected data, screen the data indicators in the two environments, and extract the main influencing factors; Further, the specific steps for extracting the main influencing factors are: S101. Collect data indicators in the software production environment and data indicators in the test environment, calculate the average value and standard deviation of each data indicator respectively, and use the average value and standard deviation of each data indicator to achieve normalization. The formula is: ; In the formula, x norm represents the normalized data indicator, x represents the collected data indicator, xp represents the average value of the data indicator, and xz represents the standard deviation of the data indicator; S102. Construct a data matrix using the normalized data metrics. Assume the number of samples in the matrix is m, and calculate the covariance matrix of the data matrix. The formula is as follows: ; In the formula, E represents the covariance matrix, Z represents the data matrix, and Z T represents the transpose matrix of the data matrix; S103. Perform eigenvalue decomposition on the calculated covariance matrix, calculate the eigenvectors and eigenvalues, and select the eigenvectors corresponding to the k largest eigenvalues as the main influencing factors. The value of k is set manually based on professional experience.
[0005] The original data in the production environment and the test environment may have inconsistent dimensions due to differences in deployment architecture, data scale, user behavior, etc. Normalization can unify the data standard and make the data in different environments comparable. By screening the main influencing factors (such as response time, error rate, throughput, etc.), redundant indicators can be eliminated, the data complexity can be reduced, interference from secondary factors in the analysis can be avoided, and the subsequent calculation efficiency can be improved. After clarifying the key indicators, in-depth analysis can be carried out on the core factors to avoid confusion in the analysis dimensions caused by excessive indicators.
[0006] S200. Calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result using the information entropy, and construct an environmental difference weight matrix; Furthermore, the specific steps for constructing the environmental difference weight matrix are as follows: S201. Assume the set of selected main influencing factors is F = {f1, f2, f3,..., f n}, where f1, f2, f3,..., f n represent the 1st, 2nd, 3rd,..., nth selected main influencing factors, and n is a positive integer; calculate the information entropy of each main influencing factor. The formula is as follows: ; In the formula, H i represents the information entropy of the i-th main influencing factor, and f ik represents the i-th main influencing factor in the k-th sample; S202. Calculate the influence weight of each main influencing factor on the software test result using the information entropy. The formula is as follows: ; In the formula, w i represents the influence weight of the i-th main influencing factor, and n represents the total number of main influencing factors; integrate the influence weights of all calculated main influencing factors to construct an environmental difference weight matrix W = [w1, w2, w3,..., w n, w1, w2, w3, ..., w n represent the influence weights of the 1st, 2nd, 3rd, ..., nth main influencing factors.
[0007] The information entropy calculates the weight through the uncertainty (disorder degree) of the data itself, avoiding the deviation of subjective assignment by humans and making the weight more in line with the actual data characteristics; the environmental difference weight matrix can reflect the importance differences of different indicators in the production / test environment, helping testers to prioritize the factors that have a greater impact on the results; the weight matrix can be iterated with the data update to adapt to the changes in the importance of indicators brought about by software version iteration or environmental changes.
[0008] S300. Extract the probability distribution of each main influencing factor in the production environment and the test environment respectively, calculate the distribution difference using the distribution probabilities of the same main influencing factor in the two environments, and analyze the distribution difference to determine whether software testing needs compensation; Furthermore, the specific steps to determine whether software testing needs compensation are as follows: S301. Discretize the continuous data for each main influencing factor, divide the data range of each main influencing factor into bins, and divide it into h equal-width intervals; extract the frequency distribution within each interval in the production environment and the test environment respectively, and calculate the probability distribution within each interval in the production environment. The formula is: ; In the formula, P(s u ) represents the probability distribution within the u-th interval, q u represents the frequency distribution within the u-th interval, represents the smoothing factor, and h represents the total number of intervals; use the same method to calculate the probability distribution within each interval in the test environment as T(s u ) S302. Calculate the distribution difference between the two environments using the probability distributions within all intervals. The formula is: ; In the formula, D KL (P||T) represents the distribution difference of the main influencing factor between the two environments. Calculate the distribution differences of all main influencing factors in the same way. Set the difference threshold as D threshold , and use the difference threshold to judge the distribution differences of all main influencing factors. When D KL (P||T) > D threshold , it is judged that the distribution difference has a significant impact on software testing, and the compensation mechanism is triggered.
[0009] By comparing the probability distributions of the same factors in two types of environments, the true differences between the test environment and the production environment can be identified. Based on the quantification results of the distribution differences, it can be objectively judged whether compensation is needed (for example, when the data distribution in the test environment is too concentrated, marginal data needs to be artificially injected) to avoid misjudgments caused by empiricism. If the distribution differences of the key factors are significant, the risk of insufficient test coverage can be warned in advance to avoid failures caused by environmental differences after going live.
[0010] S400. Calculate the compensation intensity coefficient using the distribution difference, calculate and extract the project criticality, data update frequency, and data fault tolerance threshold during software testing using professional knowledge, and calculate the data sensitivity factor; calculate the final compensation amount using the compensation intensity coefficient and the data sensitivity factor; Further, the specific steps for calculating the final compensation amount using the compensation intensity coefficient and the data sensitivity factor are as follows: S401. Calculate the compensation intensity coefficient using the distribution difference, and the formula is: ; In the formula, α represents the compensation intensity coefficient, and α base represents the basic compensation intensity; S402. For the normalized main influencing factors, set the sensitivity weight and multiply it by the normalized value to obtain the criticality of each main influencing factor. The sensitivity weight is within [0, 1]. Calculate the average value of each main influencing factor in the production environment and the test environment respectively, and calculate the difference between the average values of each main influencing factor in the two environments as the fault tolerance threshold; calculate the data sensitivity factor using the criticality and fault tolerance threshold of each main influencing factor, and the formula is: ; In the formula, Sd represents the data sensitivity factor, Cr represents the criticality, Fr represents the data update frequency, and Fa represents the fault tolerance threshold; S403. Calculate the final compensation amount using the compensation intensity coefficient and the data sensitivity factor, and the formula is: ; In the formula, △ i represents the final compensation amount of the i-th main influencing factor, p i represents the data value of the i-th main influencing factor in the production environment, t i represents the data value of the i-th main influencing factor in the test environment, α i represents the compensation intensity coefficient of the i-th main influencing factor, w i represents the influence weight of the i-th main influencing factor, Sd i represents the data sensitivity factor of the i-th main influencing factor. Calculate the final compensation amounts of all main influencing factors in the same way and construct the final compensation amount set.
[0011] The absolute value of the combined distribution difference of the compensation intensity coefficient (such as the probability difference), and the data sensitivity factor integrates business characteristics (such as higher requirements for data consistency in financial scenarios). The multiplication of the two can achieve "compensation on demand" to avoid overcompensation or undercompensation.
[0012] Quantifying sensitivity through business metrics such as criticality, update frequency, and fault tolerance threshold can prioritize compensating the data that has the greatest impact on the business in the case of limited resources, improving the cost-effectiveness of testing. S500. When it is determined that software testing requires compensation, use the final compensation amount to compensate all the main influencing factors; recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
[0013] Furthermore, the specific steps for constructing a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix are as follows: S501. After the compensation mechanism is triggered, use the final compensation amount of each main influencing factor in the final compensation amount set to perform data compensation in the test environment during software testing; recalculate the distribution difference D between the two environments based on the compensated data. new When D new ≤ D threshold , it is determined that the compensation is qualified; when it is not satisfied, it is determined that the compensation is unqualified; use D new to recalculate the compensation intensity coefficient and the environmental difference weight matrix, and update and optimize them.
[0014] After compensation, recalculate the distribution difference and update the weight matrix and compensation coefficient to form a closed loop of "analysis - compensation - verification - iteration", enabling the solution to be continuously optimized as the environment changes.
[0015] The feedback mechanism can be integrated into the test tool chain to achieve automatic adjustment of the compensation strategy and reduce the cost of manual intervention.
[0016] A software test data management system based on a cloud platform, where the software test data management system includes a data collection module, an influencing factor screening module, a weight matrix construction module, a compensation judgment module, a compensation amount calculation module, and an update verification module; The data collection module is used to collect data metrics in the software production environment and data metrics in the test environment; The influencing factor screening module is used to normalize the collected data, screen the data metrics in the two environments, and extract the main influencing factors; The weight matrix construction module is used to calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result by using the information entropy, and construct an environmental difference weight matrix; The compensation judgment module is used to calculate the distribution difference by using the distribution probability of the same main influencing factor in two environments, and analyze the distribution difference to judge whether software testing needs compensation; The compensation amount calculation module is used to calculate and extract the project criticality, data update frequency and data fault tolerance threshold during software testing by using professional knowledge, and calculate the data sensitivity factor; calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor; The update verification module is used to compensate all main influencing factors by using the final compensation amount; recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
[0017] The weight matrix construction module includes an information entropy unit and an influence weight unit; The information entropy unit is used to calculate the information entropy of each main influencing factor; The influence weight unit is used to calculate the influence weight of each main influencing factor on the software test result by using the information entropy, and integrate and construct an environmental difference weight matrix.
[0018] The compensation judgment module includes a probability distribution unit and a distribution difference unit; The probability distribution unit is used to extract the frequency distribution in each interval in the production environment and the test environment respectively, and calculate the probability distribution in each interval in the production environment; The distribution difference unit is used to calculate the distribution difference in two environments by using the probability distribution in all intervals.
[0019] The compensation amount calculation module includes a compensation intensity coefficient unit, a data sensitivity factor unit and a final compensation amount unit; The compensation intensity coefficient unit is used to calculate the compensation intensity coefficient by using the distribution difference; The data sensitivity factor unit is used to calculate the data sensitivity factor by using the criticality and fault tolerance threshold of each main influencing factor; The final compensation amount unit is used to calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor.
[0020] Compared with the prior art, the beneficial effects of the present invention are: By using mathematical tools such as information entropy and probability distribution, the present invention avoids test deviation caused by subjective experience, and enables the environmental difference analysis and compensation strategy to have a quantitative basis.
[0021] The present invention systematically solves the data difference problem between the test environment and the production environment through a complete link of "data normalization → weight quantization → difference analysis → precise compensation → dynamic iteration". Description of the Drawings
[0022] Figure 1 It is a module distribution diagram of a software test data management system based on a cloud platform according to the present invention; Figure 2 It is a step schematic diagram of a software test data management method based on a cloud platform according to the present invention. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, A software test data management method based on a cloud platform, the method includes the following steps: S100. Collect data indicators in the software production environment and data indicators in the test environment, normalize the collected data, screen the data indicators in the two environments, and extract the main influencing factors; The specific steps for extracting the main influencing factors are: S101. Collect data indicators in the software production environment and data indicators in the test environment, calculate the average value and standard deviation of each data indicator respectively, and use the average value and standard deviation of each data indicator to achieve normalization. The formula is: ; In the formula, x norm represents the normalized data indicator, x represents the collected data indicator, xp represents the average value of the data indicator, and xz represents the standard deviation of the data indicator; S102. Use the normalized data indicators to construct a data matrix. Assume that the number of samples in the matrix is m, and calculate the covariance matrix of the data matrix. The formula is: ; In the formula, E represents the covariance matrix, Z represents the data matrix, and Z T represents the transposed matrix of the data matrix; S103. Perform eigen-decomposition on the calculated covariance matrix, calculate the eigenvectors and eigenvalues, and select the eigenvectors corresponding to the k largest eigenvalues as the main influencing factors, where the value of k is set manually based on professional experience.
[0025] The original data in the production environment and the test environment may have inconsistent dimensions due to differences in deployment architecture, data scale, user behavior, etc. Normalization can unify the data standard and make the data in different environments comparable. By screening the main influencing factors (such as response time, error rate, throughput, etc.), redundant indicators are eliminated, the data complexity is reduced, interference from secondary factors to the analysis is avoided, and the subsequent calculation efficiency is improved. After clarifying the key indicators, in-depth analysis can be carried out on the core factors to avoid confusion in the analysis dimensions caused by excessive indicators.
[0026] S200. Calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result using the information entropy, and construct an environmental difference weight matrix; The specific steps for constructing the environmental difference weight matrix are as follows: S201. Let the set of selected main influencing factors be F = {f1, f2, f3,..., f n}, where f1, f2, f3,..., f n represent the 1st, 2nd, 3rd,..., nth selected main influencing factors, and n is a positive integer; calculate the information entropy of each main influencing factor, and the formula is: ; In the formula, H i represents the information entropy of the i-th main influencing factor, and f ik represents the i-th main influencing factor in the k-th sample; S202. Calculate the influence weight of each main influencing factor on the software test result using the information entropy, and the formula is: ; In the formula, w i represents the influence weight of the i-th main influencing factor, and n represents the total number of main influencing factors; integrate the influence weights of all calculated main influencing factors to construct an environmental difference weight matrix W = [w1, w2, w3,..., w n , where w1, w2, w3,..., w n represent the influence weights of the 1st, 2nd, 3rd,..., nth main influencing factors.
[0027] The information entropy calculates weights based on the uncertainty (chaos degree) of the data itself, avoiding the deviation of manually assigned subjective values and making the weights more in line with the actual data characteristics; the environmental difference weight matrix can reflect the importance differences of different indicators in the production / test environment, helping testers to prioritize the factors that have a greater impact on the results; the weight matrix can be iterated as the data is updated to adapt to the changes in the importance of indicators brought about by software version iteration or environmental changes.
[0028] S300. Extract the probability distributions of each main influencing factor in the production environment and the test environment respectively, calculate the distribution differences using the distribution probabilities of the same main influencing factors in the two environments, and analyze the distribution differences to determine whether software testing needs compensation; Further, the specific steps to determine whether software testing needs compensation are as follows: S301. Discretize the continuous data for each main influencing factor, divide the data range of each main influencing factor into bins, and divide it into h equal-width intervals; extract the frequency distributions within each interval in the production environment and the test environment respectively, and calculate the probability distribution within each interval in the production environment. The formula is: ; In the formula, P(s u ) represents the probability distribution within the u-th interval, q u represents the frequency distribution within the u-th interval, represents the smoothing factor, and h represents the total number of intervals; calculate the probability distribution within each interval in the test environment as T(s u ) using the same method S302. Calculate the distribution differences in the two environments using the probability distributions within all intervals. The formula is: ; In the formula, D KL (P||T) represents the distribution difference of the main influencing factor in the two environments. Calculate the distribution differences of all main influencing factors in the same way, set the difference threshold as D threshold , and use the difference threshold to judge the distribution differences of all main influencing factors. When D KL (P||T) > D threshold , it is judged that the distribution difference has a significant impact on software testing and the compensation mechanism is triggered.
[0029] By comparing the probability distribution of the same factors in the two types of environments, the real differences between the test environment and the production environment can be identified. Based on the quantitative results of the distribution differences, it is possible to objectively determine whether compensation is needed (for example, when the test environment data distribution is too concentrated, it is necessary to manually inject edge data) to avoid misjudgments caused by empiricism. If the distribution differences of key factors are significant, the risk of insufficient test coverage can be warned in advance to avoid failures caused by environmental differences after going online.
[0030] S400, using the distribution difference to calculate the compensation intensity coefficient, using professional knowledge to calculate and extract the project criticality, data update frequency and data fault tolerance threshold during software testing, and calculate the data sensitivity factor; using the compensation intensity coefficient and the data sensitivity factor to calculate the final compensation amount; The specific steps of calculating the final compensation amount using the compensation intensity coefficient and the data sensitivity factor are as follows: S401. Calculate the compensation strength coefficient using the distribution difference. The formula is: ; In the formula, α represents the compensation intensity coefficient, α base Indicates the basic compensation intensity; S402: For the normalized main influencing factors, set the sensitivity weight and multiply the normalized value to obtain the criticality of each main influencing factor, the sensitivity weight is in [0, 1], calculate the average value of each main influencing factor in the production environment and the test environment respectively, and calculate the difference between the average values of each main influencing factor in the two environments as the fault tolerance threshold; use the criticality and fault tolerance threshold of each main influencing factor to calculate the data sensitivity factor, the formula is: ; In the formula, Sd represents the data sensitivity factor, Cr represents the criticality, Fr represents the data update frequency, and Fa represents the fault tolerance threshold; S403, using the compensation intensity coefficient and the data sensitivity factor to calculate the final compensation amount, the formula is: ; In the formula, △ i represents the final compensation amount of the i-th main influencing factor, p i Indicates the data value of the i-th major influencing factor in the production environment, t i Indicates the data value of the i-th main influencing factor in the test environment, α i represents the compensation intensity coefficient of the i-th main influencing factor, w i represents the influence weight of the i-th main influencing factor, Sd i Represents the data sensitivity factor of the i-th main influencing factor. The final compensation amount of all main influencing factors is obtained by the same calculation, and the final compensation amount set is constructed.
[0031] The absolute value of the combined distribution difference of the compensation intensity coefficient (such as the probability difference), and the data sensitivity factor integrates business characteristics (such as higher requirements for data consistency in financial scenarios). The multiplication of the two can achieve "compensation on demand" to avoid over-compensation or under-compensation.
[0032] Quantify the sensitivity through business metrics such as criticality, update frequency, and fault tolerance threshold, so as to preferentially compensate the data that has the greatest impact on the business under limited resources, and improve the cost performance of testing; S500. When it is judged that software testing needs compensation, use the final compensation amount to compensate all the main influencing factors; recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
[0033] The specific steps for constructing a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix are as follows: S501. After the compensation mechanism is triggered, use the final compensation amount of each main influencing factor in the final compensation amount set to perform data compensation in the test environment during software testing; recalculate the distribution difference D between the two environments based on the compensated data new , when D new ≤D threshold , it is judged that the compensation is qualified; when it is not satisfied, it is judged that the compensation is unqualified; use D new to recalculate the compensation intensity coefficient and the environmental difference weight matrix, and update and optimize them.
[0034] After compensation, recalculate the distribution difference and update the weight matrix and compensation coefficient to form a closed loop of "analysis - compensation - verification - iteration", so that the solution can be continuously optimized as the environment changes.
[0035] The feedback mechanism can be integrated into the test tool chain to realize the automatic adjustment of the compensation strategy and reduce the cost of manual intervention.
[0036] A software test data management system based on a cloud platform, the software test data management system includes a data acquisition module, an influencing factor screening module, a weight matrix construction module, a compensation judgment module, a compensation amount calculation module, and an update verification module; The data acquisition module is used to collect data indicators in the software production environment and data indicators in the test environment; The influencing factor screening module is used to normalize the collected data, screen the data indicators in the two environments, and extract the main influencing factors; The weight matrix construction module is used to calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result by using the information entropy, and construct an environmental difference weight matrix; The compensation judgment module is used to calculate the distribution difference by using the distribution probabilities of the same main influencing factors in two environments, and analyze the distribution difference to determine whether software testing needs compensation; The compensation amount calculation module is used to calculate and extract the project criticality, data update frequency and data fault tolerance threshold during software testing by using professional knowledge, and calculate the data sensitivity factor; calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor; The update verification module is used to compensate all main influencing factors by using the final compensation amount; recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
[0037] The weight matrix construction module includes an information entropy unit and an influence weight unit; The information entropy unit is used to calculate the information entropy of each main influencing factor; The influence weight unit is used to calculate the influence weight of each main influencing factor on the software test result by using the information entropy, and integrate and construct the environmental difference weight matrix.
[0038] The compensation judgment module includes a probability distribution unit and a distribution difference unit; The probability distribution unit is used to extract the frequency distribution in each interval in the production environment and the test environment respectively, and calculate the probability distribution in each interval in the production environment; The distribution difference unit is used to calculate the distribution difference in the two environments by using the probability distributions in all intervals.
[0039] The compensation amount calculation module includes a compensation intensity coefficient unit, a data sensitivity factor unit and a final compensation amount unit; The compensation intensity coefficient unit is used to calculate the compensation intensity coefficient by using the distribution difference; The data sensitivity factor unit is used to calculate the data sensitivity factor by using the criticality and fault tolerance threshold of each main influencing factor; The final compensation amount unit is used to calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor.
[0040] Example: In the software testing of an e-commerce system, two main influencing factors are collected as network latency and the number of CPU cores; for network latency, it is 50ms in the test environment and 80ms in the production environment, and for the number of CPU cores, it is 8 cores in the test environment and 16 cores in the production environment; Calculate the final compensation value respectively, △ network latency = 0.8 × 0.4 × |80 - 50| / 80 × 2.5 = 0.3ms △cpu = 0.8×0.6×|16 - 8| / 16×3.0 = 0.72; Perform the compensation for the main influencing factors in the test environment, specifically: In the network layer: increase the latency by 30 ms; In the application layer: adjust the number of threads from 100 to 100×(1 + 0.75) = 172.
[0041] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A software test data management method based on a cloud platform, characterized in that: The method includes the following steps: S100. Collect data metrics in the software production environment and data metrics in the test environment, normalize the collected data, screen the data metrics in the two environments, and extract the main influencing factors; S200. Calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result by using the information entropy, and construct an environmental difference weight matrix; S300. Extract the probability distribution of each main influencing factor in the production environment and the test environment respectively, calculate the distribution difference by using the distribution probabilities of the same main influencing factor in the two environments, and analyze the distribution difference to determine whether software testing needs compensation; S400. Calculate the compensation intensity coefficient by using the distribution difference, calculate and extract the project criticality, data update frequency and data fault tolerance threshold during software testing by using professional knowledge, and calculate the data sensitivity factor; Calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor; S500. When it is determined that software testing needs compensation, compensate all the main influencing factors by using the final compensation amount; Recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
2. The software test data management method based on a cloud platform according to claim 1, wherein: The specific steps for extracting the main influencing factors in S100 are as follows: S101. Collect data metrics in the software production environment and data metrics in the test environment, calculate the average value and standard deviation of each data metric respectively, and realize normalization by using the average value and standard deviation of each data metric. The formula is: ; In the formula, x norm represents the normalized data index, x represents the collected data index, xp represents the average value of the data index, and xz represents the standard deviation of the data index; S102. Construct a data matrix by using the normalized data metrics. Assume that the number of samples in the matrix is m, and calculate the covariance matrix of the data matrix. The formula is: ; In the formula, E represents the covariance matrix, Z represents the data matrix, and Z T represents the transposed matrix of the data matrix; S103. Perform eigenvalue decomposition on the calculated covariance matrix, calculate the eigenvectors and eigenvalues, and select the eigenvectors corresponding to the k largest eigenvalues as the main influencing factors. The value of k is set manually according to professional experience.
3. A software test data management method based on a cloud platform according to claim 2, characterized in that: The specific steps for constructing the environmental difference weight matrix in S200 are as follows: S201. Let the set of the main influencing factors selected be \(F = \{f_1, f_2, f_3, \ldots, f_n\}\), where \(f_1, f_2, f_3, \ldots, f_n\) represent the 1st, 2nd, 3rd, \(\ldots\), \(n\)th main influencing factors selected, and \(n\) is a positive integer. Calculate the information entropy of each main influencing factor. The formula is as follows: n}, where \(f_1, f_2, f_3, \ldots, f_n\) n represent the 1st, 2nd, 3rd, \(\ldots\), \(n\)th main influencing factors selected, and \(n\) is a positive integer. Calculate the information entropy of each main influencing factor. The formula is as follows: ; In the formula, H i represents the information entropy of the i-th main influencing factor, and f ik represents the i-th main influencing factor in the k-th sample; S202. Calculate the influence weight of each main influencing factor on the software test result by using the information entropy. The formula is: ; In the formula, w i represents the influence weight of the i-th main influencing factor, and n represents the total number of main influencing factors; the influence weights of all calculated main influencing factors are integrated to construct an environmental difference weight matrix W = [w1, w2, w3,..., w n , where w1, w2, w3,..., w n represent the influence weights of the 1st, 2nd, 3rd,..., n-th main influencing factors.
4. A software test data management method based on a cloud platform according to claim 3, characterized in that: The specific steps for determining whether software testing needs compensation in S300 are as follows: S301. Discretize the continuous data for each main influencing factor, divide the data range of each main influencing factor into bins, and divide it into h equal-width intervals; Extract the frequency distribution within each interval in the production environment and the test environment respectively, and calculate the probability distribution within each interval in the production environment. The formula is: ; In the formula, P(s u ) represents the probability distribution within the u-th interval, q u represents the frequency distribution within the u-th interval, represents the smoothing factor, h represents the total number of intervals; the probability distribution of each interval within the test environment is calculated using the same method as T(s u ) S302. Calculate the distribution difference between the two environments by using the probability distributions in all intervals. The formula is: ; In the formula, D KL (P||T) represents the distribution difference of the main influencing factors in two environments. The distribution differences of all main influencing factors are obtained by the same calculation, and the difference threshold is set as D threshold . The difference threshold is used to judge the distribution differences of all main influencing factors. When D KL (P||T) > D threshold , it is judged that the distribution difference has a significant impact on software testing and the compensation mechanism is triggered.
5. A software test data management method based on a cloud platform according to claim 4, characterized in that: The specific steps for calculating the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor in S400 are as follows: S401. Calculate the compensation intensity coefficient by using the distribution difference. The formula is: ; In the formula, α represents the compensation intensity coefficient, and α base represents the basic compensation intensity; S402. For the normalized main influencing factors, set the sensitivity weight and multiply it by the normalized value to obtain the criticality of each main influencing factor. The sensitivity weight is within [0, 1]. Calculate the average value of each main influencing factor in the production environment and the test environment respectively, and calculate the difference between the average values of each main influencing factor in the two environments as the fault tolerance threshold. Calculate the data sensitivity factor using the criticality and fault tolerance threshold of each main influencing factor. The formula is: ; In the formula, Sd represents the data sensitivity factor, Cr represents the criticality, Fr represents the data update frequency, and Fa represents the fault tolerance threshold. S403. Calculate the final compensation amount using the compensation intensity coefficient and the data sensitivity factor. The formula is: ; In the formula, △ i represents the final compensation amount of the i-th main influencing factor, p i represents the data value of the i-th main influencing factor in the production environment, t i represents the data value of the i-th main influencing factor in the test environment, α i represents the compensation intensity coefficient of the i-th main influencing factor, w i represents the influence weight of the i-th main influencing factor, Sd i represents the data sensitivity factor of the i-th main influencing factor. The final compensation amounts of all main influencing factors are obtained through the same calculation to construct a set of final compensation amounts.
6. A software test data management method based on a cloud platform according to claim 5, characterized in that: The specific steps for constructing the feedback update mechanism in S500 to update the compensation intensity coefficient and the environmental difference weight matrix are: S501. After triggering the compensation mechanism, use the final compensation amounts of each main influencing factor in the final compensation amount set to perform data compensation in the test environment during software testing; recalculate the distribution difference D between the two environments based on the compensated data. new , when D new ≤D threshold , determine that the compensation is qualified; when not satisfied, determine that the compensation is unqualified; use D new to recalculate the compensation intensity coefficient and the environmental difference weight matrix, and update and optimize them.
7. A software test data management system based on a cloud platform, characterized in that: The software test data management system includes a data acquisition module, an influencing factor screening module, a weight matrix construction module, a compensation judgment module, a compensation amount calculation module, and an update verification module. The data acquisition module is used to collect data indicators in the software production environment and data indicators in the test environment. The influencing factor screening module is used to normalize the collected data, screen the data indicators in the two environments, and extract the main influencing factors. The weight matrix construction module is used to calculate the information entropy of each main influencing factor respectively, calculate the influence weight of each main influencing factor on the software test result using the information entropy, and construct the environmental difference weight matrix. The compensation judgment module is used to calculate the distribution difference using the distribution probabilities of the same main influencing factors in the two environments, analyze the distribution difference, and judge whether software testing requires compensation. The compensation amount calculation module is used to calculate and extract the project criticality, data update frequency, and data fault tolerance threshold during software testing using professional knowledge, and calculate the data sensitivity factor. Calculate the final compensation amount using the compensation intensity coefficient and the data sensitivity factor. The update verification module is used to compensate all the main influencing factors using the final compensation amount; recalculate the distribution difference after compensation, and construct a feedback update mechanism to update the compensation intensity coefficient and the environmental difference weight matrix.
8. A software test data management system based on a cloud platform according to claim 7, characterized in that: The weight matrix construction module includes an information entropy unit and an influence weight unit. The information entropy unit is used to calculate the information entropy of each main influencing factor. The influence weight unit is used to calculate the influence weight of each main influencing factor on the software test result using the information entropy, and integrate and construct the environmental difference weight matrix.
9. A software test data management system based on a cloud platform according to claim 7, characterized in that: The compensation judgment module includes a probability distribution unit and a distribution difference unit. The probability distribution unit is used to extract the frequency distribution in each interval in the production environment and the test environment respectively, and calculate the probability distribution in each interval in the production environment. The distribution difference unit is used to calculate the distribution difference in the two environments using the probability distributions in all intervals.
10. A software test data management system based on a cloud platform according to claim 7, characterized in that: The compensation amount calculation module includes a compensation intensity coefficient unit, a data sensitivity factor unit, and a final compensation amount unit. The compensation intensity coefficient unit is used to calculate the compensation intensity coefficient using the distribution difference. The data sensitivity factor unit is used to calculate the data sensitivity factor by using the criticality and fault tolerance threshold of each main influencing factor; The final compensation amount unit is used to calculate the final compensation amount by using the compensation intensity coefficient and the data sensitivity factor.
Citation Information
Patent Citations
Volunteer work intention prediction method based on Logistic generalized linear regression model
CN109934407A
Comprehensive judgment method and system for child brain development state
CN117438080A
Data compensation method and system of sensor
CN119043401A