Method for managing server test services and electronic device
By preprocessing data and optimizing resources for server testing, the problems of insufficient data processing capacity and unreasonable resource scheduling were solved, achieving efficient management and reasonable allocation, and improving testing efficiency and quality.
Patent Information
- Application Number
- CN202511470432.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing technologies for server testing suffer from insufficient data processing capabilities, unreasonable resource scheduling, and inadequate real-time monitoring, resulting in low testing efficiency, high costs, and difficulty in guaranteeing quality.
By preprocessing business test data, establishing business processes, conducting early warning scoring and resource demand prediction, optimizing resource scheduling using genetic algorithms, and introducing external influencing factors to improve the accuracy of resource demand prediction.
It enables efficient management of business test data, reasonable allocation of test resources, improved testing efficiency, reduced testing costs, and ensured the quality of server products.
Smart Images

Figure CN120950317B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a management method and electronic device for server testing operations. Background Technology
[0002] With the widespread adoption of emerging technologies such as 5G and the Internet of Things, server applications are becoming increasingly diverse, leading to ever-increasing performance demands on servers. This has highlighted the growing importance of server testing. Server testing involves numerous test projects, each containing a large number of test tasks and requiring specific test configurations and test cases. Furthermore, the testing process necessitates coordinating various resources, including test personnel, test materials, and test prototypes, making the entire testing process extremely complex.
[0003] Therefore, how to efficiently manage massive amounts of test data, rationally allocate limited test resources, and monitor the test process in real time are technical problems that need to be solved in related technologies. Summary of the Invention
[0004] In view of the above problems, this application provides a management method and electronic device for server testing services.
[0005] According to a first aspect of this application, a method for managing server testing services is provided, comprising: preprocessing business testing data collected from a server testing service system to obtain preprocessed data; defining business process elements and adding attribute information of the process elements to the preprocessed data according to preset process rules to obtain a business process; determining an early warning score for the business process based on the process elements and attribute information in the business process; when the early warning score triggers an early warning and a mismatch between business and resources is determined based on the business process, predicting resource requirements for the business process based on external influencing factors to obtain predicted resource requirements for the business process; and optimizing the scheduling of test resources required by the business process based on the predicted resource requirements for the business process to obtain a resource scheduling optimization strategy for the business process.
[0006] The second aspect of this application provides a management device for server testing services, comprising: a preprocessing module, an acquisition module, a determination module, a prediction module, and an optimization module. The preprocessing module is used to preprocess business test data collected from a server testing service system to obtain preprocessed data; the acquisition module is used to define business process elements and add attribute information to the preprocessed data according to preset process rules to obtain a business process; the determination module is used to determine a warning score for the business process based on the process elements and attribute information in the business process; the prediction module is used to predict the resource requirements of the business process based on external influencing factors when the warning score triggers a warning and a mismatch between business and resources is determined based on the business process, to obtain a predicted resource requirement for the business process; and the optimization module is used to optimize the scheduling of test resources required by the business process based on the predicted resource requirement for the business process, to obtain a resource scheduling optimization strategy for the business process.
[0007] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0008] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0009] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] According to the server testing service management method provided in this application, efficient management of business testing data and reasonable allocation of effective testing resources are achieved through preprocessing of business testing data, establishment of business processes, early warning, demand forecasting, and resource scheduling optimization. This helps improve testing efficiency, reduce testing costs, and ensure server product quality. Furthermore, introducing external influencing factors helps improve the accuracy of resource demand forecasting. Attached Figure Description
[0011] Figure 1 An application scenario diagram of a server testing service management method according to an embodiment of this application is shown;
[0012] Figure 2 A flowchart illustrating a method for managing server testing services according to an embodiment of this application is shown;
[0013] Figure 3A flowchart illustrating the preprocessing of business test data according to an embodiment of this application is shown;
[0014] Figure 4 A flowchart illustrating missing value processing of business test data according to an embodiment of this application is shown;
[0015] Figure 5 A flowchart illustrating the scheduling optimization of test resources for a business process according to an embodiment of this application is shown;
[0016] Figure 6 A schematic diagram of a management system for server testing services according to an embodiment of this application is shown;
[0017] Figure 7 A structural block diagram of a management device for server testing services according to an embodiment of this application is shown; and
[0018] Figure 8 A block diagram of an electronic device suitable for implementing a management method for server testing services, according to an embodiment of this application, is shown. Detailed Implementation
[0019] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0023] In the digital age, servers are the core pillars of information systems, and their performance, stability, and security directly determine the operational effectiveness of various applications. Therefore, as server applications become increasingly widespread and performance requirements continue to rise, server testing has become particularly crucial. Server testing encompasses the entire lifecycle, from product design verification to mass production quality inspection, involving multiple testing projects such as performance testing, compatibility testing, and reliability testing. Each testing project not only includes a large number of testing tasks but also requires specific test configurations and test cases as support. Furthermore, the testing process necessitates the coordination of various resources, making the entire testing process extremely complex.
[0024] The server testing business management and optimization involved in related technologies mainly include data management, business management and resource scheduling.
[0025] In terms of data management, traditional database technologies, such as relational databases, are primarily used to manage server test data, enabling data storage and basic query functions. However, traditional databases have limited capabilities for processing massive, multi-source, and heterogeneous test data. Data warehousing technology can also be used to manage server test data, integrating data from different sources through an ETL (Extract, Transform, Load) process. However, the ETL process is inefficient and struggles to handle real-time data changes.
[0026] Consequently, data management technologies suffer from insufficient data processing capabilities and low data collection efficiency, failing to meet real-time requirements. Furthermore, the lack of intelligent methods in the data ETL process prevents the full realization of data value, and a significant amount of information generated during testing is wasted, thus impacting subsequent analysis and decision-making.
[0027] In terms of business management, the system primarily employs self-developed test task management and automated task execution processes, which support task allocation and execution. However, it lacks sufficient support for complex testing processes such as iterative testing. Furthermore, it lacks the capability for real-time monitoring and dynamic adjustment of business processes, making it difficult to adapt to changes in testing operations.
[0028] Consequently, the lack of real-time monitoring of business processes in business management technologies leads to insufficient coordination between various stages, making it impossible to grasp test progress and quality in real time. This results in untimely problem detection, difficulty in dynamically adjusting and optimizing processes, and impacts test efficiency and quality.
[0029] In terms of resource scheduling, test resources mainly include test personnel, materials, and prototypes. The relevant technologies mainly adopt rule-based scheduling, such as first-come, first-served and shortest-job-first. These rules are simple and easy to use, but they do not fully consider the diversity of test resources and the complexity of test tasks, which can easily lead to low resource utilization and extended test cycles.
[0030] Consequently, resource scheduling technologies have not fully considered the diversity of test resources and the complexity of test tasks, and lack scientific allocation methods. As a result, test manpower, materials, prototypes and other resources are not used rationally, often resulting in some resources being excessively idle while others are severely insufficient, leading to increased test costs and extended project cycles.
[0031] In summary, how to efficiently manage massive amounts of test data, rationally allocate limited test resources, and monitor the test process in real time are technical problems that need to be solved in related technologies. To this end, the embodiments of this application provide a method for managing server testing operations, which can comprehensively, intelligently, and efficiently manage server testing operations, thereby improving testing efficiency, reducing testing costs, and ensuring server product quality.
[0032] Figure 1 An application scenario diagram of a server testing service management method according to an embodiment of this application is shown.
[0033] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0034] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0035] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0036] For example, server testing business systems are deployed on the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0037] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0038] For example, server 105 can collect business test data from first terminal device 101, second terminal device 102, and third terminal device 103, and preprocess the collected business test data to obtain preprocessed data. Then, according to preset process rules, business process elements are defined and attribute information of process elements is added to the preprocessed data to obtain the business process. Based on the process elements and attribute information in the business process, a warning score can be determined for the business process. When the warning score triggers a warning and a mismatch between business and resources is determined based on the business process, resource requirements for the business process can be predicted based on external influencing factors to obtain predicted resource requirements for the business process. This allows for the scheduling and optimization of the test resources required by the business process based on the predicted resource requirements, resulting in a resource scheduling optimization strategy for the business process.
[0039] It should be noted that the server testing service management method provided in this application embodiment can generally be executed by server 105. Correspondingly, the server testing service management device provided in this application embodiment can generally be located in server 105. The server testing service management method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the server testing service management device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0040] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0041] The following will be based on Figure 1 The described scene, through Figures 2-6 The management method for server testing services according to embodiments of this application will be described in detail.
[0042] Figure 2 A flowchart illustrating a method for managing server testing services according to an embodiment of this application is shown.
[0043] like Figure 2 As shown, the method 200 includes operations S210 to S250.
[0044] In operation S210, the business test data collected from the server test business system is preprocessed to obtain preprocessed data.
[0045] According to embodiments of this application, a server testing business system may include a test management system, a test execution system, a monitoring system, etc. Therefore, the server testing business system includes multiple data sources, and the collection of business test data is multi-source. Furthermore, each data source may include multiple types of data.
[0046] In one embodiment, data in the data source can be categorized based on its data structure type; that is, different acquisition methods can be used for data of different structure types in the data source. The structure type can include structured data, semi-structured data, and unstructured data.
[0047] Specifically, structured data can be obtained by direct database connection, semi-structured data can be obtained by XML (eXtensible Markup Language) / JSON (JavaScript Object Notation) parsing and schema mapping, and unstructured data can be obtained by text feature extraction and image OCR (Optical Character Recognition) mapping.
[0048] Because structured data allows us to determine its attributes and data range, we can also perform integrity checks on the structured data within the business test data before preprocessing it.
[0049] In one embodiment, during the data acquisition process, the integrity of the structured data is verified in real time. For example, the structured test task data must contain fields such as task ID and start time, and it is determined whether the test task data contains the corresponding fields. The accuracy of the structured data is also verified, such as whether the value range of the structured test data is within a reasonable range.
[0050] According to embodiments of this application, the integrity verification rules are not static and may be updated subsequently based on rule maintenance. Therefore, for data that fails an integrity verification once, it can be retried a preset number of times; that is, the data integrity verification can be retried up to a preset number of times. If the re-verification passes within the preset number of times, it can also indicate that the data integrity verification has passed; if the verification fails after the preset number of times, it indicates that the data integrity verification has failed.
[0051] Data that fails integrity checks can be placed in an exception database so that testers can easily check whether the data that failed integrity checks meets the requirements.
[0052] According to embodiments of this application, preprocessing may include outlier detection, missing value handling, and data transformation. Preprocessing business test data can improve data quality.
[0053] According to embodiments of this application, business test data can be periodically collected from the server test business system. Furthermore, since the importance of each piece of data in the business test data to the server test business varies, and the update speed of each piece of data in the server test business system also differs, the data collection frequency can be set according to the importance of the data and the speed of its update.
[0054] For example, for critical test task execution data, the collection frequency can be set to once every minute; for non-critical test configuration historical data, the collection frequency can be set to once every hour.
[0055] In operation S220, according to the preset process rules, the preprocessed data is used to define business process elements and add attribute information to the process elements to obtain the business process.
[0056] In one embodiment, preset process rules can be built based on the BPMN 2.0 (Business Process Model and Notation 2.0) standard to define detailed process elements based on the preset process rules and add attributes to the process elements.
[0057] Specifically, process elements include activity elements, gateway elements, and time elements. Activity elements include test requirement review, test case design, and test environment setup; gateway elements are used to handle process branches (such as deciding whether to proceed to regression testing or defect repair based on test results); event elements include timed events (such as test task deadline reminders) and message events (such as defect submission events).
[0058] Specifically, the attribute information added to each process element includes the role of the person in charge of the activity (test engineer, project manager, etc.), execution time constraints (such as the test case design activity must be completed within 3 days), input and output data (such as the output of the test case design activity being a test case document), etc.
[0059] According to the embodiments of this application, after defining process elements and adding attribute information of process elements to preprocessed data in accordance with preset process rules, a "standardized and executable business process model blueprint" will be obtained, which realizes the transformation of the fuzzy logic of "how to execute the test business" into a "digital process rule template" that the server can understand and monitor.
[0060] Therefore, the business process obtained by defining business process elements and adding attribute information to process elements from preprocessed data can include a visual process topology and structured process metadata.
[0061] Visualized process topology is a standard graphical flowchart that can clearly show the flow sequence, branch conditions (such as the "pass / fail" judgment of the gateway), and triggering events (such as timed reminders) of the "test requirements review → test case design → test execution → defect repair" stages. It serves as a "carrier" for the consensus reached by business and technical personnel.
[0062] The structured process metadata is a "rule database" bound to the flowchart, recording the specific attributes of each element. For example: Activity node: Test case design → Responsible person = Test engineer, execution time ≤ 3 days, output = Test case document; Event node: Task deadline reminder → Trigger time = 1 day before deadline, notification object = VM (Test Manager); Gateway node: Test result judgment → Condition 1 = All PASS (passed) → Enter "End", Condition 2 = Not all PASS → Enter "Defect Repair Verification".
[0063] According to embodiments of this application, when a business process is obtained, the business process can also be validated to verify the syntactic correctness (such as whether the connection between the gateway and the activity is reasonable) and semantic integrity (such as whether there are activities without assigned responsible persons), that is, to verify the process elements defined in the business process and the attribute information added.
[0064] In one embodiment, based on the rules set by the verification, all process elements defined in the business process can be verified, or only a part of them can be verified. For example, which process elements must be verified and which process elements can be skipped.
[0065] In operation S230, based on the process elements and attribute information in the business process, an early warning score for the business process is determined.
[0066] According to an embodiment of the present application, in the process of calculating the early warning score, the calculation of multiple indicators is involved. The process elements in the business process can be used for collecting the data required for indicator calculation, and the attribute information can be used to determine the indicator preset for the indicators. Therefore, the process elements and attribute information in the business process are the basis for calculating the early warning score.
[0067] In operation S240, when the early warning score triggers an early warning and it is determined based on the business process that there is a mismatch between the business and resources, based on the external influence factor, a resource demand prediction for the business process is performed to obtain the predicted resource demand for the business process.
[0068] According to an embodiment of the present application, the magnitude of the early warning score AlertScore is divided into four levels of early warning: AlertScore ≤ 0.2 is no early warning, 0.2 < AlertScore ≤ 0.5 is to trigger a blue early warning, 0.5 < AlertScore ≤ 0.8 is to trigger a yellow early warning, and AlertScore > 0.8 is to trigger a red early warning. Thus, when the early warning score is greater than 0.2, an early warning can be triggered.
[0069] According to an embodiment of the present application, the business process can also provide a "root cause location basis" for the early warning result. Specifically, after the early warning is triggered, it is necessary to trace back to the business process to find the reason. And when it is determined based on the business process that there is a mismatch between the business and resources, based on the external influence factor, a resource demand prediction for the business process is performed to obtain the predicted resource demand.
[0070] In operation S250, based on the predicted resource demand for the business process, the test resources required for the business process are scheduled and optimized to obtain a resource scheduling optimization strategy for the business process.
[0071] According to an embodiment of the present application, a genetic algorithm can be used to schedule and optimize the test resources required for the business process to obtain the best resource scheduling optimization strategy.
[0072] According to embodiments of this application, efficient management of business test data and reasonable allocation of effective test resources are achieved through preprocessing of business test data, establishment of business processes, early warning, demand forecasting, and resource scheduling optimization. This helps improve testing efficiency, reduce testing costs, and ensure server product quality. Furthermore, introducing external influencing factors helps improve the accuracy of resource demand forecasting.
[0073] Figure 3 A flowchart illustrating the preprocessing of business test data according to an embodiment of this application is shown.
[0074] like Figure 3 As shown, the method 300 includes operations S310 to S340.
[0075] Business test data includes data related to test tasks, test projects, and test resources. Before preprocessing the business test data, it is necessary to integrate the data, that is, to integrate the business test data into different dimension tables, such as test task table, test project dimension table, test resource dimension table, and test time dimension table.
[0076] In one embodiment, a star schema can be used to construct a data warehouse for business test data, that is, the business test data is divided into different data tables according to different dimensions. In this data warehouse, the test task table serves as the fact table, and the test task table is associated with test project dimension tables, test resource dimension tables, etc.
[0077] Because data is collected periodically, the data in each table can be processed in different ways according to the data attributes.
[0078] In one embodiment, for attributes whose data changes slowly over time, such as the skill level of testers, Type 2 processing can be used, which means storing the old data in the table into a history table and updating the original table with the new data; for attributes whose data does not change frequently, such as the name of the test project, Type 1 processing can be used, which means directly overwriting the old data in the table with the new data.
[0079] According to the embodiments of this application, only in the first collection cycle is it necessary to integrate the business test data to obtain tables of different dimensions. In subsequent collection cycles, it is only necessary to process the data in the tables based on the new data.
[0080] In operation S310, the test data of any test task in the business test data is used as a data point. The number of nearest neighbors of the data point is calculated based on the local density of the data point, the average local density of the dataset in which the data point is located, and the initial number of nearest neighbors.
[0081] The dataset includes test data from multiple test tasks within the business test data.
[0082] Based on the tables of different dimensions established during the above data integration process, since the feature dimensions of different tables are different and the data differences are large, the local density and average local density cannot be calculated across tables. The local density and average local density need to be calculated independently for each table. That is, a table of one dimension can be regarded as a dataset.
[0083] According to embodiments of this application, outlier detection can be performed on the business test data first. Specifically, test data for any test task in the business test data can be regarded as a data point, and the number of nearest neighbors of the data point can be calculated based on the local density of the data point, the average local density of the dataset in which the data point is located, and the initial number of nearest neighbors.
[0084] The nearest neighbor count of a data point can be used to determine whether the data point is in a sparse or dense region, thereby determining whether the data point is an outlier.
[0085] Specifically, the number of nearest neighbors k(X) of data point X can be calculated using the following formula (1).
[0086] k(X) = ×exp(- (X) / (1);
[0087] in, The initial nearest neighbor number, (X) represents the local density of data point X. This represents the average local density of the dataset containing the data points.
[0088] In one embodiment, the initial nearest neighbor number It is set according to needs. For example, the initial number of nearest neighbors. The value of is generally "half the square root of the total number of data points m in the dataset containing the data point". For example, when m=1000, the initial nearest neighbor number is... ≈16, this avoids the initial nearest neighbor count. Too small a value can lead to unstable local density calculations, while too large a value can cause neighboring points to include non-local points.
[0089] According to an embodiment of this application, the local density of data points is calculated. (X) and the average local density of the dataset containing the data points In the case of , we can substitute into the above formula (1) to calculate the number of nearest neighbors for each data point.
[0090] Sparse region (X) is small, exp( (X) / When k(X) is close to 1, k(X) is large; dense region (X) is large, exp( (X) / The decay is small, and k(X) is small.
[0091] In operation S320, the target data point farthest from the data point is determined from the first preset number of data points closest to the data point in the dataset. If the distance between the data point and the target data point is greater than a preset distance, the data point is determined to be an outlier and the outlier is removed.
[0092] The first preset number is the number of nearest neighbors of the data point.
[0093] According to an embodiment of this application, the preset distance is set as needed to determine whether a data point is an outlier.
[0094] In one embodiment, after calculating the number of nearest neighbors of a data point based on the above formula (1), the number of data points that are closest to the data point in the dataset that need to be determined from the dataset can be determined based on the number of nearest neighbors of the data point, i.e., the first preset number. Specifically, the first preset number of data points that are closest to the data point in the dataset can be determined by calculating the Euclidean distance between the data point and other data points in the dataset.
[0095] For example, if a data point has 8 nearest neighbors, you need to determine the 8 data points that are closest to the data point in the dataset.
[0096] According to an embodiment of this application, a target data point farthest from the data point is determined from a first preset number of data points, and it is determined whether the distance between the data point and the target data point is greater than a preset distance. If the distance between the data point and the target data point is greater than the preset distance, it indicates that the data point is relatively far from its nearest neighbor data points and is relatively isolated; therefore, the data point can be identified as an outlier. If the distance between the data point and the target data point is less than or equal to the preset distance, it indicates that the data point is not relatively isolated and therefore is not an outlier.
[0097] When operating S330, after outlier removal from the business test data, missing value processing is performed on the outlier-removed business test data based on the time decay factor to obtain business test data with missing values supplemented.
[0098] According to the embodiments of this application, since the business test data is collected periodically and involves time-related data, a time decay factor can be introduced during the missing value processing of the business test data so as to process the missing value of the business test data based on the influence of time on the missing value.
[0099] When operating S340, the business test data after missing value supplementation is transformed according to the standardized method for business test data after missing value supplementation to obtain preprocessed data.
[0100] According to embodiments of this application, for numerical data, standardization methods include min-max standardization and Z-score standardization; for categorical data, target coding can be used for standardization.
[0101] For numerical data, an appropriate standardization method can be automatically selected based on the data distribution characteristics of the numerical data.
[0102] The formula for Z-score standardization is shown in formula (2) below.
[0103] = (2);
[0104] in, The mean of the dataset. Let be the standard deviation of the dataset. For data to be standardized, This is the standardized data.
[0105] In one embodiment, the dataset used in the data transformation process may refer to the collection of all data values contained in the "feature field" currently undergoing standardization processing, rather than the entire data warehouse or data table. For example, if the numerical feature "actual execution time of test cases" is to be processed, then the "dataset" is the collection of the "actual execution time" values of all test cases.
[0106] The core logic of Z-score standardization is to convert the original data into "a multiple of the standard deviation relative to the mean", so that the standardized data follows a standard normal distribution with a mean of 0 and a standard deviation of 1. It is suitable for scenarios where the data is approximately normally distributed and there are a few extreme values.
[0107] In one embodiment, for a certain data x to be standardized, such as the defect pass rate of a certain test item being 85%; the mean μ of the dataset containing this feature, such as the average defect pass rate of all test items being 75%; and the standard deviation σ of the dataset, such as the standard deviation of the defect pass rate of all test items being 10%. Based on this, the standardized data x′ obtained through the above formula (2) is... =1, meaning that 85% of the original values are 1 standard deviation higher than the mean.
[0108] The formula for min-max standardization is shown in formula (3) below.
[0109] = (3);
[0110] in, The data is to be standardized, and min(x) and max(x) are the minimum and maximum values in the dataset, respectively. These are standardized data values.
[0111] Based on the above, for numerical data, if the data is approximately normally distributed and has a small number of extreme values, Z-score standardization should be selected; if the data distribution is non-normal, has no extreme values, and needs to be fixed to the 0-1 interval, min-max standardization should be selected.
[0112] For categorical data, target encoding is introduced. For example, for high cardinality categorical features (such as test case numbers), target encoding converts the categorical values into statistics of their corresponding target variables (such as test pass rates), reducing the curse of dimensionality.
[0113] According to embodiments of this application, preprocessing of business test data, specifically including outlier detection, missing value handling, and data transformation, aims to improve data quality and enhance the accuracy of subsequent calculations based on the business test data. Furthermore, in outlier detection, the accuracy is improved by calculating the nearest neighbors of each data point; and in missing value handling, the accuracy of identified missing values is improved by introducing a time decay factor to account for the impact of time on missing values.
[0114] According to an embodiment of this application, the local density of a data point and the average local density of the dataset containing the data point are obtained through the following operations: calculating the Euclidean distance between the data point and other data points in the dataset; determining the minimum first preset number of distances from the Euclidean distances between the data point and other data points, wherein the first preset number is the initial number of nearest neighbors; determining the local density of the data point based on the average of the first preset number of distances; and determining the average local density of the dataset containing the data point based on the local densities of multiple data points in the dataset.
[0115] According to embodiments of this application, the dataset containing the data point includes multiple data points, and each data point includes data with multiple feature dimensions. For example, a data point may involve data with multiple feature dimensions such as the execution time of test cases under a test task, the number of defects reported by the test cases, and memory usage. A data point X can be represented as X = ( , ,..., ), where p corresponds to the feature dimension of the data point. , , ..., These are the numerical values corresponding to the features of each dimension. Given that the feature dimensions of a data point are determined, the relationship between data point X and other data points in the dataset can be calculated using the following formula (4). The Euclidean distance d(X, between them) ).
[0116] d(X, )= (4);
[0117] According to embodiments of this application, the local density ρ(X) of a data point can be used to describe the "crowding" around a single data point. Local density The core logic of (X) is: the more nearest neighbors a data point X has, and the closer they are to it, then... The larger (X) is.
[0118] In one embodiment, the local density of data points can be calculated using the nearest neighbor average distance method.
[0119] Specifically, this nearest neighbor average distance method uses "X to its nearest neighbor" to calculate the distance between neighbors. The average distance between points is used as an inverse measure of density; that is, the smaller the distance, the greater the density. The specific steps are as follows: For any data point X in the dataset, calculate the Euclidean distance between X and all other data points in the dataset; select the smallest distance from all distances. The distance, i.e., X's Calculate the distances to the nearest neighbors; The average of the distances is denoted as . (X); Local density (X) is defined as the reciprocal of this average (ensuring "small distance → high density"), i.e., local density. (X)= .
[0120] For example, if X's =The distances to the 5 nearest neighbors are 2, 3, 4, 5, and 6 respectively, and the average distance is... (X)=4, then (X)= =0.25; if X's =The distances of the 5 nearest neighbors are 1, 1.5, 2, 2.5, and 3, and the average distance is... (X)=2, then (X)= =0.5 (higher density, which makes sense).
[0121] In another embodiment, the local density of data points can also be calculated using the nearest neighbor maximum distance method.
[0122] Specifically, this nearest neighbor maximum distance method uses "X to its nearest neighbor" to define the maximum distance between neighbors. The density is represented by the reciprocal of the distance to the farthest nearest neighbor, thus avoiding the influence of a single extreme nearest neighbor distance on the results.
[0123] The specific steps are as follows: Calculate the Euclidean distance between data point X and all other data points in the dataset, and select the one with the smallest distance. That distance; take this The maximum value among the distances is denoted as . Then the local density (X) is: (X)= .
[0124] In summary, the nearest neighbor average distance method is suitable for calculating the local density of data points in most cases; the nearest neighbor maximum distance method is suitable for cases where there are a few "outlier nearest neighbor points" in the dataset, that is, it is more suitable for cases with extreme values.
[0125] Average local density It is the "density benchmark" of the entire dataset, used to judge individual data points. Is (X) "above average" (dense area) or "below average" (sparse area)?
[0126] The specific calculation steps are as follows: Calculate the local density of all data points in the dataset using the method described above. ( ), ( ),...、 ( ), where m is the total number of data points in the dataset; calculate all The arithmetic mean of (X) is the average local density. : = .
[0127] For example, the dataset has 100 data points, and all the points The sum of (X) is 50, then = =0.5; If a data point X has (X) = 0.2 (< 0.5), indicating that X is in the sparse region; if (X) = 0.8 (>0.5), indicating that X is in a dense region.
[0128] According to embodiments of this application, local density (X) can be used to describe the "crowding" around a single data point, the average local density. It can be used to reflect the overall density level of the entire dataset. Based on this, the number of nearest neighbors of the data points can be calculated using the above formula (1) based on the local density of the data points and the average local density of the dataset, which helps to improve the accuracy of outlier detection.
[0129] Figure 4 A flowchart illustrating the missing value processing of business test data according to an embodiment of this application is shown.
[0130] like Figure 4 As shown, the method 400 includes operations S410 to S470.
[0131] According to embodiments of this application, the missing data in the business test data may include continuous missing data and discrete missing data. For continuous missing data, missing value processing can be performed through the following operations S410 to S470.
[0132] In operation S410, multiple missing locations were identified in the business test data after outlier removal.
[0133] According to embodiments of this application, after outlier detection and removal from the business test data, multiple missing locations can be identified in the outlier-removed business test data. These missing locations represent the positions where missing data exists.
[0134] In one embodiment, for missing parts in business test data, it is first necessary to determine whether the missing part belongs to the normal test plan. If so, the missing part does not need to be supplemented; otherwise, the missing part needs to be supplemented, and the missing part is identified as the missing location.
[0135] In operation S420, for any missing location among multiple missing locations, based on the time series, a second preset number of nearest neighbor data that are temporally adjacent to any missing location are determined from the business test data after outlier removal.
[0136] According to an embodiment of this application, the business test data contains time-series data, such as data for month Y, data for month Y+1, etc. From the business test data after outlier removal, a second preset number of nearest neighbor data points that are temporally adjacent to any missing position are determined. For example, if the missing position is data for month Y, then the nearest neighbor data for that missing position may include data for month Y-1, etc. The second preset number is set as needed.
[0137] In operation S430, for any nearest neighbor data in the second preset number of nearest neighbor data, calculate the Euclidean distance between any missing position and any nearest neighbor data.
[0138] According to an embodiment of this application, the Euclidean distance between the data corresponding to any missing position and any nearest neighbor data can be calculated using the above formula (4).
[0139] In operation S440, the time decay factor is calculated based on the time difference between any missing location and any nearest neighbor data, as well as the time decay constant.
[0140] According to an embodiment of this application, the time corresponding to any nearest neighbor data can be determined based on the time series, thereby the time difference between the time of the data corresponding to any missing position and any nearest neighbor data can be calculated.
[0141] In one embodiment, the time decay factor can characterize the negative value of the ratio of the time difference to the time decay constant. The time decay constant is set as needed. For example, the time decay constant can be set to 7 days.
[0142] In operation S450, the weight for any nearest neighbor data is calculated based on the Euclidean distance between any missing location and any nearest neighbor data, as well as the time decay factor.
[0143] According to embodiments of this application, the weights of nearest neighbor data It can be calculated using the following formula (5).
[0144] = ×exp(- (5);
[0145] in, The time difference between the i-th nearest neighbor and the current data point; T is the time decay constant; - This is the time decay factor.
[0146] Specifically, the time decay constant T can be set based on the impact of time difference on the weights. The time difference has a significant impact on the weights; a smaller time decay constant T indicates that the most recent data is given more weight, while a larger time decay constant T indicates that data with a longer time span is given more weight.
[0147] In operation S460, the missing value for any missing position is calculated based on the Euclidean distance and weight of each of the second preset number of nearest neighbor data.
[0148] According to an embodiment of this application, based on the Euclidean distance and weight of each of the second preset number of nearest neighbor data, the Euclidean distances of the second preset number of nearest neighbor data are weighted and summed, and the weighted summed value is determined as the missing value for any missing position.
[0149] In one embodiment, based on the above formula (5), it can be determined that the closer the neighboring data is in time to the missing position, the smaller the weight is assigned.
[0150] In operation S470, based on the missing values at multiple missing locations, the missing values in the business test data after outlier removal are filled in to obtain the business test data with missing values filled in.
[0151] According to an embodiment of this application, after calculating the missing values for multiple missing locations, the multiple missing locations in the business test data are supplemented with the corresponding missing values.
[0152] In one embodiment, for continuous missing data, the KNN (k-Nearest Neighbors) interpolation algorithm can be used with the introduction of a time decay factor to handle missing values in the data.
[0153] According to embodiments of this application, since business test data contains time-related data and is collected periodically, a time decay factor is introduced during the missing value processing of the business test data to account for the impact of time on missing values at the missing locations. Furthermore, during the calculation of missing values at the missing locations, due to the time decay factor, neighboring data that is further away from the missing location in time is given less weight. This makes the calculation of missing values more focused on neighboring data that is closer to the missing location in time, ensuring the accuracy of missing value calculation and improving the accuracy of missing value imputation.
[0154] For discrete missing data, a Bayesian network-based prediction imputation method can be used to predict the most likely value of the missing data by constructing a dependency model between discrete data.
[0155] In one embodiment, the dependencies between business test data are not so strong. The most likely scenario is that the execution of a certain test case depends on the execution result of another test case. Most test cases do not have many dependencies on each other, that is, test cases are independent of each other.
[0156] According to an embodiment of this application, a warning score for a business process is calculated based on process elements and attribute information in the business process, including: collecting indicator data for each indicator from the business process based on process elements; calculating the actual indicator value for each indicator based on the indicator data for each indicator and the indicator constraints determined by the attribute information; selecting multiple target indicators from the indicators, and calculating the warning score for the business process based on the actual indicator value, expected indicator value, weight, and historical trend coefficient of each of the multiple target indicators.
[0157] According to the embodiments of this application, it is recommended to have multiple levels of indicators. The first-level indicators may include multiple indicators such as test progress achievement rate and test quality pass rate; the second-level indicators include multiple indicators such as test task on-time completion rate and test case pass rate; and the third-level indicators include multiple indicators such as the execution time of a single test case and the defect repair time.
[0158] According to embodiments of this application, process elements defined in the business process can provide "collection points" for the data required for indicator calculation. Specifically, the calculation of the early warning score requires real-time data support, and the data collection must correspond to specific elements in the process. For example, to calculate the "on-time completion rate of test tasks," it is necessary to first determine "which activities belong to test tasks," and then collect the "actual completion time" of these activities (compared with the "planned time" in the attributes). That is, the indicator data required for calculating the indicators can be collected based on the process elements defined in the business process.
[0159] According to embodiments of this application, the attribute information added to process elements in the business process can determine the indicator constraints for the metrics. Specifically, the core of the metric is "measuring whether the standard is met", and the "standard of meeting the standard" comes from the attributes of the process elements. For example, the secondary indicator "on-time completion rate of test tasks" requires statistics on "the percentage of all tasks completed within the time constraint specified by the attribute". If the indicator constraint is determined by a certain task attribute to be "completed within 3 days", then exceeding the deadline will be counted as "not completed on time".
[0160] Therefore, based on the process elements defined in the business process and the attribute information added to the process elements, the business process defines "what steps, who does it, and how long it takes to complete", and the indicators are based on these definitions to "measure whether it is done well", so that the actual indicator values for each indicator can be calculated.
[0161] According to an embodiment of this application, the AlertScore for a business process can be calculated using the following formula (6).
[0162] AlertScore= ×( (6);
[0163] Where n is the number of target indicators. Let i be the weight of the i-th target indicator. The actual value of the target indicator. The expected value of the target indicator. This represents the historical trend coefficient.
[0164] Specifically, This refers to the actual result of a specific target metric calculated from the data actually collected within the current monitoring period. It reflects the actual operational status of the testing business process in terms of that metric.
[0165] Example of a primary indicator: If "Test Quality Pass Rate" is the currently calculated indicator i, and statistics show that the actual percentage of qualified test outputs is 92%, then... =92%.
[0166] Example of a secondary metric: If "test case pass rate" is metric i, and 880 out of 1000 test cases actually executed pass, then... =880 / 1000=88%.
[0167] Example of a Level 3 indicator: If "Defect Repair Time" is indicator i, and the actual repair time for a certain defect is 8 hours, then... =8 hours.
[0168] This refers to a pre-set standard or threshold that is expected to be achieved for a specific target metric. It serves as a benchmark for measuring whether the metric has been met and whether the process is functioning correctly, and is usually determined by business needs, project plans, or previous best practices.
[0169] Example of a primary indicator: If the project plan requires setting a target of 95% for the "test quality pass rate", then... =95%.
[0170] Example of a secondary indicator: If the business requirement dictates that the "test case pass rate" is no less than 90%, then... =90%.
[0171] Example of a Level 3 indicator: Based on the previous best performance, the "defect repair time" should be controlled within 6 hours. =6 hours.
[0172] Combination and The core logic of the above formula (6) is: calculate the degree to which the actual value of each target indicator deviates from the expected value, and then combine the indicator weight and historical trend coefficient to obtain the warning score.
[0173] In one embodiment, in formula (6) This represents the relative deviation rate of target indicator i. If the result is positive, it means that the actual indicator value exceeds the expected indicator value (this could be a negative deviation, such as defect repair time exceeding the time limit; or it could be a positive deviation, such as pass rate exceeding expectations, which needs to be combined with the indicator type definition); if the result is negative, it means that the actual indicator value does not reach the expected indicator value.
[0174] In formula (6) This is the historical trend coefficient, used to amplify the impact of "continuous deviations." If the indicator deviates from the expected value three consecutive times, the historical trend coefficient becomes 1.5, meaning that the deviation weight of the indicator has been increased, reflecting the emphasis on "persistent anomalies"; otherwise... =1.
[0175] In formula (6) The weights of the indicators are determined by the analytic hierarchy process and can reflect the importance of different indicators to the overall testing process (e.g., the weight of "test quality pass rate" is usually higher than that of "single test case execution time").
[0176] In formula (6), ∑ (summation) represents the summation of the "weighted deviation contribution" of all target indicators to obtain the comprehensive early warning score, and finally the early warning level is determined based on the score range.
[0177] The early warning score calculation based on multi-level indicators is a summation of indicator items, a calculation process for a specific third-level, second-level, or even first-level indicator, rather than a summation of levels. Specifically, the role of levels is "the basis for indicator classification and weight allocation," not "the dimension for summation."
[0178] In one embodiment, the relationship between hierarchy and indicator items is as follows: "History is a classification, and indicator items are calculation units." Multi-level indicators form a "tree structure": first-level indicators are "outcome indicators," second-level indicators are "process indicators," and third-level indicators are "atomic indicators." However, only third-level indicators are directly calculable "atomic items." Second-level and first-level indicators are usually "summaries or weighted results" of third-level indicators.
[0179] For example: Level 3 metrics (atomic items): Task 1 execution time, Task 2 execution time, Test Case 1 pass rate, Test Case 2 pass rate (each item has a clear definition) and The secondary indicator "On-time Completion Rate of Test Tasks" is calculated by summing the execution time of Task 1, Task 2, and other tertiary indicators (number of tasks completed on time / total number of tasks); the primary indicator "Test Progress Achievement Rate" is calculated by weighting the secondary indicators such as "On-time Completion Rate of Tasks" and "Test Case Pass Rate".
[0180] The calculation logic for the early warning score is: "Select key indicator items and sum them according to their weights." The early warning score is not calculated for all levels of indicators, but rather selects "core monitoring indicator items" (which may be third-level atomic items or second-level summary items) from each level. That is, multiple target indicators are selected from multiple indicators, weighted, and then summed to obtain the early warning score.
[0181] For example: If we select three levels of indicators for calculation, assuming we choose three key items: Task A execution time ( =0.4), Execution time of Task B ( =0.3), test case pass rate ( =0.3), then the summation is performed on these three indicator items. ×[( - ) / × Add them together; if you choose to calculate using secondary indicators, let's assume you choose two secondary indicators: on-time completion rate of tasks ( =0.6), test case pass rate ( =0.4), then the summation is the addition of the calculation results of these two second-level terms.
[0182] According to embodiments of this application, since the business process defines process elements for collecting data required for indicator calculation and attribute elements for determining indicator constraints required for indicator calculation, the indicator values of each indicator can be calculated based on the business process. Furthermore, several key target indicators are selected from multiple indicators for calculating the early warning score of the business process. Since the early warning score is obtained by weighted summation of multiple target indicators, the essence of summation is "comprehensive multi-dimensional risk." A single indicator anomaly (such as a task timeout) may have a limited impact, but multiple target indicators simultaneously anomaly (task timeout + low test case pass rate) poses a higher risk. Through weighted summation, scattered "local anomalies" can be transformed into a "global early warning score," more accurately reflecting the health status of the business.
[0183] According to an embodiment of this application, resource demand forecasting for a business process is performed based on external influencing factors to obtain the predicted resource demand for the business process. This includes: obtaining the resource demand for any test resource required by the business process for each of the d time periods prior to the current time period; calculating an external influencing factor item for any test resource based on the number of external influencing factors for any test resource, the value of each external influencing factor, and the linear coefficient of the external influencing factors; and calculating the predicted demand for any test resource in the current time period based on the resource demand for any test resource and the external influencing factor item for each of the d time periods prior to the current time period.
[0184] Where d is an integer greater than or equal to 1, and d is determined based on the difference order.
[0185] According to embodiments of this application, a comprehensive resource information database is constructed based on the test resources collected from business test data. This database stores static information about the test resources (such as the educational background and professional skills certifications of test personnel; the model and hardware configuration of test prototypes; and the specifications and suppliers of test materials) and dynamic information (such as the current workload and available time periods of test personnel; the current status of test prototypes (idle, occupied, under maintenance); and the inventory quantity and consumption rate of test materials). The test resources in the database can be updated at a preset frequency to ensure their timeliness.
[0186] According to embodiments of this application, an extended ARIMA-X model can be constructed using an ARIMA (AutoRegressive Integrated Moving Average) model and by introducing external influencing factors (such as the complexity coefficient of the test project and holiday factors) for demand forecasting. The demand for test resources can be predicted using the following formula (7).
[0187] (B) (7);
[0188] in, Let j be the value of the j-th external influence factor at time t. Let q be its coefficient, and q be the number of external influencing factors.
[0189] In formula (7), This is the random error term, representing "random fluctuations in the ARIMA-X model that cannot be explained by historical data and external influencing factors." Specifically, in time series forecasting, even considering historical trends (through the autoregressive and moving average components of ARIMA) and external influencing factors (…), the random fluctuations cannot be explained by historical data and external influencing factors. The actual data may still contain some unpredictable random disturbances (such as sudden adjustments to test tasks, temporary resource changes, etc.). It is precisely this quantitative representation of such disturbances that It is a correction for random fluctuations (random errors). These random disturbances typically satisfy the conditions of "mean is 0, variance is constant, and there is no correlation at any time," meaning that they have no fixed pattern and cannot be predicted in advance.
[0190] ARIMA-X is an extended model that introduces external influence factors (X represents external variables) into the traditional ARIMA model. Its core is through linear terms... This quantifies the impact of external factors on the prediction target and belongs to the time series regression model.
[0191] In formula (7), (B) , These are the core operators of the ARIMA model: autoregressive, differencing, and moving average.
[0192] Specifically, B stands for Backshift Operator, representing "time lag," used to simplify the expression of "past values" in time series. For any time t, the variable... The operation rule for B is: B = (Lapdown of 1 period, i.e., the value at the time preceding time t); B² =B(B )= (Lag by 2 periods); and so on. = (Lag k periods).
[0193] For example, if If B represents "the number of testers needed in month t", then B It refers to the demand in month t-1 (previous month), B² It refers to the demand in t-2 months (the month before last).
[0194] In formula (7), (B) and Both are polynomials with B as the variable, and their essence is to integrate the influence of historical data on the current value through lag operators.
[0195] Let t be the target variable for prediction at time t (e.g., the demand for testers in month t), representing "the actual value of the core indicator to be predicted at time t", which is the "dependent variable" of the time series.
[0196] For example, if t represents "the 3rd month", then This indicates the "testing staff requirement for the third month"; if categorized by staff skill level, then... It can specifically refer to "the demand for senior test engineers at time t".
[0197] d represents the differencing order. By performing "difference operations", the "trend" or "periodicity" of the time series can be eliminated, making the data stationary (ARIMA models require that the input data be stationary).
[0198] This is the mathematical expression for the difference operator: when d=1, (1-B) = - , representing "first-order difference" (current value minus previous period value), used to eliminate linear trends; when d=2, (1-B)² =(1-B)( - )= -2 + , which represents "second-order difference", is used to eliminate quadratic trends.
[0199] The value of d is usually 0, 1, or 2 (d=0 indicates that the data is stationary and no differencing is needed).
[0200] For example, if the demand for testers increases linearly over time (has a trend), the "growth trend" can be transformed into a smooth fluctuation through the first difference of d=1, thus meeting the input requirements of the ARIMA model.
[0201] For external impact factors, where, Let j be the value of the factor at time t. The linear coefficient of this factor.
[0202] According to the embodiments of this application, resource demand prediction using the above formula (7) is applicable to scenarios where "external influencing factors are clear and have a linear relationship with the target," such as: the number of project testing requirements ( ) and personnel needs ( ) are positively correlated, and can be obtained through Quantify "how much the personnel demand increases for every 10 additional units required"; holiday factor ( (e.g., 1 for holidays, 0 for non-holidays) can be accessed through Quantify the extent to which holidays reduce resource demand.
[0203] According to an embodiment of this application, based on the difference order determined in the above formula (7), the resource requirements for any test resource for each of the d time periods before the current time period can be obtained from the resource information database.
[0204] According to the embodiments of this application, the predicted demand for any test resource in the current period can be calculated using the above formula (7) based on the resource demand and external influencing factor items for each of the d time periods prior to the current time period. Since the external influencing factor is introduced in formula (7), the demand prediction is more accurate, which is beneficial to improving the effectiveness of subsequent resource scheduling optimization.
[0205] According to embodiments of this application, the sliding window technique can be used to train and predict the aforementioned ARIMA-X model. The window size is set to 6 months, and the model parameters are updated monthly to improve prediction accuracy.
[0206] For forecasting the demand for testers, considering the differences in demand at different skill levels (such as senior test engineers and junior test engineers), separate forecasting models can be established.
[0207] Based on the differences in demand between different skill levels, all core parameters of the model's final output (including autoregressive coefficients) can be indirectly affected by changing the model's data input. Moving average coefficient Difference order d, external influence factor coefficient ).
[0208] Specifically: When modeling separately for "senior test engineers" and "junior test engineers," the core difference between the two models lies in predicting the target variable. The definition and the corresponding dataset are completely different.
[0209] Model 1 (Senior Test Engineer): This represents the demand for senior test engineers at time t. The data consists of historical demand data for senior engineers at various points in time, along with corresponding external influencing factors (such as the proportion of highly complex projects, holidays, etc.).
[0210] Model 2 (Junior Test Engineer): This represents the demand for junior test engineers at time t. The data consists of historical demand data for junior engineers at various points in time, along with corresponding external influencing factors (such as the proportion of basic functional test projects, holidays, etc.).
[0211] Because the distribution, trends, and correlations with external factors of the data differ fundamentally (i.e., reflecting "differences in demand"), the solutions obtained during the training and prediction processes... (Autoregressive coefficient) (Moving average coefficient), d (difference order) The (external factor coefficients) will all be different.
[0212] For example, the demand for senior engineers may be related to the "percentage of highly complex projects" ( Strongly correlated, therefore the corresponding model The coefficient of the complexity factor will be larger; the needs of junior engineers may be more significantly affected by the "proportion of basic functional testing projects," and the corresponding factor in their model will be larger. It will stand out more.
[0213] When there are differences in requirements for the same testing resources, such as the requirements for testers, different testers have different skill levels. Therefore, it is necessary to establish multiple prediction models to distinguish between different requirements (such as different skill levels). The core of establishing independent prediction models is to adapt to the essential characteristics of the differences in requirements and avoid the decrease in prediction accuracy caused by mixing different models.
[0214] Take the differences in demand for different skill levels as an example.
[0215] The demand patterns for different skill levels are completely different, meaning that the trends, periodicity, and fluctuations in resource demand at different skill levels vary significantly over time. For example, the demand for senior test engineers is usually strongly correlated with the "launch cycle of highly complex projects" and may exhibit quarterly fluctuations; the demand for junior test engineers may be related to the "frequency of basic function iterations" and exhibit monthly fluctuations (such as routine function testing in the middle of each month).
[0216] The correlation strength between different skill levels and "external influencing factors" varies. That is, the core advantage of the ARIMA-X model is the introduction of external influencing factors, but the correlation between different skill requirements and factors varies greatly.
[0217] External influencing factors (such as project complexity, holidays, and version iteration type) have completely different weights in terms of the demand for different skills: the impact of "project complexity coefficient" on the demand for senior engineers ( The impact on junior engineers is far greater than the impact on senior engineers; the "new employee training cycle" only affects the demand for junior engineers and has almost no impact on the demand for senior engineers. If separate predictive models are not built, the model will force the use of a single set of methods. The (external factor coefficients) are adapted to two types of needs, resulting in insufficient weight estimation of strongly correlated factors and redundant weight estimation of weakly correlated factors, ultimately reducing prediction accuracy.
[0218] To avoid "averaging error" and meet the needs of refined scheduling, the goal of intelligent resource scheduling is to "precisely match supply and demand." "Skill level" is an important dimension of resource supply. If a prediction model is not established separately and only the "demand for test engineers" is predicted, the specific number of senior and junior engineers cannot be broken down during subsequent scheduling. This may lead to "a surplus of senior engineers and a shortage of junior engineers," resulting in resource waste or demand gaps.
[0219] Therefore, establishing separate prediction models can directly output the precise demand for each skill level, providing a direct basis for subsequent "resource allocation on demand" and enabling refined scheduling.
[0220] According to an embodiment of this application, based on the predicted resource requirements for a business process, the scheduling optimization of test resources required by the business process is performed to obtain a resource scheduling optimization strategy for the business process, including: determining multiple objective functions based on the predicted resource requirements for the business process; initializing genetic parameters and encoding chromosomes to generate an initial population; determining a fitness function according to the weights of the multiple objective functions; optimizing the scheduling of test resources for the business process according to the fitness function and the initial population to generate a target population; and determining a resource scheduling optimization strategy for the business process based on the global optimal solution of the target population.
[0221] According to embodiments of this application, based on the predicted resource requirements for a business process, it is possible to determine what conditions the current test resource scheduling needs to meet, thereby determining multiple objective functions for the business process.
[0222] In one embodiment, the constructed multi-objective optimization model, i.e., multiple objective functions, can be represented by the following formulas (8), (9) and (10).
[0223] Min (8);
[0224] Min (9);
[0225] Max (10);
[0226] Where h is the number of test tasks and r is the number of resources. The time for resource j to execute task i. The cost of executing task i for resource j. Let the skill matching degree between resource j and task i be denoted as . The variable is a 0-1 variable (1 indicates that resource j is allocated to task i, and 0 indicates that it is not allocated), used to represent the probability that resource j is allocated to task i.
[0227] In one embodiment, formula (8) is used to minimize the total test time. Formula (9) is used to minimize the total test resource cost. Formula (10) is used to maximize the resource skill matching degree. .
[0228] According to the embodiments of this application, constraints are also set for multiple objective functions, including resource capacity constraints (such as testers working no more than 8 hours a day), task dependency constraints (such as task B can only start after task A is completed), and resource conflict constraints (such as a test prototype cannot execute two tasks at the same time).
[0229] According to embodiments of this application, genetic parameters may include chromosome length, population size, preset crossover rate, and preset mutation rate. Two-dimensional real-number encoding can be used to encode the chromosomes to generate the initial population.
[0230] According to an embodiment of this application, the fitness function can be determined based on the weights of each objective function. The fitness function F can be determined by the following formula (11).
[0231] F= + + (11);
[0232] in, , , These are the weights corresponding to the objective function.
[0233] Therefore, based on the above formula (11), the weighted summation method is used to transform the multi-objective into a single objective, that is, into a fitness function.
[0234] According to embodiments of this application, based on a fitness function, the scheduling optimization of test resources for a business process can be achieved by optimizing individuals in the initial population. The globally optimal solution of the final generated target population can then be determined as the resource scheduling optimization strategy for the business process.
[0235] According to embodiments of this application, multiple objective functions are determined based on predicted resource demand, so that the resource scheduling optimization strategy determined based on the objective functions can meet the demand and ensure the effectiveness of the resource scheduling optimization strategy.
[0236] According to an embodiment of this application, the genetic parameters are initialized and the chromosomes are encoded to generate an initial population, including: initializing the chromosome length, the number of individuals in the population, the preset crossover rate, and the preset mutation rate; encoding the positions in the chromosomes based on the chromosome length to obtain the encoded chromosomes; and generating the initial population based on the number of individuals in the population and the encoded chromosomes.
[0237] Each position in the encoded chromosome represents the proportion of resources allocated to tasks; the number of chromosomes in the initial population is determined based on the number of individuals in the population.
[0238] According to embodiments of this application, chromosome length, population size, preset crossover rate, and preset mutation rate can be initialized.
[0239] In one embodiment, the population size can be set to 100, the preset crossover rate can be set to 0.8, and the preset mutation rate can be set to 0.1. The chromosome length, population size, preset crossover rate, and preset mutation rate are set as needed.
[0240] According to an embodiment of this application, the positions in the chromosome can be encoded based on the chromosome length, such that each position in the chromosome can represent the proportion of resource j allocated to task i.
[0241] In one embodiment, the chromosome length can be set to h×r, where each gene (location) in the chromosome... This can represent the proportion of resource j allocated to task i. The chromosome length is h×r (number of test tasks × number of resources), which represents the matching relationship between corresponding tasks and resources, and is the "matching matrix dimension" between tasks and resources. Here, h is the number of test tasks (e.g., task 1, task 2, ..., task h), and r is the number of resources (e.g., resource 1, resource 2, ..., resource r).
[0242] According to embodiments of this application, a corresponding number of chromosomes can be encoded based on the number of individuals in the population, thereby generating an initial population based on the same number of chromosomes as the number of individuals in the population.
[0243] According to an embodiment of this application, by encoding chromosomes so that each position in the chromosome represents the allocation ratio of a certain resource to a certain task, the optimal allocation scheme of resources to tasks can be obtained by subsequently optimizing individuals in the initial population, i.e., resource scheduling optimization strategy.
[0244] Figure 5 A flowchart illustrating the scheduling optimization of test resources for a business process according to an embodiment of this application is shown.
[0245] like Figure 5 As shown, the method 500 includes operations S510 to S570.
[0246] In operation S510, based on the preset crossover rate and fitness function, the second preset number of chromosomes with the largest fitness function value are selected from the first population.
[0247] In the case of a cycle of 1, the first population is the initial population.
[0248] According to an embodiment of this application, the number of chromosomes to be selected from the first population, i.e., the second preset number, can be determined based on a preset crossover rate.
[0249] According to embodiments of this application, the fitness function value of each chromosome in the first population can be calculated based on the fitness function. Therefore, based on the fitness function values of each chromosome, a second predetermined number of chromosomes with the highest fitness function value can be selected from the first population.
[0250] In one embodiment, a tournament selection method can be used to select a second preset number of chromosomes from a first population.
[0251] In operation S520, a crossover operation is performed on any two chromosomes from the second preset number of chromosomes to generate a second population.
[0252] According to embodiments of this application, a second preset number of chromosomes can be paired to obtain multiple pairs of chromosomes, each pair containing two chromosomes. Thus, a crossover operation can be performed on the two chromosomes in each pair. In one embodiment, simulated binary crossover can be used to crossover the two chromosomes in each pair to generate new individuals, thereby creating a second population.
[0253] In operation S530, the target chromosome is selected from the second population, and the individual optimal solution and the global optimal solution of the target chromosome are determined.
[0254] According to an embodiment of this application, a target chromosome can be randomly selected from a second population.
[0255] According to the embodiments of this application, the determination of the individual optimal solution (pbest) and the global optimal solution (gbest) of chromosomes in the population is based on the fitness function as the core evaluation. The determination depends on the above formula (11). The larger the fitness function value corresponding to the chromosome, the better the resource allocation scheme reflected by the chromosome.
[0256] According to embodiments of this application, the individual optimal solution pbest is the historical best solution during the iteration process of each chromosome, recording the scheduling scheme corresponding to the highest fitness ever achieved by the chromosome. The individual optimal solution is determined as follows: when generating the initial population, the initial scheduling scheme of each chromosome is its initial pbest (at this time, there is no historical data, and the current state is optimal); after each iteration, the fitness function value of the current scheme of the chromosome is calculated and compared with the fitness function value of its own historical pbest. If the current fitness is higher, pbest is updated to the current scheme; otherwise, pbest remains unchanged.
[0257] According to an embodiment of this application, the global optimal solution (gbest) is the historical best solution during the iteration process of the entire population (the set of all chromosomes), recording the scheduling scheme corresponding to the highest fitness ever achieved by all chromosomes. The global optimal solution is determined as follows: the individual with the highest fitness is selected from the initial population as the initial gbest; after each iteration, the fitness function values of all chromosomes in the current generation are summarized, along with the pbest fitness function values of all chromosomes, and the maximum value is selected. If this value exceeds the fitness of the historical gbest, then gbest is updated to the corresponding scheme; otherwise, gbest remains unchanged.
[0258] In operation S540, the target allocation ratio for any target position is determined based on the allocation ratio of any target position in the target chromosome and the allocation ratio for any target position in the individual optimal solution and the global optimal solution.
[0259] According to embodiments of this application, the target allocation ratio for any target location can be determined using the following formulas (12) and (13). That is, the allocation ratio of any target position in the next round.
[0260] = + (12);
[0261] = + (13);
[0262] in, Let be the coefficient, and pbest be the individual optimal solution. This represents the proportion of any target position in the individual optimal solution for the target chromosome, where gbest is the global optimal solution. This represents the allocation ratio of any target position in the global optimal solution of the target chromosome, where t represents the current round and t+1 represents the next round. This indicates the allocation ratio for any target position in the current round. This indicates the speed in the next round.
[0263] When operating S550, based on the target allocation ratio and preset mutation rate, the target chromosome is mutated to generate a third population.
[0264] According to embodiments of this application, given the target allocation ratios at each position in the target chromosome, the target chromosome can be mutated based on a preset mutation rate. Specifically, the allocation ratios at corresponding positions are replaced with the target allocation ratios.
[0265] In the above formula (12), Promoting individuals to adjust towards their own historical best; : It uses the "currently found global optimal solution" to guide individuals toward the global optimal direction of the population, allowing the current individual (resource allocation ratio) to... )Towards To move closer together and avoid getting trapped in local optima. This includes the allocation ratio. This represents the probability that resource j is allocated to task i.
[0266] According to embodiments of this application, the objective function in It is the final decision variable, strictly a 0-1 variable (1 means resource j is fully allocated to task i, 0 means not allocated), used to define the final state of the "optimal scheduling scheme", because the actual resource allocation is "either / or" (e.g. a test engineer cannot "partially allocate" to two manual tasks at the same time).
[0267] And in formula (12) This is an intermediate variable in the search process, representing the proportion of resource j allocated to task i (e.g., 0.3 means 30% of the resource capacity is allocated to that task), and is not the final decision. This is because the hybrid genetic algorithm needs to retain flexibility in the iterative search phase; forcing... The search can only take 0 or 1, which greatly limits the algorithm's search space (it can only choose from discrete "yes / no" options), making it prone to getting trapped in local optima; adopting a "proportional" approach... (e.g., 0.2, 0.5, 0.8) allows the algorithm to explore the optimal solution more precisely in the continuous space. After the search is close to the optimal solution, it is then transformed into the final 0-1 decision by "threshold truncation" (e.g., take 1 if the ratio is ≥0.5, otherwise take 0).
[0268] speed Its function is to guide proportional types Adjust towards a better direction. The velocities in formulas (12) and (13) above... It is a "magnitude and direction indicator" for proportional adjustment, used to optimize resource allocation ratios during iterations. Its core function is to drive... It tends to approach the individual optimal (pbest) and the global optimal (gbest). If the allocation ratio of resource j to task i is higher in pbest or gbest (e.g. =0.6, current =0.2), then For positive, to promote Increase (e.g., from 0.2 to 0.4); if the proportion is lower in pbest or gbest (e.g. If the corresponding ratio is 0.1, then... Negative, driving Decrease it (e.g., from 0.2 to 0.1). This provides a "probabilistic basis" for the final 0-1 decision.
[0269] After the iteration is complete, the proportional type The magnitude of directly reflects the "rationality of allocating resource j to task i". The higher the proportion, the better the allocation scheme is in multi-objective optimization (time, cost, matching degree), and it is more likely to be the globally optimal solution when finally transformed into a 0-1 decision.
[0270] In operation S560, determine whether the cycle number meets the termination condition.
[0271] According to an embodiment of this application, if the cycle does not meet the termination condition, the above operations S510 to S550 are performed; if the cycle meets the termination condition, operation S570 is performed.
[0272] The number of cycles is set as needed.
[0273] In operation S570, the third population obtained when the termination condition is met in the cycle rounds will be identified as the target population.
[0274] According to embodiments of this application, during the optimization of the initial population, crossover and mutation are performed on the chromosomes in the population to bring them closer to the global optimum and avoid getting trapped in local optima. Therefore, the target population generated through the above operations S510 to S570 can be used to determine resource scheduling optimization strategies.
[0275] Figure 6 A schematic diagram of a management system for server testing services according to an embodiment of this application is shown.
[0276] like Figure 6 As shown, the server testing service management system 600 includes a data integration and preprocessing module 610, a business process management and monitoring module 620, and an intelligent resource scheduling module 630. The server testing service management system 600 is used to execute management methods for server testing services.
[0277] Specifically, the data integration and preprocessing module 610 is used to perform the above operation S210, the business process management and monitoring module 620 is used to perform the above operations S220 and S230, and the intelligent resource scheduling module 630 is used to perform the above operations S240 and S250.
[0278] For the data integration and preprocessing module 610, an adaptive nearest neighbor density outlier detection algorithm and a missing value handling method that integrates domain knowledge are proposed, which significantly improves the processing accuracy of server test data. An automatic conversion mechanism based on data features is designed to achieve efficient data adaptation and solve the problem of low data processing quality in related technologies.
[0279] For the business process management and monitoring module 620, a multi-level indicator system and a weighted early warning scoring model were constructed based on the intelligent optimization system for business processes with dynamic early warning, which enabled accurate early warning of the test business process; an automatic identification and intelligent adjustment algorithm for process bottlenecks was developed, which enabled dynamic optimization of business processes and broke through the limitations of lagging monitoring and passive adjustment of related technical processes.
[0280] For the intelligent resource scheduling module 630, an intelligent resource scheduling method integrating external factors is proposed, and the ARIMA-X demand prediction model is proposed to improve the accuracy of resource demand prediction. An improved hybrid genetic algorithm is designed to achieve the global optimal allocation of resources under multi-objective constraints, solving the problem of low resource scheduling efficiency in related technologies.
[0281] Through the collaborative work of the three core modules mentioned above, the server testing business management method of this application realizes intelligent management and risk control of server testing business, comprehensively solves the problems existing in related technologies, and achieves the goals of improving testing efficiency, reducing testing costs, and ensuring server product quality.
[0282] In one embodiment, a hybrid data storage architecture can be built based on the Spring and Spring Boot frameworks, using Java, Spring Data JPA, and Elasticsearch (an open-source search engine), to establish a management system 600 for server testing services.
[0283] According to embodiments of this application, the data integration and preprocessing module 610 is used for the collection, cleaning (outlier detection and missing value handling), transformation, and integration of business test data, providing high-quality data support for subsequent analysis and decision-making. Based on the Spring Boot framework, the data integration and preprocessing module 610 is designed with scalable data collection, data processing, and data provision services to achieve unified management of multi-source heterogeneous data.
[0284] In one embodiment, the data integration and preprocessing module 610 in the server testing business management system 600 is designed with a unified data acquisition framework based on Spring Boot. It adopts a plug-in architecture, designs a multi-source data interface adaptation layer, supports data acquisition from multiple data sources such as the test management system, test execution system, and monitoring system, and provides a flexible configuration mechanism that can perform customized acquisition according to the characteristics of different data sources.
[0285] In one embodiment, based on the constructed hybrid data storage architecture, the data integration and preprocessing module 610 also designs a metadata-driven data integration method. Specifically, the data integration and preprocessing module 610 integrates the cleaned and transformed data into a relational database and a search engine by defining a unified data model and mapping rules, forming a unified data service interface for other modules to call. For example, structured data is stored in the relational database, and unstructured data and full-text indexes are stored in Elasticsearch, enabling efficient data querying and analysis.
[0286] According to an embodiment of this application, the business process management and monitoring module 620 is used for refined management and real-time monitoring of the entire process of server testing business, ensuring that the testing business is carried out efficiently and orderly.
[0287] In one embodiment, the business process management and monitoring module 620 can analyze preprocessed data under different collection cycles using CDC (Change Data Capture) technology, thereby capturing real-time changes in data in the server test business system (such as a test task status changing from "not started" to "in execution"), and transmitting the changed data to the monitoring center in real time. The business process management and monitoring module 620 also features a multi-dimensional visualization dashboard, including a project-level dashboard (displaying the overall progress, pass rate, etc. of each test project), a task-level dashboard (displaying the detailed execution status of each test task), and a resource-level dashboard (displaying the resource usage status). The business process management and monitoring module 620 can also use 3D bar charts to display defect trends for different test projects, use heatmaps to display test case coverage, and use dynamic flowcharts to display the execution status of business processes in real time.
[0288] Therefore, the business process management monitoring module 620 can be used to see the issues and progress that the management needs to pay attention to in the server testing business.
[0289] In another embodiment, if an early warning score triggers an alert and it is determined based on the business process that there is no mismatch between business and resources (i.e., it is determined based on the business process that there are other problems in the business testing process), the business process management and monitoring module 620 can also optimize and adjust the process. Specifically, the business process management and monitoring module 620 can use a development process bottleneck analysis algorithm to identify bottleneck activities by calculating indicators such as the waiting time percentage and resource utilization rate of each activity. For example, if the waiting time percentage of a certain testing activity exceeds 40%, it is determined to be a bottleneck activity. Automatic process adjustment is achieved based on the workflow engine. When a bottleneck activity is identified, adjustment strategies are automatically triggered (such as increasing resource allocation for the activity or splitting the activity into multiple sub-activities for parallel execution). For cases of test task delays, the system automatically calculates a catching-up path, compressing the time of non-critical path activities to ensure the overall project schedule.
[0290] According to an embodiment of this application, the intelligent resource scheduling module 630 is used to achieve optimal configuration of test resources, improve resource utilization, and reduce test costs when an early warning score triggers an early warning and a mismatch between business and resources is determined based on the business process.
[0291] In one embodiment, the intelligent resource scheduling module 630 can dynamically adjust resources after obtaining a resource scheduling optimization strategy for the business process.
[0292] Specifically, the intelligent resource scheduling module 630 establishes a resource scheduling monitoring mechanism to track in real time the deviation between the actual resource usage and the scheduling plan in the resource scheduling optimization strategy (e.g., a tester's actual workload exceeds the planned workload in the scheduling plan by 10%). The intelligent resource scheduling module 630 also includes dynamic adjustment trigger conditions. When the deviation exceeds a threshold (e.g., 15%) or an unexpected event occurs (e.g., a tester taking leave, equipment failure), the resource rescheduling process is triggered. The intelligent resource scheduling module 630 uses an incremental optimization algorithm for dynamic resource adjustment, re-optimizing only the affected tasks and resources, reducing computational load and improving adjustment efficiency.
[0293] Based on the above, this application, through the coordinated operation of the three modules, brings significant improvements to server testing services in multiple dimensions, as detailed below:
[0294] At the data management level, the intelligent integration and preprocessing mechanism for multi-source heterogeneous data significantly improves data quality. The improved outlier detection algorithm increases outlier identification accuracy to 92%, and the KNN interpolation method with a time decay factor controls missing value processing errors to within 5%. A unified data service interface improves data access efficiency, provides reliable data support for decision-making, and reduces test rework rates caused by inaccurate data.
[0295] In terms of business process management, standardized process templates based on BPMN 2.0 improve the standardization of business testing processes by 50%, avoiding efficiency losses caused by chaotic processes. The combination of real-time monitoring dashboards and reinforcement learning decision-makers reduces the response time to process anomalies from 2 hours to 15 minutes, improves the on-time completion rate of test tasks, and significantly shortens the average test cycle.
[0296] In terms of resource scheduling, the optimization effect is significant. The ARIMA-X model reduces the resource demand prediction error to 8%, and the improved genetic algorithm increases resource utilization by 35%. The dynamic scheduling mechanism enables flexible resource allocation, greatly reduces manpower idle time, improves prototype turnover efficiency, and reduces overall testing costs.
[0297] Overall, the server testing management method proposed in this application realizes the transformation of testing operations from passive response to proactive control, improves comprehensive testing efficiency, reduces resource costs, and provides strong support for the large-scale and refined management of server testing operations.
[0298] Based on the aforementioned management method for server testing services, this application also provides a management device for server testing services. The following will be combined with... Figure 7 The device is described in detail.
[0299] Figure 7 A structural block diagram of a server testing service management device according to an embodiment of this application is shown.
[0300] like Figure 7 As shown, the server testing service management device 700 of this embodiment includes a preprocessing module 710, an acquisition module 720, a determination module 730, a prediction module 740, and an optimization module 750.
[0301] The preprocessing module 710 is used to preprocess the business test data collected from the server test business system to obtain preprocessed data. In one embodiment, the preprocessing module 710 can be used to perform the operation S210 described above, which will not be repeated here.
[0302] The obtaining module 720 is used to define business process elements and add attribute information of process elements to the preprocessed data according to preset process rules to obtain the business process. In one embodiment, the obtaining module 720 can be used to perform the operation S220 described above, which will not be repeated here.
[0303] The determination module 730 is used to determine a warning score for the business process based on the process elements and attribute information in the business process. In one embodiment, the determination module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0304] The prediction module 740 is used to predict resource requirements for the business process based on external influencing factors when an early warning score triggers an early warning and a mismatch between business and resources is determined based on the business process. In one embodiment, the prediction module 740 can be used to execute the operation S240 described above, which will not be repeated here.
[0305] The optimization module 750 is used to optimize the scheduling of test resources required by the business process based on the predicted resource requirements for the business process, thereby obtaining a resource scheduling optimization strategy for the business process. In one embodiment, the optimization module 750 can be used to execute the operation S250 described above, which will not be repeated here.
[0306] According to an embodiment of this application, the preprocessing module 710 includes a first calculation submodule, an outlier removal submodule, a missing value processing submodule, and a data transformation submodule.
[0307] The first calculation submodule is used to determine the target data point farthest from the data point from a first preset number of data points closest to the data point in the dataset, and to determine the data point as an outlier and remove the outlier if the distance between the data point and the target data point is greater than a preset distance; wherein, the first preset number is the number of nearest neighbors of the data point.
[0308] The outlier removal submodule is used to compare the number of nearest neighbors of a data point with a preset nearest neighbor threshold. If the comparison result indicates that the number of nearest neighbors of a data point is greater than the preset nearest neighbor threshold, the data point is determined to be an outlier and the outlier is removed.
[0309] The missing value processing submodule is used to process missing values in the business test data after outlier removal, based on the time decay factor, to obtain business test data with missing values filled in.
[0310] The data transformation submodule is used to transform the business test data after missing values are filled in according to the standardized method for business test data after missing values are filled in, so as to obtain preprocessed data.
[0311] According to an embodiment of this application, the first calculation submodule includes a first calculation unit, a first determination unit, a second determination unit, and a third determination unit.
[0312] The first calculation unit is used to calculate the Euclidean distance between a data point and other data points in the dataset.
[0313] The first determining unit is used to determine a minimum first preset number of distances from the Euclidean distances between data points and other data points, wherein the first preset number is the initial number of nearest neighbors.
[0314] The second determining unit is used to determine the local density of data points based on the average value of a first preset number of distances.
[0315] The third determining unit is used to determine the average local density of the dataset to which the data point is located based on the local density of each of the multiple data points in the dataset.
[0316] According to an embodiment of this application, the missing value processing submodule includes a fourth determining unit, a fifth determining unit, a second calculation unit, a third calculation unit, a fourth calculation unit, a fifth calculation unit, and an obtaining unit.
[0317] The fourth determination unit is used to determine multiple missing locations in the business test data after outlier removal.
[0318] The fifth determining unit is used to determine, based on time series, a second preset number of nearest neighbor data that are temporally adjacent to any missing location among multiple missing locations in the business test data after outlier removal.
[0319] The second calculation unit is used to calculate the Euclidean distance between any missing position and any nearest neighbor data for any of the second preset number of nearest neighbor data.
[0320] The third calculation unit is used to calculate the time decay factor based on the time difference between any missing location and any nearest neighbor data and the time decay constant.
[0321] The fourth calculation unit is used to calculate the weight for any nearest neighbor data based on the Euclidean distance between any missing position and any nearest neighbor data and the time decay factor.
[0322] The fifth calculation unit is used to calculate the missing value for any missing position based on the Euclidean distance and weight of each of the second preset number of nearest neighbor data.
[0323] The acquisition unit is used to correct missing values in the business test data after outlier removal based on the missing values at multiple missing locations, so as to obtain business test data with missing values.
[0324] According to an embodiment of this application, the determining module 730 includes a data acquisition submodule, a second calculation submodule, and a third calculation submodule.
[0325] The data collection submodule is used to collect indicator data for various metrics from business processes based on process elements.
[0326] The second calculation submodule is used to calculate the actual indicator value for each indicator based on the indicator data for each indicator and the indicator constraints determined by the attribute information.
[0327] The third calculation submodule is used to select multiple target indicators from various indicators, and calculate the early warning score for the business process based on the actual indicator value, expected indicator value, weight and historical trend coefficient of each of the multiple target indicators.
[0328] According to an embodiment of this application, the prediction module 740 includes an acquisition submodule, a fourth calculation submodule, and a fifth calculation submodule.
[0329] The acquisition submodule is used to obtain the resource requirements for any test resource required by the business process for each of the d time periods prior to the current time period, where d is an integer greater than or equal to 1, and d is determined according to the difference order.
[0330] The fourth calculation submodule is used to calculate the external influence factor item for any test resource based on the number of external influence factors, the value of each external influence factor, and the linear coefficient of the external influence factors.
[0331] The fifth calculation submodule is used to calculate the predicted demand for any test resource in the current period based on the resource demand and external influencing factors for each of the d previous periods.
[0332] According to an embodiment of this application, the optimization module 750 includes a first determining submodule, a first generating submodule, a second determining submodule, a second generating submodule, and a third determining submodule.
[0333] The first determination submodule is used to determine multiple objective functions based on the predicted resource requirements for business processes.
[0334] The first generation submodule is used to initialize genetic parameters and encode chromosomes to generate the initial population.
[0335] The second determination submodule is used to determine the fitness function based on the weights of the multiple objective functions.
[0336] The second generation submodule is used to schedule and optimize the test resources of the business process based on the fitness function and the initial population, and generate the target population.
[0337] The third determination submodule is used to determine the resource scheduling optimization strategy for the business process based on the global optimal solution of the target population.
[0338] According to an embodiment of this application, the first generation submodule includes an initialization unit, an encoding unit, and a first generation unit.
[0339] The initialization unit is used to initialize chromosome length, population size, preset crossover rate, and preset mutation rate.
[0340] The encoding unit is used to encode positions in the chromosome based on the chromosome length to obtain an encoded chromosome, wherein each position in the encoded chromosome represents the allocation ratio of resources to tasks.
[0341] The first generation unit is used to generate an initial population based on the number of individuals in the population and the encoded chromosomes, wherein the number of chromosomes in the initial population is determined based on the number of individuals in the population.
[0342] According to an embodiment of this application, the second generation submodule includes a selection unit, a second generation unit, a sixth determination unit, a seventh determination unit, a third generation unit, and an eighth determination unit.
[0343] The selection unit is used to select a second preset number of chromosomes with the largest fitness function value from the first population when the cycle number does not meet the termination condition, based on the preset crossover rate and fitness function. In the case of cycle number 1, the first population is the initial population.
[0344] The second generation unit is used to perform a crossover operation on any two chromosomes from the second preset number of chromosomes to generate a second population.
[0345] The sixth determining unit is used to select the target chromosome from the second population and determine the individual optimal solution and the global optimal solution of the target chromosome.
[0346] The seventh determining unit is used to determine the target allocation ratio for any target position based on the allocation ratio of any target position in the target chromosome and the allocation ratio for any target position in the individual optimal solution and the global optimal solution.
[0347] The third generation unit is used to mutate the target chromosome based on the target allocation ratio and the preset mutation rate to generate a third population.
[0348] The eighth determining unit is used to determine the third population obtained when the termination condition is met in the cycle rounds as the target population.
[0349] According to embodiments of this application, any multiple modules among the preprocessing module 710, obtaining module 720, determining module 730, predicting module 740, and optimizing module 750 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the preprocessing module 710, obtaining module 720, determining module 730, predicting module 740, and optimizing module 750 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the preprocessing module 710, the obtaining module 720, the determining module 730, the predicting module 740, and the optimizing module 750 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0350] Figure 8 A block diagram of an electronic device suitable for implementing a management method for server testing services, according to an embodiment of this application, is shown.
[0351] like Figure 8 As shown, an electronic device 800 according to an embodiment of this application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0352] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0353] According to embodiments of this application, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0354] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0355] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 described above.
[0356] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.
[0357] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0358] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0359] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0360] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0361] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0362] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0363] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A management method for server testing services, characterized in that, The method includes: Preprocess the business test data collected from the server test business system to obtain preprocessed data; According to the preset process rules, the preprocessed data is used to define business process elements and add attribute information to the process elements to obtain the business process. Based on the process elements and attribute information in the business process, a warning score is determined for the business process, including: collecting indicator data for each indicator from the business process based on the process elements; calculating the actual indicator value for each indicator based on the indicator data for each indicator and the indicator constraints determined by the attribute information; selecting multiple target indicators from the indicators, and calculating the warning score for the business process based on the actual indicator value, expected indicator value, weight, and historical trend coefficient of each of the multiple target indicators. When the warning score triggers a warning and it is determined based on the business process that there is a mismatch between business and resources, the resource demand of the business process is predicted based on external influencing factors to obtain the predicted resource demand for the business process. Based on the predicted resource requirements for the business process, the test resources required by the business process are scheduled and optimized to obtain a resource scheduling optimization strategy for the business process.
2. The method according to claim 1, characterized in that, The process of preprocessing the business test data collected from the server test business system to obtain preprocessed data includes: Using the test data of any test task in the business test data as a data point, the number of nearest neighbors of the data point is calculated based on the local density of the data point, the average local density of the dataset in which the data point is located, and the initial number of nearest neighbors; wherein, the dataset includes the test data of multiple test tasks in the business test data; From a first preset number of data points in the dataset that are closest to the data point, determine the target data point that is farthest from the data point. If the distance between the data point and the target data point is greater than a preset distance, determine the data point as an outlier and remove the outlier. The first preset number is the number of nearest neighbors of the data point. After outlier removal from the business test data, missing value processing is performed on the outlier-removed business test data based on the time decay factor to obtain business test data with missing values supplemented. The preprocessed data is obtained by performing data transformation on the business test data after the missing values are filled in according to the standardized method for the business test data after the missing values are filled in.
3. The method according to claim 2, characterized in that, The local density of the data point and the average local density of the dataset to which the data point belongs are obtained through the following operations: Calculate the Euclidean distance between the data point and other data points in the dataset; Determine the minimum first preset number of distances from the Euclidean distances between the data points and other data points, where the first preset number is the initial number of nearest neighbors; The local density of the data points is determined based on the average value of the first preset number of distances; Based on the local density of each of the multiple data points in the dataset, the average local density of the dataset to which the data points are located is determined.
4. The method according to claim 2, characterized in that, The process of processing missing values in the business test data after outlier removal based on the time decay factor to obtain business test data with missing values filled in includes: Identify multiple missing locations in the business test data after outlier removal; For any missing location among the multiple missing locations, based on the time series, a second preset number of nearest neighbor data that are temporally adjacent to any missing location are determined from the business test data after outlier removal. For any one of the second preset number of nearest neighbor data, calculate the Euclidean distance between any missing position and any one of the nearest neighbor data; The time decay factor is calculated based on the time difference between any missing location and any nearest neighbor data and the time decay constant. The weight for any nearest neighbor is calculated based on the Euclidean distance between any missing location and any nearest neighbor data and the time decay factor. Based on the Euclidean distance and weight of each of the second preset number of nearest neighbor data, the missing value for any missing position is calculated; Based on the missing values at each of the multiple missing locations, the missing values in the business test data after the outlier removal are supplemented to obtain the business test data with the missing values supplemented.
5. The method according to claim 1, characterized in that, The method of predicting resource requirements for the business process based on external influencing factors, resulting in predicted resource requirements for the business process, includes: any test resources required for the business process. Obtain the resource requirements for any test resource for each of the d time periods prior to the current time period, where d is an integer greater than or equal to 1, and d is determined based on the difference order; Based on the number of external influence factors for any test resource, the value of each external influence factor, and the linear coefficient of the external influence factors, the external influence factor item for any test resource is calculated. Based on the resource requirements for any test resource in the previous d time periods and the external influencing factor, the predicted demand for any test resource in the current time period is calculated.
6. The method according to claim 1, characterized in that, The step of optimizing the scheduling of test resources required by the business process based on the predicted resource requirements for the business process, to obtain a resource scheduling optimization strategy for the business process, includes: Based on the predicted resource requirements for the business process, several objective functions are determined; Initialize genetic parameters and encode chromosomes to generate an initial population; The fitness function is determined based on the weights of the multiple objective functions. Based on the fitness function and the initial population, the test resources of the business process are scheduled and optimized to generate the target population; Based on the global optimal solution of the target population, a resource scheduling optimization strategy for the business process is determined.
7. The method according to claim 6, characterized in that, The initialization of genetic parameters and encoding of chromosomes to generate an initial population includes: Initialize chromosome length, population size, preset crossover rate, and preset mutation rate; Based on the chromosome length, the positions in the chromosome are encoded to obtain an encoded chromosome, wherein each position in the encoded chromosome represents the allocation ratio of resources to tasks. The initial population is generated based on the number of individuals in the population and the encoded chromosomes, wherein the number of chromosomes in the initial population is determined based on the number of individuals in the population.
8. The method according to claim 7, characterized in that, The step of scheduling and optimizing the test resources of the business process based on the fitness function and the initial population to generate a target population includes repeatedly performing the following operations until the termination condition is met: If it is determined that the termination condition is not met in the current cycle round. Based on the preset crossover rate and the fitness function, select a second preset number of chromosomes with the largest fitness function value from the first population, wherein, in the case of 1 cycle, the first population is the initial population; Perform a crossover operation on any two chromosomes from the second preset number of chromosomes to generate a second population; Select the target chromosome from the second population, and determine the individual optimal solution and the global optimal solution of the target chromosome; Based on the allocation ratio of any target position in the target chromosome and the allocation ratio of any target position in the individual optimal solution and the global optimal solution, determine the target allocation ratio for any target position. Based on the target allocation ratio and the preset mutation rate, the target chromosome is mutated to generate a third population; The third population obtained when the termination condition is met in the cycle is determined as the target population.
9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Business system test method, device and equipment
CN111625458A
Enterprise whole process service resource scheduling optimization method and system based on big data
CN118674123A