Workload prediction method and device, equipment, storage medium and product
By processing and reducing the dimensionality of variables in software engineering projects, and utilizing intelligent optimization algorithms and random forest models, the problem of low accuracy in workload prediction in existing technologies has been solved, achieving higher-precision workload prediction.
Patent Information
- Application Number
- CN202410561996.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies have low accuracy in predicting workload in software engineering projects. Traditional function point methods are highly subjective, while deep learning models cannot be effectively trained due to a lack of sufficient historical data, resulting in poor prediction accuracy.
By acquiring the variables of the target software engineering project, handling outliers and missing variables, a scree plot is constructed, the number of common factors is determined using an intelligent optimization algorithm, a loading coefficient matrix is constructed and coordinate transformation is performed, and after normalization, it is input into a pre-trained random forest model for prediction.
It improves the accuracy of workload prediction, avoids the subjectivity of traditional methods and the poor interpretability of deep learning models, and achieves more accurate workload prediction.
Smart Images

Figure CN120930836A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of workload assessment, and in particular to a workload prediction method, apparatus, device, storage medium, and product. Background Technology
[0002] Existing digital engineering workload prediction methods can be roughly divided into two categories: (1) workload assessment methods based on function point method, which measure the scale of software by quantifying system functions from the user's perspective; and (2) workload assessment methods based on neural networks, which use deep learning technology to predict the workload of software engineering.
[0003] Because traditional function point evaluation methods are highly subjective and have low prediction accuracy, often only user requirements documents are available in the early stages of a project, and a complete software system specification document is lacking. Furthermore, neural networks require a large amount of sample data for training, and historical workload evaluation data is often scarce, causing the trained model to fail to converge to a satisfactory level. This results in poor interpretability and low prediction accuracy for deep learning models.
[0004] Therefore, there is an urgent need for a workload prediction method that can improve the accuracy of workload prediction in software engineering. Summary of the Invention
[0005] The purpose of the embodiments in this specification is to provide a workload prediction method, apparatus, device, storage medium, and product to improve the accuracy of workload prediction in software engineering.
[0006] To achieve the above objectives, in one aspect, embodiments of this specification provide a workload prediction method for at least one target software engineering project, including:
[0007] Obtain all variables of the target software engineering project, which are used to characterize the basic situation of the target software engineering project;
[0008] The variables are processed by handling outlier and missing variables to obtain the processed variables;
[0009] The processed variables are tested to obtain test values;
[0010] When the test value meets the set standard, a scree map is established based on the processed variables;
[0011] The scree map is processed using an intelligent optimization algorithm to obtain the number of common factors in the processed variables;
[0012] The loading coefficient matrix is determined based on the number of processed variables and common factors.
[0013] The processed variables are subjected to coordinate transformation based on the load coefficient matrix to obtain the values of the common factors;
[0014] The values of the common factors are normalized to obtain the standard factors of the target software engineering project.
[0015] The standard factors are input into a pre-trained random forest model to predict the workload of the target software engineering project.
[0016] Preferably, the step of processing the scree map using an intelligent optimization algorithm to obtain the number of common factors in the processed variables further includes:
[0017] The following steps are iterated N times, where N is a set number of iterations:
[0018] Randomly select any number from a set range;
[0019] When the random number is the horizontal coordinate value in the gravel map, the vertical coordinate value corresponding to the point in the gravel map that matches the horizontal coordinate value is obtained;
[0020] Calculate the product of the x-coordinate value and the y-coordinate value;
[0021] After the loop iteration ends, the random number corresponding to the smallest of all product values obtained from N loop iterations is taken as the number of common factors.
[0022] Preferably, determining the loading coefficient matrix based on the number of processed variables and common factors further includes:
[0023] Based on any one of the principal component method, maximum likelihood method, and principal axis factor method, analyze the number of variables and common factors after processing to determine the loading coefficient matrix.
[0024] Preferably, the step of performing coordinate transformation on the processed variables based on the loading coefficient matrix to obtain the values of the common factors further includes:
[0025] The processed variables are orthogonally rotated based on the loading coefficient matrix to obtain the values of the common factors.
[0026] Preferably, the training process of the random forest model includes:
[0027] Obtain all known variables from multiple known software engineering projects, where the workload of each known software engineering project is the known workload.
[0028] The known variables are processed by handling missing and outlier variables to obtain the processed known variables;
[0029] The processed known variables are tested to obtain known test values;
[0030] When the test value meets the set standard, a known gravel map is established based on the processed known variables;
[0031] The known scree map is processed using an intelligent optimization algorithm to obtain the number of known common factors among the known variables after processing.
[0032] Based on the number of known variables and known common factors after processing, determine the known loading coefficient matrix;
[0033] Based on the known load coefficient matrix, the processed known variables are subjected to coordinate transformation to obtain the values of the known common factors;
[0034] The values of the known common factors are normalized to obtain the known standard factors of the known software engineering project;
[0035] The known standard factors of the multiple known software engineering projects are divided into training set and test set;
[0036] The training set is input into the random forest model to predict the predicted workload of the known software engineering project;
[0037] The predicted workload is compared with the known workload, and the random forest model is optimized and trained to obtain a well-trained random forest model.
[0038] The random forest model is tested using the test set to obtain the performance of the trained random forest model.
[0039] Preferably, the hyperparameter tuning method for the random forest model includes:
[0040] Set the value range for each hyperparameter;
[0041] Set the initial mean squared error loss function to positive infinity, and set each hyperparameter in the initial hyperparameter set to a random value;
[0042] The following steps are iterated M times, where M is a preset number of iterations;
[0043] The hyperparameter set is obtained by randomly selecting values from the range of each hyperparameter.
[0044] The random forest model is trained using the hyperparameter set and the training set to obtain the trained random forest model;
[0045] The trained random forest model is subjected to multi-fold cross-validation using the training set to obtain the mean squared error loss function.
[0046] If the mean squared error loss function is less than the initial mean squared error loss function, then the mean squared error loss function is assigned to the initial mean squared error loss function, and the hyperparameter set is assigned to the initial hyperparameter set.
[0047] If the mean squared error loss function is not less than the initial mean squared error loss function, then continue to the next iteration;
[0048] After the loop iteration is completed, the initial hyperparameter set is output.
[0049] On the other hand, embodiments of this specification provide a workload prediction apparatus for at least one target software engineering project, the apparatus comprising:
[0050] The acquisition module is used to acquire all variables of the target software engineering project, and the variables are used to characterize the basic situation of the target software engineering project.
[0051] The variable processing module is used to handle abnormal variables and missing variables to obtain the processed variables;
[0052] The testing module is used to test the processed variables and obtain test values;
[0053] The scree plot creation module is used to create a scree plot based on the processed variables when the test value meets the set standard.
[0054] The scree graph processing module is used to process the scree graph using an intelligent optimization algorithm to obtain the number of common factors in the processed variables.
[0055] The determination module is used to determine the loading coefficient matrix based on the number of processed variables and common factors;
[0056] The transformation module is used to perform coordinate transformation on the processed variables according to the load coefficient matrix to obtain the values of the common factors.
[0057] The normalization module is used to normalize the values of the common factors to obtain the standard factors of the target software engineering project.
[0058] The prediction module is used to input the standard factors into a pre-trained random forest model to predict the workload of the target software engineering project.
[0059] In another aspect, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the computer program, when executed by the processor, performs instructions of any of the methods described above.
[0060] In another aspect, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of a computer device, performs instructions according to any one of the methods described above.
[0061] In another aspect, embodiments of this specification provide a computer program product that, when run by the processor of a computer device, executes instructions for any of the methods described above.
[0062] As can be seen from the technical solutions provided in the embodiments of this specification above, these embodiments enable the acquisition of all variables of the target software engineering project. After processing the variables, dimensionality reduction is performed. Specifically, an intelligent optimization algorithm is used to process the scree graph to obtain the number of common factors, and the loading coefficient matrix is further determined. This yields the values of the common factors after dimensionality reduction. The values of the common factors are then normalized to obtain standard factors, which are input into a pre-trained random forest model to predict the workload. This method improves the accuracy of workload prediction by avoiding the problems of strong subjectivity and poor model interpretability in traditional evaluation methods.
[0063] To make the above and other objects, features and advantages of this specification more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 A flowchart illustrating a workload prediction method provided in an embodiment of this specification is shown.
[0066] Figure 2 A diagram of crushed stone provided in the embodiments of this specification is shown;
[0067] Figure 3 This document illustrates a flowchart of an embodiment of the process of using an intelligent optimization algorithm to process a scree map and obtain the number of common factors.
[0068] Figure 4 A flowchart illustrating the training process of the random forest model provided in the embodiments of this specification is shown.
[0069] Figure 5A flowchart illustrating the hyperparameter tuning method for the random forest model provided in the embodiments of this specification is shown.
[0070] Figure 6 This specification shows a schematic diagram of the module structure of a workload prediction device provided in an embodiment.
[0071] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this specification is shown.
[0072] Explanation of symbols in the attached drawings:
[0073] 100. Acquisition Module;
[0074] 200. Variable processing module;
[0075] 300. Inspection module;
[0076] 400. Gravel Map Creation Module;
[0077] 500. Gravel Map Processing Module;
[0078] 600. Determine the module;
[0079] 700, Transformation Module;
[0080] 800, Normalization module;
[0081] 900. Prediction module;
[0082] 702. Computer equipment;
[0083] 704, Processor;
[0084] 706. Memory;
[0085] 708. Drive mechanism;
[0086] 710. Input / Output Module;
[0087] 712. Input devices;
[0088] 714. Output devices;
[0089] 716. Presentation equipment;
[0090] 718. Graphical User Interface;
[0091] 720. Network interface;
[0092] 722. Communication link;
[0093] 724. Communication bus. Detailed Implementation
[0094] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the embodiments of this specification.
[0095] Existing digital engineering workload prediction methods can be roughly divided into two categories: (1) workload assessment methods based on function point method, which measure the scale of software by quantifying system functions from the user's perspective; and (2) workload assessment methods based on neural networks, which use deep learning technology to predict the workload of software engineering.
[0096] Because traditional function point evaluation methods are highly subjective and have low prediction accuracy, often only user requirements documents are available in the early stages of a project, and a complete software system specification document is lacking. Furthermore, neural networks require a large amount of sample data for training, and historical workload evaluation data is often scarce, causing the trained model to fail to converge to a satisfactory level. This results in poor interpretability and low prediction accuracy for deep learning models.
[0097] To address the aforementioned issues, this specification provides a workload prediction method through its embodiments. Figure 1 This is a flowchart illustrating a workload prediction method provided in an embodiment of this specification. This specification provides the operational steps of the method described in the embodiment or flowchart, but based on conventional or non-inventive labor, it may include more or fewer operational steps. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the methods shown in the embodiment or drawings can be executed sequentially or in parallel.
[0098] It should be noted that the terms "first," "second," etc., in the description, claims, and accompanying drawings of the embodiments in this specification are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0099] Reference Figure 1This specification discloses a workload prediction method for at least one target software engineering project, comprising:
[0100] S101: Obtain all variables of the target software engineering project, the variables being used to characterize the basic situation of the target software engineering project;
[0101] S102: Perform outlier and missing variable processing on the variables to obtain the processed variables;
[0102] S103: Test the processed variables to obtain test values;
[0103] S104: When the test value meets the set standard, a scree map is established based on the processed variables;
[0104] S105: The scree map is processed using an intelligent optimization algorithm to obtain the number of common factors in the processed variables;
[0105] S106: Determine the loading coefficient matrix based on the number of processed variables and common factors;
[0106] S107: Perform coordinate transformation on the processed variables according to the loading coefficient matrix to obtain the values of the common factors;
[0107] S108: Normalize the values of the common factors to obtain the standard factors of the target software engineering project;
[0108] S109: Input the standard factors into the pre-trained random forest model to predict the workload of the target software engineering project.
[0109] The basic information of the target software engineering project includes the number of developers, their programming skills, their analytical skills, the development cycle, memory requirements, etc. Each of these basic information items corresponds to a variable, and all variables of the target software engineering project can be obtained.
[0110] For variables, there may be missing and outlier variables. It is necessary to further remove outlier variables and impute missing variables. The box plot method can be used to check for outlier variables; if outlier variables are found, they are removed, resulting in missing variables. Whether the missing variables are due to removal or were originally present, they need to be imputed. Random forest regression can be used to impute missing variables. Each target software engineering project corresponds to processed variables. Subsequent steps S103-S108, such as verifying the processed variables and building scree plots, are based on all processed variables corresponding to all target software engineering projects.
[0111] After handling outliers and missing variables to obtain the processed variables, they can be tested, specifically using the KMO test or Bartlett's test of sphericity. The purpose is to test the correlation between the processed variables. Only if a correlation exists can subsequent steps be performed. Specifically, the KMO test value or Bartlett's test of sphericity value corresponding to the processed variables can be calculated. When the KMO test value is greater than its corresponding threshold, or the Bartlett's test of sphericity value is less than its corresponding threshold, the test value is considered to meet the set standard, that is, there is a correlation between the processed variables, and subsequent steps can be performed.
[0112] Reference Figure 2 When the test value meets the set criteria, a scree plot is constructed based on the processed variables. The horizontal axis of the scree plot ranges from [0, a], representing the number of processed variables corresponding to each target software engineering project. For example, each target software engineering project has a total of 16 processed variables, with the horizontal axis ranging from [0, 16]. The vertical axis ranges from [0, b], representing the magnitude of the eigenvalues. By reading the scree plot constructed based on the processed variables, it can be seen that as the number of processed variables increases, the eigenvalues become smaller. Smaller eigenvalues indicate that the newly added variables contribute less and are less important. Therefore, variables with small contributions are ignored through the scree plot.
[0113] To reduce the dimensionality of the processed variables and improve the prediction accuracy of the subsequent random forest model, an intelligent optimization algorithm can be used to process the scree graph to obtain the number of common factors in the processed variables. Common factors are those variables with high correlation that are grouped into the same category, and each category of variables becomes a common factor. A small number of common factors can be used to reflect most of the information of the target software engineering project.
[0114] Among them, reference Figure 3 The step of processing the scree map using an intelligent optimization algorithm to obtain the number of common factors in the processed variables further includes:
[0115] S201: Iterate N times using the following steps S202-S204, where N is a set number of iterations:
[0116] S202: Randomly select any random number from the set range;
[0117] S203: When the random number is the horizontal coordinate value in the gravel map, obtain the vertical coordinate value corresponding to the point in the gravel map that matches the horizontal coordinate value;
[0118] S204: Calculate the product of the horizontal coordinate value and the vertical coordinate value;
[0119] S205: After the loop iteration ends, the random number corresponding to the smallest of all product values obtained from N loop iterations is taken as the number of common factors.
[0120] The number of iterations (N) can be set according to actual conditions. The range of the number is defined as the range of the horizontal coordinates in the scree plot, representing the number of processed variables corresponding to each target software engineering project. After randomly selecting any number within the set range, this random number should be the horizontal coordinate value in the scree plot. In the scree plot, there is exactly one point matching this horizontal coordinate value. The vertical coordinate value of this point can be determined by reading the graph. Then, the product of the horizontal and vertical coordinate values is calculated. Each iteration yields a product value. Finally, after the iteration ends, the random number (i.e., the horizontal coordinate value) corresponding to the smallest product value is the number of common factors.
[0121] Furthermore, determining the loading coefficient matrix based on the number of processed variables and common factors includes:
[0122] Based on any one of the principal component method, maximum likelihood method, and principal axis factor method, analyze the number of variables and common factors after processing to determine the loading coefficient matrix.
[0123] Furthermore, the step of performing coordinate transformation on the processed variables based on the loading coefficient matrix to obtain the values of the common factors includes:
[0124] The processed variables are orthogonally rotated based on the loading coefficient matrix to obtain the values of the common factors.
[0125] Suppose that each target software engineering project has 16 processed variables before orthogonal rotation. These 16 processed variables include the number of developers, their programming skills, their analytical skills, development cycle, memory requirements, etc. Suppose that there are 6 common factors. That is, we need to reduce the dimensionality of these 16 processed variables to 6 common factors. The values of the 6 common factors cover the values of the above 16 processed variables. The specific dimensionality reduction needs to be achieved through orthogonal rotation.
[0126] It should be noted that the loading coefficient matrix has 16 rows, corresponding to 16 processed variables, and 6 columns, corresponding to 6 common factors. The value of each element in the loading coefficient matrix represents the contribution of the corresponding processed variable to the corresponding common factor; the larger the value, the higher the contribution. For further evaluation of the target software engineering project, the 6 common factors can be named. When naming, all element values in the column corresponding to the common factor in the loading coefficient matrix are obtained, sorted from largest to smallest, and the processed variables corresponding to the first X element values are selected. X can be set according to the actual situation. The common factors are named based on the commonalities of the selected processed variables. For example, it can include 6 names: F1-F6, where F1 represents software performance requirements, F2 platform development difficulty, F3 personnel development capabilities, F4 teamwork capabilities, F5 software operation and maintenance difficulty, and F6 software data scale. The meanings of the 6 names are shown in Table 1 below.
[0127] Table 1
[0128]
[0129] In the embodiments of this specification, it is also necessary to normalize the values of the common factors to obtain the standard factors of the target software engineering project. The purpose of normalization is to unify all common factors into the 0-1 range to facilitate subsequent model prediction. Specifically, normalization can be achieved through the z-score model.
[0130] Finally, the standard factors are input into the pre-trained random forest model to predict the workload of the target software engineering project. In software engineering workload prediction, workload is generally represented by the number of lines of code, and the random forest model should ultimately predict the specific number of lines of code.
[0131] The embodiments described in this specification enable the acquisition of all variables in a target software engineering project. After processing these variables, dimensionality reduction is performed. Specifically, an intelligent optimization algorithm is used to process the scree graph to obtain the number of common factors. The loading coefficient matrix is then determined, yielding the values of the common factors after dimensionality reduction. These values are then normalized to obtain standard factors, which are input into a pre-trained random forest model to predict the workload. This method improves the accuracy of workload prediction by avoiding the problems of strong subjectivity and poor model interpretability in traditional evaluation methods.
[0132] In the embodiments described in this specification, reference is made to Figure 4 The training process of the random forest model includes:
[0133] S301: Obtain all known variables of multiple known software engineering projects, where the workload of the known software engineering projects is the known workload;
[0134] S302: Perform missing variable handling and outlier variable handling on the known variables to obtain the processed known variables;
[0135] S303: Test the processed known variables to obtain known test values;
[0136] S304: When the test value meets the set standard, a known gravel map is established based on the processed known variables;
[0137] S305: The known scree map is processed using an intelligent optimization algorithm to obtain the number of known common factors among the known variables after processing;
[0138] S306: Determine the known loading coefficient matrix based on the number of known variables and known common factors after the processing;
[0139] S307: Perform coordinate transformation on the processed known variables based on the known load coefficient matrix to obtain the values of the known common factors;
[0140] S308: Normalize the values of the known common factors to obtain the known standard factors of the known software engineering project;
[0141] S309: Divide the known standard factors of the plurality of known software engineering projects into a training set and a test set;
[0142] S310: Input the training set into the random forest model to predict the predicted workload of the known software engineering project;
[0143] S311: Compare the predicted workload with the known workload, optimize the training of the random forest model, and obtain a trained random forest model;
[0144] S312: Test the random forest model using the test set to obtain the performance of the trained random forest model.
[0145] Steps S301-S308 follow the same processing logic as the target software engineering project and can be referenced from the steps described above. When training and testing the model, known software engineering projects are required. The workload of these projects is the known workload, which is used as the actual value. The predicted workload of the known software engineering projects, obtained from the random forest model, is used as the predicted value for training the random forest model. Furthermore, the known standard factors of multiple known software engineering projects need to be divided into training and testing sets. The training set is used to train the random forest model, and the testing set is used to test the trained random forest model. Specifically, the predicted workload can be obtained by inputting the testing set into the trained random forest model. Based on the known workload and the predicted workload, the median relative error of MMRE, the median relative error of MdMRE, and PRED (0.25) are calculated to evaluate the model's performance.
[0146] In the embodiments described in this specification, reference is made to Figure 5 The hyperparameters of the random forest model need to be tuned. The methods for tuning the hyperparameters of the random forest model include:
[0147] S401: Set the value range of each hyperparameter;
[0148] S402: Set the initial mean squared error loss function to positive infinity, and set each hyperparameter in the initial hyperparameter set to a random value;
[0149] S403: Iterate M times as follows: steps S404-S408, where M is the preset number of iterations;
[0150] S404: Randomly select values from the range of each hyperparameter to obtain a set of hyperparameters;
[0151] S405: Use the hyperparameter set and training set to train the random forest model and obtain the trained random forest model;
[0152] S406: Perform multi-fold cross-validation on the training set to obtain the mean squared error loss function of the trained random forest model;
[0153] S407: If the mean squared error loss function is less than the initial mean squared error loss function, then the mean squared error loss function is assigned to the initial mean squared error loss function, and the hyperparameter set is assigned to the initial hyperparameter set.
[0154] S408: If the mean squared error loss function is not less than the initial mean squared error loss function, then continue to the next iteration;
[0155] S409: After the loop iteration is completed, output the initial hyperparameter set.
[0156] Each hyperparameter includes the number of regression trees in the random forest model, the depth of each regression tree, the training rate, etc. The hyperparameter set is a collection of each hyperparameter, including the value of each hyperparameter. The hyperparameters in the initial hyperparameter set are random values.
[0157] During the iterative process, values are randomly selected from the range of each hyperparameter to obtain a set of hyperparameters. A random forest model is constructed based on the hyperparameter values in the set of hyperparameters, and the training set is used to obtain the trained random forest model.
[0158] Further multi-fold cross-validation is needed to obtain the mean squared error loss function. The initial mean squared error loss function is positive infinity. If the mean squared error loss function is less than the initial mean squared error loss function, the initial mean squared error loss function and the initial hyperparameter set are reassigned. Otherwise, the next iteration continues. After the final iteration, the initial hyperparameter set is output. The hyperparameter values in the current initial hyperparameter set can be used as the values of each hyperparameter in the random forest model.
[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse.
[0160] This application provides users with access to relevant big data analysis (such as personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), allowing users to choose to agree to or reject automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0161] Based on the workload prediction method described above, this specification also provides a workload prediction device. The device may include a system (including a distributed system), software (application), module, component, server, client, etc., using the method described in this specification, combined with necessary hardware implementation. Based on the same innovative concept, the devices in one or more embodiments provided in this specification are as described in the following embodiments. Since the implementation schemes and methods for solving the problem are similar, the implementation of specific devices in this specification can refer to the implementation of the aforementioned method, and repeated details will not be repeated. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0162] Specifically, Figure 6 This is a schematic diagram of the module structure of one embodiment of the workload prediction device provided in this specification, with reference to... Figure 6 As shown in the embodiments of this specification, a workload prediction device is provided for at least one target software engineering project, including: an acquisition module 100, a variable processing module 200, a verification module 300, a scree map establishment module 400, a scree map processing module 500, a determination module 600, a transformation module 700, a normalization module 800, and a prediction module 900.
[0163] The acquisition module 100 is used to acquire all variables of the target software engineering project, and the variables are used to characterize the basic situation of the target software engineering project.
[0164] The variable processing module 200 is used to process the variables for abnormal variables and missing variables, and obtain the processed variables.
[0165] The testing module 300 is used to test the processed variables and obtain test values;
[0166] The scree plot creation module 400 is used to create a scree plot based on the processed variables when the test value meets the set standard.
[0167] The scree map processing module 500 is used to process the scree map using an intelligent optimization algorithm to obtain the number of common factors in the processed variables.
[0168] The determination module 600 is used to determine the loading coefficient matrix based on the number of processed variables and common factors.
[0169] Transformation module 700 is used to perform coordinate transformation on the processed variables according to the load coefficient matrix to obtain the values of common factors;
[0170] The normalization module 800 is used to normalize the values of the common factors to obtain the standard factors of the target software engineering project.
[0171] The prediction module 900 is used to input the standard factors into a pre-trained random forest model to predict the workload of the target software engineering project.
[0172] Reference Figure 7 As shown, based on the workload prediction method described above, one embodiment of this specification also provides a computer device 702, wherein the above method is run on the computer device 702. The computer device 702 may include one or more processors 704, such as one or more central processing units (CPUs) or graphics processing units (GPUs), each processing unit implementing one or more hardware threads. The computer device 702 may also include any memory 706 for storing any kind of information such as code, settings, data, etc. In one specific embodiment, a computer program is stored on the memory 706 and can run on the processor 704. When the computer program is run by the processor 704, it can execute instructions according to the above method. Non-limitingly, for example, the memory 706 may include any type of RAM, any type of ROM, flash memory device, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 702. In one scenario, when processor 704 executes associated instructions stored in any memory or combination of memories, computer device 702 can perform any operation of the associated instructions. Computer device 702 also includes one or more drive mechanisms 708 for interacting with any memory, such as hard disk drive mechanisms, optical disk drive mechanisms, etc.
[0173] Computer device 702 may also include an input / output module 710 (I / O) for receiving various inputs (via input device 712) and providing various outputs (via output device 714). A specific output mechanism may include a presentation device 716 and an associated graphical user interface 718 (GUI). In other embodiments, the input / output module 710 (I / O), input device 712, and output device 714 may be omitted, and the device may function solely as a computer device within a network. Computer device 702 may also include one or more network interfaces 720 for exchanging data with other devices via one or more communication links 722. One or more communication buses 724 couple the components described above together.
[0174] Communication link 722 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 722 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0175] Corresponding to Figures 1-5 In addition to the methods described above, embodiments of this specification also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the methods described above.
[0176] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the following... Figures 1 to 5 The method shown.
[0177] This specification also provides a computer program product, which, when executed by the processor of a computer device, performs the following... Figures 1 to 5 The method shown.
[0178] The computer program product described in this specification is a software product that mainly implements the methods described in this specification through a computer program.
[0179] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0180] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the embodiments of this specification, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0181] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments in this specification.
[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0183] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0184] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described in this specification, depending on actual needs.
[0185] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0186] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this specification, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0187] This specification uses specific embodiments to illustrate the principles and implementation methods of the embodiments. The above description of the embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments in this specification. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments in this specification. Therefore, the content of this specification should not be construed as a limitation on the embodiments in this specification.
Claims
1. A workload prediction method, characterized in that, For use in at least one target software engineering project, including: Obtain all variables of the target software engineering project, which are used to characterize the basic situation of the target software engineering project; The variables are processed by handling outlier and missing variables to obtain the processed variables; The processed variables are tested to obtain test values; When the test value meets the set standard, a scree map is established based on the processed variables; The scree map is processed using an intelligent optimization algorithm to obtain the number of common factors in the processed variables; The loading coefficient matrix is determined based on the number of processed variables and common factors. The processed variables are subjected to coordinate transformation based on the load coefficient matrix to obtain the values of the common factors; The values of the common factors are normalized to obtain the standard factors of the target software engineering project. The standard factors are input into a pre-trained random forest model to predict the workload of the target software engineering project.
2. The method according to claim 1, characterized in that, The step of processing the scree map using an intelligent optimization algorithm to obtain the number of common factors in the processed variables further includes: The following steps are iterated N times, where N is a set number of iterations: Randomly select any number from a set range; When the random number is the horizontal coordinate value in the gravel map, the vertical coordinate value corresponding to the point in the gravel map that matches the horizontal coordinate value is obtained; Calculate the product of the x-coordinate value and the y-coordinate value; After the loop iteration ends, the random number corresponding to the smallest of all product values obtained from N loop iterations is taken as the number of common factors.
3. The method according to claim 1, characterized in that, Determining the loading coefficient matrix based on the number of processed variables and common factors further includes: Based on any one of the principal component method, maximum likelihood method, and principal axis factor method, analyze the number of variables and common factors after processing to determine the loading coefficient matrix.
4. The method according to claim 1, characterized in that, The step of performing coordinate transformation on the processed variables based on the loading coefficient matrix to obtain the values of the common factors further includes: The processed variables are orthogonally rotated based on the loading coefficient matrix to obtain the values of the common factors.
5. The method according to claim 1, characterized in that, The training process of the random forest model includes: Obtain all known variables from multiple known software engineering projects, where the workload of each known software engineering project is the known workload. The known variables are processed by handling missing and outlier variables to obtain the processed known variables; The processed known variables are tested to obtain known test values; When the test value meets the set standard, a known gravel map is established based on the processed known variables; The known scree map is processed using an intelligent optimization algorithm to obtain the number of known common factors among the known variables after processing. Based on the number of known variables and known common factors after processing, determine the known loading coefficient matrix; Based on the known load coefficient matrix, the processed known variables are subjected to coordinate transformation to obtain the values of the known common factors; The values of the known common factors are normalized to obtain the known standard factors of the known software engineering project; The known standard factors of the multiple known software engineering projects are divided into training set and test set; The training set is input into the random forest model to predict the predicted workload of the known software engineering project; The predicted workload is compared with the known workload, and the random forest model is optimized and trained to obtain a well-trained random forest model. The random forest model is tested using the test set to obtain the performance of the trained random forest model.
6. The method according to claim 5, characterized in that, The hyperparameter tuning methods for the random forest model include: Set the value range for each hyperparameter; Set the initial mean squared error loss function to positive infinity, and set each hyperparameter in the initial hyperparameter set to a random value; The following steps are iterated M times, where M is a preset number of iterations; The hyperparameter set is obtained by randomly selecting values from the range of each hyperparameter. The random forest model is trained using the hyperparameter set and the training set to obtain the trained random forest model; The trained random forest model is subjected to multi-fold cross-validation using the training set to obtain the mean squared error loss function. If the mean squared error loss function is less than the initial mean squared error loss function, then the mean squared error loss function is assigned to the initial mean squared error loss function, and the hyperparameter set is assigned to the initial hyperparameter set. If the mean squared error loss function is not less than the initial mean squared error loss function, then continue to the next iteration; After the loop iteration is completed, the initial hyperparameter set is output.
7. A workload prediction device, characterized in that, For use in at least one target software engineering project, the apparatus includes: The acquisition module is used to acquire all variables of the target software engineering project, and the variables are used to characterize the basic situation of the target software engineering project. The variable processing module is used to handle abnormal variables and missing variables to obtain the processed variables; The testing module is used to test the processed variables and obtain test values; The scree plot creation module is used to create a scree plot based on the processed variables when the test value meets the set standard. The scree graph processing module is used to process the scree graph using an intelligent optimization algorithm to obtain the number of common factors in the processed variables. The determination module is used to determine the loading coefficient matrix based on the number of processed variables and common factors; The transformation module is used to perform coordinate transformation on the processed variables according to the load coefficient matrix to obtain the values of the common factors. The normalization module is used to normalize the values of the common factors to obtain the standard factors of the target software engineering project. The prediction module is used to input the standard factors into a pre-trained random forest model to predict the workload of the target software engineering project.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the computer program is run by the processor, it executes the instructions of the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor of the computer device, it executes the instructions of the method according to any one of claims 1-6.
10. A computer program product, characterized in that, When the computer program product is run by the processor of a computer device, it executes the instructions of the method according to any one of claims 1-6.