Internet-based performance prediction system for software development
By fusing multi-source heterogeneous data and using an adaptive performance prediction model, the problem of time lag and accuracy degradation in performance prediction during software development is solved. This achieves efficient and accurate performance prediction and model adaptability, supporting cross-project and cross-team applications.
Patent Information
- Application Number
- CN202510801180.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-18
AI Technical Summary
In the current software development process, performance prediction methods rely on data from a single source, making it difficult to comprehensively characterize the performance of software modules. They are also time-consuming and lack timeliness. Furthermore, traditional models are difficult to adapt to version evolution and changes in data distribution, leading to a decrease in prediction accuracy.
By employing a multi-source heterogeneous data fusion mechanism, an adaptive performance prediction system is constructed through data acquisition, feature extraction, performance prediction models, and feedback optimization modules. Combining deep neural networks and ensemble regression algorithms, the system optimizes model parameters in real time and supports multi-dimensional performance indicator prediction.
It improves the foresight and accuracy of performance prediction, enhances the model's adaptability to complex software environments, reduces the burden of later operation and maintenance, and supports modular deployment and multi-scenario application.
Smart Images

Figure CN120973560A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer system performance evaluation and fault prediction technology, specifically relating to a performance prediction system for software development based on the Internet. Background Technology
[0002] In modern software development, performance metrics (such as response time, resource utilization, and system reliability) are key factors in measuring software quality and user experience. However, existing methods often rely on performance testing and monitoring feedback after system deployment, which frequently suffers from time lags and an inability to provide effective early warnings, leading to high costs and significant impacts in fixing performance defects later on.
[0003] Most current performance analysis methods rely on data from a single source, such as runtime logs or static code metrics, making it difficult to comprehensively depict the true performance of software modules. Especially in microservice and continuous integration environments, the relationships between multi-source heterogeneous information such as log data, source code features, and defect tracking are complex, making it difficult for traditional methods to efficiently extract, model, and use for prediction tasks, resulting in poor accuracy and generalization ability.
[0004] Traditional performance prediction systems mostly employ static modeling strategies, where the model remains fixed after training, making it difficult to adapt to version evolution and changes in data distribution. In practical applications, there is often a deviation between predicted and true values. Without an effective error feedback optimization mechanism, model performance will gradually degrade over time, making it difficult to maintain long-term stability and prediction accuracy. Summary of the Invention
[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solution: a performance prediction system for software development based on the Internet, comprising the following modules:
[0006] The data acquisition module is used to acquire multi-dimensional performance-related raw data from multiple stages of the software development lifecycle, including source code metric data D. src , Run log data D log Historical defect data D def ;
[0007] The feature extraction module is used to construct a high-dimensional feature vector set F based on the collected raw data. vec The feature vector set includes the code complexity index C. com Calling the relation graph G call Test coverage T cov Abnormal mode E pat and defect density D den ;
[0008] The performance prediction model module is used to predict the performance based on the high-dimensional feature vector set F.vec Construct a predictive model P for software module-level performance prediction. mod Output predictive performance index P ind The predictive performance index P ind Including response time R tim Resource utilization rate R res Reliability R rel ;
[0009] The feedback optimization module is used to optimize the predicted performance metric P. ind Compared with actual operating performance index A ind Perform difference analysis and extract the error vector E. vec The prediction model P is optimized using the error backpropagation algorithm. mod ;
[0010] The visualization analysis module is used to graphically present performance prediction results, error feedback trends, and key performance bottlenecks. This module provides a performance trend chart T. perf Error convergence plot T err Module risk scoring function S risk .
[0011] Preferably, the data acquisition module includes:
[0012] The source code static analysis unit is used to analyze the code repository during the development process, extracting code lines, cyclomatic complexity, and class inheritance depth to form source code metrics data D. src ;
[0013] The log processing unit is used to extract timestamps, exception stack traces, and service call chain information from the runtime log to form runtime log data D. log ;
[0014] The defect tracking interface unit is used to obtain historical defect tags, closure times, and severity, forming historical defect data D. def .
[0015] Preferably, the feature extraction module includes constructing a high-dimensional feature vector set F using a graph embedding optimization-based structure learning algorithm. vec The high-dimensional feature vector set F vec include:
[0016] Code complexity metric C com It is calculated based on the Halstead index fusion function;
[0017] Calling the relation graph G call Dependencies between functions are extracted through static code analysis and graph convolution.
[0018] Test coverage Tcov The calculation is based on the frequency of execution of the test path and the branch coverage ratio;
[0019] Abnormal mode E pat Log data D is analyzed using a deep sequence model. log Perform abnormal cluster identification and quantization;
[0020] Defect density D den The calculation is based on the frequency of defect records and the number of lines of code in the module.
[0021] Preferably, the performance prediction model module includes:
[0022] Hierarchical modeling units are used to process the high-dimensional feature vector set F output by the feature extraction module. vec The input is fed into a multi-layer coupled prediction network, and a performance prediction model P is constructed using a feature channel grouping + residual connection strategy. mod ;
[0023] Multi-objective prediction unit for use with the trained prediction model P mod Output prediction performance index P ind The predictive performance index P ind Includes: response time R tim Resource utilization rate R res Reliability R rel The results are obtained through the following coupling prediction functions:
[0024] R tim =α1·log(W) tim ·F vec +b tim );
[0025] R res =α2·tanh(W) res ·F vec +b res );
[0026] R rel =α3·σ(W rel ·F vec +b rel );
[0027] Among them, W tim W is the weight matrix for response time. res The weight matrix for resource utilization, W rel b is the reliability weight matrix. tim b is the bias vector for the response time. res Let b be the bias vector for resource utilization. relLet α1, α2, and α3 be the reliability bias vector, and let α1, α2, and α3 be the correlation adjustment coefficients, satisfying α1 2 +α2 2 +α3 2 =1, σ(·) is the Sigmoid function, used to model the confidence distribution on reliability;
[0028] Temporal memory units are used to combine historical version feature sequences with the current feature vector set F. vec Sliding window prediction optimization is performed to improve the continuity and generalization ability of performance prediction during version evolution. The temporal memory unit uses a gated recurrent unit structure to model the sequence dependency relationship of feature vectors of each version and outputs a temporal enhancement vector F. seq :
[0029]
[0030] Where t is the current time and t-1 is the previous time.
[0031] Preferably, the feedback optimization module includes:
[0032] The performance error calculation unit is used to compare the predicted performance index P. ind Compared with actual operating performance index A ind Construct the performance error vector E vec The error vector is calculated according to the following formula:
[0033] E vec =[|R tim -A tim |,|R res -A res |,|R rel -A rel |];
[0034] Among them, A tim A res A rel These are the actual performance indicators monitored.
[0035] Model gradient feedback unit, used for based on error vector E vec Calculate the loss function L err And for the prediction model P mod Perform backpropagation update, the loss function L err The calculation formula is as follows:
[0036] L err =∑ j∈{tim,res,rel} λE j 2 ;
[0037] Where λ is the weighting coefficient.
[0038] The model gradient feedback unit uses the following algorithm to predict model P. mod Update:
[0039] P mod ′=β1·P mod ·(1-L err )+β2·L err ;
[0040] Among them, P mod ′ represents the updated prediction model, L err Let β1 and β2 be the loss function and β1 and β2 be the dynamic adjustment coefficients.
[0041] Preferably, the visualization analysis module includes:
[0042] The trend analysis unit is used to analyze the performance index P within consecutive iteration cycles. ind Perform sequence aggregation to generate a performance trend chart T. perf It intuitively reflects the response time R tim Resource utilization rate R res Reliability R rel The trend analysis unit supports displaying performance prediction time series at the module level to reflect dynamic fluctuations with version changes. The graphical interface can dynamically select and display indicator combinations, including: indicator multi-selection control, sliding time axis window, version switching tab, and provides comparative analysis functions.
[0043] The error convergence display unit is used to display the error vector E in the feedback optimization module. vec The changes during the training process are plotted as an error convergence plot T. err It helps identify whether the training has converged and abnormal oscillation cycles;
[0044] Module risk scoring unit, used for predictive reliability R rel With defect density D den Integration Module Risk Scoring Function S risk The function is defined as follows:
[0045]
[0046] Where β1 and β2 are weighting parameters, reflecting the reliability R rel With defect density D den Contribution to risk scoring.
[0047] An internet-based software development performance prediction system, deployed in a cloud-based distributed container environment, supports cross-project and cross-team sharing of model training results and visualization reports. The system also includes:
[0048] The model service interface module is used to encapsulate the prediction model P. mod And provide RESTful API calls;
[0049] The data synchronization module is used to synchronize logs, code, and defect data to the data acquisition module in real time, ensuring the consistency and real-time performance of inputs from each module.
[0050] The security encryption module is used to encrypt the high-dimensional feature vector set F during transmission. vec With performance index P ind Data is used to ensure the confidentiality and integrity of the model's inputs and outputs.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] Improving the foresight and accuracy of performance prediction: This invention introduces a multi-source heterogeneous data fusion mechanism to comprehensively collect multi-dimensional features such as code static indicators, log data, and historical defect records. It combines deep neural networks and ensemble regression algorithms to build a performance prediction model, which can accurately predict the performance of future software during the development stage, effectively improving the ability to predict in advance and reducing the burden of later operation and maintenance.
[0053] Enhancing the model's adaptability to complex software environments: This invention adopts a dynamic model update strategy with adaptive capabilities. By introducing a feedback correction module, it compares the prediction results with the actual running results in real time and automatically adjusts the model parameters, enabling the prediction system to adapt to complex software iteration environments such as version evolution and changes in data distribution, and maintain long-term prediction accuracy and stability.
[0054] Supports modular deployment and multi-scenario application expansion: The system structure of this invention adopts a modular design. Each functional module, such as the data processing module, feature extraction module, predictive modeling module, and feedback optimization module, can be flexibly combined and deployed. It can be used independently in a single project or integrated into a CI / CD pipeline or DevOps platform, providing multi-scenario and multi-dimensional support for software performance management in different projects and at different stages, thereby improving development efficiency and quality assurance capabilities. Attached Figure Description
[0055] Figure 1 The system module flowchart provided for this application;
[0056] Figure 2 A schematic diagram of the system modules provided in this application. Detailed Implementation
[0057] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0058] refer to Figures 1-2 This invention provides an Internet-based software development performance prediction system, comprising the following modules:
[0059] The data acquisition module is used to acquire multi-dimensional performance-related raw data from multiple stages of the software development lifecycle, including source code metric data D. src , Run log data D log Historical defect data D def .
[0060] First, the source code static analysis unit is primarily deployed on the code hosting platform (GitHub) used by the development team. This unit employs static analysis tools to perform non-intrusive analysis on a specified version of the source code, extracting multiple static structure metrics. In practical application, this unit automatically identifies various types of source files by traversing and scanning the project's code repository, and calculates the number of lines of code (including valid lines and comment lines), cyclomatic complexity (used to measure the complexity of code logic branches), and class inheritance depth (used to reflect the hierarchical complexity of object-oriented structures) for each file. Finally, the above analysis results are integrated into a structured source code metric dataset D. src It is indexed according to metadata such as file path, author and submission time, which facilitates timeline correlation analysis with subsequent logs or defect information.
[0061] Secondly, the log processing unit is deployed in the test environment or online environment node where the target system runs. This unit mainly performs structured preprocessing on the large amount of log data generated during service runtime by connecting to the log collection system. In this embodiment, the unit uses regular expression matching and template parsing technology to accurately extract runtime timestamps, exception stack traces, and service call chain information for service log content of different formats, and generates a unified structured runtime log dataset D. log This dataset is stored in units of service modules and supports fast retrieval by time period, anomaly type, and other dimensions, providing basic data support for performance bottleneck identification and runtime modeling.
[0062] Secondly, the defect tracking interface unit interfaces with a commonly used defect tracking system (JIRA) in software projects, retrieving defect record information corresponding to a specified project version by calling its open API. In this embodiment, this unit periodically pulls historical defect data and extracts the defect tag, closure time, and severity level for each defect record, forming a historical defect dataset D. def The dataset also supports mapping to source code version information, which facilitates subsequent models inferring risk factors that may lead to performance degradation based on the defect density and severity at different stages.
[0063] The feature extraction module is used to construct a high-dimensional feature vector set F based on the collected raw data. vec The feature vector set includes the code complexity index C. com Calling the relation graph G call Test coverage T cov Abnormal mode E pat and defect density D den .
[0064] The feature extraction module is mainly used to extract features from multi-source heterogeneous data (including source code indicator data D) acquired by the data acquisition module. src , Run log data D log and historical defect data D def Transform it into a unified high-dimensional feature vector set F that can be used for subsequent modeling and training. vec This module employs a graph embedding optimization-based structure learning algorithm, combining static structural information, dynamic behavioral patterns, and historical quality features to achieve a deep representation of performance influencing factors across multiple dimensions. The feature extraction module primarily includes the following five core feature generation processes: code complexity index C... com Calling the relation graph G call Test coverage T cov Abnormal mode E pat and defect density D den .
[0065] First, the code complexity metric C com The extraction is based on the Halstead complexity metric, and its input data comes from the D output of the source code static analysis unit. src Dataset. In the specific implementation, the system first parses the operators and operands inside each function, then performs statistical analysis on the computational logic of the fusion function according to the Halstead formula, generating a series of indicators describing the cognitive complexity of the code (such as program size, program difficulty, etc.), and finally sums them up to obtain a unified code complexity vector C. comBy standardizing the complexity distribution of each module or function, this feature can effectively reflect the implementation complexity of different functional modules, providing a static structural basis for performance bottleneck analysis.
[0066] Secondly, call the relation graph G call The system is built based on static code analysis results, achieved through the extraction of inter-function call information and feature learning using a graph convolutional network. In this embodiment, the system first uses static analysis tools to construct call graphs within modules and between functions across modules. Then, it represents the call graph as a directed graph structure and extracts the structural embedding features of each node (function) in the call dependencies using a graph convolutional network, ultimately forming a vectorized representation G of the inter-function dependencies. call This process not only captures direct calls between functions, but also captures potential indirect dependencies through graph neural networks, which is of great value for modeling the overall structural stability of the system.
[0067] Next, test coverage T cov The extraction is based on the execution results of unit tests and integration tests during the development process. The system uses a testing framework to collect metrics such as the code paths reached by each test case during execution, execution frequency, and branch coverage, constructing a test execution matrix. Based on dimensions such as coverage ratio of different modules and path execution frequency, the system calculates the test coverage characteristic T. cov In this embodiment, the feature pays particular attention to the intersection of low-coverage areas and high-risk functions, thereby revealing potential performance anomalies and improving the model's ability to identify test blind spots.
[0068] Subsequently, abnormal mode E pat The main approach is to use deep sequence models to analyze the runtime log data D. log The system learns and clusters data. In practical applications, it first encodes structured log information into time-series sequences and then trains a deep learning model (Transformer) on these sequences to learn the behavioral boundaries between normal and abnormal events. After training, the system performs cluster analysis on abnormal events and uses the center vector of each cluster as a pattern feature representing the abnormal cluster. Finally, the structural information of the abnormal clusters is converted into E using vectorized encoding. pat Vectors reflect abnormal trends and change patterns in logs.
[0069] Finally, the defect density D den The calculation relies on historical defect data D def With source code static analysis data D src The system combines [various methods]. In this embodiment, the system counts the frequency of defect records for each module within a certain time window, and performs normalization calculations based on the number of lines of code for that module to form a defect density index D. denThis metric can effectively identify modules that frequently cause quality issues in historical versions and, to some extent, reflect their potential impact on performance stability. The system will also include D... den With C com T cov Features are used in conjunction with modeling to enhance the overall expressive power of the model.
[0070] The performance prediction model module is used to predict the performance based on the high-dimensional feature vector set F. vec Construct a predictive model P for software module-level performance prediction. mod Output predictive performance index P ind The predictive performance index P ind Including response time R tim Resource utilization rate R res Reliability R rel .
[0071] The performance prediction model module is used to predict the high-dimensional feature vector set F output by the feature extraction module. vec Build and train the performance prediction model P mod This module enables automated prediction of multi-dimensional performance indicators. It comprises three key functional units: a hierarchical modeling unit, a multi-objective prediction unit, and a temporal memory unit.
[0072] First, the hierarchical modeling unit receives the high-dimensional feature vector set F generated by the feature extraction module. vec Construct a performance prediction model P with deep expressive capabilities mod In its implementation, this unit employs a modeling structure based on a multi-layer coupled prediction network. To prevent the degradation of high-dimensional features during multi-layer propagation, this network structure introduces a feature channel grouping strategy, which groups F... vec The model is divided into multiple feature sub-channels, each undergoing convolution or fully connected operations to effectively avoid information loss caused by feature redundancy. Simultaneously, a residual connection mechanism is introduced between each layer to ensure that important features are preserved in the deep network, thereby enhancing the model's ability to fit complex nonlinear relationships. During the model training phase, the system... mod Multiple rounds of iterative training were conducted to optimize its performance prediction capabilities on the sample set.
[0073] Next, the multi-object prediction unit is responsible for processing the trained P... mod The model is used to infer actual performance indicators. Specifically, this unit is based on model P. mod The intermediate feature representation results are used to calculate three key performance indicators P. ind That is, response time R tim Resource utilization rate R res Reliability R relEach metric is modeled using a dedicated coupled prediction function to better simulate the nonlinear interactions between performance metrics across different dimensions, where:
[0074] R tim From function R tim =α1·log(W) tim ·F vec +b tim Modeling is performed by taking the logarithmic function form of the result after linear transformation to simulate the logarithmic distribution characteristics of the response time;
[0075] R res Then through R res =α2·tanh(W) res ·F vec +b res The expression uses the hyperbolic tangent function to constrain the boundary conditions of resource utilization.
[0076] R rel Then use R rel =α3·σ(W rel ·F vec +b rel ), where σ represents the Sigmoid function, which is suitable for probabilistic modeling of reliability indicators.
[0077] In the above function, W tim W res W rel These are the weight matrices corresponding to different indicators, b tim b res b rel α1, α2, and α3 are the corresponding bias vectors; α1, α2, and α3 are the correlation adjustment coefficients among the indicators, and the sum of their squares is 1 to ensure the weight balance of multiple indicators in joint prediction. This design can improve the model's expressive ability when facing multi-objective optimization tasks, and is especially suitable for industrial software application scenarios that require trade-offs between performance, resources, and reliability.
[0078] Finally, to address the significant performance fluctuations between versions, the temporal memory unit introduces a sliding window prediction optimization mechanism, which models the feature sequences of historical versions and the current F... vec This enhances the predictive continuity of the model over time by establishing dependencies. Specifically, the unit employs a gated recurrent unit structure as its core modeling framework, with multiple consecutive versions of F as input. vec Sequence. During operation, the gated recurrent unit structure encodes and updates historical feature states, retains long-term dependency information, and combines it with the current version of F. vec Generate time-series augmentation vector F seq This is to improve the predictive model's ability to perceive long-term evolution trends.seq The generating formula is Where t is the current time, t-1 is the previous time, and F vec (t) Let represent the feature vector of the t-th version. This structure ensures that the model retains sufficient temporal context information when evaluating the performance trend during version evolution, thereby reducing the abruptness of prediction results and improving the model's generalization ability in version transition scenarios.
[0079] The feedback optimization module is used to optimize the predicted performance metric P. ind Compared with actual operating performance index A ind Perform difference analysis and extract the error vector E. vec The prediction model P is optimized using the error backpropagation algorithm. mod .
[0080] The feedback optimization module mainly includes two functional units: a performance error calculation unit and a model gradient feedback unit.
[0081] First, the performance error calculation unit is used to monitor and evaluate the performance of the prediction model in real-time during practical applications. After system deployment, in each version or task execution cycle, the performance prediction model P... mod It will output the predictive performance metric P ind Including response time R tim Resource utilization rate R res Reliability R rel Meanwhile, the system monitoring module will collect actual performance indicators A from the operating environment. ind , respectively corresponding to A tim A res A rel This unit will P ind With A ind A one-to-one comparison is performed, and the performance error vector E is constructed using the absolute difference method. vec The specific calculation method is as follows:
[0082] E vec =[|R tim -A tim |,|R res -A res |,|R rel -A rel |];
[0083] The error vector E vec It fully expresses the magnitude of the deviation between the predicted results and the actual data, providing a clear indicator and facilitating subsequent model adjustments. This mechanism can adapt to the dynamic changes in performance data during actual operation, ensuring that the model has the ability to provide real-time feedback to the real operating environment.
[0084] Next, the model gradient feedback unit uses the aforementioned error vector E vec Further construct a unified loss function L err This is used to measure the overall prediction accuracy of the current model. The loss function is designed to consider the differences in importance among various performance metrics, and is summed using a weighted sum of squares. Its calculation formula is as follows:
[0085] L err =∑ j∈{tim,res,rel} λE j 2 ;
[0086] Among them, E j Let be the components of the error vector, corresponding to the prediction errors of the three performance dimensions. λ is the weighting coefficient for each error term, used to balance the influence of different indicators on the overall loss. Through this loss function, the system can quantify the current model bias, serving as a basis for subsequent optimization and updates.
[0087] During the model update phase, the feedback optimization module further optimizes the loss function L. err Applied to prediction model P mod In adjusting the structural parameters, a lightweight backpropagation optimization strategy is adopted to maintain a balance between system real-time performance and stability. The specific update formula is as follows:
[0088] P mod ′=β1·P mod ·(1-L err )+β2·L err ;
[0089] Among them, P mod ' represents the updated prediction model, and β1 and β2 are adjustment coefficients that control the proportional relationship between maintaining the original model structure and adapting to new error feedback. This update method does not directly reconstruct the model structure, but rather introduces new error information by weighting the existing model state, thereby achieving continuous and gradual model evolution while retaining existing predictive capabilities. This method is more efficient and practical than the traditional full retraining mechanism, and is especially suitable for software deployment scenarios with frequent version changes and unstable operating environments.
[0090] In addition, to prevent the model from making drastic adjustments when the error is small, the system can set a threshold mechanism, when L err Pause P when the level falls below the set level. mod This module updates the model to reduce unnecessary calculations and model oscillations. It also periodically stores the model state, creating model snapshots for quick recovery in case of erroneous updates or system rollbacks.
[0091] The visualization analysis module is used to graphically present performance prediction results, error feedback trends, and key performance bottlenecks. This module provides a performance trend chart T. perf Error convergence plot T err Module risk scoring function S risk .
[0092] The visualization analysis module mainly includes three functional units: a trend analysis unit, an error convergence display unit, and a module risk scoring unit. First, the trend analysis unit focuses on the performance prediction index P within a continuous iteration cycle. ind This unit performs time-series aggregation and dynamic display. It collects response time R from the performance prediction model module. tim Resource utilization rate R res Reliability R rel The metric data is organized into a time series according to version order, and a trend chart T reflecting performance changes over time is generated. perf This trend chart not only displays the overall performance trend but also supports segmented views at the module level. Users can flexibly select and combine performance metrics of interest through the graphical interface. For example, the multi-selection control allows users to freely combine and display one or more performance metrics according to their actual needs, enabling personalized analysis; the sliding timeline window allows users to dynamically adjust the time range they are interested in, facilitating observation of performance changes between different versions; and the version switching tab allows users to quickly switch between performance data from different versions for horizontal comparison. This feature design greatly enhances the interactivity and readability of performance data, supporting technical teams in deeply exploring performance fluctuation patterns and potential optimization points during version evolution.
[0093] Secondly, the error convergence demonstration unit aims to assist in monitoring the feedback optimization process of the performance prediction model during the training phase. This unit uses the error vector E provided by the feedback optimization module. vec Based on this, the error variation curve during the training process is plotted to form an error convergence graph T. err This graph clearly shows the decreasing trend of error during training iterations, helping to determine whether the model has reached convergence or whether there are abnormal oscillations. By observing the error convergence plot, researchers can promptly identify unstable factors in model training, adjust training strategies or hyperparameters, and ensure stable improvement in the performance of the prediction model. The intuitive presentation of the error convergence plot helps improve the transparency and controllability of model training, promoting the effective implementation of feedback optimization mechanisms.
[0094] Finally, the module risk scoring unit combines the predicted reliability index R rel With historical defect density D den Construct a quantitative risk scoring function S risk This unit is used to assess the risk level of a software module. It is calculated using the following scoring formula:
[0095]
[0096] Here, β1 and β2 are weighting parameters, measuring the contribution of prediction reliability and defect density to the overall risk score, respectively. Through this risk scoring function, the system can integrate complex performance and defect information into a single risk value, facilitating risk level ranking and focused attention. The risk scoring results can not only guide test prioritization and resource allocation but also serve as an important basis for iterative optimization. The module risk scoring unit supports graphical display of risk results and provides multi-dimensional views combined with version and module information, helping users comprehensively understand the current risk status and evolution trend of each module.
[0097] An internet-based software development performance prediction system, deployed in a cloud-based distributed container environment, supports cross-project and cross-team sharing of model training results and visualization reports. The system also includes:
[0098] The model service interface module is used to encapsulate the prediction model P. mod And provide RESTful API calls;
[0099] The data synchronization module is used to synchronize logs, code, and defect data to the data acquisition module in real time, ensuring the consistency and real-time performance of inputs from each module.
[0100] The security encryption module is used to encrypt the high-dimensional feature vector set F during transmission. vec With performance index P ind Data is used to ensure the confidentiality and integrity of the model's inputs and outputs.
[0101] It should be noted that, unless otherwise specified, the embodiments and features and technical solutions in the present invention can be combined with each other.
[0102] Obviously, the embodiments described above are merely some embodiments of the present invention, not all embodiments. The accompanying drawings show preferred embodiments of the present invention, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the patent protection scope of this invention.
Claims
1. A performance prediction system for software development based on the Internet, characterized in that, Includes the following modules: The data acquisition module is used to acquire multi-dimensional performance-related raw data from multiple stages of the software development lifecycle, including source code metric data D. src , Run log data D log Historical defect data D def ; The feature extraction module is used to construct a high-dimensional feature vector set F based on the collected raw data. vec The feature vector set includes the code complexity index C. com Calling the relation graph G call Test coverage T cov Abnormal mode E pat and defect density D den ; The performance prediction model module is used to predict the performance based on the high-dimensional feature vector set F. vec Construct a predictive model P for software module-level performance prediction. mod Output predictive performance index P ind The predictive performance index P ind Including response time R tim Resource utilization rate R res Reliability R rel ; The feedback optimization module is used to optimize the predicted performance metric P. ind Compared with actual operating performance index A ind Perform difference analysis and extract the error vector E. vec The prediction model P is optimized using the error backpropagation algorithm. mod ; The visualization analysis module is used to graphically present performance prediction results, error feedback trends, and key performance bottlenecks. This module provides a performance trend chart T. perf Error convergence plot T err Module risk scoring function S risk .
2. The performance prediction system for Internet-based software development according to claim 1, characterized in that, The data acquisition module includes: The source code static analysis unit is used to analyze the code repository during the development process, extracting code lines, cyclomatic complexity, and class inheritance depth to form source code metrics data D. src ; The log processing unit is used to extract timestamps, exception stack traces, and service call chain information from the runtime log to form runtime log data D. log ; The defect tracking interface unit is used to obtain historical defect tags, closure times, and severity, forming historical defect data D. def .
3. The performance prediction system for Internet-based software development according to claim 1, characterized in that, The feature extraction module includes constructing a high-dimensional feature vector set F using a graph embedding optimization-based structure learning algorithm. vec The high-dimensional feature vector set F vec include: Code complexity metric C com It is calculated based on the Halstead index fusion function; Calling the relation graph G call Dependencies between functions are extracted through static code analysis and graph convolution. Test coverage T cov The calculation is based on the frequency of execution of the test path and the branch coverage ratio; Abnormal mode E pat Log data D is analyzed using a deep sequence model. log Perform abnormal cluster identification and quantization; Defect density D den The calculation is based on the frequency of defect records and the number of lines of code in the module.
4. The performance prediction system for Internet-based software development according to claim 1, characterized in that, The performance prediction model module includes: Hierarchical modeling units are used to process the high-dimensional feature vector set F output by the feature extraction module. vec The input is fed into a multi-layer coupled prediction network, and a performance prediction model P is constructed using a feature channel grouping + residual connection strategy. mod ; Multi-objective prediction unit, used for prediction based on trained prediction model P mod Output prediction performance index P ind The predictive performance index P ind Includes: response time R tim Resource utilization rate R red Reliability R rel The results are obtained through the following coupling prediction functions: R tim =α1·log(W tim ·F vec +b tim ); R res =α2·tanh(W res ·F vec +b res ); R rel =α3·σ(W rel ·F vec +b rel ); Among them, W tim W is the weight matrix for response time. res The weight matrix for resource utilization, W rel b is the reliability weight matrix. tim b is the bias vector for the response time. res Let b be the bias vector of resource utilization. rel Let α1, α2, and α3 be the reliability bias vector, and let α1, α2, and α3 be the correlation adjustment coefficients, satisfying α1 2 +α2 2 +α3 2 =1, σ(·) is the Sigmoid function, used to model the confidence distribution on reliability; Temporal memory units are used to combine historical version feature sequences with the current feature vector set F. vec We optimize sliding window prediction to improve the continuity and generalization of performance prediction during version evolution.
5. The performance prediction system for Internet-based software development according to claim 1, characterized in that, The feedback optimization module includes: The performance error calculation unit is used to compare the predicted performance index P. ind Compared with actual operating performance index A ind Construct the performance error vector E vec The error vector is calculated according to the following formula: E vec =[|R tim -A tim |,|R res -A res |,|R rel -A rel |]; Among them, A tim A res A rel These are the actual performance indicators monitored. Model gradient feedback unit, used for based on error vector E vec Calculate the loss function L err And for the prediction model P mod Perform backpropagation update, the loss function L err The calculation formula is as follows: L err =∑ j∈{tim,res,rel} λE j 2 ; Where λ is the weighting coefficient.
6. The performance prediction system for Internet-based software development according to claim 1, characterized in that, The visualization analysis module includes: The trend analysis unit is used to analyze the performance index P within consecutive iteration cycles. ind Perform sequence aggregation to generate a performance trend chart T. pref It intuitively reflects the response time R tim Resource utilization rate R res Reliability R rel Dynamic fluctuations as versions change; The error convergence display unit is used to display the error vector E in the feedback optimization module. vec The changes during the training process are plotted as an error convergence plot T. err It helps identify whether the training has converged and abnormal oscillation periods; Module risk scoring unit, used for predictive reliability R rel With defect density D den Integration Module Risk Scoring Function S risk The function is defined as follows: Where β1 and β2 are weighting parameters, reflecting the reliability R rel With defect density D den Contribution to risk scoring.
7. The performance prediction system for Internet-based software development according to claim 4, characterized in that, The temporal memory unit uses a gated recurrent unit structure to model the sequence dependencies of the feature vectors of each version and outputs a temporal enhancement vector F. seq : Where t is the current time and t-1 is the previous time.
8. A performance prediction system for Internet-based software development according to claim 5, characterized in that, The model gradient feedback unit uses the following algorithm to predict model P. mod Update: P mod ′=β1·P mod ·(1-L err )+β2· Lerr ; Among them, P mod For the updated prediction model, L err Let β1 and β2 be the loss function and β1 and β2 be the dynamic adjustment coefficients.
9. A performance prediction system for Internet-based software development according to claim 6, characterized in that, The trend analysis unit supports displaying performance forecast time series at the module granularity, and the graphical interface allows for dynamic selection of display indicator combinations, including: Multi-select indicator control; Slide the timeline window; The version switching tab provides comparative analysis functionality.
10. A performance prediction system for Internet-based software development according to any one of claims 1 to 9, characterized in that, The system is deployed in a cloud-based distributed container environment, supporting cross-project and cross-team sharing of model training results and visualization reports. The system also includes: The model service interface module is used to encapsulate the prediction model P. mod And provide RESTful API calls; The data synchronization module is used to synchronize logs, code, and defect data to the data acquisition module in real time, ensuring the consistency and real-time performance of inputs from each module. The security encryption module is used to encrypt the high-dimensional feature vector set F during transmission. vec With performance index P ind Data is used to ensure the confidentiality and integrity of the model's inputs and outputs.
Citation Information
Cited By
AI-based game performance analysis and operational stability prediction system
CN122489445A