Machine learning-based APP performance multi-dimensional analysis method and system
By constructing a machine learning-based multi-dimensional analysis method for APP performance, this method solves the problems of existing technologies that cannot predict the performance impact of code changes and lack long-term evolution analysis. It achieves accurate prediction and identification of complex related paths during the code change stage, thereby improving the performance optimization and maintainability of the APP.
Patent Information
- Application Number
- CN202511222627.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-29
AI Technical Summary
In existing technologies, APP performance optimization lacks predictive capabilities during the code change phase and cannot identify complex relationships, resulting in performance problems being discovered only after code deployment. Furthermore, it lacks long-term performance evolution analysis, making it difficult to achieve a balance between code maintainability and performance optimization.
We construct a multi-dimensional analysis method for APP performance based on machine learning. Through code semantic understanding model, performance evolution analysis model and code performance knowledge graph, we can predict the long-term and short-term impacts of code changes on performance and identify complex correlation paths.
It enables accurate prediction of potential performance issues during code change phases, reduces online incidents, provides problem-solving suggestions, supports informed technical decisions, improves the overall score of code quality and performance, and identifies code changes that are not directly related but have long-term impact.
Smart Images

Figure CN120723614B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile application performance analysis technology, and more specifically, to a machine learning-based multi-dimensional analysis method and system for APP performance. Background Technology
[0002] With the rapid development of the mobile application market and the increasing performance demands of users, app performance optimization has become a crucial aspect of the software development process. In traditional app development and maintenance processes, code review and performance analysis are typically handled separately. This approach exposes numerous technical limitations when dealing with increasingly complex app systems.
[0003] Existing code review technologies primarily focus on the functional correctness, security, and maintainability of code, identifying potential code defects through static code analysis and manual code review. However, these methods struggle to effectively predict the impact of code changes on app performance during the code modification phase, often resulting in performance issues being discovered only after code deployment. Regarding performance monitoring, current technologies typically rely on the collection and analysis of runtime performance data, including monitoring metrics such as CPU utilization, memory usage, and response time. For example, traditional application performance monitoring tools can monitor the app's running status in real time, but this reactive monitoring approach has a significant lag. By the time performance issues are discovered, they have often already negatively impacted the user experience, leading to high repair costs. Furthermore, existing performance analysis methods mostly focus on the immediate performance impact of a single code change, lacking an understanding of performance changes over the long-term evolution of the app. A lack of in-depth analysis of trends reveals that this short-term perspective makes it difficult for development teams to identify performance degradation issues with cumulative effects, and also makes it impossible to predict the potential performance impact of code modifications during long-term iterations. In complex app systems, there are intricate dependencies and calling relationships between code modules. Existing technologies are often limited to analyzing the performance impact of directly related code changes. For code changes that seem unrelated but actually affect system performance through complex paths, there is a lack of effective identification and early warning mechanisms. This limitation makes it impossible to detect and handle potential performance risks in a timely manner. In software engineering practice, development teams often face the trade-off between code maintainability and performance optimization. Due to the lack of objective and quantitative decision support tools, teams often have to rely on experience to make judgments. This subjective decision-making approach can easily lead to suboptimal technology choices, affecting the long-term development quality of the app.
[0004] Therefore, there is an urgent need for an APP performance analysis method that can predict performance impact during code change phases, support long-term performance evolution analysis, and identify complex relationships, in order to solve the above-mentioned technical problems. Summary of the Invention
[0005] This invention provides a machine learning-based method and system for multi-dimensional analysis of APP performance, which solves the technical problems of code review and performance analysis separation, inability to predict the impact of code changes on performance, lack of long-term evolution analysis and identification of complex related paths in related technologies.
[0006] This invention provides a multi-dimensional analysis method for app performance based on machine learning, including:
[0007] Build a code semantic understanding model to analyze the APP source code and extract the structured features and semantic information of the code;
[0008] Based on the extracted code structure features and semantic information, a performance evolution analysis model is constructed. By collecting and analyzing the performance data of the APP in different versions, the long-term change patterns of performance indicators are identified.
[0009] Based on the code entity information output by the code semantic understanding model and the performance index data output by the performance evolution analysis model, a code performance knowledge graph is constructed, with code entities and performance indicators as nodes and the influence relationships between them as edges, to realize the systematic modeling of the relationship between code changes and performance changes.
[0010] Based on the node relationships and edge weights in the code performance knowledge graph, a dual-state mapping prediction model is constructed. It takes the code state and performance state at the current time point as input and predicts the long-term and short-term impacts of code changes on APP performance.
[0011] Furthermore, the steps for constructing the code semantic understanding model include:
[0012] Code entity extraction identifies and extracts key code entities, including classes, methods, functions, and modules, and generates a code entity set.
[0013] Code syntax structure analysis converts source code into an abstract syntax tree, capturing the syntactic structure relationships of the code;
[0014] Code semantic vectorization applies graph neural networks to process the abstract syntax tree of code, generating semantic vector representations of code entities;
[0015] Code change feature extraction: Compare two versions of the code and extract the structured features of code changes.
[0016] Furthermore, the steps for constructing the performance evolution analysis model include:
[0017] Performance metrics collection: Collect key performance metrics data for the app, including response time, memory usage, CPU utilization, battery consumption, startup time, and other multi-dimensional metrics.
[0018] Performance time series construction involves organizing the collected performance metric data into a time series by version, forming a performance evolution dataset.
[0019] Performance evolution pattern identification: Long Short-Term Memory (LSTM) networks are used to analyze performance time series data to identify long-term trends, periodic patterns, and outliers in performance metrics.
[0020] Performance status vectorization involves comprehensively processing multi-dimensional performance metrics for each version to generate a vector representation of the app's performance status.
[0021] Furthermore, the steps for constructing the code performance knowledge graph include:
[0022] Node construction uses code entities and performance metrics as nodes in the knowledge graph, with each node containing its corresponding vector representation and attribute information;
[0023] Edge relationship mining analyzes the statistical correlation between code changes and performance changes in historical versions, identifies potential causal relationships, and uses these relationships as edges in the knowledge graph.
[0024] Relationship weight calculation: Analyze the conditional probability relationship between code entities and performance indicators using a Bayesian network model, and calculate the weight value of the edges;
[0025] The knowledge graph is updated incrementally. With the release of new versions and the collection of new data, the knowledge graph is continuously updated using an incremental learning method, and the relationships and weights between nodes are dynamically adjusted.
[0026] Furthermore, the step of constructing the bi-state mapping prediction model includes:
[0027] The construction of a two-state continuous mapping model is based on the theory of code performance co-evolution. The model takes the code state and performance state at the current time point as input and outputs the predicted code state and performance state at the future time point.
[0028] Evolutionary mapping function training utilizes historical versions of code and performance data to train evolutionary mapping functions, including code state evolution functions and performance state evolution functions;
[0029] Path impact analysis, based on knowledge graphs, analyzes the impact of code changes on performance metrics through different paths, calculates the comprehensive impact score, and identifies key impact paths;
[0030] The performance prediction results are generated by taking the current code status and planned code changes as input. The prediction system generates a performance impact assessment report, including short-term direct impact and long-term cumulative impact, as well as the confidence interval of the impact.
[0031] Furthermore, the graph neural network includes:
[0032] The input layer receives node and edge information from the abstract syntax tree, with each node corresponding to a code element;
[0033] The embedding layer converts the input code elements into initial feature vectors, which are generated using a pre-trained code word embedding model.
[0034] The message passing layer contains multiple graph convolutional layers, which update the representation of the central node by recursively aggregating neighbor node information.
[0035] Attention mechanism layer: Introduces an attention mechanism to calculate the importance weights of different neighboring nodes;
[0036] The output layer generates the final semantic vector representation of the code entities.
[0037] Furthermore, the Long Short-Term Memory network includes:
[0038] The input layer receives time-series data of multi-dimensional performance metrics, with each time point corresponding to the performance status of a version.
[0039] An LSTM unit contains three gated units: an input gate, a forget gate, and an output gate, which can selectively memorize and forget information.
[0040] The multi-layer stacked structure uses a 3-layer LSTM unit stack to enhance the model's ability to capture long-term dependencies.
[0041] The attention layer introduces a temporal attention mechanism, which adaptively focuses on historical time points in the sequence where the attention weight value exceeds a preset threshold by calculating attention weights.
[0042] The fully connected layer performs a nonlinear transformation on the LSTM output to generate the final performance trend prediction result.
[0043] Furthermore, the code performance knowledge graph includes:
[0044] Node types include code entity nodes and performance metric nodes;
[0045] Node attributes: Each node has multi-dimensional attributes. Code entity nodes include semantic vectors, complexity metrics, and change frequency, while performance metric nodes include baseline values, fluctuation ranges, and importance weights.
[0046] Edge types include three types: direct impact, indirect impact, and potential impact, which represent different degrees of certainty in the relationship.
[0047] Edge attributes: Each edge has a weight, confidence level, and delay period, describing the strength of the impact, the degree of confidence, and the time characteristics.
[0048] Furthermore, the dual-state mapping prediction model includes:
[0049] The code state encoder encodes code entities and their relationships into fixed-dimensional vector representations, implemented using a graph convolutional network.
[0050] The performance status encoder encodes multidimensional performance metrics into fixed-dimensional vector representations, and is implemented using an autoencoder network.
[0051] The feature extractor is modified to extract structured features from code differences, and this is achieved by combining convolutional neural networks and recurrent neural networks.
[0052] The evolution predictor receives the current state and change information, predicts the future state, and is implemented using a sequence-to-sequence model based on the Transformer architecture.
[0053] A multi-task learning framework that simultaneously optimizes two tasks: code state prediction and performance state prediction, and improves the model's generalization ability by sharing parameters.
[0054] This invention provides a machine learning-based multi-dimensional analysis system for app performance, used to execute the aforementioned machine learning-based multi-dimensional analysis method for app performance, comprising:
[0055] The code semantic understanding module is used to analyze the APP source code and extract the structural features and semantic information of the code;
[0056] The performance evolution analysis module is used to identify long-term change patterns in performance metrics by collecting and analyzing performance data of the APP in different versions.
[0057] The code performance knowledge graph construction module is used to model the relationship between code changes and performance changes by using code entities and performance metrics as nodes and the influence relationships between them as edges.
[0058] The dual-state mapping prediction module is used to receive the code state and performance state at the current time as input to predict the short-term and long-term impacts of code changes on APP performance.
[0059] The beneficial effects of this invention are: by deeply understanding code semantics and long-term performance evolution patterns, it can accurately predict the impact of code changes on performance, which is an improvement over traditional rule-based performance evaluation methods;
[0060] This invention can predict potential performance issues during the code change phase, shifting the problem discovery time from post-deployment to the code review stage, reducing most online performance incidents. The system can also provide problem-solving suggestions to help developers fix potential performance risks more quickly.
[0061] By using code performance knowledge graph technology, we can intuitively show the long-term impact path of code changes on performance, enabling development teams to understand complex cause-and-effect relationships and support more informed technology decisions.
[0062] This invention provides a quantitative relationship model between code maintainability and performance optimization, helping teams achieve the best balance between the two, and improving the overall score of code quality and performance compared to traditional methods;
[0063] Compared with separate code review and performance monitoring systems, this invention reduces data fragmentation and analysis redundancy through a unified code performance knowledge graph framework. The system's incremental learning capability ensures that the knowledge graph can be continuously updated to adapt to constantly changing code.
[0064] This invention can identify the performance evolution differences caused by different code modification paths, providing development teams with a deeper perspective on performance optimization. The system can also predict code changes that are not directly related but have long-term impact, demonstrating a systematic understanding ability that goes beyond simple mapping relationships. Attached Figure Description
[0065] Figure 1 This is a flowchart of a multi-dimensional analysis method for APP performance based on machine learning, as described in this invention.
[0066] Figure 2 It is a line graph showing the evolution of APP performance metrics over version versions;
[0067] Figure 3 This is a radar chart comparing the performance of the method of this invention with that of traditional methods in five dimensions;
[0068] Figure 4 This is a Sankey diagram illustrating the path of code changes' impact on performance metrics.
[0069] Figure 5 This is a bar chart comparing the performance prediction accuracy of the method of this invention with that of traditional methods at different development stages;
[0070] Figure 6 It is a scatter plot showing the relationship between code change complexity and prediction accuracy;
[0071] Figure 7 This is a pie chart showing the distribution of the root causes of performance problems identified by this invention. Detailed Implementation
[0072] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0073] At least one embodiment of the present invention discloses a multi-dimensional analysis method for APP performance based on machine learning, such as... Figure 1 As shown, it includes:
[0074] Step 1: Build a code semantic understanding model, analyze the APP source code, and extract the structured features and semantic information of the code;
[0075] According to embodiments of this application, this step utilizes deep learning technology to construct a code semantic understanding model, analyzes the APP source code, and extracts the structured features and semantic information of the code. Specifically, it includes:
[0076] Step 1.1, code entity extraction;
[0077] Analyze the app's source code to identify and extract key code entities, including classes, methods, functions, and modules, generating a set of code entities. Optionally, during the code entity extraction process, filtering can be performed based on the importance and performance sensitivity of the code, prioritizing those code entities that may impact performance.
[0078] Step 1.2, code syntax structure analysis;
[0079] The source code is converted into an Abstract Syntax Tree (AST) to capture the syntactic structural relationships of the code, including call relationships, inheritance relationships, dependency relationships, etc. In some implementations, in addition to the AST, control flow graphs and data flow graphs can also be constructed to more comprehensively represent the structural features of the code.
[0080] Step 1.3, code semantic vectorization;
[0081] Graph Neural Networks (GNNs) are used to process the abstract syntax tree of code, generating semantic vector representations of code entities that capture the functional semantics and contextual relationships of the code.
[0082] The specific implementation of the graph neural network model includes the following structure:
[0083] Input layer: Receives node and edge information from the abstract syntax tree, with each node corresponding to a code element (such as a class, method, variable, etc.).
[0084] Embedding layer: Converts the input code elements into initial feature vectors, generated using a pre-trained code word embedding model;
[0085] Message passing layer: Contains multiple graph convolutional layers, which recursively aggregate neighbor node information to update the representation of the central node;
[0086] Attention Mechanism Layer: An attention mechanism is introduced to calculate the importance weights of different neighbor nodes, enhancing the model's ability to perceive key dependencies;
[0087] Output layer: Generates the final semantic vector representation of the code entity.
[0088] It should be noted that the neural network model in this graph uses a contrastive learning method during training, which makes code entities with similar functions closer together in the semantic space, and code entities with large functional differences farther apart, thereby achieving effective modeling of code semantics.
[0089] Step 1.4, code change feature extraction;
[0090] Compare the two versions of the code, extract the structured features of the code changes, including change type, scope, and complexity, and generate a code change feature vector. Optionally, when extracting change features, further analysis of the contextual information and potential impact of the changes can enhance the understanding of their effects.
[0091] Therefore, the key innovation of this step lies in using graph neural networks to process code structure, which achieves a deep understanding of code semantics, surpasses traditional rule-based static code analysis methods, and can capture complex relationships between code entities.
[0092] Step 2: Based on the extracted code structure features and semantic information, construct a performance evolution analysis model. By collecting and analyzing the performance data of the APP in different versions, identify the long-term change patterns of performance indicators.
[0093] According to one embodiment of this application, this step involves collecting and analyzing performance data of the app in different versions to construct a performance evolution analysis model and identify long-term change patterns in performance metrics. Specifically, this includes:
[0094] Step 2.1, performance metrics collection;
[0095] Collect key performance metrics data for the app across multiple versions, including response time, memory usage, CPU utilization, battery consumption, startup time, and other multi-dimensional metrics. It should be noted that in practice, different combinations of performance metrics can be flexibly selected based on the app's characteristics and priorities. For example, for interaction-intensive apps, more attention can be paid to metrics such as interaction response time and frame rate.
[0096] Step 2.2, Performance time series construction;
[0097] The collected performance metric data is organized into time series by version to form a performance evolution dataset. Optionally, the raw data can be preprocessed, such as denoising, standardization, and outlier handling, to improve the accuracy of subsequent analysis.
[0098] Step 2.3, performance evolution pattern identification;
[0099] Long Short-Term Memory (LSTM) networks are used to analyze performance time series data to identify long-term trends, periodic patterns, and outliers in performance metrics.
[0100] The specific implementation of the Long Short-Term Memory (LSTM) network model includes the following structure:
[0101] Input layer: Receives time-series data of multi-dimensional performance metrics, with each time point corresponding to the performance status of a version;
[0102] LSTM unit: It contains three gated units: input gate, forget gate and output gate, which can selectively remember and forget information;
[0103] Multi-layer stacked structure: A 3-layer LSTM unit stack is used to enhance the model's ability to capture long-term dependencies;
[0104] Attention layer: Introduces a temporal attention mechanism, which adaptively focuses on historical time points in the sequence where the attention weight value exceeds a preset threshold by calculating attention weights. Through statistical analysis of the attention weights in the training dataset, the mean of the attention weights plus a standard deviation is used as the preset threshold.
[0105] Fully connected layer: Performs a nonlinear transformation on the LSTM output to generate the final performance trend prediction result.
[0106] It should be understood that this Long Short-Term Memory (LSTM) network model uses the sliding window method during training, storing historical data... Using performance data from one version as input, the system predicts the performance of the next version and optimizes it using the mean squared error loss function, thereby achieving accurate modeling of performance evolution trends. In some implementations, ensemble learning methods can also be used to combine the prediction results of multiple models to further improve the accuracy and stability of the prediction.
[0107] Step 2.4, performance state vectorization;
[0108] The multi-dimensional performance metrics for each version are comprehensively processed to generate a vector representation of the app's performance status. This vector can comprehensively reflect the app's performance in a specific version. Optionally, a weighting mechanism can be introduced to assign different weights based on the degree of impact of different performance metrics on user experience, constructing a performance status representation that is more in line with user perception.
[0109] Therefore, the key innovation of this step lies in using long short-term memory networks to process performance time series data, which enables the modeling of long-term performance evolution trends and captures the gradual patterns and periodic characteristics of performance changes, laying the foundation for subsequent correlation analysis.
[0110] like Figure 2 As shown in the figure, the changes in two key performance indicators, response time and memory usage, of the APP during multiple version iterations are illustrated. This figure intuitively reflects the performance evolution trajectory and verifies the system's ability to model long-term performance evolution trends.
[0111] like Figure 3 As shown, the radar chart compares the prediction accuracy of this system with traditional methods from five dimensions: response time, memory usage, CPU utilization, battery consumption, and startup time, demonstrating the comprehensive advantages of the system in multi-dimensional performance analysis.
[0112] Step 3: Based on the code entity information output by the code semantic understanding model and the performance index data output by the performance evolution analysis model, construct a code performance knowledge graph, with code entities and performance indicators as nodes and the influence relationships between them as edges, to realize the systematic modeling of the relationship between code changes and performance changes.
[0113] According to embodiments of this application, this step uses code entities and performance metrics as nodes, and the influence relationships between them as edges, to construct a code performance knowledge graph, thereby achieving systematic modeling of the relationship between code changes and performance changes. Specifically, it includes:
[0114] Step 3.1, Node Construction;
[0115] The code entities obtained in step 1 and the performance metrics obtained in step 2 are used as nodes in the knowledge graph. Each node contains its corresponding vector representation and attribute information. It should be noted that during the node construction process, code entities of different granularities can be selected according to actual needs, ranging from coarse-grained modules to fine-grained functions or methods.
[0116] Step 3.2, edge relationship mining;
[0117] Analyzing the statistical correlation between code changes and performance variations in historical versions identifies potential causal relationships, which are then used as edges in a knowledge graph. In some implementations, in addition to statistical correlation, expert knowledge and heuristic rules can be combined to help discover more reliable causal relationships.
[0118] Step 3.3, Calculate relation weights;
[0119] Using a Bayesian network model, we analyze the conditional probability relationship between code entities and performance metrics, and calculate the weights of the edges. This reflects the strength and confidence level of the impact.
[0120] The specific implementation of the code performance knowledge graph includes:
[0121] Node types include code entity nodes (classes, methods, functions, modules, etc.) and performance metric nodes (response time, memory usage, CPU utilization, etc.).
[0122] Node attributes: Each node It has multi-dimensional attributes; code entity nodes include semantic vectors, complexity indicators, change frequency, etc.; performance indicator nodes include benchmark values, fluctuation range, importance weight, etc.
[0123] Edge types: There are three types: direct impact, indirect impact, and potential impact, which represent different degrees of certainty in the relationship.
[0124] Edge attributes: Each edge has a weight. Attributes such as confidence level and delay period describe the strength of the impact, the degree of credibility, and the time characteristics.
[0125] It should be understood that the knowledge graph construction process employs a semi-supervised learning method, combining expert rules and data-driven approaches to iteratively optimize the graph structure and parameters. Simultaneously, the graph supports multi-level representations, capable of describing code performance relationships at different levels of abstraction, from coarse-grained module-level to fine-grained function-level. Optionally, knowledge distillation techniques can be introduced to extract concise key knowledge from large-scale code performance data, reducing graph complexity while retaining core information.
[0126] Step 3.4, Incremental update of the knowledge graph;
[0127] With the release of new versions and the collection of new data, the knowledge graph is continuously updated using incremental learning methods, dynamically adjusting the relationships and weights between nodes. Optionally, an aging mechanism can be set to gradually reduce the weight of data that is too old, making the graph more closely reflect the actual state of the current codebase.
[0128] Therefore, the key innovation of this step lies in using knowledge graph technology to explicitly model the relationship between code performance, realizing the representation and analysis of complex relational networks, capturing direct and indirect influence paths, and providing a structured knowledge foundation for performance prediction.
[0129] like Figure 4 As shown, the Sankey diagram, which analyzes the impact path of code changes on performance metrics, demonstrates how code changes affect different performance metrics through various intermediate paths (data structure changes, algorithm complexity changes, concurrency model changes, and resource management changes), ultimately verifying the system's ability to analyze complex impact paths.
[0130] Step 4: Based on the node relationships and edge weight information in the code performance knowledge graph, construct a dual-state mapping prediction model, which takes the code state and performance state at the current time point as input, and predicts the short-term and long-term impacts of code changes on APP performance.
[0131] According to one embodiment of this application, this step, based on the model and knowledge graph constructed in the previous three steps, implements a bi-state mapping prediction model, which can accurately predict the short- and long-term impacts of code changes on APP performance. Specifically, it includes:
[0132] Step 4.1, Construction of the two-state continuous mapping model;
[0133] Based on the "code performance co-evolution theory," a two-state continuous mapping model is constructed. This model takes the code state and performance state at the current time point as input and outputs the predicted code state and performance state at future time points, thus describing the co-evolution of code state and performance state over time. It should be noted that this model not only considers the impact of code changes on performance but also the feedback effect of performance state on subsequent code evolution, forming a closed-loop two-way influence model.
[0134] The dual-state continuous mapping model is specifically implemented as a composite function system composed of multiple neural networks. This system first transforms the code state and performance state into representations in the latent space using an encoder, and then predicts future states through a series of nonlinear transformations. Specifically, the system uses an attention mechanism to capture the interaction between the code state and performance state, and then generates predictions of future states through a multi-layer feedforward neural network. This mapping function not only considers the current state but also incorporates historical evolution information, making the predictions more accurate.
[0135] Step 4.2, Evolutionary mapping function training;
[0136] Using historical code and performance data, evolutionary mapping functions are trained, including the code state evolution function (CSE, CodeStateEvolution) and the performance state evolution function (PSE, PerformanceStateEvolution). The code state evolution function takes the current code state, the current performance state, and code changes as input and outputs the predicted future code state; the performance state evolution function also takes the current code state, the current performance state, and code changes as input and outputs the predicted future performance state.
[0137] The specific implementation of the evolutionary mapping function is as follows:
[0138] For the code state evolution function, the current code state and the performance state are first connected and their features are fused through a fully connected layer to obtain a fused state representation. Then, the code change is converted into a change representation through a feature extraction network. Next, an attention mechanism is used to calculate the part of the fused state representation that is most relevant to the change. Finally, the impact of the change is combined with the original code state through residual connections to generate a prediction of the future code state.
[0139] For the performance state evolution function, we first analyze the propagation path of code changes on the knowledge graph using a graph neural network (GNN) to identify the affected performance nodes; then we calculate the weight of each influencing path. The system then considers latency characteristics; next, it predicts performance changes at future time points based on the current performance state and path impact; finally, it integrates the predictions at each time point into a complete representation of the future performance state.
[0140] The specific implementation of the bi-state mapping prediction model includes the following components:
[0141] Code state encoder: Encodes code entities and their relationships into fixed-dimensional vector representations, implemented using graph convolutional networks;
[0142] Performance Status Encoder: Encodes multidimensional performance metrics into fixed-dimensional vector representations, implemented using an autoencoder network;
[0143] Modify the feature extractor: extract structured features from code differences, using a combination of convolutional neural networks and recurrent neural networks;
[0144] Evolution predictor: Receives current state and change information, predicts future state, and is implemented using a sequence-to-sequence model based on the Transformer architecture;
[0145] Multi-task learning framework: Simultaneously optimizes two tasks: code state prediction and performance state prediction, and improves the model's generalization ability by sharing parameters.
[0146] It should be understood that the prediction model is trained in an end-to-end manner, simultaneously minimizing both code state prediction errors and performance state prediction errors to achieve accurate modeling of the co-evolution of code and performance. The model also incorporates an uncertainty estimation mechanism, providing a confidence interval for each prediction to reflect its reliability. In some implementations, an active learning strategy can be employed to identify predictions with high uncertainty, guiding the system to collect more relevant data to improve the prediction accuracy in these areas.
[0147] Step 4.3, Path Impact Analysis;
[0148] Based on knowledge graphs, the impact of code changes on performance metrics through different paths is analyzed, a comprehensive impact score is calculated, and key impact paths are identified. Optionally, explainable AI technology can be used to generate visual explanations of path impacts, helping developers understand complex causal relationships.
[0149] Step 4.4, Performance prediction results are generated;
[0150] Given the current code state and planned code changes, the prediction system generates a performance impact assessment report, including short-term direct impacts and long-term cumulative impacts, as well as the confidence intervals for the impacts. In some implementations, it can also generate multiple possible evolution scenarios to help the development team assess the potential risks and benefits of different decisions.
[0151] Therefore, the key innovation of this step lies in realizing the bi-state mapping relationship between code changes and performance impacts, which can simulate the co-evolution process of code and performance states, predict the long-term performance impact of code changes, and in particular, discover code modifications that appear unrelated but actually have long-term effects.
[0152] like Figure 5As shown, the performance prediction accuracy of this system and traditional methods were compared at different development stages (code review, deployment testing, online operation), verifying the system's predictive advantage in the early stages of code changes.
[0153] like Figure 6 As shown, the relationship between code change complexity and performance prediction accuracy is illustrated. As code change complexity increases, prediction accuracy gradually decreases, but this system still maintains a high accuracy under high complexity conditions.
[0154] like Figure 7 As shown, the distribution of root causes of performance problems identified by the system is illustrated, including code quality issues, algorithm efficiency issues, resource management issues, and concurrency issues, verifying the system's ability to identify different types of performance problems.
[0155] A machine learning-based multi-dimensional performance analysis system for apps, used to execute the aforementioned machine learning-based multi-dimensional performance analysis method for apps, includes:
[0156] The code semantic understanding module is used to analyze the APP source code and extract the structural features and semantic information of the code;
[0157] The performance evolution analysis module is used to identify long-term change patterns in performance metrics by collecting and analyzing performance data of the APP in different versions.
[0158] The code performance knowledge graph construction module is used to model the relationship between code changes and performance changes by using code entities and performance metrics as nodes and the influence relationships between them as edges.
[0159] The dual-state mapping prediction module is used to receive the code state and performance state at the current time as input to predict the short-term and long-term impacts of code changes on APP performance.
[0160] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A multi-dimensional analysis method for APP performance based on machine learning, characterized in that, include: Build a code semantic understanding model to analyze the APP source code and extract the structured features and semantic information of the code; Based on the extracted code structure features and semantic information, a performance evolution analysis model is constructed. By collecting and analyzing the performance data of the APP in different versions, the long-term change patterns of performance indicators are identified. The steps involved in constructing a performance evolution analysis model include: Performance metrics collection: Collect key performance metrics data for the app, including response time, memory usage, CPU utilization, battery consumption, startup time, and other multi-dimensional metrics. Performance time series construction involves organizing the collected performance metric data into a time series by version, forming a performance evolution dataset. Performance evolution pattern identification: Long Short-Term Memory (LSTM) networks are used to analyze performance time series data to identify long-term trends, periodic patterns, and outliers in performance metrics. Performance status vectorization involves comprehensively processing multi-dimensional performance metrics for each version to generate a vector representation of the app's performance status. Based on the code entity information output by the code semantic understanding model and the performance index data output by the performance evolution analysis model, a code performance knowledge graph is constructed, with code entities and performance indicators as nodes and the influence relationships between them as edges, to realize the systematic modeling of the relationship between code changes and performance changes. Based on the node relationships and edge weight information in the code performance knowledge graph, a dual-state mapping prediction model is constructed. It takes the code state and performance state at the current time point as input and predicts the short-term and long-term impacts of code changes on APP performance. The steps to construct a bi-state mapping prediction model include: The construction of a two-state continuous mapping model is based on the theory of code performance co-evolution. The model takes the code state and performance state at the current time point as input and outputs the predicted code state and performance state at the future time point. Evolutionary mapping function training utilizes historical versions of code and performance data to train evolutionary mapping functions, including code state evolution functions and performance state evolution functions; Path impact analysis, based on knowledge graphs, analyzes the impact of code changes on performance metrics through different paths, calculates the comprehensive impact score, and identifies key impact paths; The performance prediction results are generated by taking the current code status and planned code changes as input. The prediction system generates a performance impact assessment report, including short-term direct impact and long-term cumulative impact, as well as the confidence interval of the impact. The specific implementation of the evolutionary mapping function includes: For the code state evolution function, the current code state and the performance state are first connected and their features are fused through a fully connected layer to obtain a fused state representation. Then, the code change is converted into a change representation through a feature extraction network. Next, an attention mechanism is used to calculate the part of the fused state representation that is most relevant to the change. Finally, the impact of the change is combined with the original code state through residual connections to generate a prediction of the future code state. For the performance state evolution function, firstly, the propagation path of code changes on the knowledge graph is analyzed through a graph neural network to identify the affected performance nodes; then, the weight and latency characteristics of each influencing path are calculated; next, the performance changes at each future time point are predicted based on the current performance state and the influence of the paths; finally, the predictions at each time point are combined into a complete representation of the future performance state. Two-state mapping prediction models include: The code state encoder encodes code entities and their relationships into fixed-dimensional vector representations, implemented using a graph convolutional network. The performance status encoder encodes multidimensional performance metrics into fixed-dimensional vector representations, and is implemented using an autoencoder network. The feature extractor is modified to extract structured features from code differences, and this is achieved by combining convolutional neural networks and recurrent neural networks. The evolution predictor receives the current state and change information, predicts the future state, and is implemented using a sequence-to-sequence model based on the Transformer architecture. A multi-task learning framework that simultaneously optimizes two tasks: code state prediction and performance state prediction, and improves the model's generalization ability by sharing parameters.
2. The method for multi-dimensional analysis of APP performance based on machine learning according to claim 1, characterized in that, The steps for constructing the code semantic understanding model include: Code entity extraction identifies and extracts key code entities, including classes, methods, functions, and modules, and generates a code entity set. Code syntax structure analysis converts source code into an abstract syntax tree, capturing the syntactic structure relationships of the code; Code semantic vectorization applies graph neural networks to process the abstract syntax tree of code, generating semantic vector representations of code entities; Code change feature extraction: Compare two versions of the code and extract the structured features of code changes.
3. The method for multi-dimensional analysis of APP performance based on machine learning according to claim 1, characterized in that, The steps for constructing the code performance knowledge graph include: Node construction uses code entities and performance metrics as nodes in the knowledge graph, with each node containing its corresponding vector representation and attribute information; Edge relationship mining analyzes the statistical correlation between code changes and performance changes in historical versions, identifies potential causal relationships, and uses these relationships as edges in the knowledge graph. Relationship weight calculation: Analyze the conditional probability relationship between code entities and performance indicators using a Bayesian network model, and calculate the weight value of the edges; The knowledge graph is updated incrementally. With the release of new versions and the collection of new data, the knowledge graph is continuously updated using an incremental learning method, and the relationships and weights between nodes are dynamically adjusted.
4. The method for multi-dimensional analysis of APP performance based on machine learning according to claim 1, characterized in that, The long short-term memory network includes: The input layer receives time-series data of multi-dimensional performance metrics, with each time point corresponding to the performance status of a version. An LSTM unit contains three gated units: an input gate, a forget gate, and an output gate, which can selectively memorize and forget information. The multi-layer stacked structure uses a 3-layer LSTM unit stack to enhance the model's ability to capture long-term dependencies. The attention layer introduces a temporal attention mechanism to adaptively focus on the historical time points in the sequence that contribute the most to the prediction results. The fully connected layer performs a nonlinear transformation on the LSTM output to generate the final performance trend prediction result.
5. The method for multi-dimensional analysis of APP performance based on machine learning according to claim 2, characterized in that, The graph neural network includes: The input layer receives node and edge information from the abstract syntax tree, with each node corresponding to a code element; The embedding layer converts the input code elements into initial feature vectors, which are generated using a pre-trained code word embedding model. The message passing layer contains multiple graph convolutional layers, which update the representation of the central node by recursively aggregating neighbor node information. Attention mechanism layer: Introduces an attention mechanism to calculate the importance weights of different neighboring nodes; The output layer generates the final semantic vector representation of the code entities.
6. The method for multi-dimensional analysis of APP performance based on machine learning according to claim 3, characterized in that, The code performance knowledge graph includes: Node types include code entity nodes and performance metric nodes; Node attributes: Each node has multi-dimensional attributes. Code entity nodes include semantic vectors, complexity metrics, and change frequency, while performance metric nodes include baseline values, fluctuation ranges, and importance weights. Edge types include three types: direct impact, indirect impact, and potential impact, which represent different degrees of certainty in the relationship. Edge attributes: Each edge has a weight, confidence level, and delay period, describing the strength of the impact, the degree of confidence, and the time characteristics.
7. A multi-dimensional performance analysis system for apps based on machine learning, characterized in that, A method for performing a multi-dimensional analysis of APP performance based on machine learning as described in any one of claims 1-6 includes: The code semantic understanding module is used to analyze the APP source code and extract the structural features and semantic information of the code; The performance evolution analysis module is used to identify long-term change patterns in performance metrics by collecting and analyzing performance data of the APP in different versions. The code performance knowledge graph construction module is used to model the relationship between code changes and performance changes by using code entities and performance metrics as nodes and the influence relationships between them as edges. The dual-state mapping prediction module is used to receive the code state and performance state at the current time as input to predict the short-term and long-term impacts of code changes on APP performance.
Citation Information
Patent Citations
Software development data processing system based on data analysis
CN118152221A
Nursing risk intervention decision-making system and method based on knowledge graph
CN119560121A