APP Response Time Analysis Method Based on Deep Neural Network and Distributed Computing
The method leverages deep learning and distributed computing to improve APP performance analysis accuracy and efficiency by frame-wise image processing and model-based analysis, addressing limitations of existing methods and enabling automated, comprehensive performance evaluation.
Patent Information
- Application Number
- CN202510572269.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing APP response time analysis methods have problems such as insufficient accuracy, low efficiency and excessive human resources investment. Especially in dynamic content loading scenarios, it is difficult to accurately judge the full loading status of complex UI elements, and it is not suitable for competitive product analysis.
Using a method based on deep neural network and distributed computing, the video is loaded by recording APP pages and allocated to different computing nodes for pre-processing and post-processing, the trained classification model is used for frame-by-frame analysis, and data processing and result summary are combined with distributed CPU and GPU server.
It improves the accuracy and efficiency of APP performance analysis, can more truly reflect user perception experience, automates complex data analysis, reduces the need for manual intervention, and is suitable for APP analysis of your own and competitors.
Smart Images

Figure CN120091126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for analyzing the response time of an APP, specifically a method for analyzing the response time of an APP based on a deep neural network and distributed computing, belonging to the technical field of software testing. Background Art
[0002] With the wide popularization of mobile applications (APPs) and the increasing requirements for user experience, their application performance has become a key factor affecting user retention rate and brand value. Poor-performing APPs will not only lead to user loss and damage to brand reputation, but also directly affect commercial conversion rate and user satisfaction.
[0003] In the field of APP page loading performance testing, there are currently the following several mainstream methods:
[0004] Android Device Monitor. This method directly depends on the resources provided by the development tools and is suitable for developers to quickly evaluate the startup time. However, in actual use, data deviation may occur due to factors such as system load.
[0005] ADB Shell commands. This method is intuitive and easy to use, but the output time only represents the recording time point of the system log, and there is a significant difference from the visually perceived loading completion time of users, especially in the case of dynamic content loading, the error is more obvious.
[0006] Code instrumentation. Although this method is simple to implement and can capture the accurate internal execution time, it has three major limitations: 1) It is only applicable to its own APP with access to the source code and cannot be used for competitive product analysis; 2) It is often difficult to accurately judge the fully loaded state of complex UI elements, especially in asynchronous loading and dynamic rendering scenarios; 3) The instrumentation process itself will generate additional overhead on the application performance, affecting the measurement accuracy. Therefore, its application scope is limited and it is difficult to meet the comprehensive performance evaluation requirements.
[0007] High-speed camera or screen recording analysis. The traditional pixel-based analysis method faces severe challenges when processing dynamic pages: 1) It is difficult to determine the optimal comparison threshold, resulting in inaccurate judgment; 2) The computational complexity is high, and the efficiency is extremely low when processing a large number of video files; 3) It lacks adaptability to applications with different visual styles and requires frequent manual intervention to adjust parameters. Summary of the Invention
[0008] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a method for analyzing the response time of an APP based on a deep neural network and distributed computing, which splits the preprocessing and postprocessing parts and distributes them to different computing nodes for processing to achieve more accurate, efficient and automated APP performance analysis.
[0009] Technical solution: The APP response time analysis method based on deep neural network and distributed computing provided by the present invention on the one hand includes the following steps:
[0010] Record the video of the APP page loading and upload it to the file server;
[0011] Use the distributed CPU server to obtain the video from the file server, split the video into frame-by-frame pictures, preprocess the pictures, generate a data format file for classification model inference, and upload the data format file to the file server;
[0012] Obtain the arrival timestamp information of each frame of the image, associate the arrival timestamp information with the corresponding image data; store the associated arrival timestamp information in the local storage medium of the distributed CPU server in a structured data format;
[0013] Use the central GPU server to receive the analysis task request of the APP response time, obtain the preprocessed data format file from the file server, use the trained classification model to perform inference on the data format file, and upload the inference result to each distributed CPU server;
[0014] Use the distributed CPU server to post-process the received inference result, obtain the start frame and end frame of each page loading, calculate the page loading time, and simultaneously obtain the start screenshot and end screenshot during each page loading;
[0015] Use the central GPU server to summarize the analysis results of the distributed CPU servers, store the analysis results in the database, and provide an interface for query.
[0016] Further, before the step of using the trained classification model to perform inference on the data format file includes:
[0017] Use the data set and labels to train the classification model.
[0018] Further, the step of using the data set and labels to train the classification model includes:
[0019] Construct a data set, where the data set is composed of videos obtained by disassembling the APP page loading, and each frame of the video is divided into different page states, including stable state, loading state and click state pictures, and the page states constitute different class labels;
[0020] Construct a classification model, where the classification model is a deep neural network;
[0021] Construct a loss function;
[0022] Use the images in the constructed dataset as the input of the classification model. The output of the classification model is the class probability score corresponding to each image. Use the loss function to measure the difference between the model prediction and the actual label, and use gradient descent to optimize the classification model.
[0023] Furthermore, the steps of using the distributed CPU server to post-process the received inference results include:
[0024] Obtain the start frame number and end frame number of all click states, and record them as a list , denote the i-th tuple. The list includes n tuples, and each tuple is denoted as , and in each tuple denotes the frame number when the i-th click state starts to click, denotes the frame number when the i-th click state releases the click;
[0025] Traverse each tuple in the list in sequence, obtain the count of different state frames between the end frame of the current click state and the start frame of the next click state, select the loading state with the largest number of occurrences and at least 3 frames, and use this loading state as the loading state triggered by the current click operation; if not found, skip the current tuple, traverse the next tuple, and calculate the start frame number and end frame number when the current state appears after finding it;
[0026] Regard all blank loading states as the same class, and reset the start frame of the loading state to the next frame of the end frame of the current click state;
[0027] Summarize the start frame number and end frame number of all loading states, use the timestamp information file to query the arrival time of the start frame and end frame, calculate the actual page loading time, and obtain the images when each page starts loading and ends loading.
[0028] Furthermore, the steps of preprocessing the images include:
[0029] Use the image processing library to read the video and obtain each frame of the image frame by frame;
[0030] Perform a scaling operation on each frame of the image;
[0031] Perform a normalization operation on the scaled image;
[0032] Perform standardization processing with the mean and standard deviation;
[0033] Concatenate the processed frame data;
[0034] Perform lossless compression on the concatenated data and store the compressed data in a structured format.
[0035] On the other hand, the present invention provides an APP response time analysis system based on a deep neural network and distributed computing, including:
[0036] A file server for storing the original video files, preprocessed data, and processing results;
[0037] Multiple distributed CPU servers for preprocessing video files and post-processing the inference results of the classification model;
[0038] A central GPU server for creating analysis tasks, performing classification model inference, summarizing processing results, and providing a query interface.
[0039] Furthermore, the system further includes a task scheduling module for coordinating the communication and data transmission among the distributed CPU servers, the central GPU server, and the file server.
[0040] Furthermore, the system adopts an asynchronous communication mechanism.
[0041] Furthermore, the system adopts a fault tolerance mechanism.
[0042] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0043] 1. Improve the accuracy of APP performance analysis: The present invention uses a deep neural network for frame-by-frame analysis, which can more truly reflect the user's perceptual experience, thereby improving the accuracy of page loading time measurement. This means that it can objectively analyze not only its own applications but also the APPs of competitors.
[0044] 2. Enhance processing efficiency: The present invention adopts a distributed computing method, which can distribute the preprocessing and post-processing tasks to different computing nodes, make full use of the computing resources of the servers, and speed up the data processing speed. And the ability of multi-node parallel processing of videos greatly shortens the overall analysis time, meeting the needs of large-scale data processing.
[0045] 3. Save human resource investment: The present invention can automatically process complex data analysis tasks, reduce the need for manual intervention, and thus reduce the labor cost of testing and analysis. This automated analysis method not only improves the work efficiency of the team but also can more quickly discover performance problems during the testing process. Description of the Drawings
[0046] Figure 1 is the overall flowchart of the method in the embodiment of the present invention;
[0047] Figure 2 is the UML sequence diagram of distributed computing in the embodiment of the present invention;
[0048] Figure 3 It is a flowchart for preprocessing a video in an embodiment of the present invention;
[0049] Figure 4 It is a flowchart for post-processing the model inference result in an embodiment of the present invention. Specific embodiments
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0051] Embodiment 1
[0052] The method for analyzing the APP response time based on deep neural network and distributed computing described in this embodiment has a flowchart as Figure 1 shown, combined with Figure 2 the UML sequence diagram of the distributed computing shown, the method includes the following steps:
[0053] Step 1, record the video of the APP page loading and upload it to the file server.
[0054] Step 2, use the distributed CPU server to obtain the video from the file server, split the video into frame-by-frame pictures, preprocess the pictures, generate a data format file for classification model inference, and upload the data format file to the file server.
[0055] The number of the distributed CPU servers can be dynamically adjusted according to the processing requirements, and the load balancing is achieved through the task queue to ensure the efficient processing of a large number of video files. When a certain distributed CPU server fails to process, the system can automatically transfer the task to other available distributed CPU servers to ensure the continuity and reliability of the analysis task. Use the stress testing server to record the video and upload the video file to the file server, download the video stored in the file server to the distributed CPU server, read the video using the image processing library, and then split the video into frame-by-frame pictures, and obtain each frame picture frame by frame, where the image processing library includes but is not limited to OpenCV, FFmpeg, etc.
[0056] Combined with Figure 3 shown, further, the steps of preprocessing the pictures include:
[0057] Use an image processing library such as OpenCV to read the video and obtain each frame picture frame by frame;
[0058] Perform a scaling operation on each frame picture, for example, it can be uniformly scaled to the size of;
[0059] Perform a normalization operation on the scaled picture to make its pixel value between 0 and 1;
[0060] Normalize with a mean and standard deviation of 0.5;
[0061] Concatenate the processed frame data in batches of 256;
[0062] Use the DEFLATE algorithm to losslessly compress the concatenated data and store the compressed data in a structured format, such as an npz file.
[0063] Step 3: Obtain the arrival timestamp information for each frame of the image, associate the arrival timestamp information with the corresponding image data; store the associated arrival timestamp information in a structured data format in the local storage medium of the distributed CPU server.
[0064] The arrival timestamp of each frame of the image can be obtained using the FFmpeg image processing library and saved as a JSON file in the local area of each distributed CPU server.
[0065] Step 4: Use the central GPU server to receive the analysis task request for the APP response time, obtain the preprocessed data format file from the file server, use the trained classification model to perform inference on the data format file, and upload the inference result to each distributed CPU server.
[0066] Dataset preparation: To enhance the classification model's ability to judge the starting frame of page loading, it is necessary to enable the pointer position in the developer mode on the mobile phone. After triggering a click operation, the touch point and its coordinates will be displayed on the screen, providing accurate visual features for the classification model as the basis for judgment.
[0067] Furthermore, before the step of using the trained classification model to perform inference on the data format file, it includes:
[0068] Train the classification model using the dataset and labels.
[0069] Furthermore, the steps of training the classification model using the dataset and labels include:
[0070] Construct a dataset, where the dataset consists of videos of disassembled APP page loading. Each frame of the video is divided into different page states, including stable state, loading state, and click state pictures. The page states constitute different class labels;
[0071] Construct a classification model, where the classification model is a deep neural network;
[0072] Construct a loss function;
[0073] Use the images in the constructed dataset as the input of the classification model. The output of the classification model is the class probability score corresponding to each image. Use the loss function to measure the difference between the model prediction and the actual label, and use gradient descent to optimize the classification model.
[0074] In one example, the deep neural network can adopt the ResNet18 network model. The ResNet18 network model is a deep residual network. Due to its efficient classification ability and low computational overhead, it becomes an ideal choice for classification tasks. Use the cross-entropy loss function to measure the difference between the model prediction and the actual label. This is a loss function commonly used in multi-class classification problems. The optimizer uses Adam, which is an adaptive learning rate optimization algorithm that can efficiently perform gradient descent optimization and quickly converge to the optimal solution.
[0075] Since there may be large blank areas during the page loading process, and the blank pages of different pages may be extremely similar, it is necessary to uniformly classify all blank loading pages into the same category to reduce classification errors. The training dataset is obtained by disassembling the APP page loading video. Then, manually divide the video frames into different page states, including stable state, loading state, and click state images, to form different category annotations.
[0076] After the classification model inference is completed, the category to which each frame of the video belongs can be obtained. According to the classification results of the classification model, all states can be summarized into three categories: loading state, stable state, and click state. The specific categories included in each major category are shown in Table 1.
[0077] Table 1
[0078]
[0079] Step 5, use the distributed CPU server to post-process the received inference results to obtain the start frame and end frame of each page loading, calculate the page loading time, and at the same time obtain the start screenshot and end screenshot during each page loading process.
[0080] Combined with Figure 4 Furthermore, the steps of using the distributed CPU server to post-process the received inference results include:
[0081] Obtain the start frame numbers and end frame numbers of all click states, denoted as the list , represents the i-th tuple. The list includes n tuples, and each tuple is denoted as , in each tuple represents the frame number at which the i-th click state starts to click, represents the frame number at which the i-th click state releases the click;
[0082] Traverse each tuple in the list in sequence, obtain the count of different state frames between the end frame of the current click state and the start frame of the next click state, select the loading state with the largest number of occurrences and at least 3 frames as the loading state triggered by the current click operation; if not found, skip the current tuple and traverse the next tuple, and calculate the start frame number and end frame number when the current state appears after finding it;
[0083] Regard all blank loading states as the same class, and reset the start frame of the loading state to the next frame of the end frame of the current click state;
[0084] Summarize the start frame numbers and end frame numbers of all loading states, query the arrival times of the start frame and end frame using the timestamp JSON file, calculate the actual page loading time, and obtain the pictures when each page starts loading and ends loading.
[0085] In an example, first, obtain the start and end frame numbers of all click states from it, such as [(10, 20), (30, 40), (60, 70)], where the list contains 3 tuples, and each tuple represents the start click frame number and the release click frame number of a click state. Traverse each tuple in sequence, obtain the count of different state frames between the end frame of the current click state and the start frame of the next click state, select the loading state with the largest number of occurrences and at least 3 frames as the loading state triggered by the current click operation, and if not found, skip and enter the next tuple loop. After finding it, calculate the start and end frames when the current state appears. Due to possible model classification errors and other situations, there may be multiple segments of loading states between two click frames. At the same time, some pages are blank when loading, and it is impossible to distinguish which loading state it is. During training, all blank loading states are regarded as the same class. Therefore, multiple loading states of the same class and blank loading states will be merged. And because there is no change in the page within a few frames before and after the user triggers a click operation, most cases within a few frames after the start frame of the click state will be classified as the previous state of the click state. Therefore, the start frame of the loading state will be reset to the next frame of the end frame of the current click state. Finally, summarize the start frame numbers and end frame numbers of all loading states, query the arrival times of the start frame and end frame using the timestamp JSON file, calculate the actual page loading time, and obtain the pictures when each page starts loading and ends loading.
[0086] Step 6: Use the central GPU server to summarize the analysis results of the distributed CPU servers, store the analysis results in the database, and provide an interface for querying.
[0087] The central GPU server is used to aggregate the analysis results of distributed CPU servers, store the analysis results in a database, and provide an interface for querying. The analysis results include the name of each loading state of the APP, the start frame number, the start frame screenshot, the start frame timestamp, the end frame number, the end frame screenshot, the end frame timestamp, and the loading time.
[0088] In summary, this example provides an APP response time analysis method based on deep neural networks and distributed computing. By using deep neural networks for frame-by-frame analysis, this method can more realistically reflect the user's perceptual experience, thereby improving the accuracy of page loading time measurement. This means that it is not limited to its own application programs, but can also objectively analyze the APPs of competitors; by adopting a distributed computing strategy, the preprocessing and postprocessing tasks can be assigned to different computing nodes, making full use of the computing resources of the servers and accelerating the data processing speed. The ability of multiple nodes to process videos in parallel greatly shortens the overall analysis time and meets the needs of large-scale data processing; it can automatically process complex data analysis tasks, reduce the need for manual intervention, and thus reduce the labor cost of testing and analysis.
[0089] Embodiment 2
[0090] The APP response time analysis system based on deep neural networks and distributed computing described in this embodiment includes:
[0091] A file server for storing the original video files, preprocessed data, and processing results;
[0092] Multiple distributed CPU servers for preprocessing video files and postprocessing the model inference results;
[0093] A central GPU server for creating analysis tasks, performing classification model inference, aggregating processing results, and providing a query interface.
[0094] Furthermore, the system further includes a task scheduling module for coordinating the communication and data transmission between the distributed CPU servers, the central GPU server, and the file server to ensure that each component works together according to a predefined timing process.
[0095] Furthermore, the system adopts an asynchronous communication mechanism, enabling each server component to process tasks in parallel and improving the overall throughput and performance of the system.
[0096] Furthermore, the system adopts a fault tolerance mechanism. When a certain distributed CPU server fails to process, the system can automatically transfer the task to other available servers to ensure the continuity and reliability of the analysis task.
[0097] The system uses distributed computing to accelerate the data processing speed and utilizes the data flow of distributed computing, including:
[0098] a) The stress testing server is responsible for recording videos and uploading the video files to the file server;
[0099] b) The file server is responsible for storing the original video files, preprocessed data, and processing results;
[0100] c) The distributed CPU server is responsible for preprocessing the video files, receiving and responding to asynchronous processing requests, and post-processing the model inference results;
[0101] d) The central GPU server is responsible for creating analysis tasks, performing deep neural network model inference, summarizing the processing results, and providing a query interface;
[0102] Furthermore, the system realizes distributed computing through the following timing process:
[0103] S1, The stress testing server uploads the recorded video file to the file server;
[0104] S2, The file server downloads the video file to the distributed CPU server;
[0105] S3, The stress testing server requests the asynchronous processing interface from the distributed CPU server;
[0106] S4, The distributed CPU server responds to the asynchronous processing interface callback;
[0107] S5, The distributed CPU server uploads the preprocessing result file to the file server;
[0108] S6, The central GPU server downloads the preprocessing file from the file server;
[0109] S7, The stress testing server creates an analysis task for the central GPU server;
[0110] S8, The central GPU server returns the model inference result to the distributed CPU server;
[0111] S9, The distributed CPU server returns the post-processing result to the central GPU server;
[0112] S10, The stress testing server queries the analysis task result from the central GPU server.
Claims
1. A method for analyzing the response time of an APP based on a deep neural network and distributed computing, characterized in that, It includes the following steps: Record the video of the APP page loading and upload it to the file server; Use the distributed CPU server to obtain the video from the file server, split the video into frame-by-frame pictures, preprocess the pictures, generate the data format file for the classification model inference, and upload the data format file to the file server; Obtain the arrival timestamp information of each frame of the image, and associate the arrival timestamp information with the corresponding image data; Store the associated arrival timestamp information in the local storage medium of the distributed CPU server in a structured data format; Use the central GPU server to receive the analysis task request of the APP response time, obtain the preprocessed data format file from the file server, use the trained classification model to perform inference on the data format file, and upload the inference result to each distributed CPU server; Use the distributed CPU server to post-process the received inference result to obtain the start frame and end frame of each page loading, calculate the page loading time, and at the same time obtain the start screenshot and end screenshot during each page loading; Use the central GPU server to summarize the analysis results of the distributed CPU server, store the analysis results in the database, and provide an interface for query; The steps of using the distributed CPU server to post-process the received inference result include: Obtain the start frame number and end frame number of all click states, and record them as a list , represents the i-th tuple. There are n tuples in the list, and each tuple is recorded as . In each tuple represents the frame number when the i-th click state starts to click, represents the frame number when the i-th click state releases the click; Traverse each tuple in the list in turn, obtain the different state frame counts between the end frame of the current click state and the start frame of the next click state, select the loading state with the largest number of occurrences and at least 3 frames or more, and use this loading state as the loading state triggered by the current click operation; if not found, skip the current tuple, traverse the next tuple, and calculate the start frame number and end frame number when the current state appears after finding; Regard all blank loading states as the same class, and reset the start frame of the loading state to the next frame of the end frame of the current click state; Summarize the start frame numbers and end frame numbers of all loading states, use the timestamp information file to query the arrival times of the start frame and end frame, calculate the actual page loading time, and obtain the pictures when each page starts loading and ends loading.
2. The APP response time analysis method based on a deep neural network and distributed computing according to claim 1, wherein Before the step of using the trained classification model to perform inference on the data format file includes: Train the classification model using the data set and labels.
3. The method for analyzing the response time of an APP based on a deep neural network and distributed computing according to claim 2, wherein The steps of training the classification model using the data set and labels include: Construct a data set, where the data set consists of the videos of the disassembled APP page loading, divide each frame of the video into different page states, including stable state, loading state, and click state pictures, and the page states constitute different class labels; Construct a classification model, where the classification model is a deep neural network; Construct a loss function; Use the pictures in the constructed data set as the input of the classification model, the output of the classification model is the class probability score corresponding to each picture, use the loss function to measure the difference between the model prediction and the actual label, and use gradient descent to optimize the classification model.
4. The method for analyzing the APP response time based on a deep neural network and distributed computing according to any one of claims 1 to 3, characterized in that And the steps of preprocessing the pictures include: Use the image processing library to read the video and obtain each frame of picture frame by frame; Perform a scaling operation on each frame of the picture; Perform a normalization operation on the scaled picture; Standardize with the mean and standard deviation; Concatenate the processed frame data; Perform lossless compression on the concatenated data and store the compressed data in a structured format.
5. The APP response time analysis system based on a deep neural network and distributed computing implemented by the method according to claim 1, characterized in that, Include: A file server for storing the original video file, the preprocessed data, and the processing results; Multiple distributed CPU servers for preprocessing the video file and postprocessing the classification model inference results; A central GPU server for creating analysis tasks, performing classification model inference, aggregating the processing results, and providing a query interface.
6. The APP response time analysis system based on deep neural network and distributed computing according to claim 5, characterized in that, The system further includes a task scheduling module for coordinating the communication and data transfer among the distributed CPU servers, the central GPU server, and the file server.
7. The APP response time analysis system based on a deep neural network and distributed computing according to claim 6, characterized in that, The system adopts an asynchronous communication mechanism.
8. The APP response time analysis system based on a deep neural network and distributed computing according to claim 6 or 7, characterized in that The system adopts a fault tolerance mechanism.
Citation Information
Patent Citations
A method and device for determining response time
CN110704294A