Cloud mobile phone end-to-end performance tracking method and related equipment
By obtaining client touch event data and network event data, combining server resource event data, hardware-level clock synchronization and software compensation algorithm are used for timestamp calibration, and event alignment and causality calculation are performed based on dynamic time regularization algorithm and Bayesian network model, the problems of poor data correlation, low time accuracy, and large resource consumption in traditional cloud mobile phone performance tracking methods are solved, and problem positioning and resource optimization of high-precision cloud mobile phone end-to-end performance tracking and automation are achieved.
Patent Information
- Application Number
- CN202510367378.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional cloud mobile phone performance tracking methods have problems such as poor data correlation, low time accuracy, and large resource consumption, which is difficult to meet the needs of end-to-end performance tracking of cloud mobile phones.
By obtaining client touch event data and network event data, generating global unique identifiers based on preset hash algorithms, combining server resource event data, hardware-level clock synchronization and software compensation algorithm are used for timestamp calibration, event alignment and causality calculation are performed based on dynamic time alignment algorithm and Bayesian network model, a priority-sorted root cause list is generated and the self-healing strategy execution is triggered.
It realizes efficient association and tracking of cross-end events, improves the accuracy of event alignment, quickly calculates the causal relationship weights between resource events, realizes automated problem positioning and resource optimization, improves the performance tracking accuracy and response speed of cloud mobile phones, reduces resource consumption, and improves user experience and system stability.
Smart Images

Figure CN120151231A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to an end-to-end performance tracking method and related devices for cloud mobile phones. Background Art
[0002] With the popularization of cloud mobile phone technology, the demand for its performance tracking is increasing day by day. Traditional performance tracking methods have problems such as poor data correlation, low time accuracy, and high resource consumption, and it is difficult to meet the requirements of end-to-end performance tracking for cloud mobile phones. For example, traditional solutions usually analyze front-end and back-end data independently, and the time accuracy can only reach the level of 10 - 100 ms, and the resource consumption is large, and the positioning speed is slow.
[0003] Traditional solutions cannot efficiently correlate cross-end events, resulting in difficult root cause location. For example, traditional methods cannot effectively align the time axes of client touch events and server resource events, and it is difficult to establish a causal relationship, so the root cause of performance problems cannot be located in a timely and accurate manner. In addition, traditional solutions lack an automated self-healing mechanism and cannot automatically adjust resource allocation according to performance problems, resulting in high operation and maintenance costs and poor user experience. Therefore, there is an urgent need for an end-to-end performance tracking method for cloud mobile phones to solve the above-mentioned technical problems. Summary of the Invention
[0004] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further detailed in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0005] In a first aspect, this application provides a gray box testing method for a cloud mobile phone server, and the method includes:
[0006] Obtain client touch event data and network event data, and generate a globally unique identifier based on a preset hash algorithm;
[0007] Obtain server resource event data, where the server resource event data includes system call trace data, CPU scheduling trace data, and resource consumption monitoring data;
[0008] Based on a hardware-level clock synchronization and software compensation algorithm, calibrate the timestamps of the client touch event data and the server resource event data to generate synchronized timestamps;
[0009] Based on the synchronized timestamps, align the client and server event sequences through a dynamic time warping algorithm to generate an aligned event sequence;
[0010] Based on a Bayesian network model, calculate the causal probability of the aligned event sequence to determine the causal relationship weight between resource events;
[0011] Match the root cause according to the causal relationship weight and the preset abnormal pattern library, generate a root cause list with priority sorting, and trigger the execution of the self-healing strategy.
[0012] In some embodiments, obtain client touch event data and network event data, and generate a globally unique identifier based on a preset hash algorithm, including:
[0013] Intercept touch events in the input subsystem at the kernel layer through eBPF probes, and capture touch coordinates, pressure values, and sliding trajectories;
[0014] Inject trace points at the TCP protocol stack state migration point, and record round-trip delay, window size, and socket operation parameters;
[0015] Transmit the client touch event data and network event data to the user space through a zero-copy channel and bind them to the globally unique identifier.
[0016] In some embodiments, the hardware-level clock synchronization includes:
[0017] Implement a precision time protocol transparent clock through a programmable logic device, and correct the link layer timestamp error to a first preset threshold;
[0018] The software compensation algorithm includes:
[0019] Adopt a state equation model to correct the logical clock drift, and calibrate based on the cross-validation of the timestamp counter, so that the cross-terminal time error is less than a second preset threshold.
[0020] In some embodiments, based on the synchronized timestamp, align the client and server event sequences through the dynamic time warping algorithm to generate an aligned event sequence, including:
[0021] Constrain the window width to a preset proportional value of the lengths of the client and server event sequences, and match the resource fluctuation events within the time window through a sliding window;
[0022] Interpolate and complete the out-of-order or lost events to generate a continuous aligned event sequence.
[0023] In some embodiments, calculate the causal probability of the aligned event sequence based on the Bayesian network model to determine the causal relationship weight of the resource events, including:
[0024] Accelerate the sampling of the event posterior probability based on a parallel computing framework;
[0025] Construct a device dependency graph, and generate the probability weight between events through the Markov chain Monte Carlo algorithm.
[0026] In some embodiments, root cause matching and self-healing strategy triggering include:
[0027] When it is detected that the waiting time of a high-priority thread exceeds a first preset duration or the memory allocation difference exceeds a third preset threshold, match the thread priority inversion or memory leak mode;
[0028] According to the root cause confidence ranking result, dynamically adjust the virtual CPU binding, video memory pre-allocation, or input / output current limiting strategy, and send it to the kernel for execution in real time through the kernel communication channel;
[0029] Among them, dynamically adjusting the virtual CPU binding is to obtain the physical core load status based on the control group nested monitoring model, and bind the target process to the low-load physical core through the affinity setting command;
[0030] Video memory pre-allocation is to reserve video memory space based on historical peak data and lock the pre-allocated area through the memory mapping system call.
[0031] In some embodiments, it further includes:
[0032] Through the ring buffer dynamic adjustment mechanism, adjust the size of the shared memory area between the kernel mode and the user mode according to the data throughput;
[0033] Adopt boundary check instructions and loop unrolling compilation directives to prevent the probe program from accessing out-of-bounds memory.
[0034] In a second aspect, the present application proposes an end-to-end performance tracking device for cloud mobile phones, and the device includes:
[0035] A data collection identification unit, which obtains client touch event data and network event data, and generates a globally unique identifier based on a preset hash algorithm;
[0036] A service resource monitoring unit, which is used to obtain server resource event data, where the server resource event data includes system call tracking data, CPU scheduling trace data, and resource consumption monitoring data;
[0037] A time synchronization and calibration unit, based on the hardware-level clock synchronization and software compensation algorithm, calibrates the timestamps of the client touch event data and the server resource event data to generate synchronized timestamps;
[0038] An event alignment processing unit, based on the synchronized timestamps, aligns the client and server event sequences through the dynamic time warping algorithm to generate an aligned event sequence;
[0039] A causal inference analysis unit, based on the Bayesian network model, calculates the causal probability of the aligned event sequence to determine the causal relationship weight between resource events;
[0040] The root cause localization and self-healing unit matches the root cause according to the causal relationship weight and the preset abnormal pattern library, generates a root cause list with priority sorting, and triggers the execution of the self-healing strategy.
[0041] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, it implements the steps of the cloud mobile phone end-to-end performance tracking method according to any one of the first aspects.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the cloud mobile phone end-to-end performance tracking method according to any one of the first aspects.
[0043] In summary, the present application obtains client touch event data and network event data, generates a globally unique identifier based on a preset hash algorithm, and combines the collection of server-side resource event data to achieve efficient association and tracking of cross-terminal events. Through hardware-level clock synchronization and software compensation algorithms, it ensures that the timestamp accuracy of client and server events reaches ±0.3 ms, significantly improving the accuracy of event alignment. Based on the dynamic time warping algorithm and the Bayesian network model, it can quickly calculate the causal relationship weight between resource events, and through root cause matching and self-healing strategy execution, it realizes automated problem localization and resource optimization. The present application effectively improves the performance tracking accuracy and response speed of cloud mobile phones, reduces resource consumption, and improves user experience and system stability. Description of the Drawings
[0044] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this specification. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0045] Figure 1 It is a schematic flowchart of the cloud mobile phone end-to-end performance tracking method provided by the embodiment of the present application;
[0046] Figure 2 It is a schematic structural diagram of the cloud mobile phone end-to-end performance tracking device provided by the embodiment of the present application;
[0047] Figure 3 It is a structural diagram of the cloud mobile phone end-to-end performance tracking electronic device provided by the embodiment of the present application. Detailed Embodiments
[0048] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims, and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0049] Please refer to Figure 1 , which is a schematic flowchart of an end-to-end performance tracking method for a cloud mobile phone provided by an embodiment of this application, and specifically may include:
[0050] S110. Obtain client touch event data and network event data, and generate a globally unique identifier based on a preset hash algorithm;
[0051] Exemplarily, the client touch event data is captured in real time through the kernel layer probe technology, including interaction information such as touch coordinates, pressure values, and sliding trajectories. At the same time, the network event data injects tracking logic at key nodes in the transport layer (such as TCP state migration points) to record network performance metrics such as round-trip delay and window parameters. The two types of data are efficiently transmitted to the user state through the zero-copy mechanism, avoiding the memory copy overhead in traditional data collection.
[0052] To establish accurate association of cross-end events, a preset hash algorithm (such as SHA-1) is used to encrypt the touch event timestamp, device fingerprint, and random salt to generate a globally unique trace_id. This identifier runs through the client and server event links, ensuring that touch operations and backend resource fluctuations are traceable in the space-time dimension and providing a unified data identification basis for subsequent causal analysis.
[0053] S120. Obtain server resource event data, where the server resource event data includes system call tracking data, CPU scheduling tracking data, and resource consumption monitoring data;
[0054] Exemplarily, the server - side resource event data is collected through kernel - level probes, including three core metrics: system call tracing, CPU scheduling tracking, and resource consumption monitoring. System call tracing intercepts file operation events at the virtual file system (VFS) layer, recording the operation path and latency; CPU scheduling tracking constructs a running state map based on thread state transition events (such as sched_switch) to detect high - priority thread blocking; resource consumption monitoring quantifies the usage of hardware resources such as GPU video memory, CPU / IO pressure, etc. through the cgroup nested model to form a multi - dimensional performance baseline.
[0055] The above - mentioned data is associated and integrated with a preset time window (such as 500 ms) through hardware - level time synchronization (PTP protocol), and cross - event association is achieved by combining the process PID, thread TID, and cgroup identifier. Metrics such as system call latency, thread blocking duration, and resource usage peak together constitute an evidence chain for server - side performance analysis, providing inputs for the causal inference engine to trigger resource chain reactions (such as touch → rendering → video memory allocation).
[0056] S130. Based on hardware - level clock synchronization and software compensation algorithms, calibrate the timestamps of client - side touch event data and server - side resource event data to generate synchronized timestamps;
[0057] Exemplarily, the timestamp calibration between the client and the server is achieved relying on hardware - level clock synchronization technology. By deploying programmable logic devices (such as FPGA) that support the Precision Time Protocol (PTP), the timestamp error in link transmission is corrected at the physical layer. Hardware synchronization eliminates the millisecond - level deviation of the traditional Network Time Protocol (NTP), ensuring the unity of the time reference for cross - end events and providing sub - millisecond - level accuracy guarantee for subsequent event alignment. The hardware - level clock synchronization adopts the PTP transparent clock technology implemented by FPGA, following the IEEE1588 - 2008 standard. Its core mechanism is: at the physical layer, the FPGA corrects the link - layer timestamp of network packets to eliminate the influence of transmission delay and clock jitter.
[0058] On the basis of hardware synchronization, the dynamic drift of the logical clock is corrected through the Kalman filter model, and combined with the cross - verification mechanism of the timestamp counter (TSC), the clock deviation is further calibrated. The software compensation algorithm can adapt to clock jitter and transmission delay, converging the global time error within a preset threshold, thereby generating highly consistent synchronized timestamps to support the reliability of cross - end event causal association.
[0059] S140. Based on the synchronized timestamps, align the client - side and server - side event sequences through the dynamic time warping algorithm to generate an aligned event sequence;
[0060] Exemplarily, after the timestamps of the client touch event and the server resource event are hardware-synchronized and software-compensated, a globally consistent timing benchmark is formed. Based on this, the Dynamic Time Warping (DTW) algorithm uses a sliding window matching mechanism to elastically align the two event sequences on the time axis, eliminating time offsets caused by network jitter or sampling frequency differences, and ensuring the association of touch operations and resource fluctuation events in a unified time dimension.
[0061] The DTW algorithm restricts the path search range by constraining the window width (such as a preset proportional value of the lengths of the two sequences), reducing the computational complexity while allowing local stretching and alignment of the event sequences. For the situation of event disorder or partial loss, the algorithm generates a continuous alignment sequence through interpolation completion, ensuring the integrity of the cross-terminal causal chain and laying a data foundation for the accurate analysis of subsequent resource chain reactions.
[0062] S150. Calculate the causal probability of the aligned event sequence based on the Bayesian network model to determine the causal relationship weight between resource events;
[0063] Exemplarily, based on the aligned event sequence, the Bayesian network model quantifies the interaction logic of device and thread states by constructing a dependency graph between resource events (such as the touch event → GPU rendering → video memory allocation chain relationship). The model combines prior knowledge and real-time event data, and uses the Markov Chain Monte Carlo (MCMC) method to perform parallel probability sampling on millions of events, calculating the conditional probability distribution between events, thereby revealing the potential impact of resource contention or scheduling anomalies on performance issues.
[0064] Efficiently solve the posterior probability through a parallel computing framework (such as CUDA acceleration) to generate the causal relationship weight between events (such as the probability of rendering delay caused by insufficient video memory). The weight value reflects the contribution degree of resource events to performance issues and is dynamically updated with the system state, providing a quantifiable decision basis for root cause location and ensuring the accuracy of abnormal pattern matching and self-healing strategy triggering.
[0065] S160. Perform root cause matching according to the causal relationship weight and the preset abnormal pattern library, generate a root cause list with priority ranking, and trigger the execution of the self-healing strategy.
[0066] Exemplarily, based on the causal relationship weight, match the resource events with the preset abnormal pattern library (such as thread priority inversion, memory leak, etc.), and screen out candidate root causes that meet the threshold conditions. Use a probability sorting engine to sort the confidence levels of the candidate root causes in descending order to generate a root cause list with priority ranking, ensuring that high-probability abnormal events trigger self-healing actions first.
[0067] According to the sorting result of the root cause list, map the root causes to a predefined self-healing policy library (such as video memory pre-allocation, CPU core binding), and send the policies to the target module for execution in real time through the kernel communication channel (such as netlink). At the same time, the execution effect of the policies is fed back to the analysis system through the resource metric change data, forming a closed-loop optimization link to continuously improve the accuracy of root cause location and self-healing decision-making.
[0068] In summary, the embodiment of the present application realizes the unique identification and cross-terminal association of client events by obtaining client touch event data and network event data and generating a globally unique identifier based on a preset hash algorithm; through hardware-level clock synchronization and software compensation algorithms, calibrate the timestamps of the client touch event data and the server resource event data to generate synchronized timestamps, ensuring the time alignment accuracy of cross-terminal events; based on the synchronized timestamps, align the client and server event sequences through the dynamic time warping algorithm to generate an aligned event sequence, improving the accuracy of event association; calculate the causal probability of the aligned event sequence based on the Bayesian network model to determine the causal relationship weights between resource events, realizing the quantitative analysis of the causal relationship between resource events; match the root causes according to the causal relationship weights and a preset abnormal pattern library to generate a root cause list with priority sorting, and trigger the execution of self-healing policies, realizing the automatic location and self-healing of performance problems. In summary, the present invention realizes the high-precision tracking of the end-to-end performance of the cloud phone, the accurate analysis of the causal relationship, and the automatic optimization of performance problems, significantly improving the performance and user experience of the cloud phone and reducing the operation and maintenance costs.
[0069] In some instances, obtaining client touch event data and network event data and generating a globally unique identifier based on a preset hash algorithm includes:
[0070] Intercept input subsystem touch events at the kernel layer through eBPF probes to capture touch coordinates, pressure values, and sliding trajectories;
[0071] Inject trace points at the TCP protocol stack state transition points to record round-trip delays, window sizes, and socket operation parameters;
[0072] Transmit the client touch event data and network event data to the user space through a zero-copy channel and bind them to the globally unique identifier.
[0073] Exemplarily, the embodiments of the present application achieve efficient collection of touch events through the eBPF probe technology at the kernel layer. Specifically, in the input subsystem of the Android kernel, the input_handler function is dynamically hooked through Kprobe to intercept user touch operations (such as clicks and swipes) in real time. This mechanism accurately obtains raw data such as touch coordinates (X / Y axes, 0 - 1080p precision), pressure values (0 - 255 levels), and swipe trajectories through the event capture module in the kernel state. At the same time, the CLOCK_MONOTONIC clock source is used to generate nanosecond-level timestamps to ensure the high stability and anti-drift characteristics of event timestamps. The captured touch data is transmitted with low latency between the kernel state and the user state through the eBPF double-ring buffer (BPF_MAP_TYPE_RINGBUF). This buffer supports dynamic capacity adjustment from 4KB to 16MB, and a zero-copy channel is created through the ring_buffer__new() API to avoid redundant copying of data between the kernel state and the user state, reducing the memory overhead by 3 times.
[0074] At the TCP protocol stack level, eBPF injects trace points at state transition points (such as SYN_SENT → ESTABLISHED) to monitor the network connection status and performance metrics in real time. Specifically, it includes: recording the TCP round-trip time (RTT), the receive window size (sk_rcv_wnd), and the target IP, port, and data size of socket operations (connect / send / recv). The meta-information of network packets is obtained by parsing the sk_buff structure, and key metrics are output using bpf_trace_printk. The network event data and the touch event data are cross-layer associated through the shared trace_id: the trace_id is generated by the SHA-1 hash algorithm, with the input being the timestamp, device fingerprint (such as IMEI), and random salt value, generating a 128-bit unique identifier (collision probability < 1e-18). This identifier is embedded in the event data through memory mapping or protocol headers (such as the TCP_OPT_TRACE_ID field) to ensure cross-terminal consistency of client events.
[0075] The client data is efficiently transmitted through a zero-copy channel. The kernel-mode buffer is mapped to the user-mode address space through mmap. The user-mode program asynchronously reads the data through ring_buffer_poll to avoid data replication. After the touch events and network events are aggregated in the user mode, they are encapsulated into a unified format (such as Protocol Buffers) and UDP data packets are batch-sent through sendmmsg to reduce the number of system calls. Sensitive data (such as touch coordinates) is encrypted and transmitted using TLS1.3 to protect data privacy. The binding mechanism of the globally unique identifier ensures the integrity of the client event identifiers during transmission, providing a basis for subsequent cross-terminal event alignment and causal reasoning. The embodiments of this application achieve the characteristics of low loss and high timeliness of data collection through eBPF technology, meeting the real-time requirements of end-to-end performance tracking for cloud mobile phones.
[0076] In some instances, the hardware-level clock synchronization includes:
[0077] Implementing a Precision Time Protocol transparent clock through a programmable logic device to correct the link-layer timestamp error to a first preset threshold;
[0078] The software compensation algorithm includes:
[0079] Using a state equation model to correct the logical clock drift and calibrating based on the cross-validation of the timestamp counter to make the cross-terminal time error less than a second preset threshold.
[0080] Exemplarily, the hardware-level time synchronization is implemented using the PTP transparent clock technology based on FPGA, following the IEEE1588-2008 standard. Specifically, at the physical layer, the FPGA corrects the link-layer timestamp of the network packet in real time: when the packet enters the FPGA, the entry timestamp is recorded, and when it leaves, the exit timestamp is recorded, and the transmission path delay is calculated through hardware logic to compensate for the timestamp error caused by link jitter and transmission distance. This technology compresses the physical layer timestamp error to the first preset threshold (±300 microseconds), providing a high-precision time reference for cross-terminal events. The hardware implementation is based on a PHY chip (such as Intel I210) that supports IEEE1588-2008, combined with the programmable logic characteristics of the FPGA, to implement the transparent clock function, ensuring that the time references of the client and the server are aligned at the sub-millisecond level at the physical layer.
[0081] The software compensation algorithm realizes logical clock calibration through Kalman filtering and TSC cross-verification. Kalman filtering is based on the state equation model to dynamically estimate the clock offset and drift rate, and correct the logical clock drift caused by the crystal oscillator frequency deviation. At the same time, the TSC counter value of the CPU is read through the rdtscp instruction, and its high-frequency stability (usually at the GHz level) is used to cross-verify the software timestamp. The combination of the two further optimizes the cross-terminal time error to the second preset threshold (±0.3 ms). Specifically, in implementation, the Kalman filtering module predicts the current clock offset through the historical timestamp sequence, and the TSC counter provides a hardware-independent time reference to eliminate the errors introduced by CPU frequency fluctuations or system load changes, ensuring the spatio-temporal consistency of the client and server timestamps in causal reasoning.
[0082] The hardware-level synchronization and the software compensation algorithm form a hierarchical collaborative architecture. The hardware layer corrects the timestamp error of the physical layer through the FPGA transparent clock, and the software layer compensates for the logical clock drift through Kalman filtering and TSC verification. This hybrid synchronization scheme optimizes the 2 ms-level error of the traditional NTP protocol to ±0.3 ms, meeting the sub-millisecond alignment requirement for end-to-end performance tracking of cloud mobile phones. The synchronized timestamp provides a reference for the dynamic time warping (DTW) algorithm, supporting the non-linear alignment of client touch events (such as 120 Hz sampling) and server resource events (such as 1 kHz GPU metrics), ensuring that the causal reasoning engine can accurately associate cross-terminal events and locate performance bottlenecks. This design reduces the physical layer error through hardware and compensates for the logical drift through software, constructing an end-to-end high-precision time synchronization system, providing a time reference for subsequent causal analysis and self-healing decision-making.
[0083] In some instances, based on the synchronized timestamp, the dynamic time warping algorithm is used to align the client and server event sequences to generate an aligned event sequence, including:
[0084] The constraint window width is a preset proportional value of the lengths of the client and server event sequences, and the resource fluctuation events within the time window are matched through a sliding window;
[0085] Interpolate and complete the out-of-order or missing events to generate a continuous aligned event sequence.
[0086] Exemplarily, the DTW algorithm restricts the path search range by constraining the window width (W = floor(0.1 × max(m, n)), where m and n are the lengths of the client and server event sequences), reducing the time complexity of the alignment calculation. This constrained window is based on the Sakoe-Chiba Band theory, reducing the complexity of the traditional DTW algorithm from O(mn) to O(0.2mn), while ensuring that the path deviation does not exceed a preset ratio, avoiding alignment distortion caused by excessive differences in the lengths of event sequences. The setting of the window width takes into account both computational efficiency and alignment accuracy, adapting to the real-time processing requirements of high-frequency touch events and resource collection data in the cloud phone scenario.
[0087] Driven by synchronized timestamps, the sliding window mechanism is used to filter server resource fluctuation events (such as sudden increases in CPU load and video memory allocation failures) within a specific time window (e.g., 0ms, 500ms) after client touch operations. The events within the window calculate the similarity using the Euclidean distance, elastically aligning the local features of the two sequences at both ends, and eliminating the timing misalignment caused by network latency or inconsistent sampling frequencies (such as 120Hz touch data on the client side and 1kHz GPU data on the server side). This mechanism can accurately correlate touch operations with backend resource responses. For example, it aligns the user's sliding event with the GPU rendering delay on the timeline, providing a spatio-temporally consistent data basis for causal chain analysis.
[0088] To address the possible out-of-order or partial loss problems during event transmission, linear interpolation is used to complement the missing events. For example, if server resource events are detected as missing within a certain timestamp range, interpolation filling is performed based on the resource metrics (such as video memory usage rate, thread status) of adjacent events in the front and back windows, generating a continuous aligned event sequence. The interpolation and complementation mechanism allows a certain proportion (e.g., 5%) of event tolerance, ensuring that the causal inference engine can still construct a complete resource chain reaction model (such as touch → rendering → video memory allocation link) in the scenario of incomplete data, avoiding misjudgment of the root cause due to single-point data loss.
[0089] In some instances, based on the Bayesian network model, causal probability calculations are performed on the aligned event sequence to determine the causal relationship weights of resource events, including:
[0090] Accelerating sampling of the event posterior probability based on a parallel computing framework;
[0091] Constructing a device dependency graph and generating probability weights between events through the Markov chain Monte Carlo algorithm.
[0092] Exemplarily, a Bayesian network model is used as the core tool for causal inference, and its structure constructs a directed acyclic graph (DAG) based on the dependency relationships between resource events. Specifically, network nodes represent event types (such as GPU rendering latency, insufficient video memory), edges represent causal relationships, and the edge weights are causal probabilities. The network structure is defined according to a preset resource chain reaction (such as touch event → SurfaceFlinger composition → GPU rendering → video memory allocation) and a device dependency map (GPU-CPU-IO DAG). Parameter estimation uses the Markov chain Monte Carlo (MCMC) method, and samples the posterior probability distribution through the Metropolis-Hastings algorithm. This algorithm constructs a Markov chain and gradually approximates the target probability distribution to quantify the strength of the causal relationship between events. For example, when calculating the probability that insufficient video memory causes rendering latency, MCMC sampling will comprehensively consider historical data, the current resource state, and the spatio-temporal correlation between events.
[0093] To cope with the complex computing requirements of millions of events, the present invention uses the CUDA parallel computing framework to accelerate posterior probability sampling. Specifically, when implementing, multi-threaded parallel processing is started on NVIDIA GPUs: each thread independently calculates the likelihood values of some events, and optimizes the data access efficiency through shared memory and thread cooperation. For example, for 1 million events, 4096 threads are started to execute in parallel, reducing the single-iteration time from the second level of a single-core CPU to the millisecond level of the GPU. This parallel design compresses the overall processing time within 15 ms, significantly improving the real-time performance of causal inference. At the same time, the efficient memory management of CUDA (such as the hierarchical optimization of global memory and shared memory) and the thread synchronization mechanism ensure the stability and accuracy of parallel computing.
[0094] The generation process of the causal relationship weight includes the following steps: First, based on the timestamps (accuracy ±0.3 ms) of the aligned event sequences and the resource chain reaction path, filter out resource fluctuation events within 500 ms after the front-end operation; Second, filter out non-directly related events through resource topology modeling (such as GPU-CPU-IO dependencies) to ensure the physical rationality of the causal relationship; Finally, combine the posterior probability distribution obtained by MCMC sampling to calculate the causal weight between events. For example, if the touch event is associated with GPU rendering latency within the time window, and the probability weight of video memory allocation failure and rendering latency is 0.87, then it is determined that insufficient video memory is the potential root cause. The generated weight matrix ensures the spatio-temporal consistency of the causal relationship through time window matching ([0 ms, 500 ms]) and resource dependency verification, providing a quantitative basis for subsequent root cause matching. This mechanism combines parallel computing with domain knowledge to achieve efficient modeling and verification of causal relationships in complex systems.
[0095] In some instances, root cause matching and self-healing strategy triggering include:
[0096] When it is detected that the waiting time of a high-priority thread exceeds the first preset duration or the memory allocation difference exceeds the third preset threshold, match the thread priority inversion or memory leak mode;
[0097] According to the root cause confidence ranking result, dynamically adjust the virtual CPU binding, video memory pre-allocation, or input / output rate limiting strategy, and send it to the kernel for execution in real time through the kernel communication channel;
[0098] Among them, dynamically adjusting the virtual CPU binding is to obtain the physical core load status based on the control group nested monitoring model, and bind the target process to the low-load physical core through the affinity setting command;
[0099] Video memory pre-allocation is to reserve video memory space based on historical peak data and lock the pre-allocated area through the memory mapping system call.
[0100] Exemplarily, based on the sched_switch event to monitor the thread state migration, when the high-priority thread is in the TASK_UNINTERRUPTIBLE state for more than the first preset duration (such as 50 ms), it is determined as thread priority inversion. At the same time, monitor the difference histogram of the number of kmalloc and kfree calls within the same control group (cgroup). If the difference continuously exceeds the third preset threshold (such as 100 times / second), it is determined as the memory leak mode. The root cause confidence is calculated by the formula Confidence = causal probability × α + pattern matching degree / β (where α = 0.7 and β = 1.2), and a root cause list sorted by priority is generated in descending order to ensure that high-confidence exceptions are processed first.
[0101] Based on the control group (cgroup v2) nested monitoring model, the load status of the physical core (such as CPU utilization rate, run queue length) is collected in real time. Bind the target process (such as the rendering thread) to the low-load physical core (such as cores 12 - 15) through the taskset - pc command, and set the CPU affinity to eliminate cross-core resource contention. In the specific implementation, the kernel dynamically updates the CPU binding relationship by parsing the / sys / fs / cgroup / cpuset.cpus file, and combines the pressure stall information (PSI) to predict the load trend of the physical core to achieve adaptive adjustment of the binding strategy.
[0102] The video memory pre-allocation strategy dynamically calculates the reserved space (e.g., 110% of the peak value) based on historical peak data (such as the maximum video memory usage in the past 24 hours). The pre-allocated video memory area is locked through the mmap system call, and the cgroup video memory upper limit is set by writing to the / sys / fs / cgroup / gpu / memory.high file. The pre-allocated area is mapped to the virtual address space at the startup of the process, avoiding rendering delays or process crashes caused by insufficient video memory. At the same time, combined with the Sliding Window Binning (SWAB) encoding technology to compress historical data, and real-time predict the peak value of video memory requirements to optimize the utilization rate of the reserved space.
[0103] The self-healing strategy is sent to the target module in real time through the kernel communication protocol (such as netlink), and the message header type is defined as 0xA0E0 to distinguish policy instructions. For example, the IO rate limiting strategy configures the token bucket algorithm through the tc qdisc command to limit the bandwidth or IOPS of the specified device. After the policy is executed, the causal inference engine is fed back by collecting the changed data of resource metrics (such as the CPU load decrease rate, the video memory usage fluctuation), and the prior probability of the Bayesian network and the state-action value table (Q-table) of the Q-learning model are dynamically updated, forming a closed-loop link of "detection → execution → verification → optimization" to continuously improve the accuracy and response speed of self-healing decisions.
[0104] In some instances, it further includes:
[0105] Dynamically adjust the size of the shared memory area between the kernel state and the user state through the circular buffer dynamic adjustment mechanism according to the data throughput;
[0106] Adopt boundary check instructions and loop unrolling pragmas to prevent out-of-bounds memory access of the probe program.
[0107] Exemplarily, the embodiment of the present application realizes efficient data transmission between the kernel state and the user state through the circular buffer dynamic adjustment mechanism. The circular buffer adopts a double-ring structure design. The kernel state establishes a shared memory area through the eBPF map (type BPF_MAP_TYPE_RINGBUF), and the user state creates a zero-copy channel by calling the ring_buffer__new() API. The buffer capacity supports dynamic adjustment, and the adjustment range is between 4KB and 16MB. The specific adjustment algorithm is:
[0108] size = min(16MB, 4KB × 2n) (n = log2(throughput / 80%))
[0109] This algorithm dynamically adjusts the buffer size based on real-time data throughput, ensuring that the buffer capacity is expanded in high-load scenarios to improve throughput, and reduced in low-load scenarios to reduce memory occupancy. For example, when the throughput reaches 80% of the current buffer capacity, exponential expansion is used to avoid data overflow; conversely, when the throughput is below the threshold, the buffer is gradually reduced to release resources. The dynamic adjustment mechanism realizes hot update through the sysctl interface (kernel.ebpf_ringbuf_size), adapting to load fluctuations without restarting the service. In addition, zero-copy transmission directly accesses the kernel buffer through memory mapping (mmap), eliminating the data copy overhead, tripling the transmission efficiency, and significantly reducing the CPU occupancy rate to below 1%.
[0110] To ensure the security and execution efficiency of the eBPF probe program, the embodiments of this application adopt a dual protection mechanism of boundary check instructions and loop unrolling compilation directives. During the kernel-state data collection process, the eBPF program accesses memory through the bpf_probe_read_kernel instruction, which automatically verifies the legality of the target address before execution to prevent out-of-bounds access from causing kernel crashes or data tampering. For example, when parsing network packets (sk_buff), boundary checks are used to ensure that the reading of the sk_rcv_wnd field does not exceed the structure range. At the same time, for loop logic, the loop body is forcibly unrolled through the #pragma unroll compilation directive to eliminate the risk of branch prediction errors and improve instruction-level parallelism. For example, when counting the difference between kmalloc / kfree, loop unrolling converts a fixed number of iterations into a linear instruction sequence, avoiding loop control overhead and reducing potential execution latency. The above measures, combined with the static analysis of the eBPF validator, ensure that the probe program complies with the kernel security specifications, and at the same time improve the collection performance by optimizing the instruction pipeline, achieving a balance between high security and low resource consumption.
[0111] In some instances, it also includes:
[0112] Predicting bottlenecks in computing resources, storage resources, and input / output resources based on the pressure blocking information model;
[0113] When detecting an abnormal resource chain reaction, trace back the touch event to the complete call chain of the rendering delay and mark the associated resource nodes.
[0114] Exemplarily, this application uses the Pressure Stall Information (PSI) model to predict bottlenecks in computing resources, storage resources, and input / output resources. The PSI model monitors the pressure status of CPU, memory, and IO resources in real time by quantifying the proportion of latency time caused by resource contention. Specifically, the PSI model is based on the cgroup v2 nested monitoring architecture and collects the following key metrics: CPU pressure reflects the proportion of latency time caused by CPU resource contention, memory pressure monitors the memory reclaim frequency and blocking time, and IO pressure quantifies the proportion of IO operation blocking time. For example, when the memory pressure metric shows that the page cache reclaim frequency exceeds 100 times per second, it is determined as memory thrashing; when the IO pressure metric shows that the disk latency exceeds 200 ms, it is determined as an IO bottleneck. The PSI model uses a sliding window mechanism (such as a 10 ms window) to statistically calculate the maximum, minimum, and average values of resource pressure, generating a multi-dimensional performance baseline. Combining historical data with real-time metrics, the PSI model can predict resource bottlenecks in advance, providing a quantitative basis for self-healing strategies. For example, when the video memory usage rate reaches 90%, a video memory pre-allocation strategy is triggered to avoid performance degradation caused by resource exhaustion.
[0115] When an abnormal resource chain reaction is detected, this application accurately locates associated resource nodes by tracing back the complete call chain from the touch event to the rendering latency. The specific process is as follows: First, align the time axes of the client touch event and the server resource event based on the Dynamic Time Warping (DTW) algorithm, with the window width constrained to 10% of the lengths of both sequences (i.e., W = floor(0.1 × max(m, n))), ensuring the spatio-temporal consistency between the touch event and the resource fluctuation event. Second, construct a GPU-CPU-IO device dependency graph (DAG) through resource topology modeling to clarify the resource chain reaction path, such as touch event → SurfaceFlinger composition → GPU rendering → video memory allocation. When a sudden increase in rendering latency is detected, trace back and analyze the complete call chain to identify abnormal nodes (such as video memory allocation failure or thread blocking). For example, if the GPU rendering latency exceeds the threshold and the video memory usage rate reaches 95%, the video memory node is marked as abnormal; if the waiting time of high-priority threads exceeds 50 ms, the CPU scheduling node is marked as abnormal. Calculate the causal probability weights between events through a Bayesian network model to generate a root cause list sorted by probability, providing an accurate decision-making basis for self-healing strategies. For example, when the probability of rendering latency caused by insufficient video memory is 87%, a video memory pre-allocation strategy is triggered to ensure the stability of the resource chain reaction.
[0116] In some instances, after the self-healing strategy is executed, it further includes:
[0117] Collecting data on changes in resource metrics after the execution of the strategy through a custom communication protocol;
[0118] Update the state space and value table parameters based on the reinforcement learning model to optimize the response efficiency of subsequent self-healing decisions.
[0119] Exemplarily, real-time collect the changed data of resource metrics after the execution of the self-healing strategy through a custom communication protocol (such as the netlink channel) to form a closed-loop feedback mechanism. Specifically, the netlink packet header type is defined as 0xA0E0 to distinguish policy instructions from feedback data. After the policy is executed, the system collects the changed data of resource metrics such as CPU utilization, memory pressure (PSI metric), and IO latency through the cgroup interface and encapsulates them into a feedback packet in a unified format. For example, after the video memory pre-allocation policy is executed, read the / sys / fs / cgroup / gpu / memory.usage_in_bytes file in real time to monitor the fluctuation of the video memory usage rate; after the CPU core binding policy is executed, collect the changed data of the CPU load through the / proc / stat interface. The feedback packet is sent back to the causal inference engine through the netlink channel to ensure the real-time and low latency of data collection. For example, when the CPU load drops by 20% or the fluctuation of the video memory usage rate is less than 5%, the feedback data will be used as the quantitative basis for the policy execution effect to provide data support for subsequent optimization.
[0120] Adopt a reinforcement learning model (such as Q-learning) to dynamically optimize the response efficiency of self-healing decisions. Specifically, the feedback data after the policy is executed is used to update the state space and value table (Q-table) parameters of the Q-learning model. The state space s = {CPU load, memory pressure, IO latency} is a three-dimensional vector quantifying the system load, and the action space a = {cpuset adjustment, memory compression, IO rate limiting} defines the optional optimization actions. The Q-table update formula is:
[0121] Q(s,a)←Q(s,a)+α[r+γmax a′ Q(s′,a′)-Q(s,a)]
[0122] where α = 0.1 is the learning rate, γ = 0.9 is the discount factor, and r is the immediate reward after the policy is executed (such as the CPU load drop rate). For example, when the video memory pre-allocation policy effectively reduces the rendering latency, the Q value of the corresponding state-action pair in the Q-table will increase, improving the priority of this policy in subsequent decisions. At the same time, the exploration rate ε decays according to ε = 1 / √t (t is the number of training iterations) to balance exploration and exploitation, ensuring that the model takes into account both stability and adaptability during the optimization process. By dynamically updating the state space and Q-table parameters, the reinforcement learning model can continuously optimize the response efficiency of self-healing decisions, forming a closed-loop link of "detection → execution → verification → optimization" to improve system performance and user experience.
[0123] Please refer to Figure 2 , which is a schematic structural diagram of an end-to-end performance tracking device for a cloud mobile phone provided by an embodiment of the present application. The device includes:
[0124] A data acquisition identification unit 21, which acquires client touch event data and network event data, and generates a globally unique identifier based on a preset hash algorithm;
[0125] A service resource monitoring unit 22, which is used to acquire server resource event data. Among them, the server resource event data includes system call tracing data, CPU scheduling trace data, and resource consumption monitoring data;
[0126] A time synchronization and calibration unit 23, which calibrates the timestamps of the client touch event data and the server resource event data based on a hardware-level clock synchronization and software compensation algorithm, and generates a synchronized timestamp;
[0127] An event alignment processing unit 24, which aligns the client and server event sequences through a dynamic time warping algorithm based on the synchronized timestamp, and generates an aligned event sequence;
[0128] A causal inference and analysis unit 25, which calculates the causal probability of the aligned event sequence based on a Bayesian network model to determine the causal relationship weight between resource events;
[0129] A root cause location and self-healing unit 26, which performs root cause matching according to the causal relationship weight and a preset abnormal pattern library, generates a root cause list with priority sorting, and triggers the execution of a self-healing strategy.
[0130] Please refer to Figure 3 , an embodiment of the present application also provides an electronic device 300, which includes a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for end-to-end performance tracking of a cloud mobile phone.
[0131] Since the electronic device introduced in this embodiment is the device adopted for implementing an end-to-end performance tracking device for a cloud mobile phone in an embodiment of the present application, based on the method introduced in an embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in an embodiment of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in an embodiment of the present application belongs to the scope protected by the present application.
[0132] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.
[0133] It should be noted that in the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0134] Those skilled in the art should understand that the embodiments of the present application can provide methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0138] The embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a cloud mobile phone end-to-end performance tracking method in the corresponding embodiment.
[0139] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are wholly or partly generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by the computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.
[0140] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0141] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in an electrical, mechanical, or other form.
[0142] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware and / or software functional units.
[0144] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods of the various embodiments of the present application.
[0145] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
[0146] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications that fall within the scope of this specification.
[0147] Obviously, those skilled in the art can make various changes and deformations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification is also intended to include these modifications and deformations.
Claims
1. A cloud phone end-to-end performance tracking method, characterized in that: The method comprises: Obtain client touch event data and network event data, and generate a globally unique identifier based on a preset hash algorithm; Acquire server-side resource event data, wherein the server-side resource event data includes system call tracking data, CPU scheduling tracking data, and resource consumption monitoring data; Based on the hardware-level clock synchronization and software compensation algorithm, the client touch event data and the server resource event data are timestamped to generate a synchronization timestamp; Based on the synchronization timestamp, align the client and server event sequences through a dynamic time warping algorithm to generate an aligned event sequence; Calculate the causal probability of the alignment event sequence based on the Bayesian network model to determine the causal relationship weights between resource events; The root causes are matched with a preset abnormal pattern library according to the causal relationship weights, a prioritized root cause list is generated, and the execution of the self-healing strategy is triggered.
2. The method according to claim 1, characterized in that The acquiring of the client touch event data and the network event data and generating a globally unique identifier based on a preset hash algorithm includes: Intercept the input subsystem touch events at the kernel layer through the eBPF probe to capture touch coordinates, pressure values, and sliding trajectories; Inject trace points at the TCP protocol stack state transition points to record round-trip delay, window size, and socket operation parameters; The client touch event data and network event data are transmitted to the user state through a zero-copy channel and bound to a globally unique identifier.
3. The method according to claim 1, characterized in that The hardware-level clock synchronization includes: Implementing a precision time protocol transparent clock through a programmable logic device to correct a link layer timestamp error to a first preset threshold; The software compensation algorithm includes: The state equation model is used to correct the logical clock drift, and the calibration is cross-validated based on the timestamp counter so that the cross-end time error is less than a second preset threshold.
4. The method according to claim 1, characterized in that: The step of aligning the client and server event sequences based on the synchronization timestamp by a dynamic time warping algorithm to generate an aligned event sequence includes: The window width is constrained to be a preset ratio of the length of the client and server event sequences, and a sliding window is used to match resource fluctuation events within the time window; Interpolate out-of-order or missing events to generate a continuous aligned event sequence.
5. The method according to claim 1, characterized in that The causal probability calculation of the aligned event sequence based on the Bayesian network model to determine the causal relationship weight of the resource event includes: Accelerate sampling of event posterior probabilities based on a parallel computing framework; Construct a device dependency graph and generate probability weights between events using the Markov Chain Monte Carlo algorithm.
6. The method according to claim 1, characterized in that The root cause matching and self-healing strategy triggering include: When it is detected that the waiting time of the high priority thread exceeds a first preset time or the memory allocation difference exceeds a third preset threshold, matching the thread priority inversion or memory leak mode; According to the root cause confidence ranking results, dynamically adjust the virtual CPU binding, video memory pre-allocation or input and output current limiting strategy, and send it to the kernel for execution in real time through the kernel communication channel; The dynamic adjustment of virtual CPU binding is to obtain the physical core load status based on the control group nested monitoring model, and bind the target process to the low-load physical core through the affinity setting command; The video memory pre-allocation is to reserve video memory space based on historical peak data, and lock the pre-allocated area through a memory mapping system call.
7. The method according to claim 1, characterized in that Also includes: Through the dynamic adjustment mechanism of the ring buffer, the size of the kernel-mode and user-mode shared memory areas is adjusted according to the data throughput; Bounds checking instructions and loop unrolling compilation instructions are used to prevent probe programs from accessing memory out of bounds.
8. A cloud phone end-to-end performance tracking device, characterized in that: The device comprises: A data collection and identification unit obtains client touch event data and network event data, and generates a globally unique identifier based on a preset hash algorithm; A service resource monitoring unit, used to obtain server resource event data, wherein the server resource event data includes system call tracking data, CPU scheduling tracking data and resource consumption monitoring data; A time synchronization calibration unit, based on hardware-level clock synchronization and software compensation algorithm, performs timestamp calibration on the client touch event data and the server resource event data to generate a synchronization timestamp; An event alignment processing unit, based on the synchronization timestamp, aligns the client and server event sequences by a dynamic time warping algorithm to generate an aligned event sequence; A causal reasoning analysis unit, which performs causal probability calculation on the alignment event sequence based on a Bayesian network model to determine the causal relationship weights between resource events; The root cause positioning self-healing unit matches the root cause with the preset abnormal pattern library according to the causal relationship weight, generates a priority-ordered root cause list, and triggers the execution of the self-healing strategy.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the cloud phone end-to-end performance tracking method as described in any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the end-to-end performance tracking method of a cloud phone as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Marketing risk control system based on big data
CN120338881A
A marketing risk control system based on big data
CN120338881B
File-oriented IO strength visualization and evaluation system construction method and device and medium
CN120371780A
Method, device and medium for constructing file-oriented IO intensity visualization and evaluation system
CN120371780B
PCIE power consumption testing device
CN120492275A