Cloud mobile phone feedback intelligent evaluation method and related equipment

Through the spatiotemporal convolution neural network model, the network quality and hardware resource data of the cloud mobile cluster are integrated with the spatiotemporal characteristics, and the network configuration is dynamically adjusted, which solves the accuracy and real-time problems of user feedback priority evaluation in the cloud mobile system, and improves the system's resource utilization and stability.

CN120416083APending Publication Date: 2025-08-01启朔(深圳)科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510492958.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The priority evaluation of user feedback in existing cloud mobile systems relies on static rules and lacks multi-dimensional dynamic data analysis, resulting in insufficient accuracy of evaluation results, rigid optimization of hardware resources and network quality, and inability to adapt to real-time business loads, affecting user experience and system stability.

Method used

The spatiotemporal convolutional neural network model is used to fuse the network quality and hardware resource data in spatiotemporal features, generate feedback priority weight coefficients, dynamically adjust network configuration parameters, and optimize bandwidth allocation and protocol adaptive switching through intelligent network cards to realize hardware-level service quality strategy.

Benefits of technology

It improves the accuracy and timeliness of feedback priority evaluation, reduces end-to-end latency, improves the resource utilization and service stability of cloud mobile clusters, and adapts to high concurrency and low latency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416083A_ABST
    Figure CN120416083A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud mobile phone feedback intelligent evaluation method and related equipment, and relates to the technical field of cloud mobile phones, the method comprises the following steps: obtaining network quality indexes and hardware resource usage data of a cloud mobile phone cluster, the network quality indexes comprising a network interruption rate and a key frame packet loss rate, and the hardware resource usage data being used by the cloud mobile phone cluster; the hardware resource use data comprises a GPU video memory occupancy rate and a computing resource utilization rate; performing spatio-temporal feature fusion on the network quality index and the hardware resource use data based on a spatio-temporal convolutional neural network model to generate a feedback priority weight coefficient; determining a target priority score of the target feedback according to the feedback priority weight coefficient; and triggering a hardware-level service quality strategy based on the target priority score, and dynamically adjusting network configuration parameters to optimize the response performance of the cloud mobile phone cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cloud mobile phones, and in particular, to a method and related devices for intelligent evaluation of cloud mobile phone feedback. Background Art

[0002] In the current cloud mobile phone system, the evaluation of user feedback priority mostly relies on static rules or single indicators, lacking comprehensive analysis of multi-dimensional dynamic data, resulting in insufficient accuracy of evaluation results. Existing methods have significant defects in the collaborative optimization of hardware resource usage and network quality indicators, with a rigid weight allocation mechanism that cannot adapt to real-time changing business loads. In addition, traditional solutions rely on software layer monitoring, with low data collection efficiency, high latency, and lack of linkage between hardware resource scheduling strategies and network configuration adjustments, resulting in significant delays in responding to high-priority issues and affecting user experience and system stability. Therefore, there is an urgent need for a method for intelligent evaluation of cloud mobile phone feedback to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a method for intelligent evaluation of cloud mobile phone feedback, the method including:

[0005] Obtain the network quality indicators and hardware resource usage data of the cloud mobile phone cluster, where the network quality indicators include network interruption rate and key frame packet loss rate, and the hardware resource usage data includes GPU video memory occupancy rate and computing resource utilization rate;

[0006] Based on a spatio-temporal convolutional neural network model, perform spatio-temporal feature fusion on the network quality indicators and hardware resource usage data to generate a feedback priority weight coefficient;

[0007] According to the feedback priority weight coefficient, determine the target priority score of the target feedback;

[0008] Based on the target priority score, trigger a hardware-level service quality policy to dynamically adjust network configuration parameters to optimize the response performance of the cloud mobile phone cluster.

[0009] In some embodiments, obtaining the network quality indicators and hardware resource usage data of the cloud mobile phone cluster includes:

[0010] Collect the network interruption rate through a dedicated probe deployed in the cloud mobile phone container. Among them, the dedicated probe calculates the TCP retransmission rate by monitoring the depth of the network card DMA queue through bypass, and realizes data collection based on the user-state zero-copy technology;

[0011] Real-time monitor the key frame loss rate of the cloud mobile phone video stream. Among them, the key frame loss rate is obtained by parsing the SEI metadata of the video encoder;

[0012] Collect the time series data of the GPU video memory occupancy rate curve and the computing resource utilization rate through the GPU resource monitor, and format the time series data into a three-dimensional vector. Among them, the dimensions of the three-dimensional vector include the microsecond-level timestamp, the GPU video memory occupancy rate, and the computing resource utilization rate.

[0013] In some embodiments, perform spatio-temporal feature fusion on the network quality metrics and the hardware resource usage data based on the spatio-temporal convolutional neural network model to generate the feedback priority weight coefficient, including:

[0014] Input the network interruption rate, the key frame loss rate, and the three-dimensional vector into the spatio-temporal convolutional neural network model for spatio-temporal feature extraction. Among them, the spatial dimension convolutional kernel performs weight allocation based on the correlation between the GPU video memory occupancy rate and the computing resource utilization rate; the time dimension convolutional kernel performs temporal modeling on the network interruption event based on the microsecond-level timestamp in the three-dimensional vector.

[0015] Normalize the output result of the spatio-temporal convolutional neural network model into the feedback priority weight coefficient. Among them, the weight coefficient has a dynamic non-linear association with the network interruption rate, the key frame loss rate, and the computing resource utilization rate.

[0016] In some embodiments, determine the target priority score of the target feedback according to the feedback priority weight coefficient, including:

[0017] Generate an initial priority score based on the feedback priority weight coefficient;

[0018] When it is detected that the user identifier associated with the target feedback is a high-priority user, multiply the initial priority score by a preset gain coefficient to generate a gain-adjusted priority score;

[0019] Based on the gain-adjusted priority score, perform weighted average calculation on the priority scores at the current moment and the historical moments through the sliding window algorithm, and output the target priority score. Among them, the window length of the sliding window is negatively correlated with the change frequency of the network interruption rate.

[0020] In some embodiments, trigger the hardware-level service quality policy based on the target priority score and dynamically adjust the network configuration parameters, including:

[0021] Determine the protocol priority mark of the target network transmission queue according to the target priority score, where the target priority score and the protocol priority mark are positively correlated;

[0022] Based on the protocol priority mark, adjust the transmission queue bandwidth allocation strategy of the smart network card, where the bandwidth allocation weight of the high-priority queue is linearly mapped and generated by the target priority score;

[0023] Real-time collect the adjusted network quality metrics and hardware resource usage data;

[0024] Input the network quality metrics and hardware resource usage data into the spatio-temporal convolutional neural network model, and iteratively update the feedback priority weight coefficient to form a closed-loop optimization link.

[0025] In some embodiments, it further includes:

[0026] When it is detected that the computing resource utilization rate exceeds the preset energy efficiency threshold, switch the spatio-temporal convolutional neural network model to the low-power computing mode;

[0027] In the low-power computing mode, enable the sparse matrix acceleration architecture to extract features from the input data, and reduce the GPU core voltage through dynamic voltage and frequency adjustment;

[0028] Real-time monitor the model inference accuracy, and when the inference accuracy is lower than the preset fault tolerance threshold, automatically fallback to the normal computing mode and trigger a computing power compensation instruction.

[0029] In some embodiments, it further includes:

[0030] Based on the leaf-spine network topology structure, real-time monitor the bandwidth utilization rate of the cloud phone cluster;

[0031] When the bandwidth utilization rate is less than or equal to the preset protocol switching threshold, use the RoCEv2 protocol for data transmission;

[0032] When a network congestion event is detected and the bandwidth utilization rate is greater than the preset protocol switching threshold, switch to the TCP / IP protocol and enable the BBRv2 congestion control algorithm, where the protocol switching process is implemented through the P4 programmable pipeline, and the switching delay is less than or equal to 100 microseconds.

[0033] In a second aspect, the present application proposes a cloud phone feedback intelligent evaluation device, and the device includes:

[0034] A cluster data acquisition unit, configured to acquire the network quality metrics and hardware resource usage data of the cloud phone cluster, where the network quality metrics include the network interruption rate and the key frame packet loss rate, and the hardware resource usage data includes the GPU video memory occupancy rate and the computing resource utilization rate;

[0035] A weight coefficient generation unit that performs spatio-temporal feature fusion on network quality metrics and hardware resource usage data based on a spatio-temporal convolutional neural network model to generate feedback priority weight coefficients;

[0036] A weight score determination unit for determining the target priority score of the target feedback according to the feedback priority weight coefficients;

[0037] A parameter configuration adjustment unit triggers a hardware-level service quality policy based on the target priority score and dynamically adjusts network configuration parameters to optimize the response performance of the cloud phone cluster.

[0038] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to implement the steps of the cloud phone feedback intelligent evaluation method according to any one of the first aspects when executing the computer program stored in the memory.

[0039] In a fourth aspect, the present application provides a computer-readable storage medium with a computer program stored thereon. The computer program, when executed by a processor, implements the cloud phone feedback intelligent evaluation method according to any one of the first aspects.

[0040] In summary, the present application uses a dedicated probe to collect multi-dimensional data such as network interruption rate and key frame packet loss rate in real time, and adopts a hardware-accelerated spatio-temporal convolutional neural network to achieve spatio-temporal feature fusion, significantly improving the accuracy and timeliness of feedback priority evaluation. Specifically, the spatio-temporal convolutional neural network model can accurately capture the real-time state changes of the cloud phone cluster by dynamically and non-linearly associating the network interruption rate, key frame packet loss rate, and computing resource utilization rate, combined with the time series modeling of microsecond-level timestamps. Based on the dynamic adjustment of the feedback priority weight coefficients, the present application realizes a millisecond-level response of the hardware-level QoS policy, and controls the end-to-end delay within 50 μs through the bandwidth allocation optimization and protocol adaptive switching of the intelligent network card. In addition, the closed-loop optimization link continuously optimizes the network configuration by iteratively updating the model parameters, effectively improving the resource utilization rate and service stability of the cloud phone cluster, and providing a reliable solution for user feedback processing in high-concurrency and low-latency scenarios. Description of the Drawings

[0041] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of this specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0042] Figure 1 It is a schematic flowchart of a cloud phone feedback intelligent evaluation method provided by an embodiment of the present application;

[0043] Figure 2 Schematic diagram of the structure of an intelligent cloud phone feedback evaluation device provided by an embodiment of the present application;

[0044] Figure 3 Structural diagram of an electronic device for intelligent cloud phone feedback evaluation provided by an embodiment of the present application. Specific implementation manners

[0045] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0046] Please refer to Figure 1 , which is a schematic diagram of the process of an intelligent cloud phone feedback evaluation method provided by an embodiment of the present application, and specifically may include:

[0047] S110. Obtain the network quality indicators and hardware resource usage data of the cloud phone cluster. Among them, the network quality indicators include the network interruption rate and the key frame packet loss rate, and the hardware resource usage data includes the GPU video memory occupancy rate and the computing resource utilization rate;

[0048] Exemplarily, the acquisition of the network quality indicators of the cloud phone cluster is the basis for evaluating the feedback priority. Among them, the network interruption rate reflects the stability of data transmission and is collected in real time by a dedicated probe deployed in the container. Its principle is to indirectly calculate the TCP retransmission rate by bypass monitoring the depth of the network card DMA queue to ensure the efficiency and accuracy of data collection; the key frame packet loss rate is directly related to the video stream service quality and is obtained by parsing the SEI metadata of the video encoder, which can accurately capture the deterioration of video quality caused by network fluctuations. These two indicators jointly depict the network health status of the cloud phone cluster and provide the core input for subsequent dynamic evaluation.

[0049] The collection of hardware resource usage data focuses on the GPU memory occupancy rate and the computing resource utilization rate. The former records the memory usage curve in real time through a GPU resource monitor, and the latter quantifies the computing load of the GPU core. These data are formatted as three-dimensional vectors based on microsecond-level timestamps, not only retaining the temporal characteristics, but also laying a data foundation for the spatio-temporal feature fusion of the spatio-temporal convolutional neural network model through the correlation mapping between the memory and computing resources. The dynamic monitoring of such hardware metrics ensures a deep adaptation of the evaluation strategy to the actual operating state of the cloud phone cluster.

[0050] S120. Perform spatio-temporal feature fusion on the network quality metrics and hardware resource usage data based on the spatio-temporal convolutional neural network model to generate feedback priority weight coefficients;

[0051] Exemplarily, the spatio-temporal convolutional neural network (ST-CNN) jointly models the network quality metrics (network interruption rate, key frame packet loss rate) and hardware resource usage data (GPU memory occupancy rate, computing resource utilization rate) through a multi-dimensional feature fusion mechanism. Its core lies in simultaneously capturing the dynamic associations in the spatio-temporal dimensions: in the spatial dimension, the convolutional kernel analyzes the synergy between the GPU memory occupancy and the computing resource utilization rate, reflecting the impact of resource distribution on system performance; in the time dimension, based on the microsecond-level timestamp, it models the temporal evolution of network interruption events to identify the periodic or sudden characteristics of network fluctuations. This spatio-temporal fusion mechanism can extract key features from a global perspective, providing a dynamic and multi-dimensional data representation for priority evaluation.

[0052] The model transforms the extracted spatio-temporal features into feedback priority weight coefficients through non-linear mapping. The weight coefficients not only reflect the immediate state of the current network quality and hardware load, but also dynamically adjust their allocation ratios through the learning of historical data. For example, when the network interruption rate and GPU utilization rate increase simultaneously, the weight coefficient will adaptively enhance the evaluation weight of the key frame packet loss rate to ensure the accurate determination of the user experience priority in high-load scenarios. This process avoids the subjectivity of static rules and realizes a deep adaptation of the evaluation strategy to the real-time system state.

[0053] S130. Determine the target priority score of the target feedback according to the feedback priority weight coefficients;

[0054] Exemplarily, the determination of the target priority score is based on the dynamic mapping of the feedback priority weight coefficients, and the weight coefficients reflect the comprehensive influence of the network interruption rate, key frame packet loss rate, and hardware resource utilization rate. The weight coefficient is converted into an initial score through a non-linear function. For example, when the high network interruption rate and GPU memory occupancy rate are superimposed, the weight coefficient will significantly increase the value of the initial score, thereby indicating the urgency of the feedback. This process ensures that the score can respond to changes in the system state in real time through the dynamic correlation of the model output.

[0055] To further adapt to complex business scenarios, the target priority score smooths the historical score data through a sliding window algorithm and adjusts the final result by combining a dynamic correction factor for the real-time network interruption event frequency. For example, when network interruption events occur intensively, the correction factor reduces the initial score to avoid misjudgment; while during low-load periods, the scores of high-priority users are amplified to ensure an accurate match between resource allocation and business requirements. This dynamic adjustment mechanism achieves a balance between the real-time and stability of the evaluation results, providing a reliable basis for triggering hardware-level policies.

[0056] S140. Trigger hardware-level service quality policies based on the target priority score, and dynamically adjust network configuration parameters to optimize the response performance of the cloud phone cluster.

[0057] Exemplarily, the core logic of triggering hardware-level service quality policies by the target priority score lies in directly mapping the evaluation results to the physical layer resource configuration. Specifically, based on the high or low priority score, the RoCEv2 protocol priority mark of the RDMA network card and the transmission queue bandwidth allocation of the intelligent network card are dynamically adjusted. For example, a high priority score corresponds to a higher transmission queue weight and exclusive bandwidth for low-latency channels. Such adjustments are achieved through hardware offloading techniques (such as direct writing to MMIO registers) to achieve microsecond-level response, ensuring a strict match between network configuration and real-time business requirements, thereby avoiding resource contention problems caused by scheduling delays in traditional software solutions.

[0058] The optimization effect of dynamically adjusting network configuration parameters is reflected in a significant improvement in end-to-end response performance. Through a closed-loop feedback mechanism, the adjusted network quality metrics (such as key frame loss rate, end-to-end delay) and hardware resource status are collected in real time and input into the spatio-temporal convolutional neural network model to iteratively update the weight coefficients, forming a continuous optimization link of "evaluation - adjustment - verification". For example, when it is detected that the delay exceeds the threshold, the GPU resource reallocation or protocol switching strategy is immediately triggered to ensure that the system maintains a stable throughput in high-concurrency and low-latency scenarios, while strictly controlling the P99 response delay within 50 μs, improving the concurrent processing ability and response efficiency of the cloud phone cluster.

[0059] In summary, the embodiment of the present application collects multi-dimensional data such as network interruption rate and key frame packet loss rate in real time through dedicated probes, and uses hardware-accelerated ST-CNN to realize spatiotemporal feature fusion, thereby improving the accuracy and timeliness of feedback priority evaluation. Specifically, the ST-CNN model can accurately capture the real-time state changes of the cloud phone cluster by dynamically nonlinearly associating network interruption rate, key frame packet loss rate and computing resource utilization, combined with time series modeling of microsecond timestamps. Based on the dynamic adjustment of the feedback priority weight coefficient, the present invention realizes the millisecond-level response of the hardware-level QoS policy, and controls the end-to-end delay within 50μs (P99 value) through the bandwidth allocation optimization and protocol adaptive switching of the smart network card, which is more than 60% lower than the traditional solution. In addition, the closed-loop optimization link continuously optimizes the network configuration by iteratively updating the model parameters, effectively improving the resource utilization and service stability of the cloud phone cluster, and providing a reliable solution for user feedback processing in high concurrency and low latency scenarios.

[0060] In some instances, obtain cloud phone cluster network quality indicators and hardware resource usage data, including:

[0061] The network interruption rate is collected through a dedicated probe deployed in the cloud phone container. The dedicated probe calculates the TCP retransmission rate by monitoring the DMA queue depth of the network card through bypassing, and implements data collection based on user-mode zero-copy technology.

[0062] Real-time monitoring of the key frame packet loss rate of cloud mobile video streams, where the key frame packet loss rate is obtained by parsing the SEI metadata of the video encoder;

[0063] The GPU resource monitor is used to collect time series data of GPU memory occupancy curves and computing resource utilization, and the time series data is formatted into a three-dimensional vector. The dimensions of the three-dimensional vector include microsecond timestamps, GPU memory occupancy, and computing resource utilization.

[0064] Exemplarily, the acquisition of network interruption rate is achieved through a dedicated probe deployed in the cloud phone container. Its core technology is to bypass the monitoring of the network card DMA (direct memory access) queue depth, directly obtain the queue status data at the hardware level, and avoid the performance loss caused by the container virtualization layer. The probe reads the queue status of the network card DMA buffer in real time, counts the number of TCP retransmission packets, and calculates the network interruption rate based on the ratio of the number of retransmissions to the total number of transmissions. In order to avoid the performance loss of traditional software monitoring, the probe uses user-mode zero-copy technology to directly access the network card data through memory mapping (AF_XDP Socket) bypassing the kernel protocol stack to achieve microsecond data collection accuracy. This method significantly improves data collection efficiency through hardware-level monitoring and user-mode optimization, ensuring the real-time and accuracy of the network interruption rate indicator.

[0065] The monitoring of the key frame packet loss rate relies on the parsing of the SEI (Supplemental Enhancement Information) metadata of the video encoder. The cloud mobile phone video stream adopts the H.265 coding standard, and the SEI metadata embeds the identification and transmission status information of the key frame (I frame). The system intercepts the video stream data packets in real time, extracts the key frame sequence number and reception status in the SEI metadata, counts the number of key frames that are not successfully received, and calculates the proportion of the total number of expected key frames as the packet loss rate. Since the loss of key frames will cause subsequent frames to be unable to be decoded, this indicator directly reflects the quality of the user experience and provides the core basis for priority evaluation.

[0066] The acquisition of the GPU video memory occupancy rate and computing resource utilization rate is achieved through a GPU resource monitor. The video memory occupancy rate is obtained in real time by monitoring the GPU video memory allocation interface (such as the NVIDIA CUDA API), and the change curve of the video memory usage is recorded; the computing resource utilization rate is obtained by reading the SM (Streaming Multiprocessor) activity counter of the GPU core and counting the proportion of the time when it is in the active state. The collected time series data is aligned through microsecond-level timestamps and formatted into a three-dimensional vector, whose dimensions include the timestamp (accurate to 0.1 μs), the video memory occupancy rate (percentage), and the computing resource utilization rate (percentage). The structured data not only retains the temporal correlation but also provides a standardized input format for the spatio-temporal convolutional neural network.

[0067] The acquisition schemes of the above three types of data achieve the collaborative monitoring of multi-dimensional indicators through hardware-level optimization and standardized processing. The bypass monitoring mechanism of the dedicated probe avoids the performance loss of the virtualization layer and ensures the low-latency acquisition of the network interruption rate data; the SEI metadata parsing makes full use of the embedded information of the video coding standard without additional computational overhead; the GPU resource monitor directly reads the hardware status through the native interface, ensuring the authenticity and timeliness of the data. The formatted three-dimensional vector is transmitted to the GPU computing cluster through the PCIe Gen4 x16 link and stored in the HBM2 high-speed memory, providing structured input for the spatio-temporal feature extraction of the spatio-temporal convolutional neural network, providing comprehensive and high-precision data support for the real-time status of the cloud mobile phone cluster, and laying a reliable data foundation for subsequent dynamic evaluation and resource scheduling.

[0068] In some instances, based on the spatio-temporal convolutional neural network model, the spatio-temporal feature fusion of the network quality indicators and the hardware resource usage data is performed to generate the feedback priority weight coefficients, including:

[0069] The network interruption rate, key frame packet loss rate, and three-dimensional vector are input into a spatio-temporal convolutional neural network model for spatio-temporal feature extraction. Among them, the spatial dimension convolutional kernel performs weight allocation based on the correlation between GPU video memory occupancy and computing resource utilization rate; the temporal dimension convolutional kernel performs temporal modeling on network interruption events based on the microsecond-level timestamps in the three-dimensional vector.

[0070] The output result of the spatio-temporal convolutional neural network model is normalized into a feedback priority weight coefficient, where the weight coefficient has a dynamic non-linear correlation with the network interruption rate, key frame packet loss rate, and computing resource utilization rate.

[0071] Exemplarily, a spatio-temporal convolutional neural network (ST-CNN) model takes the network interruption rate, key frame packet loss rate, and three-dimensional vector as input data. The three-dimensional vector consists of a microsecond-level timestamp, GPU video memory occupancy, and computing resource utilization rate, and its dimension design ensures the synchronous embedding of temporal and spatial information. The network interruption rate is calculated through the TCP retransmission rate collected by a dedicated probe, and the key frame packet loss rate is derived from the parsing result of the video encoder SEI metadata. After the input data undergoes standardized preprocessing (such as normalization and time alignment), a spatio-temporal feature tensor is formed, providing a structured input for the multi-dimensional feature extraction of the model.

[0072] The spatial dimension convolutional kernel dynamically allocates feature weights by analyzing the correlation between GPU video memory occupancy and computing resource utilization rate. Specifically, when the video memory occupancy is high, if the computing resource utilization rate rises synchronously, the convolutional kernel will enhance the feature extraction ability sensitive to hardware load. For example, when video memory resources are scarce, it will preferentially identify abnormal events related to video stream processing. The weight matrix of the convolutional kernel is optimized through the backpropagation algorithm to learn the non-linear coupling relationship between video memory and computing resources, thereby capturing the impact of hardware status on network performance in the spatial dimension.

[0073] The temporal dimension convolutional kernel performs temporal modeling on network interruption events based on the microsecond-level timestamps in the three-dimensional vector. The convolutional kernel captures the periodic fluctuation characteristics of the network interruption rate through a sliding window (window size 50 - 200 ms), such as the short-term dense distribution of sudden interruptions or the steady state under low load. The weight allocation of temporal convolution adopts weighted moving average (EMA coefficient 0.2), and dynamically adjusts the temporal sensitivity in combination with the time interval of network interruption events (such as the minimum RTO of 200 ms). In addition, the intelligent network card hardware counters (such as the number of retransmissions, out-of-order packets) are read in real time through a P4 programmable pipeline, associating the original temporal data with hardware events, and enhancing the model's forward-looking prediction ability for network fluctuations. The temporal convolutional kernel enhances the sensitivity to high-frequency abnormal events through long short-term memory (LSTM) units or temporal attention mechanisms, thereby accurately predicting the network stability trend.

[0074] The output result of the spatio-temporal convolutional neural network model is normalized by the Softmax function to generate the feedback priority weight coefficient. The normalization process introduces a dynamic scaling factor, whose value is determined by the real-time ratios of the network interruption rate, the key frame packet loss rate, and the GPU resource utilization rate. For example, when the key frame packet loss rate exceeds the preset threshold, the scaling factor automatically amplifies the contribution weight of the network interruption rate to ensure that video quality degradation events trigger policy adjustments first. The normalized weight coefficient is stored in the GPU video memory in FP16 format and directly written into the RDMA network card configuration engine through the PCIe P2P technology to achieve microsecond-level weight synchronization and avoid data transmission bottlenecks in traditional solutions.

[0075] The feedback priority weight coefficient has a dynamic non-linear relationship with the network interruption rate, the key frame packet loss rate, and the computing resource utilization rate, and its relationship is continuously optimized through the model's backpropagation algorithm. Specifically, the loss function is designed as a weighted combination of the end-to-end delay (P99 value) and the resource utilization rate (such as a delay weight of 70% and a resource utilization rate weight of 30%). During the training process, the convolution kernel parameters are adjusted through gradient descent. For example, in a high-concurrency scenario, the model preferentially reduces the weight deviation of delay-sensitive tasks; while in a low-load situation, it optimizes the resource utilization rate and energy efficiency ratio. In addition, a federated learning architecture (differential privacy ε = 0.3) is used to perform distributed fine-tuning on the model, and the bsdiff algorithm is combined to compress the parameter update amount (compression rate ≥ 85%) to ensure the global consistency and privacy security of the weight coefficient. This mechanism enables the weight coefficient to adapt to the dynamic load changes of the cloud mobile phone cluster and improves the evaluation accuracy and system stability.

[0076] In some instances, according to the feedback priority weight coefficient, determine the target priority score of the target feedback, including:

[0077] Generate an initial priority score based on the feedback priority weight coefficient;

[0078] When it is detected that the user identifier associated with the target feedback is a high-priority user, multiply the initial priority score by a preset gain coefficient to generate a gain-adjusted priority score;

[0079] Based on the gain-adjusted priority score, perform a weighted average calculation on the priority scores at the current moment and historical moments through a sliding window algorithm, and output the target priority score, where the window length of the sliding window is negatively correlated with the change frequency of the network interruption rate.

[0080] Exemplarily, the initial priority score is generated by successively multiplying each feedback priority weight coefficient by the standardized value of its corresponding network quality or hardware resource metric and then summing them up, and then adding a preset baseline priority threshold. Specifically, the weight coefficients are output by a spatio-temporal convolutional neural network model, representing the influence degree of each metric on the priority. The network quality metrics and hardware resource metrics need to be normalized to eliminate the dimension difference. The final score is obtained by accumulating the product of each weight coefficient and the corresponding metric value and adding a baseline offset, thereby objectively quantifying the priority level of the current system state.

[0081] When the system identifies that the target feedback is associated with a high-priority user (such as a VIP user authenticated by a national cryptography algorithm), the initial priority score will be dynamically amplified by multiplying it by a preset gain coefficient. The gain coefficient is set hierarchically according to the user level. For example, the coefficient for ordinary users remains 1.0, and the coefficient for VIP users is increased to 1.2. The adjusted score is achieved by directly multiplying the initial score by the gain coefficient, ensuring that high-priority user feedback occupies a higher weight in resource allocation and preferentially obtains low-latency channels and sufficient computing resources.

[0082] The priority score after gain adjustment is smoothed in time series by a sliding window algorithm. The window length of the sliding window is dynamically adjusted according to the change frequency of the network interruption rate: when the network fluctuates frequently, the window is shortened to quickly respond to sudden problems; when the network is stable, the window is extended to improve the result stability. The scores at each moment within the window are weighted and averaged according to the time decay weight, that is, the score weight closer to the current moment is higher, and the decay weight is calculated with the base of the natural logarithm and the exponent of the product of the negative decay coefficient and the time interval. The final target priority score is obtained by weighted summation and dividing by the sum of the weights, effectively suppressing instantaneous noise interference while retaining the timeliness of the score.

[0083] The final priority score calibrates the weighted average result through a dynamic correction factor. The correction factor is calculated based on the real-time network interruption event frequency. The higher the interruption frequency, the smaller the correction factor, to suppress the overestimation of the score caused by network instability; otherwise, the score is increased to strengthen the advantage of high-priority users. The calibrated score takes effect through the hardware-level priority marking module of the smart network card, supports multi-level transmission queue configuration, and realizes nanosecond-level delay through direct register writing, ensuring that the end-to-end response performance meets strict latency requirements.

[0084] The length of the sliding window is inversely related to the change frequency of the network interruption rate. The specific rule is as follows: the more frequent the fluctuation of the network interruption rate, the shorter the window length; conversely, the longer it is. The window length is calculated by the ratio of a constant to the sum of the change frequency of the network interruption rate and a smoothing factor, ensuring that the denominator is not zero and the result is smooth. For example, when the number of fluctuations per second of the network interruption rate increases from 5 times to 20 times, the window length adaptively shortens from 200 milliseconds to 50 milliseconds. This mechanism enables the target priority score to accurately guide resource scheduling decisions under different network states by dynamically balancing real-time performance and stability.

[0085] In some instances, trigger the hardware-level quality of service policy based on the target priority score, and dynamically adjust network configuration parameters, including:

[0086] Determine the protocol priority mark of the target network transmission queue according to the target priority score, where the target priority score and the protocol priority mark are positively correlated;

[0087] Based on the protocol priority mark, adjust the transmission queue bandwidth allocation policy of the intelligent network card, where the bandwidth allocation weight of the high-priority queue is linearly mapped and generated by the target priority score;

[0088] Real-time collect the adjusted network quality metrics and hardware resource usage data;

[0089] Input the network quality metrics and hardware resource usage data into the spatio-temporal convolutional neural network model, and iteratively update the feedback priority weight coefficient to form a closed-loop optimization link.

[0090] Exemplarily, the target priority score is mapped to the priority mark (range 0 to 7) of the RoCEv2 protocol through a linear proportional relationship. The higher the score, the higher the mark level. Specifically, the score interval is divided into eight levels. For example, when the score is higher than or equal to 90%, the highest priority mark 7 is assigned. For every 10% decrease in the score, the mark level decreases by 1 in sequence. The configuration of the priority mark is directly written through the hardware register of the intelligent network card and is updated by the hardware offloading engine within 100 nanoseconds, avoiding the scheduling delay of the traditional software protocol stack. At the same time, based on the programmable network pipeline, real-time monitor the network congestion status. If it is detected that the queue depth exceeds the dynamic threshold (such as more than 1024 packets backlogged in each queue), automatically enhance the transmission preemption ability of high-priority traffic to ensure lossless transmission of critical services (packet loss rate lower than 0.01%).

[0091] The bandwidth allocation strategy for the transmission queue of the intelligent network card is dynamically generated based on the protocol priority marking. The bandwidth weights of the high-priority queues (marking levels of 5 and above) are calculated through a linear relationship, and the weight values are composed of the product of the target priority score and the preset slope parameter plus the fixed intercept parameter. The slope and intercept are dynamically adjusted according to the real-time link bandwidth utilization rate. For example, when the bandwidth utilization rate exceeds 80%, the slope is increased to strengthen the bandwidth advantage of high-priority traffic. The minimum granularity of bandwidth allocation is 1 Mbps, and burst traffic buffering is supported (maximum buffer capacity of 128 frames). The configuration parameters are sent in real time through the network card performance monitoring module, combined with the dynamic congestion control algorithm to optimize the transmission efficiency, ensuring that the end-to-end delay (P99 value) of high-priority traffic is controlled within 50 microseconds, reducing by at least 60% compared with the traditional TCP / IP solution.

[0092] The adjusted network quality metrics (such as network interruption rate, key frame loss rate) and hardware resource data (such as GPU video memory occupancy rate, computing resource utilization rate) are collected in real time through dedicated probes and resource monitors. The data is based on a microsecond-level timestamp (accuracy error not exceeding 0.1 microsecond) and is formatted into a vector containing five dimensions: timestamp, network interruption rate, key frame loss rate, GPU video memory occupancy rate, and computing resource utilization rate. The data is transmitted to the storage unit of the GPU cluster through a high-speed PCIe interface, and an efficient compression algorithm (compression rate not less than 85%) and cyclic redundancy check (CRC) are used during the transmission process to ensure data integrity and transmission efficiency (bandwidth reaches over 28.9 GB / s in actual tests). The formatted data is further aligned in time series and filtered for noise through a sliding window (window size from 50 to 200 milliseconds) to provide high-quality input for model iteration.

[0093] The collected data is input into a spatio-temporal convolutional neural network model and iteratively trained through a distributed federated learning framework (differential privacy parameter ε = 0.3) to dynamically update the feedback priority weight coefficient. During the training process, the loss function combines three objectives: end-to-end delay (accounting for 70%), resource utilization rate (accounting for 20%), and energy efficiency ratio (accounting for 10%), and uses the gradient descent algorithm to adjust the model parameters. The model weights are synchronized within the GPU cluster through a high-speed communication library (supporting GPUDirect RDMA technology), and the synchronization delay does not exceed 2 microseconds. The closed-loop optimization link completes a global update every 5 minutes and uses a differential compression algorithm (compression rate exceeding 92%) to reduce the communication data volume. After testing, this mechanism improves the resource utilization rate from 68% of the traditional solution to 92%, and reduces the standard deviation of the end-to-end delay by 76%, significantly improving the system stability and response consistency.

[0094] In some instances, it also includes:

[0095] When the detected computing resource utilization exceeds the preset energy efficiency threshold, switch the spatio-temporal convolutional neural network model to the low-power computing mode;

[0096] In the low-power computing mode, enable the sparse matrix acceleration architecture to extract features from the input data, and reduce the GPU core voltage through dynamic voltage and frequency adjustment;

[0097] Monitor the model inference accuracy in real time. When the inference accuracy is lower than the preset fault tolerance threshold, automatically fallback to the normal computing mode and trigger the computing power compensation instruction.

[0098] Exemplarily, when the GPU computing resource utilization exceeds the preset energy efficiency threshold (e.g., 70%), the system automatically switches the ST-CNN model to the low-power computing mode. The energy efficiency threshold is dynamically calibrated based on historical load data and the hardware performance curve. For example, by analyzing the power consumption-utilization relationship of the GPU under typical loads, the critical point that balances energy efficiency and performance is determined. The switching process is achieved by interrupting the current computing task and reloading the low-power model parameters, ensuring that the mode switching latency is less than 10 milliseconds and minimizing the impact on real-time services.

[0099] In the low-power mode, enable the sparse matrix acceleration architecture (sparsity ≥ 40%) to extract features from the input data by block-structured compression (4:2 mode), and only perform convolution operations on non-zero values, increasing the computing density to 8.4 TFLOPs / mm 2 . At the same time, reduce the GPU core voltage to 0.85V through dynamic voltage and frequency adjustment (DVFS), and adjust the frequency in 12 levels (step size 50 MHz), reducing the maximum power consumption by 40%. The voltage adjustment signal is generated by a digital PWM controller (switching frequency 2 MHz) and takes effect in real time through the power management unit (PMU) of the HBM2 memory, ensuring that the energy efficiency is increased by ≥ 30% compared to the normal mode.

[0100] In the low-power mode, monitor the model inference accuracy in real time (such as classification accuracy or regression error), and use the sliding window algorithm (window size 50 frames) to statistically analyze the accuracy fluctuations. If the accuracy is lower than the preset fault tolerance threshold (e.g., 95%), immediately trigger a two-stage fallback strategy:

[0101] Mode fallback: Call the MIG instance of the spare GPU resource pool through the CXL 2.0 interconnection technology to restore the normal computing mode within ≤ 5 μs;

[0102] Computing power compensation: Enable the FPGA preprocessing unit (such as Xilinx Alveo U280) to recalculate the historical low-precision inference data, and transmit the corrected results back through the PCIe Gen4 x16 link (bandwidth ≥ 28.9 GB / s) to ensure business continuity.

[0103] Distributed fine-tuning of model parameters in the low-power mode is carried out through a federated learning architecture (differential privacy ε = 0.3). The bsdiff algorithm is used to compress the parameter update amount (compression rate ≥ 92%), and cross-GPU synchronization is achieved through the NCCL communication library (latency ≤ 2 μs). Verified by the MLPerf benchmark, the end-to-end response latency in the low-power mode is ≤ 50 μs (P99 value), the energy efficiency ratio reaches 15.8 TOPS / W, which is 35% higher than that in the conventional mode, and the accuracy regression rate < 0.5%. In addition, the resource fragmentation rate is reduced from 32% in the traditional solution to 8%, and the GPU utilization rate is stable above 90%, significantly optimizing the long-term operation stability of the cloud mobile phone cluster.

[0104] In some instances, it also includes:

[0105] Based on the leaf-spine network topology, the bandwidth utilization rate of the cloud mobile phone cluster is monitored in real time;

[0106] When the bandwidth utilization rate is less than or equal to the preset protocol switching threshold, the RoCEv2 protocol is used for data transmission;

[0107] When a network congestion event is detected and the bandwidth utilization rate is greater than the preset protocol switching threshold, switch to the TCP / IP protocol and enable the BBRv2 congestion control algorithm. Among them, the protocol switching process is implemented through the P4 programmable pipeline, and the switching latency is less than or equal to 100 microseconds.

[0108] Exemplarily, based on the Leaf-Spine network topology, the link bandwidth utilization data is collected in real time through the P4 programmable pipeline deployed on the Leaf switch. Specifically, the counter module in the pipeline counts the incoming and outgoing traffic of each port at a fixed sampling period (such as 5 seconds) and calculates the instantaneous bandwidth utilization rate (formula: utilization rate = (actual traffic / link bandwidth) × 100%). The monitoring data is reported to the central controller in real time through the DMA queue depth monitoring module (accuracy 0.1 μs), and the protocol switching threshold (default threshold 40%) is dynamically adjusted in combination with the historical traffic pattern (such as the daily peak period) to ensure that the policy adapts to the periodic changes of the service load.

[0109] When the bandwidth utilization rate ≤ the preset threshold, the RoCEv2 protocol is used for data transmission. RoCEv2 achieves lossless transmission through the DCQCN (Priority-based Dynamic Congestion Notification) algorithm, including: dividing 8 levels of traffic priorities (0 - 7) according to the target priority score, and high-priority traffic (mark ≥ 5) has the bandwidth preemption right; dynamically adjusting the sending rate through ECN (Explicit Congestion Notification) marking and rate limiter (minimum granularity 1Mbps) to ensure zero-packet-loss transmission triggered by PFC (Priority Flow Control) frames; the smart network card (such as Mellanox ConnectX-6DX) directly processes the RoCEv2 protocol stack, and the end-to-end delay ≤ 50μs (P99 value), which is reduced by ≥ 60% compared with the traditional TCP / IP solution.

[0110] When it is detected that the bandwidth utilization rate > the threshold and network congestion occurs (such as the queue depth > 1024 packets / queue), the system switches to the TCP / IP protocol and enables the BBRv2 congestion control algorithm. BBRv2 optimizes transmission through the following mechanisms: dynamically estimating the bottleneck bandwidth and link delay based on the ACK arrival time and packet interval; using the gradient descent algorithm to adjust the congestion window to avoid the aggressive speed reduction of the traditional AIMD (Additive Increase / Multiplicative Decrease) strategy; implementing protocol stack switching (such as from RoCEv2 to TCP / IP) through the P4 programmable pipeline, and the switching process includes protocol header rewriting, flow table entry update, and queue priority reset, with the whole process delay ≤ 100μs.

[0111] After the protocol switch, network quality indicators (packet loss rate, delay, throughput) are collected in real time and the effectiveness of the strategy is evaluated through a spatio-temporal convolutional neural network model. If the delay after switching still exceeds the limit (such as P99 > 50μs), the protocol priority weight coefficient in the model is iteratively updated according to the real-time data to strengthen the adaptability to high-load scenarios; through the Kubernetes DevicePlugin interface, the GPU video memory pages (4KB granularity) are preemptively reclaimed from low-priority tasks to preferentially guarantee the computing resources of TCP / IP traffic; the protocol switch decision model is updated daily through reinforcement learning (PPO algorithm), and the threshold setting is optimized in combination with the MLPerf benchmark test results. After actual measurement, this mechanism increases the throughput in high-load scenarios by 35%, reduces the resource fragmentation rate from 32% to 8%, and reduces the standard deviation of the end-to-end delay by 76%, significantly enhancing the robustness of the cloud phone cluster.

[0112] Please refer to Figure 2 , which is a schematic structural diagram of a cloud phone feedback intelligent evaluation device provided by an embodiment of the present application, including:

[0113] The cluster data acquisition unit 21 is used to acquire the network quality indicators and hardware resource usage data of the cloud mobile phone cluster. Among them, the network quality indicators include the network interruption rate and the key frame packet loss rate, and the hardware resource usage data includes the GPU video memory occupancy rate and the computing resource utilization rate;

[0114] The weight coefficient generation unit 22 performs spatio-temporal feature fusion on the network quality indicators and hardware resource usage data based on the spatio-temporal convolutional neural network model to generate feedback priority weight coefficients;

[0115] The weight score determination unit 23 is used to determine the target priority score of the target feedback according to the feedback priority weight coefficient;

[0116] The parameter configuration adjustment unit 24 triggers the hardware-level service quality policy based on the target priority score, and dynamically adjusts the network configuration parameters to optimize the response performance of the cloud mobile phone cluster.

[0117] Please refer to Figure 3 , this application embodiment also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for intelligent evaluation of cloud mobile phone feedback.

[0118] Since the electronic device introduced in this embodiment is the device adopted for implementing a cloud mobile phone feedback intelligent evaluation device in this application embodiment, based on the method introduced in this application embodiment, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in this application embodiment will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in this application embodiment belongs to the scope of protection of this application.

[0119] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0120] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0121] Those skilled in the art should understand that the embodiments of the present application can provide a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-readable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0125] The embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a cloud mobile phone feedback intelligent evaluation method in the corresponding embodiment.

[0126] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or by wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium (such as a solid-state drive), etc.

[0127] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0128] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of devices or units may be in electrical, mechanical, or other forms.

[0129] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0130] In addition, the functional units in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0131] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.

[0132] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

[0133] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0134] Obviously, those skilled in the art can make various changes and modifications to this specification without departing from the spirit and scope of this specification. In this way, if these modifications and variations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification is also intended to include these modifications and variations.

Claims

1. A method for intelligent evaluation of cloud mobile phone feedback, characterized in that, The method comprises: Obtain network quality indicators and hardware resource usage data of the cloud phone cluster, wherein the network quality indicators include network interruption rate and key frame packet loss rate, and the hardware resource usage data includes GPU memory occupancy rate and computing resource utilization rate; Based on the spatiotemporal convolutional neural network model, the network quality indicator and the hardware resource usage data are subjected to spatiotemporal feature fusion to generate a feedback priority weight coefficient; Determining a target priority score for the target feedback according to the feedback priority weight coefficient; A hardware-level quality of service policy is triggered based on the target priority score, and network configuration parameters are dynamically adjusted to optimize the response performance of the cloud phone cluster.

2. The method according to claim 1, characterized in that, The acquisition of network quality indicators and hardware resource usage data of the cloud phone cluster includes: The network interruption rate is collected through a dedicated probe deployed in the cloud phone container. The dedicated probe calculates the TCP retransmission rate by bypassing the network card DMA queue depth and implements data collection based on user-mode zero-copy technology. Real-time monitoring of the key frame packet loss rate of the cloud phone video stream, wherein the key frame packet loss rate is obtained by parsing the SEI metadata of the video encoder; The GPU resource monitor is used to collect time series data of GPU memory occupancy curve and computing resource utilization, and the time series data is formatted into a three-dimensional vector, wherein the dimensions of the three-dimensional vector include microsecond timestamp, GPU memory occupancy and computing resource utilization.

3. The method according to claim 2, wherein The spatiotemporal feature fusion of the network quality indicator and the hardware resource usage data based on the spatiotemporal convolutional neural network model to generate a feedback priority weight coefficient includes: The network interruption rate, the key frame packet loss rate, and the three-dimensional vector are input into a spatiotemporal convolutional neural network model for spatiotemporal feature extraction, wherein the spatial dimension convolution kernel performs weight assignment based on the correlation between GPU memory occupancy and computing resource utilization; the temporal dimension convolution kernel performs time series modeling of network interruption events based on the microsecond timestamp in the three-dimensional vector; The output result of the spatiotemporal convolutional neural network model is normalized into a feedback priority weight coefficient, wherein the weight coefficient is dynamically nonlinearly associated with the network interruption rate, the key frame packet loss rate, and the computing resource utilization rate.

4. The method according to claim 1, characterized in that, Determining the target priority score of the target feedback according to the feedback priority weight coefficient includes: generating an initial priority score based on the feedback priority weight coefficient; When it is detected that the user identifier associated with the target feedback is a high-priority user, multiplying the initial priority score by a preset gain coefficient to generate a gain-adjusted priority score; Based on the gain-adjusted priority score, a weighted average calculation is performed on the priority scores of the current moment and the historical moment through a sliding window algorithm to output a target priority score, wherein the window length of the sliding window is negatively correlated with the frequency of change of the network interruption rate.

5. The method according to claim 1, wherein The triggering of a hardware-level quality of service policy based on the target priority score and dynamically adjusting network configuration parameters includes: Determine the protocol priority tag of the target network transmission queue according to the target priority score, where the target priority score and the protocol priority tag are positively correlated; Based on the protocol priority tag, adjust the transmission queue bandwidth allocation strategy of the smart network card, where the bandwidth allocation weight of the high-priority queue is linearly mapped by the target priority score; Real-time collect the adjusted network quality indicators and hardware resource usage data; Input the network quality indicators and the hardware resource usage data into the spatio-temporal convolutional neural network model, and iteratively update the feedback priority weight coefficient to form a closed-loop optimization link.

6. The method according to claim 1, wherein It also includes: When it is detected that the computing resource utilization rate exceeds the preset energy efficiency threshold, switch the spatio-temporal convolutional neural network model to the low-power computing mode; In the low-power computing mode, enable the sparse matrix acceleration architecture to extract features from the input data, and reduce the GPU core voltage through dynamic voltage and frequency adjustment; Real-time monitor the model inference accuracy, and when the inference accuracy is lower than the preset fault tolerance threshold, automatically fallback to the normal computing mode and trigger a computing power compensation instruction.

7. The method according to claim 1, wherein It also includes: Based on the leaf-spine network topology, real-time monitor the bandwidth utilization rate of the cloud phone cluster; When the bandwidth utilization rate is less than or equal to the preset protocol switching threshold, use the RoCEv2 protocol for data transmission; When a network congestion event is detected and the bandwidth utilization rate is greater than the preset protocol switching threshold, switch to the TCP / IP protocol and enable the BBRv2 congestion control algorithm, where the protocol switching process is implemented through the P4 programmable pipeline, and the switching delay is less than or equal to 100 microseconds.

8. A cloud mobile phone feedback intelligent evaluation device, characterized in that, The device includes: A cluster data acquisition unit for acquiring the network quality indicators and hardware resource usage data of the cloud phone cluster, where the network quality indicators include the network interruption rate and the key frame loss rate, and the hardware resource usage data includes the GPU video memory occupancy rate and the computing resource utilization rate; A weight coefficient generation unit that generates a feedback priority weight coefficient based on the spatio-temporal convolutional neural network model for spatio-temporal feature fusion of the network quality indicators and the hardware resource usage data; A weight score determination unit for determining the target priority score of the target feedback according to the feedback priority weight coefficient; A parameter configuration adjustment unit that triggers a hardware-level service quality policy based on the target priority score and dynamically adjusts network configuration parameters to optimize the response performance of the cloud phone cluster.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the cloud phone feedback intelligent evaluation method according to any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the cloud phone feedback intelligent evaluation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent Kubernetes node maintenance method and device

    CN120896908A