E-commerce live broadcast real-time interaction quality evaluation system based on edge calculation

Through edge computing technology, multimodal interactive data of e-commerce live broadcasts is collected and analyzed in real time, combined with lightweight neural networks and deep learning models, and dynamically adjusting parameters, solving the delay and accuracy of e-commerce live broadcast interaction quality evaluation, and improving live broadcast quality and user experience.

CN120343291AActive Publication Date: 2025-07-18WUHAN QISHI MEDIA CO LTD

Patent Information

Application Number
CN202510520967.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-18
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing e-commerce live broadcast interaction quality evaluation methods cannot comprehensively and accurately reflect the interaction quality, ignore user experience, and cannot meet the real-time evaluation needs, and traditional computing models lead to delays.

Method used

A real-time interactive quality evaluation system based on edge computing is adopted, multimodal interactive data is collected in real time through a distributed edge computing node cluster, and data fusion analysis is performed using a lightweight spatiotemporal attention neural network model, combining with a deep reinforcement learning network generation optimization strategy, and dynamic parameter adjustment is performed through an adaptive fuzzy inference system.

Benefits of technology

It realizes real-time, accurate and comprehensive evaluation of the quality of e-commerce live broadcast interaction, improves the fluency, clarity and interactivity of live broadcasts, provides a better user experience, and provides reliable quality evaluation and optimization basis for live broadcast platforms and merchants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343291A_ABST
    Figure CN120343291A_ABST
Patent Text Reader

Abstract

The invention discloses an e-commerce live broadcast real-time interaction quality evaluation system based on edge calculation, and relates to the technical field of e-commerce live broadcast, and the system comprises a multi-modal interaction data collection module which collects multi-modal interaction data of a live broadcast stream in real time through a distributed edge calculation node cluster, and constructs a multi-dimensional quality feature vector, the multi-modal interaction data comprises a video coding parameter, an audio quality index, user interaction behavior data and network transmission state data; according to the invention, the distributed edge computing node cluster collects the multi-modal interaction data of the live stream in real time, the lightweight space-time attention neural network model carries out data fusion processing, the deep reinforcement learning network generates a quality optimization scheme, and the adaptive fuzzy inference system corrects the optimization scheme in real time. And dynamic parameter adjustment is carried out in combination with network bandwidth fluctuation and a terminal device resource state, so that the effect of accurately and comprehensively evaluating the e-commerce live broadcast interaction quality in real time is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of e-commerce live streaming, and particularly to a real-time interactive quality evaluation system for e-commerce live streaming based on edge computing. Background Art

[0002] With the rapid development of Internet technology and the change of consumers' shopping habits, e-commerce live streaming has become a new and popular marketing mode. In the process of e-commerce live streaming, the real-time interactive quality plays a crucial role in enhancing the user experience and promoting commodity sales. More and more merchants and brands rely on e-commerce live streaming platforms to promote products and attract consumers to purchase.

[0003] However, the existing methods for evaluating the interactive quality of e-commerce live streaming have many deficiencies: on the one hand, most of them use a single index for evaluation, such as the interaction frequency, the number of comments and the number of likes, etc., which cannot comprehensively and accurately reflect the real situation of the interactive quality and cannot comprehensively measure the interactive quality, because a single index cannot comprehensively consider various factors such as the effectiveness of interaction and the depth of user participation; on the other hand, the existing evaluation methods often consider the user experience insufficiently, ignoring the impact of factors such as the smoothness of the live stream, the picture quality, and the performance of the anchor on the interactive quality, and cannot meet the personalized needs of users, resulting in a deviation between the evaluation result and the actual user feeling; in addition, the data generated by e-commerce live streaming is huge in quantity and has high real-time requirements, and the traditional centralized computing mode is prone to delay when processing these data and cannot meet the requirements of real-time evaluation.

[0004] Therefore, there is an urgent need to provide a real-time interactive quality evaluation system for e-commerce live streaming based on edge computing, which can comprehensively and accurately evaluate the interactive quality of e-commerce live streaming in real time.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a real-time interactive quality evaluation system for e-commerce live streaming based on edge computing. The present invention collects multi-modal interaction data of the live stream in real time through a distributed edge computing node cluster, and through the fusion processing of a spatio-temporal attention neural network model, evaluates and adaptively dynamically adjusts the live interactive quality to solve the problems in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solutions: An e-commerce live streaming real-time interaction quality evaluation system based on edge computing, including a multimodal interaction data acquisition module: Through a distributed edge computing node cluster, it collects multimodal interaction data of the live stream in real time and constructs a multi-dimensional quality feature vector. Among them, the multimodal interaction data includes video coding parameters, audio quality indicators, user interaction behavior data, and network transmission status data;

[0008] An interaction quality analysis module: Using a lightweight spatio-temporal attention neural network model built into the edge node, it performs real-time fusion analysis on the multimodal interaction data to generate real-time interaction quality evaluation indicators;

[0009] An evaluation strategy generation module: Adopting a generator model of a deep reinforcement learning network, it combines real-time interaction quality evaluation indicators with historical optimization strategies to generate a quality optimization plan. The discriminator model evaluates the effectiveness of the quality optimization plan according to the feedback information of the user terminal;

[0010] A dynamic adjustment strategy module: Through an adaptive fuzzy inference system, it makes real-time corrections to the quality optimization plan, and dynamically adjusts the parameters of the quality optimization plan generated by the deep reinforcement learning network in combination with the current network bandwidth fluctuation and the resource status of the terminal device;

[0011] A feedback execution module: Based on the quality optimization plan executed after dynamic parameter adjustment, in combination with the distributed decision-making mechanism of the edge node, it collaboratively optimizes video coding parameters, transmission protocols, and resource allocation strategies to dynamically control the live streaming quality. The edge computing node then performs real-time evaluation on the e-commerce live streaming quality and feeds back the evaluation results to the live streaming platform and merchants.

[0012] Optionally, the multimodal interaction data acquisition module includes an edge deployment unit: Deploy lightweight multimodal data acquisition arrays at both the content delivery network (CDN) edge nodes and user terminals to form a distributed data acquisition network;

[0013] A multimodal synchronization unit: The distributed data acquisition network synchronizes and collects multi-dimensional data such as video frame rate, encoding bit rate, audio delay, user likes, comment frequency, and network jitter rate through a timestamp alignment mechanism to generate raw data that comprehensively reflects the interaction situation in the live broadcast room;

[0014] A preprocessing unit: It performs filtering, denoising, and normalization processing on the raw data to construct a structured interaction quality feature matrix;

[0015] A feature extraction unit: Using edge computing resources, it performs real-time feature extraction on the interaction quality feature matrix to generate a multi-dimensional quality feature vector containing time-domain and space-domain features;

[0016] Edge transmission unit: After compressing the multi-dimensional quality feature vectors processed by feature extraction using hierarchical compression technology, it is transmitted to the analysis node through the CDN edge network.

[0017] Optionally, the interactive quality analysis module includes a data fusion unit: spatio-temporally aligning and fusing video quality metrics, audio synchronization, interaction response latency, and network QoS parameters;

[0018] Model input unit: Inputting the fused multi-dimensional features into a pre-trained spatio-temporal attention neural network, where the spatio-temporal attention neural network includes a video quality attention branch and an interaction experience attention branch;

[0019] Metric generation unit: The spatio-temporal attention neural network outputs an evaluation metric set including picture smoothness score, audio-visual synchronization, interaction response index, and comprehensive QoE score after learning;

[0020] Threshold comparison unit: Comparing and analyzing the real-time output evaluation metrics with preset dynamic quality thresholds to generate a quality anomaly feature map;

[0021] Result output unit: Performing data analysis on the interactive quality according to the evaluation metric set, generating an analysis result, and encoding the analysis result into a quality state feature vector recognizable by the deep reinforcement learning network.

[0022] Optionally, the evaluation strategy generation module includes a strategy generation unit: The generator model in the deep reinforcement learning network receives the quality state feature vector and historical optimization records, and generates a set of candidate strategies including bitrate adjustment, frame rate optimization, and transmission protocol selection, where the historical optimization records include the execution effects of bitrate adjustment strategies within the past t0 time;

[0023] Strategy evaluation unit: The discriminator model calculates the expected quality improvement benefits of each candidate strategy in the set of candidate strategies based on the actual experience data reported by the user terminal, where the actual experience data includes the stuttering rate, interaction success rate, and decoding time;

[0024] Optimization decision unit: Using the reinforcement learning decision ε-greedy algorithm to balance exploration and exploitation, and selecting the optimal quality optimization strategy;

[0025] Experience replay unit: Storing the decision-making process in the experience pool of the edge node for online model update;

[0026] Model update unit: Regularly synchronously updating the model parameters between edge nodes to maintain the consistency of the evaluation strategy.

[0027] Optionally, the dynamic adjustment strategy module includes a strategy parsing unit: Decomposing the optimal quality optimization strategy output by the reinforcement learning network into executable parameter adjustment instructions;

[0028] Encoding adjustment unit: Dynamically adjust H.265 / AV1 encoding parameters according to network bandwidth prediction to achieve Pareto optimization of bitrate - image quality;

[0029] Transmission optimization unit: Adopt adaptive multi - path transmission technology and dynamically select the optimal transmission protocol combination according to the network state;

[0030] Resource allocation unit: Through the edge computing resource scheduling algorithm, allocate priority computing resources for key quality index guarantee;

[0031] Fault - tolerance processing unit: When detecting sudden network fluctuations, start the degradation strategy to ensure the basic interaction experience.

[0032] Optionally, the feedback execution module includes a quality monitoring unit: Continuously monitor the changes in actual quality indicators after the execution of the optimal quality optimization strategy;

[0033] Parameter adjustment unit: Dynamically adjust the quantization parameters and GOP structure of the video encoder through a fuzzy PID controller;

[0034] Transmission update unit: Dynamically update the forward error correction strategy and re - transmission mechanism according to the real - time network diagnosis results;

[0035] Resource re - allocation unit: Dynamically adjust the distribution of video processing tasks based on the load status of edge nodes;

[0036] Closed - loop control unit: Feed back the execution effect to the reinforcement learning network to form a continuously optimized control loop.

[0037] Optionally, the spatio - temporal attention neural network includes:

[0038] Spatio - temporal feature extraction and fusion component: In the e - commerce live - streaming scenario, extract spatio - temporal features during the live - streaming process, and use dynamic time warping and affine transformation for data alignment and integration in time and space. Among them, the spatio - temporal features include heterogeneous data such as video frame rate, audio delay, and user like / comment / order hotspots;

[0039] Interactive quality prediction component: By learning the relationship between spatio - temporal features and interactive quality in historical data, predict the real - time interactive quality of e - commerce live - streaming according to the fused spatio - temporal features, and output a quality assessment;

[0040] Anomaly detection component: Identify abnormal situations in high - latency time periods or picture areas during the live - streaming process through the attention mechanism, and take timely measures to ensure the live - streaming quality.

[0041] Optionally, the reinforcement learning model includes:

[0042] Quality Optimization Strategy Exploration and Generation Component: According to the current network status and the live broadcast status of user interaction data, explore the adjustment of new parameters during bandwidth fluctuations to generate quality optimization strategies, so as to dynamically adjust the bit rate, frame rate, and transmission protocol;

[0043] Model Agent Optimization Strategy Component: The environment sets the reward function, based on the Actor-Critic framework, and adopts Proximal Policy Optimization (PPO), and realizes stable update through importance quality index sampling;

[0044] Resource Allocation to Maximize Revenue Component: According to Prioritized Experience Replay (PER) and Edge Node Periodic Twin Delayed Deep Deterministic Policy Gradient (TD3), the agent allocates computing resources and network bandwidth according to the real-time needs of the live broadcast and the resource status of the edge node, so as to maximize the live broadcast interaction quality and resource utilization efficiency, and continuously improves the evaluation accuracy of the spatio-temporal network by using execution feedback.

[0045] A computer device, comprising: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned e-commerce live broadcast real-time interaction quality evaluation system based on edge computing are realized.

[0046] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned e-commerce live broadcast real-time interaction quality evaluation system based on edge computing are realized.

[0047] In the above technical solution, the technical effects and advantages provided by the present invention:

[0048] The present invention collects multi-modal interaction data of the live stream in real time through a distributed edge computing node cluster, the lightweight spatio-temporal attention neural network model performs data fusion processing, and the generated quality optimization scheme of the deep reinforcement learning network, and the adaptive fuzzy inference system corrects the optimization scheme in real time, combines network bandwidth fluctuations and terminal device resource status for dynamic parameter adjustment, achieving the effect of real-time, accurate and comprehensive evaluation of e-commerce live broadcast interaction quality, can dynamically generate and adjust quality optimization schemes according to actual situations, and thus effectively cope with network bandwidth fluctuations and terminal device resource changes, not only improving the fluency, clarity and interactivity of the live broadcast, but also bringing a better viewing experience to users, and at the same time providing a more reliable quality evaluation and optimization basis for live broadcast platforms and merchants, promoting the healthy development of e-commerce live broadcast business. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a block diagram of an e-commerce live broadcast real-time interaction quality evaluation system based on edge computing of the present invention.

[0051] Figure 2 It is a flowchart of an evaluation method for an e-commerce live broadcast real-time interaction quality evaluation system based on edge computing of the present invention. Specific Embodiments

[0052] Now, the exemplary embodiments will be described more comprehensively with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0053] The present invention provides an Figure 1-2 e-commerce live broadcast real-time interaction quality evaluation system based on edge computing as shown, including a multimodal interaction data collection module: through a distributed edge computing node cluster, it collects multimodal interaction data of the live stream in real time and constructs a multi-dimensional quality feature vector, where the multimodal interaction data includes video coding parameters, audio quality indicators, user interaction behavior data, and network transmission status data;

[0054] An interaction quality analysis module: using a lightweight spatio-temporal attention neural network model built in the edge node, it performs real-time fusion analysis on the multimodal interaction data to generate real-time interaction quality evaluation indicators;

[0055] An evaluation strategy generation module: using a generator model of a deep reinforcement learning network to combine the real-time interaction quality evaluation indicators with historical optimization strategies to generate a quality optimization plan, and a discriminator model evaluates the effectiveness of the quality optimization plan based on user terminal feedback information;

[0056] A dynamic adjustment strategy module: it performs real-time correction on the quality optimization plan through an adaptive fuzzy inference system, and dynamically adjusts the parameters of the quality optimization plan generated by the deep reinforcement learning network in combination with the current network bandwidth fluctuation and the terminal device resource status;

[0057] Feedback Execution Module: Based on the quality optimization scheme executed after dynamic parameter adjustment, combined with the distributed decision-making mechanism of edge nodes, it collaboratively optimizes video coding parameters, transmission protocols, and resource allocation strategies to dynamically control the live broadcast quality. The edge computing node then real-time evaluates the quality of the e-commerce live broadcast and feeds back the evaluation results to the live broadcast platform and merchants.

[0058] The working principle of the above technical solution is as follows: First, through the distributed edge computing node cluster, a lightweight data acquisition proxy matrix is deployed at the CDN edge node and the user terminal, which can collect multi-modal interaction data of the live stream in real time with low latency, including video coding parameters, audio quality indicators, user interaction behavior data, and network transmission status data, to ensure that the collected data can effectively describe the interaction quality. At the same time, a timestamp alignment mechanism is used to synchronize the time of multi-dimensional data, and denoising and normalization preprocessing operations are performed to construct a structured interaction quality feature matrix. Then, edge computing resources are used to extract time-domain and space-domain features to generate multi-dimensional quality feature vectors, which are transmitted to the analysis node after being processed by hierarchical compression technology. Then, the spatio-temporal fusion analysis of video fluency, audio-visual synchronization, interaction response latency, and network QoS parameters is input into the pre-trained lightweight spatio-temporal attention neural network model to generate real-time interaction quality evaluation indicators that can evaluate the live broadcast interaction quality. The real-time interaction quality evaluation indicators are compared with the dynamic quality threshold, and a quality anomaly feature map is output and encoded into a quality state feature vector recognizable by reinforcement learning. Secondly, the deep reinforcement learning network generator receives the quality state feature vector, generates a set of candidate optimization strategies including bitrate adjustment, frame rate optimization, and transmission protocol selection. The discriminator calculates the expected benefits of each strategy based on the actual experience data reported by the user terminal, selects the optimal strategy by combining the ε-greedy algorithm, and stores it in the edge experience pool for online learning and updating. Finally, through the adaptive fuzzy inference system, the optimization strategy output by reinforcement learning is parsed into executable parameter adjustment instructions. At the same time, combined with the real-time network bandwidth and terminal resource status, the controller dynamically adjusts the coding parameters and transmission protocols, which can achieve the Pareto optimization of bitrate and image quality. When the network fluctuates, the multi-path transmission fault tolerance mechanism and the edge resource dynamic allocation strategy are started to preferentially guarantee the key indicators for evaluating the interaction quality, and the distribution of video processing tasks is adjusted through the distributed decision-making mechanism, including the dynamic adjustment of the GOP structure by the video encoder, the enabling of forward error correction in the transport layer, and the allocation of computing resources according to priority to achieve the collaborative execution optimization of edge nodes. Continuously monitor the changes in the actual quality indicators after execution and feedback the evaluation effect to the reinforcement learning model.

[0059] The effects of the above technical solutions are as follows: Through the distributed edge node cluster and lightweight data acquisition agents, low-latency acquisition and preprocessing of multi-modal data of live streams are achieved. At the same time, a spatio-temporal attention neural network model is used to complete multi-dimensional feature fusion analysis on the edge side, reducing resource consumption while ensuring evaluation accuracy and improving the response efficiency of data transmission. By combining the deep reinforcement learning network with historical optimization strategies and real-time interactive quality evaluation metrics, the audio-visual synchronization effect can be improved, stuttering can be reduced, and interactive quality can be enhanced. Through the correction of the adaptive fuzzy inference system, dynamic adjustment of bandwidth fluctuations can optimize resource utilization. The distributed resource scheduling algorithm based on edge nodes improves the load balancing degree, and by adjusting the encoding parameters and transmission strategies in real-time feedback, the range of live broadcast quality fluctuations is reduced, and the stability of live broadcast interaction is improved. The lightweight model deployment on the edge side reduces the pressure on the central server, not only reducing bandwidth costs but also supporting real-time quality evaluation of tens of thousands of concurrent live streams, thereby ensuring that live broadcast interaction can support rapid response, optimize user experience, and ensure the reliability, stability, and scalability of evaluation accuracy in high-concurrency, low-latency, and strong-interaction scenarios.

[0060] In a specific embodiment for detailed description, the multi-modal interaction data acquisition module includes:

[0061] Edge deployment unit: Lightweight multi-modal data acquisition arrays are deployed at both the edge nodes of the content delivery network (CDN) and user terminals to form a distributed data acquisition network. Among them, the lightweight multi-modal data acquisition array includes, but is not limited to, comment monitors, like counters, share recorders, and viewing duration trackers.

[0062] Multi-modal synchronization unit: The distributed data acquisition network synchronizes multi-dimensional data such as video frame rate, encoding bit rate, audio delay, user likes, comment frequency, and network jitter rate through a timestamp alignment mechanism to generate raw data that comprehensively reflects the interaction situation in the live broadcast room.

[0063] Preprocessing unit: Filter and denoise the raw data and perform normalization processing to construct a structured interactive quality feature matrix.

[0064] Feature extraction unit: Use edge computing resources to perform real-time feature extraction on the interactive quality feature matrix to generate a multi-dimensional quality feature vector containing time-domain and space-domain features.

[0065] Edge transmission unit: After compressing the multi-dimensional quality feature vector processed by feature extraction using hierarchical compression technology, it is transmitted to the analysis node through the CDN edge network.

[0066] The working principle of the above technical solution is as follows: First, lightweight multi-modal data acquisition arrays are deployed at the CDN edge nodes and user mobile terminals / PC clients. For example, by using a comment monitor to capture user comment content and send timestamps in real time, a like counter to count the number of likes per second and click hot zone coordinates, a sharing recorder to track sharing behavior types and conversion rates, and a viewing duration tracker to record user stay durations and page scrolling behaviors, a distributed acquisition network can be constructed, and thus various interaction data in the live broadcast room can be comprehensively and timely collected. For example, the CDN node covers the push stream data of the anchor end video / audio encoding parameters, and the user terminal collects experience data such as the client-side frame rate and rendering delay. Moreover, the two-way data is interconnected with low latency through the QUIC proprietary tunnel protocol. Then, the distributed data acquisition network adopts a hybrid clock synchronization mechanism. At the hardware level, GPS / PTP clocks are used, and at the software level, the NTP protocol of the user terminal is used for time synchronization calibration. By integrating and aligning multi-dimensional data, it can provide comprehensive consideration of live broadcast quality, user behavior, and network environment for subsequent analysis. Secondly, a Kalman filter is used to eliminate the transient noise of the original data, and median filtering is used to process abnormal signals. Then, the heterogeneous data is mapped to a unified dimension to form a structured interaction quality feature matrix. By edge computing, the data transmission delay is reduced, and the time-domain and space-domain features of the interaction quality feature matrix are extracted to construct a user behavior map to more effectively represent the interaction quality in the live broadcast room. Finally, since the amount of data of the extracted multi-dimensional quality feature vectors is still large, in order to reduce the data transmission pressure and cost, hierarchical compression technology is used to compress to different degrees on the premise of ensuring that the key information of the evaluation data is not lost.

[0067] The effects of the above technical solution are as follows: The distributed data acquisition network covers the CDN edge nodes and user terminals. The multi-modal data acquisition array can collect multi-dimensional data such as live broadcast technical indicators, user interaction behaviors, and network status, thus breaking through the data acquisition dimension, not only comprehensively and accurately reflecting the interaction situation in the live broadcast room, but also providing a rich and reliable data basis for subsequent analysis and evaluation; The application of edge computing reduces the delay of data transmission to the central server. At the same time, the timestamp alignment mechanism ensures the synchronous acquisition of multi-modal data, enabling real-time acquisition and preprocessing of interaction data, improving the quality and usability of the data, contributing to subsequent feature extraction and analysis, and improving the accuracy and reliability of the analysis results; The hierarchical compression technology reduces the data volume without losing key information, reduces the data transmission pressure and cost. The use of the CDN edge network further optimizes the data transmission path, improves the data transmission efficiency, ensures that the feature vectors can be quickly and stably transmitted to the analysis node, and enables timely discovery of problems in the live broadcast room and response.

[0068] In a specific embodiment for detailed description, the interaction quality analysis module includes:

[0069] Data fusion unit: Spatially and temporally aligns and fuses video quality metrics, audio synchronization, interactive response latency, and network QoS parameters;

[0070] Model input unit: Inputs the fused multi-dimensional features into a pre-trained spatio-temporal attention neural network, where the spatio-temporal attention neural network includes a video quality attention branch and an interactive experience attention branch;

[0071] Metric generation unit: The spatio-temporal attention neural network outputs an evaluation metric set including picture smoothness score, audio-visual synchronization, interactive response index, and comprehensive QoE score after learning;

[0072] Threshold comparison unit: Compares and analyzes the real-time output evaluation metrics with preset dynamic quality thresholds to generate a quality anomaly feature map;

[0073] Result output unit: Performs data analysis on the interactive quality according to the evaluation metric set, generates an analysis result, and encodes the analysis result into a quality state feature vector recognizable by a deep reinforcement learning network.

[0074] The working principle of the above technical solution is as follows: First, the dynamic time warping (DTW) method is used to spatially and temporally align and fuse video quality metrics, audio synchronization, interactive response latency, and network QoS parameters. Among them, in the time mapping, microsecond-level alignment is achieved through interpolation based on the PTS / DTS video frame timestamps and the AAC frame audio sampling points, and user interaction events are synchronized with the audio-visual stream through the NTP protocol; as for the spatial mapping, the interactive coordinates of different resolution terminals are normalized to a unified coordinate system; at the same time, the spatio-temporally aligned multi-dimensional feature data is fused by constructing a four-dimensional tensor. Then, using the attention mechanism in the spatio-temporal attention neural network architecture, 3D convolutional attention is used for the video quality branch, and LSTM and self-attention learning processes are used for the interactive experience branch to output evaluation metrics such as picture smoothness score, audio-visual synchronization, interactive response index, and comprehensive QoE score. Among them, the expression of the picture smoothness score is In the formula, S fluency represents the picture smoothness score, FR t represents the frame rate of the t-th time segment, and t and T represent the t-th time segment and the total number of time segments respectively; the expression of the audio-visual synchronization is In the formula, Δ sync represents the audio-visual synchronization, n and N represent the n-th sample and the total number of samples for audio-visual synchronization detection respectively, represents the video timestamp of the n-th sample, represents the audio timestamp of the n-th sample, and | | represents the absolute value symbol; the expression of the interactive response index is Wherein, R response represents the interactive response index, and τ represents the average delay of interactive requests; the expression for the comprehensive QoE score is QoE = ω1·S fluency +ω2·Δ sync +ω3·R response , and ω1 + ω2 + ω3 = 1. Wherein, QoE represents, and ω1, ω2, and ω3 respectively represent the weight coefficients corresponding to the screen smoothness score S fluency , the audio-visual synchronization degree Δ sync , and the interactive response index R response . Secondly, based on the exponential weighted moving average (EWMA) method, the dynamic threshold is calculated, and then compared and analyzed with the evaluation index set output by the spatio-temporal attention neural network. The heat map coding is used to mark the abnormal spatio-temporal regions to generate the quality anomaly feature map. Finally, the multi-dimensional feature data is reduced to a 32-dimensional vector coding, and the state feature vector is output by adding the time series differential features.

[0075] The effects of the above technical solutions are as follows: By aligning and fusing multi-dimensional feature data in time and space through DTW, combined with the training and learning of the spatio-temporal attention neural network, the live interactive quality can be comprehensively and accurately evaluated, the accuracy of audio-visual synchronization is improved during live interaction, and a clear quality portrait is provided for users and operators; The quality anomaly feature map generated by calculating the dynamic threshold through the EWMA method for comparison and analysis can quickly locate the specific aspects where the interactive quality is abnormal, enabling the operator to conduct targeted problem troubleshooting and repair, improving the efficiency of problem solving. At the same time, the setting of the dynamic threshold can be adjusted according to different scenarios and requirements, with strong adaptability and flexibility; The reinforcement learning model can continuously adjust and optimize the parameters and strategies of live interaction according to the multi-dimensional feature vector to improve the interactive quality, realize the adaptive optimization of the system, and improve the real-time performance of evaluating interactive quality and the efficiency of information transmission.

[0076] In a specific embodiment for detailed description, the evaluation strategy generation module includes:

[0077] Strategy generation unit: The generator model in the deep reinforcement learning network receives the quality state feature vector and the historical optimization record, and generates a candidate strategy set including bitrate adjustment, frame rate optimization, and transmission protocol selection. Among them, the historical optimization record includes the execution effect of the bitrate adjustment strategy within the past t0 time;

[0078] Strategy evaluation unit: The discriminator model calculates the expected quality improvement benefits of each candidate strategy in the candidate strategy set based on the actual experience data reported by the user terminal. Among them, the actual experience data includes the stuttering rate, the interactive success rate, and the decoding time;

[0079] Optimized decision-making unit: The reinforcement learning decision-making ε-greedy algorithm is adopted to balance exploration and exploitation, and the optimal quality optimization strategy is selected.

[0080] Experience replay unit: The decision-making process is stored in the experience pool of the edge node for online model update.

[0081] Model update unit: Regularly synchronously update the model parameters between edge nodes to maintain the consistency of the evaluation strategy.

[0082] The working principle of the above technical solution is as follows: First, the recognizable quality state feature vector and the historical optimization record of the execution effect of the bitrate adjustment strategy in the past unit time are used as inputs and transmitted to the generator model. Based on the Actor-Critic framework, by introducing the Lagrange multiplier method to ensure that the strategy meets the bandwidth constraint conditions, the generator model outputs a candidate policy set that can dynamically adjust the H.265 encoding bitrate, adaptively switch the frame rate mode, and intelligently switch between TCP / UDP / QUIC protocols through the policy network Actor. Then, the actual experience data of the user terminal reporting the stuttering rate, interaction success rate, and decoding time consumption is input into the discriminator model, and the double-delay deep deterministic policy gradient TD3 algorithm is used to calculate the expected quality improvement benefits that each candidate policy may bring after implementation, that is, after designing the reward function, the expected benefits of the candidate policies are evaluated through the value function, and then the candidate policies are sorted in descending order according to the expected quality improvement benefits. For example, a certain bitrate adjustment policy may improve the clarity of the picture, but at the same time may also increase the demand for network bandwidth. The discriminator model will evaluate the results of the influence of these factors on the expected quality improvement benefits of the policy according to the reward function. Secondly, the reinforcement learning decision-making ε-greedy algorithm is used to find a balance between exploring new strategies and exploiting known effective strategies. Specifically, a candidate policy is randomly selected with a certain probability ε for exploration to discover potentially better optimization strategies; and a policy with the highest current expected quality improvement benefit is selected with a probability of 1 - ε for exploitation, so that while constantly trying new strategies, the existing experience can also be fully utilized. Among them, a dynamic decay mechanism is set in the ε-greedy algorithm, and finally the optimal quality optimization strategy is selected. When there are parameter conflicts among multiple strategies, the Pareto optimal solution is used for screening. Finally, a ring buffer storage experience pool is deployed at the edge node, and the entire decision-making process of the quality state feature vector, candidate policy set, selected optimal quality optimization strategy, and actual generated effect information is stored in the experience pool. Online model update is triggered according to priority sampling and batch learning to minimize the Bellman error, and the model can learn more effective optimization strategies, and the model parameters between edge nodes are regularly synchronously updated. The Byzantine fault-tolerant algorithm is adopted to filter malicious nodes to ensure the consistency of the evaluation strategies on each edge node and improve its own stability and reliability.

[0083] The effects of the above technical solution are as follows: By combining the deep reinforcement learning generator and discriminator models, multiple candidate strategies can be quickly generated and evaluated based on real-time quality status and historical experience, providing multiple options for optimizing the quality of live interaction. Using the ε-greedy algorithm helps to discover better quality optimization strategies, thereby improving the picture quality, smoothness, and user interaction experience of the live broadcast. The actual experience data reported by the user terminal accurately reflects the real feelings of the users. By evaluating the strategies through the discriminator model, the user-centered experience is realized, ensuring that the selected optimization strategies can truly enhance the live broadcast experience of the users. The experience replay unit and model update unit enable the model to have the ability of online learning and continuous optimization. The historical decision process data in the experience pool can be used for online updating of the model, allowing the model to continuously learn new optimization strategies and adapt to different network environments. Regularly synchronizing and updating the model parameters between edge nodes not only clearly ensures the consistency and stability of the evaluation strategies on different nodes, but also enables the entire evaluation strategy system to continuously evolve and improve. Furthermore, in different live broadcast scenarios and network conditions, it is possible to dynamically adjust the strategies of bit rate, frame rate, and transmission protocol, with adaptability and flexibility, ensuring the overall performance and reliability of the live broadcast quality and user experience.

[0084] In a specific embodiment for detailed description, the dynamic adjustment strategy module includes:

[0085] Policy parsing unit: Decompose the optimal quality optimization strategy output by the reinforcement learning network into executable parameter adjustment instructions;

[0086] Encoding adjustment unit: Dynamically adjust the H.265 / AV1 encoding parameters according to the network bandwidth prediction to achieve the Pareto optimization of bit rate - picture quality;

[0087] Transmission optimization unit: Adopt the adaptive multi-path transmission technology to dynamically select the optimal transmission protocol combination according to the network status;

[0088] Resource allocation unit: Through the edge computing resource scheduling algorithm, allocate priority computing resources for key quality index guarantee;

[0089] Fault tolerance processing unit: When detecting sudden network fluctuations, start the degradation strategy to ensure the basic interaction experience.

[0090] The working principle of the above technical solution is as follows: First, the optimal quality optimization strategy output by the reinforcement learning network is decomposed into executable parameter adjustment instructions, including video encoding parameter adjustment instructions, transmission protocol priority lists, and resource allocation weights. For example, if the optimal quality optimization strategy is to improve the picture quality stability of live interaction while balancing the control of the performance of the client device, the policy parsing unit will dynamically degrade the encoding according to the network bandwidth, preferentially adopt the UDP combined with the FEC transmission protocol and allocate multi-path transmission weights to achieve dynamic balance between picture quality and bitrate while improving transmission reliability and resource utilization efficiency, and then establish a mapping table between the parameter adjustment instructions of the policy and the underlying API. Then, based on the LSTM network, the bandwidth in the next short period of time is predicted, and the quantization step size and coding block size parameters of H.265 / AV1 encoding are dynamically adjusted according to the prediction results, which can find the best balance point between bitrate and picture quality, so as to reduce the bitrate without reducing the picture quality, or improve the picture quality when the bitrate remains unchanged, and achieve the Pareto optimization of finding the optimal solution on the bitrate-picture quality curve. Secondly, by adopting the adaptive multi-path transmission technology, according to the network status of real-time monitoring of packet loss rate, latency, and bandwidth, the optimal transmission protocol combination is dynamically selected. For example, when the network condition is good, a protocol with a fast transmission speed is selected; when the network is unstable, a more fault-tolerant protocol is selected to ensure that data can be transmitted efficiently and stably, and the traffic allocation weights of each path are adjusted in real time based on the Kalman filter to balance the load. At the same time, by prioritizing the key quality indicators of picture smoothness, audio-visual synchronization, and interactive response latency, an improved Best-Fit algorithm is used to allocate computing resources preferentially for the key quality indicators, and a burst task preemption mechanism is adopted. For example, for interactive links with high real-time requirements, more computing resources are immediately preempted to ensure their response speed and stability. Finally, by real-time detecting the network status, when a sudden network fluctuation is detected, the hierarchical degradation strategy is immediately started. Among them, the hierarchy includes turning off the background blurring effect, pausing the frame rate mode and locking the frame rate mode, and the minimalist mode with low resolution, and the degradation includes reducing the picture quality and reducing interactive functions.

[0091] The effects of the above technical solutions are as follows: By dynamically adjusting encoding parameters, optimizing the transmission protocol, and reasonably allocating computing resources, it is possible to maintain a high live broadcast quality in different network environments. At the same time, Pareto optimization of the bitrate-picture quality is achieved, which not only ensures the clarity of the picture but also controls the bitrate, reducing the pressure on network bandwidth. The adaptive transmission protocol selection ensures the stability and efficiency of data transmission, reduces the impact of packet loss and latency, and dynamically adjusts the network state by predicting the network bandwidth to adapt to various complex and changeable network environments. While ensuring the quality of the live broadcast and the user experience, it improves the adaptability and reliability of the live broadcast interaction quality. The edge computing resource scheduling algorithm allocates priority computing resources for key quality indicator guarantee. Even in the case of poor network conditions, users can still perform some basic interaction operations, avoiding the situation where the live broadcast cannot proceed normally due to network problems, enhancing the anti-interference ability of the live broadcast interaction link, and thus ensuring that the live broadcast interaction link can meet the overall user experience.

[0092] In a specific embodiment for detailed description, the feedback execution module includes:

[0093] Quality monitoring unit: Continuously monitor the changes in actual quality indicators after the execution of the optimal quality optimization strategy;

[0094] Parameter adjustment unit: Dynamically adjust the quantization parameters and GOP structure of the video encoder through a fuzzy PID controller;

[0095] Transmission update unit: Dynamically update the forward error correction strategy and retransmission mechanism according to the real-time network diagnosis results;

[0096] Resource reallocation unit: Dynamically adjust the distribution of video processing tasks based on the load status of edge nodes;

[0097] Closed-loop control unit: Feed back the execution effect to the reinforcement learning network to form a continuously optimized control loop.

[0098] The working principle of the above technical solution is as follows: First, the quality indicators of video freezing rate, audio-visual synchronization error, and interactive response delay after implementing the optimal quality optimization strategy are collected in real time. A sliding window is defined to calculate the moving average of the indicators, which serves as the dynamic quality benchmark for anomaly detection. By using a pre-trained isolation forest model to learn and predict the input quality indicators, the output for identifying deviations from the baseline is compared with the dynamic quality benchmark for anomaly detection using the 3σ rule to identify abnormal indicator data points, enabling an intuitive understanding of the actual effect after the implementation of the optimization strategy and providing an optimization basis for subsequent dynamic adjustment. Then, the quantization parameters of the video encoder directly affect the picture quality-bit rate of the live video, while the GOP structure affects the encoding efficiency and playback smoothness of the video. Based on the actual changes in the quality indicators of the quality monitoring unit, the centroid method is used to calculate the precise adjustment amount, and a fuzzy PID controller is used to automatically and flexibly adjust the quantization parameters and GOP structure of the video encoder to improve the picture clarity and video playback smoothness. Secondly, through real-time network diagnosis, the status information of the current network bandwidth, delay, and packet loss rate is obtained. The forward error correction strategy can add redundant information during data transmission, dynamically calculate the redundant packet ratio based on the real-time packet loss rate, and use RaptorQ coding to recover the packet loss rate as much as possible so that a certain number of errors can be corrected at the receiving end; the retransmission mechanism is to retransmit the lost data when data loss is detected. A three-time retransmission strategy is adopted for key I-frames, and a one-time retransmission combined with a discard strategy is adopted for non-key B / P frames, thereby improving the reliability and efficiency of data transmission. Finally, by monitoring the load conditions of edge nodes in real time and using K-means clustering to identify overloaded nodes and construct a load heat map, according to the results of the load heat map, an improved Hungarian algorithm is used to match tasks with nodes, and the video transcoding tasks of overloaded nodes are migrated to idle nodes for processing to dynamically adjust the distribution of video processing tasks, thereby achieving the balanced utilization of resources, avoiding the impact on the overall performance due to overloading of local nodes, and feeding back the effect after the execution of the strategy to the reinforcement learning network, encoding it into a state vector, adding metadata such as the node ID, timestamp, and protocol version of the execution environment, setting quantitative feedback data to trigger online learning of the model, and updating the policy network parameters using the PPO algorithm with adaptive adjustment of the learning rate to adjust and optimize subsequent strategies.

[0099] The effects of the above technical solutions are as follows: According to the changes in the actual quality indicators monitored in real time, the fuzzy PID controller can timely adjust the video coding parameters, transmission strategies or resource allocation according to the changes in the network conditions and quality indicators, ensuring that the live broadcast always maintains a good quality level; Through the forward error correction strategy and retransmission mechanism, it can effectively cope with network fluctuations and data loss problems, improve the reliability of data transmission and enhance the user viewing experience; Dynamically adjust the distribution of video processing tasks according to the load status of edge nodes, avoiding the situation of some nodes being overloaded while some nodes are idle, realizing the balanced utilization of load resources and the task migration efficiency, thus not only improving the overall performance of the system, but also reducing the operation cost; The reinforcement learning network continuously adjusts the strategy according to the feedback information, adapts to different network environments and user needs, improves the adaptability and stability of live broadcast interaction, and provides users with a high-quality and stable live broadcast experience.

[0100] In a specific embodiment for detailed description, the spatio-temporal attention neural network includes:

[0101] Spatio-temporal feature extraction and fusion component: In the e-commerce live broadcast scenario, extract the spatio-temporal features during the live broadcast process, and use dynamic time warping and affine transformation for alignment and integration of data in time and space. Among them, the spatio-temporal features include heterogeneous data such as video frame rate, audio delay, and user like / comment / order hotspots.

[0102] Interactive quality prediction component: By learning the relationship between spatio-temporal features and interactive quality in historical data, predict the real-time interactive quality of the e-commerce live broadcast according to the fused spatio-temporal features, and output a quality assessment.

[0103] Anomaly detection component: Identify abnormal situations in high-latency periods or picture areas during the live broadcast process through the attention mechanism, and take timely measures to ensure the live broadcast quality.

[0104] The working principle of the above technical solution is as follows: First, the heterogeneous data collected by e-commerce live broadcast is cleaned and normalized for preprocessing, and the feature data in the time and space dimensions are extracted through the structure of convolutional layer and loop layer. At the same time, DTW is used to align the time series of video frame rate and audio delay to eliminate the asynchrony of audio and video, and the coordinates of the user click hot zone are mapped to a unified coordinate system to unify the resolution difference of user-side devices, and a four-dimensional tensor is constructed to fuse the spatiotemporal feature data. Then, a large amount of historical data is input into the model for training to learn the mapping relationship between spatiotemporal features and interactive quality. The model is used to make real-time predictions of the picture smoothness, audio and video synchronization error, and interactive response index of the evaluation quality indicators, and the fine-grained indicators are strengthened, so that the model can predict the real-time interactive quality of the current e-commerce live broadcast and output the quality evaluation results. Finally, in the e-commerce live broadcast, the attention mechanism can focus on the high-incidence period or picture area during the live broadcast by calculating the abnormal weight of each time slice, and more keenly capture abnormal situations. After the dynamic threshold is compared and analyzed, corresponding measures are taken in time to adjust the encoding parameters, enhance the regional bit rate ratio, and optimize the transmission strategy.

[0105] The effects of the above technical solution are: by comprehensively extracting spatiotemporal features such as video frame rate, audio delay, user interaction hot spots, and performing fine alignment and integration, it can more accurately reflect the actual situation of e-commerce live broadcast, thereby improving the accuracy of the prediction, allowing live broadcast operators to more clearly understand the effect of the live broadcast and adjust the live broadcast strategy in time; using the attention mechanism to focus on the time periods or screen areas with high incidence of freezes, so that abnormal situations in the live broadcast process can be quickly and accurately identified, and measures can be taken quickly to deal with them, effectively reducing the impact of abnormal situations on the quality of the live broadcast and ensuring the audience's viewing experience; and when abnormal situations are detected, the live broadcast parameters of specific areas are optimized in a targeted manner to avoid waste of resources and improve resource utilization efficiency. At the same time, user participation and satisfaction are improved, and user stickiness is enhanced.

[0106] In a specific embodiment, the reinforcement learning model includes:

[0107] Quality optimization strategy exploration and generation component: Based on the current network status and the live broadcast status of user interaction data, it explores the adjustment of new parameters when bandwidth fluctuates and generates quality optimization strategies to dynamically adjust the bit rate, frame rate, and transmission protocol;

[0108] Model agent optimization strategy component: The environment sets the reward function, based on the Actor-Critic framework, and uses the proximal strategy to optimize the PPO, and achieves stable updates through sampling of importance quality indicators;

[0109] Resource Allocation Maximizing Revenue Component: Based on Prioritized Experience Replay (PER) and Edge Node Regularized Twin Delayed Deep Deterministic Policy Gradient (TD3), the agent allocates computing resources and network bandwidth according to the real-time requirements of the live broadcast and the resource status of the edge nodes to maximize the quality of live broadcast interaction and resource utilization efficiency, and continuously improves the evaluation accuracy of the spatio-temporal network using execution feedback.

[0110] The working principle of the above technical solution is as follows: First, the current network status and user interaction data are obtained in real time to construct a dynamic action space for adjusting transmission protocol selection, frame rate mode switching, bit rate, and FEC redundancy. When bandwidth fluctuations are detected, the deep exploration mode is triggered, and the component attempts to explore a combination of strategies for adjusting new parameters. Based on the exploration results, corresponding quality optimization strategies are generated to dynamically adjust the bit rate, frame rate, and transmission protocol, thereby adapting to changes in the network status and maintaining good live broadcast quality. Then, the agent is optimized based on the Actor-Critic framework and Proximal Policy Optimization (PPO) algorithm. By setting a reward function in the environment, the behavior of the agent is evaluated and feedback is provided according to the quality metrics of the clarity, smoothness, and interaction activity of the live broadcast, enabling the Actor to generate actions for quality optimization strategies and the Critic to evaluate the value of the actions generated by the Actor, thus stably updating the quality metrics using the PPO algorithm. Finally, optimal resource allocation is performed according to the real-time requirements of the live broadcast and the resource status of the edge nodes. First, the PER mechanism is used to replay the agent's experience to assign priorities, and the edge node regularized twin delayed deep deterministic policy gradient algorithm enables the agent to allocate computing resources and network bandwidth according to the real-time requirements of the live broadcast and the resource status of the edge nodes, and continuously improves the evaluation accuracy of the spatio-temporal network using execution feedback information.

[0111] The effects of the above technical solution are as follows: Adjusting the bit rate, frame rate, and transmission protocol according to changes in network status bandwidth fluctuations ensures that the live broadcast can maintain good quality under different network conditions, reduces phenomena such as stuttering and frozen screens, and improves the viewing experience of users; adopting the Actor-Critic framework and PPO algorithm, combined with the feedback of the reward function, can stably and efficiently optimize the agent's strategy, enabling the agent to quickly learn the optimal actions in different live broadcast states and improving the effect of live broadcast quality optimization; and using the PER and TD3 algorithms to allocate resources according to the real-time requirements of the live broadcast and the resource status of the edge nodes realizes the optimal utilization of computing resources and network bandwidth, thus avoiding resource waste and overload and improving the overall performance and resource utilization efficiency of the system.

[0112] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0113] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0114] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0115] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0116] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An e-commerce live streaming real-time interaction quality evaluation system based on edge computing, characterized in that It includes a multi-modal interaction data acquisition module: Through a distributed edge computing node cluster, it collects the multi-modal interaction data of the live stream in real time and constructs a multi-dimensional quality feature vector. Among them, the multi-modal interaction data includes video coding parameters, audio quality indicators, user interaction behavior data, and network transmission status data; An interactive quality analysis module: Using the lightweight spatio-temporal attention neural network model built into the edge node, it performs real-time fusion analysis on the multi-modal interaction data and generates real-time interactive quality evaluation indicators; An evaluation strategy generation module: Adopting the generator model of the deep reinforcement learning network, it combines the real-time interactive quality evaluation indicators with the historical optimization strategy to generate a quality optimization plan. The discriminator model evaluates the effectiveness of the quality optimization plan according to the user terminal feedback information; A dynamic adjustment strategy module: It performs real-time correction on the quality optimization plan through an adaptive fuzzy inference system, and dynamically adjusts the parameters of the quality optimization plan generated by the deep reinforcement learning network in combination with the current network bandwidth fluctuation and terminal device resource status; A feedback execution module: Based on the quality optimization plan executed after dynamic parameter adjustment, combined with the distributed decision-making mechanism of the edge node, it collaboratively optimizes the video coding parameters, transmission protocol, and resource allocation strategy to dynamically control the live broadcast quality. The edge computing node then performs real-time evaluation on the e-commerce live broadcast quality and feeds back the evaluation results to the live broadcast platform and merchants.

2. The real-time interactive quality evaluation system for e-commerce live streaming based on edge computing according to claim 1, wherein The multi-modal interaction data acquisition module includes an edge deployment unit: Lightweight multi-modal data acquisition arrays are deployed at both the content delivery network (CDN) edge nodes and user terminals to form a distributed data acquisition network; A multi-modal synchronization unit: The distributed data acquisition network synchronously collects multi-dimensional data such as video frame rate, coding bit rate, audio delay, user likes, comment frequency, and network jitter rate through a timestamp alignment mechanism to generate a raw data that comprehensively reflects the interaction situation in the live broadcast room; A preprocessing unit: It performs filtering, denoising, and normalization processing on the raw data to construct a structured interactive quality feature matrix; A feature extraction unit: Using edge computing resources, it performs real-time feature extraction on the interactive quality feature matrix to generate a multi-dimensional quality feature vector containing time-domain and spatial-domain features; An edge transmission unit: After compressing the multi-dimensional quality feature vector processed by feature extraction using hierarchical compression technology, it transmits it to the analysis node through the CDN edge network.

3. The e-commerce live streaming real-time interaction quality evaluation system based on edge computing according to claim 2, wherein The interactive quality analysis module includes a data fusion unit: It performs spatio-temporal alignment and fusion on video quality indicators, audio synchronization, interaction response delay, and network QoS parameters; A model input unit: It inputs the fused multi-dimensional features into a pre-trained spatio-temporal attention neural network, where the spatio-temporal attention neural network includes a video quality attention branch and an interaction experience attention branch; An index generation unit: The spatio-temporal attention neural network outputs an evaluation index set including picture smoothness score, audio-visual synchronization, interaction response index, and comprehensive QoE score after learning; A threshold comparison unit: It compares and analyzes the real-time output evaluation indicators with preset dynamic quality thresholds to generate a quality anomaly feature map; Result Output Unit: Perform data analysis on the interaction quality according to the evaluation metric set, generate an analysis result, and encode the analysis result into a quality status feature vector recognizable by the deep reinforcement learning network.

4. The e-commerce live streaming real-time interaction quality evaluation system based on edge computing according to claim 3, wherein, The said evaluation strategy generation module includes a strategy generation unit: The generator model in the deep reinforcement learning network receives the quality status feature vector and the historical optimization record, and generates a set of candidate strategies including bitrate adjustment, frame rate optimization, and transmission protocol selection. Among them, the historical optimization record includes the execution effect of the bitrate adjustment strategy within the past t0 time. Strategy Evaluation Unit: The discriminator model calculates the expected quality improvement benefits of each candidate strategy in the set of candidate strategies based on the actual experience data reported by the user terminal. Among them, the actual experience data includes the stuttering rate, interaction success rate, and decoding time. Optimization Decision Unit: Adopt the reinforcement learning decision ε-greedy algorithm to balance exploration and exploitation, and select the optimal quality optimization strategy. Experience Replay Unit: Store the decision-making process in the experience pool of the edge node for online model update. Model Update Unit: Regularly synchronously update the model parameters between edge nodes to maintain the consistency of the evaluation strategy.

5. The real-time interactive quality evaluation system for e-commerce live streaming based on edge computing according to claim 4, wherein The said dynamic adjustment strategy module includes a strategy parsing unit: Decompose the optimal quality optimization strategy output by the reinforcement learning network into executable parameter adjustment instructions. Encoding Adjustment Unit: Dynamically adjust the H.265 / AV1 encoding parameters according to the network bandwidth prediction to achieve the Pareto optimization of bitrate - image quality. Transmission Optimization Unit: Adopt the adaptive multi-path transmission technology to dynamically select the optimal transmission protocol combination according to the network state. Resource Allocation Unit: Through the edge computing resource scheduling algorithm, allocate priority computing resources for key quality index guarantee. Fault Tolerance Processing Unit: When a sudden network fluctuation is detected, start the degradation strategy to ensure the basic interaction experience.

6. The e-commerce live streaming real-time interaction quality evaluation system based on edge computing according to claim 5, wherein The said feedback execution module includes a quality monitoring unit: Continuously monitor the change of the actual quality index after the execution of the optimal quality optimization strategy. Parameter Adjustment Unit: Dynamically adjust the quantization parameters and GOP structure of the video encoder through a fuzzy PID controller. Transmission Update Unit: Dynamically update the forward error correction strategy and retransmission mechanism according to the real-time network diagnosis result. Resource Reallocation Unit: Dynamically adjust the distribution of video processing tasks based on the load status of the edge node. Closed-loop Control Unit: Feed back the execution effect to the deep reinforcement learning network to form a continuously optimized control loop.

7. The real-time interactive quality evaluation system for e-commerce live streaming based on edge computing according to claim 6, characterized in that, The said spatio-temporal attention neural network includes: Spatio-temporal Feature Extraction and Fusion Component: In the e-commerce live broadcast scenario, extract the spatio-temporal features during the live broadcast process, and use dynamic time warping and affine transformation to align and integrate the data in time and space. Among them, the spatio-temporal features include heterogeneous data such as video frame rate, audio delay, and user like / comment / order hotspots. Interaction Quality Prediction Component: By learning the relationship between spatio-temporal features and interaction quality in historical data, predict the real-time interaction quality of the e-commerce live broadcast according to the fused spatio-temporal features, and output a quality assessment. Abnormal Detection Component: Identify abnormal situations in high-stuttering time periods or picture areas during the live broadcast process through the attention mechanism, and take timely measures to ensure the live broadcast quality.

8. The real-time interactive quality evaluation system for e-commerce live streaming based on edge computing according to claim 7, characterized in that, The reinforcement learning model includes: A quality optimization strategy exploration and generation component: According to the current network state and the live broadcast state of user interaction data, explore the adjustment of new parameters during bandwidth fluctuations, generate a quality optimization strategy to dynamically adjust the bitrate, frame rate, and transmission protocol; A model agent optimization strategy component: The environment sets a reward function, based on the Actor-Critic framework, and adopts proximal policy optimization (ppo), and realizes stable update through importance quality index sampling; A resource allocation maximum revenue component: According to the prioritized experience replay (PER) and the edge node's periodic twin-delayed deep deterministic policy gradient (TD3), enable the agent to allocate computing resources and network bandwidth according to the real-time needs of the live broadcast and the resource status of the edge node, so as to maximize the live broadcast interaction quality and resource utilization efficiency, and continuously improve the evaluation accuracy of the spatio-temporal network using execution feedback.

9. A computer device, comprising: A memory and a processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the edge computing-based e-commerce live broadcast real-time interaction quality evaluation system according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the edge computing-based e-commerce live broadcast real-time interaction quality evaluation system according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent large-screen interactive broadcast control system based on multi-modal fusion

    CN118981296A

  • Private domain live broadcast peak hot spot prediction and content scheduling method based on deep learning

    CN119450099A

  • Television broadcast quality monitoring system based on multi-modal model

    CN119728959A

  • Intelligent video coding optimization system and method based on computer vision

    CN119865608A

  • Edge deployment of cloud-originated machine learning and artificial intelligence workloads

    US12033006B1

Cited By

  • Multifunctional integrated digital network broadcasting system based on intelligent algorithm

    CN120751346A

  • Video conference equipment inspection method, device, equipment, medium and product

    CN120825572A

  • Video conference equipment inspection method, device, equipment, medium and product

    CN120825572B

  • Multi-modal deep learning garment live broadcast real-time conversion rate prediction method and system

    CN120996861A

  • Multi-terminal live broadcast interactive data transmission method and system based on frequency modulation transfer

    CN121000708A