User experience optimization method and system for audio and video transmission system

By obtaining the transmission quality data of the audio and video transmission system, performing feature extraction and reinforcement learning model processing, and generating dynamic transmission optimization strategies, the existing system's lack of audio and video synchronization and real-time performance under complex network conditions is solved, and the user experience and system robustness are improved.

CN120263777AActive Publication Date: 2025-07-04SHENZHEN ZIDOO TECH CO LTD

Patent Information

Application Number
CN202510735372.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

It is difficult for existing audio and video transmission systems to maintain stable audio and video synchronization and real-time under complex network conditions. The existing quality evaluation system cannot fully reflect the complex relationship between user subjective experience and multi-dimensional technical parameters, lacks real-time perception and adaptability, resulting in disconnection of optimization strategies from actual needs, and the user feedback mechanism is vague and inefficient.

Method used

By obtaining the transmission quality data set, the transmission feature extraction is performed, and the pre-trained reinforcement learning optimization model is used to dynamically aggregate the audio-visual synchronization features, dynamic code rate features and decoded delay features to generate transmission quality evaluation results, and based on this, the dynamic transmission optimization strategy is determined and the transmission parameters are adjusted in real time.

Benefits of technology

Dynamic calibration and forward-looking prediction of transmission quality are achieved, which significantly reduces the probability of problems such as audio and video loss, lag and delay, and improves the robustness and user satisfaction of the audio and video transmission system in dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263777A_ABST
    Figure CN120263777A_ABST
Patent Text Reader

Abstract

The invention provides a user experience optimization method and system for an audio and video transmission system, and the method comprises the steps: firstly obtaining a transmission quality data set in the audio and video transmission process of a target terminal, the transmission quality data set comprises a plurality of transmission session sequences, and each transmission session sequence is composed of network fluctuation, coding and decoding parameters, user feedback and other features; performing transmission feature extraction on the data set to obtain audio and video synchronization, dynamic code rate and decoding delay features, calling a pre-trained reinforcement learning optimization model, performing dynamic aggregation processing on the features to generate a transmission quality evaluation result, and finally determining a dynamic transmission optimization strategy based on the transmission quality evaluation result. And the real-time transmission parameter adjustment operation is triggered, so that the user experience of the audio and video transmission system is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video systems, and more particularly, to a method and system for optimizing the user experience of an audio - video transmission system. Background Art

[0002] In the field of audio - video transmission technology, with the popularization of applications such as streaming media services, remote meetings, and real - time communication, users' experience requirements for transmission quality are becoming increasingly stringent. Existing audio - video transmission systems are difficult to maintain stable audio - video synchronization and real - time performance under complex network conditions. In addition, existing quality assessment systems mostly rely on single technical indicators and cannot comprehensively reflect the complex relationship between users' subjective experience and multi - dimensional technical parameters, resulting in the disconnection between optimization strategies and actual needs. Moreover, existing solutions mostly use static threshold judgment or preset rule adjustment, lacking real - time perception and adaptive capabilities for environmental changes and user feedback, and are difficult to cope with sudden quality degradation problems in dynamic scenarios.

[0003] Furthermore, the user feedback mechanisms of related solutions are mostly limited to post - event scoring or simple complaint records, not deeply associated with transmission technical parameters, resulting in fuzzy problem positioning and low optimization efficiency. Summary of the Invention

[0004] In view of the above - mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for optimizing the user experience of an audio - video transmission system, the method comprising: Obtaining a set of transmission quality data generated during the audio - video transmission of a target terminal, the set of transmission quality data including a plurality of transmission session sequences, and each transmission session sequence consisting of network fluctuation characteristics, codec parameter characteristics, and user feedback characteristics of at least one transmission period; Performing transmission feature extraction processing on the set of transmission quality data to obtain audio - video synchronization characteristics, dynamic bit - rate characteristics, and decoding delay characteristics of each transmission session sequence; Invoking a pre - trained reinforcement learning optimization model to perform dynamic aggregation processing on the audio - video synchronization characteristics, dynamic bit - rate characteristics, and decoding delay characteristics to generate a transmission quality assessment result of the transmission session sequence; Determining a dynamic transmission optimization strategy based on the transmission quality assessment result, and feeding back the dynamic transmission optimization strategy to the audio - video transmission system to trigger real - time transmission parameter adjustment operations.

[0005] In another aspect, embodiments of the present invention further provide a user experience optimization system for an audio - video transmission system, including a processor and a machine - readable storage medium, the machine - readable storage medium being connected to the processor, the machine - readable storage medium being used to store programs, instructions, or codes, and the processor being used to execute the programs, instructions, or codes in the machine - readable storage medium to implement the above - mentioned method.

[0006] Based on the above aspects, the present invention systematically integrates multi-dimensional heterogeneous data such as network fluctuations, codec parameters, and user feedback, enabling the transmission quality analysis to comprehensively reflect the dynamic coupling relationship of audio-visual synchronization, bitrate adaptability, and decoding real-time performance during the audio-visual transmission process. In the transmission feature extraction stage, by analyzing core features such as audio-visual synchronization deviation, dynamic bitrate fluctuation, and decoding delay distribution, not only the potential causes of transmission quality degradation are revealed, but also a quantitative mapping model between user perception and underlying technical parameters is established, significantly improving the accuracy of problem diagnosis. The dynamic aggregation processing mechanism based on the pre-trained reinforcement learning model effectively overcomes the static adaptation defects of traditional rule-driven methods. By online learning the complex non-linear relationship between transmission features and user experience, dynamic calibration and forward-looking prediction of transmission quality assessment are achieved. The finally generated dynamic transmission optimization strategy not only significantly reduces the occurrence probability of typical quality problems such as audio-visual desynchronization, stuttering, and delay by real-time adjusting key parameters such as bitrate control, buffer management, and decoding priority, but also achieves the Pareto optimum of bandwidth utilization and user experience quality in complex network environments through an adaptive parameter adjustment mechanism. Thus, through the deep integration of data intelligence and real-time control, the robustness and user satisfaction of the audio-visual transmission system in dynamic scenarios are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 FIG. is a schematic flowchart of an execution process of a user experience optimization method for an audio-visual transmission system provided by an embodiment of the present invention.

[0008] Figure 2 FIG. is a schematic diagram of exemplary hardware and software components of a user experience optimization system for an audio-visual transmission system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 FIG. is a flowchart of a user experience optimization method for an audio-visual transmission system provided by an embodiment of the present invention. The user experience optimization method for the audio-visual transmission system will be introduced in detail below.

[0010] Step S110: Obtain a set of transmission quality data generated during the audio-visual transmission of a target terminal, where the set of transmission quality data includes a plurality of transmission session sequences, and each transmission session sequence is composed of network fluctuation features, codec parameter features, and user feedback features of at least one transmission period.

[0011] In this embodiment, taking the user terminal of a certain video playback platform as an example, this video playback platform has a large number of users, and there are a huge amount of audio and video transmission activities every day. Assuming that the target terminal concerned in this embodiment is the device of user A, then during the process of user A watching a popular TV drama, the transmission quality data is collected. Among them, the transmission session sequence is divided at a certain time interval, for example, each 5 minutes is a transmission session sequence.

[0012] Regarding the network fluctuation characteristics in each transmission session sequence, relevant data can be collected. For example, in a certain transmission session sequence, the network fluctuation characteristics include the packet loss rate sequence, the jitter rate sequence, and the bandwidth change rate sequence. Assuming that the packet loss rate sequence shows [0.01, 0.02, 0.03, 0.02, 0.01] during this period, which represents the packet loss rate at different moments during this period; the jitter rate sequence is [5, 8, 6, 7, 5], with the unit of millisecond, reflecting the degree of network jitter; the bandwidth change rate sequence is [0.1, 0.2, -0.1, 0.1, 0.05], reflecting the change of the bandwidth during this period.

[0013] In terms of the codec parameter characteristics, it can include the initial bitrate setting value, the dynamic bitrate adjustment threshold, and the buffer capacity parameter. Assuming that the initial bitrate setting value is 2000 kbps, the dynamic bitrate adjustment threshold is ±200 kbps, and the buffer capacity parameter is 500 KB. This means that the system initially sets the bitrate to 2000 kbps, and can adjust the bitrate between 1800 kbps and 2200 kbps during network fluctuations, and at the same time, the terminal device has a 500 KB buffer to process data.

[0014] The user feedback characteristics are collected through the operations and evaluations of the user during the viewing process. For example, user A gives feedback during the viewing process, and thus the user feedback text is obtained. The user feedback text contains information such as the user's evaluation of the audio-visual synchronization situation and the feeling of latency, which is used to extract the subjective scoring characteristics of the user's audio-visual synchronization and the sensitivity characteristics to latency.

[0015] Step S120: Perform transmission feature extraction processing on the transmission quality data set to obtain the audio-visual synchronization feature, the dynamic bitrate feature, and the decoding latency feature of each transmission session sequence.

[0016] In this embodiment, step S120 may include: Step S121: Perform time series segmentation processing on the network fluctuation characteristics in the transmission session sequence to obtain multiple network fluctuation sub-segments.

[0017] In this scenario, for the network fluctuation characteristics mentioned above, each transmission session sequence can be further processed by time series segmentation. For example, a 5-minute transmission session sequence can be divided into segments every minute, resulting in 5 network fluctuation sub-segments. Taking the first network fluctuation sub-segment as an example, it contains a packet loss rate sequence [0.01, 0.02], a jitter rate sequence [5, 8], and a bandwidth change rate sequence [0.1, 0.2] within that minute, which helps to analyze the characteristics of network fluctuations in different time periods more meticulously.

[0018] Step S122: Invoke the pre-trained encoder model to encode each network fluctuation sub-segment and generate a fluctuation intensity vector for each network fluctuation sub-segment.

[0019] In this embodiment, step S122 may include: Step S1221: Obtain the packet loss rate sequence, jitter rate sequence, and bandwidth change rate sequence included in each network fluctuation sub-segment.

[0020] For example, taking the first network fluctuation sub-segment as an example, the obtained packet loss rate sequence is [0.01, 0.02], the jitter rate sequence is [5, 8], and the bandwidth change rate sequence is [0.1, 0.2].

[0021] Step S1222: Perform normalization processing on the packet loss rate sequence to obtain the packet loss rate distribution feature.

[0022] In this step, the purpose of performing normalization processing on the packet loss rate sequence is to map the data to a specific range for better comparison and analysis. Assuming that the normalization method used is to divide each value in the packet loss rate sequence by the maximum value in the packet loss rate sequence, for the packet loss rate sequence [0.01, 0.02], the maximum value is 0.02. After normalization processing, the obtained packet loss rate distribution feature is [0.5, 1], indicating the relative distribution of the packet loss rate within this network fluctuation sub-segment.

[0023] Step S1223: Perform sliding window statistical processing on the jitter rate sequence to generate a jitter peak-valley difference feature and a jitter duration feature.

[0024] In this embodiment, for the jitter rate sequence [5, 8], sliding window statistical processing is adopted. Assuming that the sliding window size is 2, that is, two consecutive jitter rate values are considered each time. In this jitter rate sequence, the first window is [5, 8], and the calculated peak-valley difference is 8 - 5 = 3, which is the jitter peak-valley difference feature within this sliding window. The jitter duration feature is determined according to the number of window movements and the window size. For example, since the window size is 2 and it moves 1 time, the jitter duration feature is 2 time units (assuming each time point represents 1 time unit here).

[0025] Step S1224: Calculate the gradient of the bandwidth change rate sequence to obtain the bandwidth change trend feature.

[0026] In this embodiment, for the bandwidth change rate sequence [0.1, 0.2], calculate the gradient. The gradient reflects the change trend of the data. The calculation method is to subtract the previous value from the next value, that is, 0.2 - 0.1 = 0.1. So the bandwidth change trend feature is 0.1, indicating that the bandwidth shows an upward trend within this network fluctuation sub-segment.

[0027] Step S1225: Input the packet loss rate distribution feature, jitter peak-valley difference feature, jitter duration feature, and bandwidth change trend feature into the encoder model, and generate the fluctuation intensity vector through multi-layer non-linear transformation.

[0028] In this embodiment, the packet loss rate distribution feature [0.5, 1], jitter peak-valley difference feature 3, jitter duration feature 2, and bandwidth change trend feature 0.1 obtained previously can be input into the pre-trained encoder model. The encoder model processes the above features through multi-layer non-linear transformation. For example, the first layer of the encoder model may perform a weighted sum on the packet loss rate distribution feature. Assuming the weights are 0.3 and 0.7 respectively, the calculation result is 0.5×0.3 + 1×0.7 = 0.85. The second layer may perform a certain non-linear combination of this result with the jitter peak-valley difference feature, such as through a non-linear function f(x, y) = x*y + 1, and the calculation result is f(0.85, 3) = 0.85×3 + 1 = 3.55. The subsequent layers continue to process the data in a similar manner, and finally generate the fluctuation intensity vector. Assuming that after multi-layer processing, the obtained fluctuation intensity vector is [2.5, 3.0, 1.5], this fluctuation intensity vector comprehensively reflects the fluctuation intensity of this network fluctuation sub-segment.

[0029] Step S123: Determine the code rate adaptation deviation value for each transmission period based on the calculation of the correlation between the fluctuation intensity vector and the codec parameter feature.

[0030] In this embodiment, step S123 may include: Step S1231: Extract the initial code rate setting value, dynamic code rate adjustment threshold, and buffer capacity parameter in the codec parameter feature.

[0031] For example, recalling the codec parameter feature mentioned above, the initial code rate setting value is 2000 kbps, the dynamic code rate adjustment threshold is ±200 kbps, and the buffer capacity parameter is 500 KB.

[0032] Step S1232: Calculate the cosine similarity between the fluctuation intensity vector and the dynamic code rate adjustment threshold to obtain the influence coefficient of network fluctuation on the code rate.

[0033] For example, assume that the fluctuation intensity vector is [2.5, 3.0, 1.5], and the vector corresponding to the dynamic code rate adjustment threshold can be assumed to be [1, 1, 1] (this is just a simplified assumption for calculating the cosine similarity, and there may be more complex representations in practice). The calculation method of the cosine similarity is the dot product of the two vectors divided by the product of their magnitudes. Specifically, first calculate the dot product: 2.5×1 + 3.0×1 + 1.5×1 = 7. Then calculate the magnitude of the fluctuation intensity vector: √(2.5² + 3.0² + 1.5²) = √(6.25 + 9 + 2.25) = √17.5 ≈ 4.18. The magnitude of the dynamic code rate adjustment threshold vector is √(1² + 1² + 1²) = √3 ≈ 1.73. Then the cosine similarity is 7 / (4.18×1.73) ≈ 0.96. This value is the influence coefficient of network fluctuation on the code rate, indicating a relatively high degree of correlation between network fluctuation and code rate adjustment.

[0034] Step S1233: Determine the delay tolerance of code rate adjustment according to the buffer capacity parameter, and generate a code rate adaptation weight in combination with the influence coefficient.

[0035] For example, it is known that the buffer capacity parameter is 500KB. Assume that according to experience or pre-set rules, the code rate adjustment delay tolerance corresponding to a 500KB buffer is 100 kbps / min (indicating the maximum allowable value of code rate adjustment per minute). Combine the previously obtained influence coefficient of 0.96 to generate a code rate adaptation weight. Assume that the calculation method of the weight is the influence coefficient multiplied by the delay tolerance, that is, 0.96×100 = 96. This code rate adaptation weight will be used for subsequent calculation of the theoretical code rate adaptation value.

[0036] Step S1234: Determine the theoretical code rate adaptation value based on the product of the initial code rate setting value and the code rate adaptation weight.

[0037] For example, the initial code rate setting value is 2000 kbps, and the code rate adaptation weight is 96. The calculation method of the theoretical code rate adaptation value is the initial code rate setting value multiplied by the code rate adaptation weight, that is, 2000×96 = 192000 kbps² (the unit here is only for explaining the calculation relationship, and the actual meaning needs to be understood in combination with the specific scenario).

[0038] Step S1235: Calculate the code rate adaptation deviation value according to the difference between the theoretical code rate adaptation value and the actual code rate change value.

[0039] For example, assume that the actual bitrate change value is 1800 kbps during this period. Then the bitrate adaptation deviation value is the theoretical bitrate adaptation value minus the actual bitrate change value, that is, 192000 - 1800 = 190200 kbps² (the unit here is only for explaining the calculation relationship, and the actual meaning needs to be understood in combination with specific scenarios). This bitrate adaptation deviation value reflects the difference between the theoretical bitrate and the actual bitrate.

[0040] Step S124: Perform semantic parsing processing on the user feedback features to extract the subjective scoring features of the user's perception of audio-visual synchronization and the sensitivity features to latency.

[0041] In this embodiment, step S124 may include: Step S1241: Obtain the set of keywords included in the user feedback text, where the set of keywords includes the first type of keywords related to audio-visual synchronization and the second type of keywords related to latency.

[0042] In this embodiment, in the feedback text of user A, a set of keywords is extracted. For example, the first type of keywords related to audio-visual synchronization are "audio-visual out of sync", "synchronization problem", etc.; the second type of keywords related to latency are "obvious latency", "lag", etc. The above keywords will be used as key information for subsequent analysis.

[0043] Step S1242: Perform sentiment polarity analysis on the first type of keywords to determine the user's satisfaction score for audio-visual synchronization, and map the satisfaction score to a numerical subjective scoring feature.

[0044] For example, perform sentiment polarity analysis on the first type of keywords "audio-visual out of sync" and "synchronization problem". Assume that the sentiment analysis algorithm used determines that "audio-visual out of sync" is a negative sentiment, and "synchronization problem" is also a negative sentiment. Therefore, according to the preset scoring rules, the satisfaction score corresponding to the negative sentiment is 3 points (out of 10), and then map this satisfaction score to a numerical subjective scoring feature. Assume the mapping rule is to divide the score by 10, and the obtained subjective scoring feature is 0.3.

[0045] Step S1243: Perform frequency statistics processing on the second type of keywords to determine the occurrence frequency and context association strength of the latency-related keywords.

[0046] For example, the occurrence frequencies of the second type of keywords, such as "obvious delay" and "lag", can be counted. Suppose "obvious delay" appears 2 times and "lag" appears 3 times, and the total occurrence frequency is 5 times. For the context association strength, it is determined by analyzing factors such as the position of the keyword in the text and its collocation with other words. For example, "obvious delay" appears in a position describing a critical moment during video playback, so the context association strength is relatively high, set to 0.8; "lag" appears in a relatively less critical description position, and the context association strength is 0.6.

[0047] Step S1244: Generate the sensitivity feature based on the weighted sum of the occurrence frequency and the context association strength.

[0048] For example, assume that the weight of the occurrence frequency is 0.6 and the weight of the context association strength is 0.4. Then the calculation method of the sensitivity feature is the occurrence frequency multiplied by its weight plus the context association strength multiplied by its weight, that is, 5×0.6+(0.8×2+0.6×3)×0.4 = 3+(1.6+1.8)×0.4 = 3+3.4×0.4 = 3+1.36 = 4.36. This sensitivity feature reflects the user's sensitivity to delay.

[0049] Step S125: Perform feature fusion on the bitrate adaptation deviation value, the subjective scoring feature, and the sensitivity feature to generate the audio-visual synchronization feature, the dynamic bitrate feature, and the decoding delay feature.

[0050] In this step, the bitrate adaptation deviation value of 190200 kbps² obtained previously (the unit here is only for explaining the calculation relationship, and the actual meaning needs to be understood in combination with the specific scenario), the subjective scoring feature of 0.3, and the sensitivity feature of 4.36 can be fused. Assume that the fusion method is to splice the above features according to a certain weight. For example, the weight of the bitrate adaptation deviation value is 0.5, the weight of the subjective scoring feature is 0.2, and the weight of the sensitivity feature is 0.3. First, normalize the bitrate adaptation deviation value. Assume it is normalized to the range of 0 - 1, and the result is 0.9 (the specific normalization method is determined according to the actual situation). Then splice according to the weights to obtain the audio - video synchronization feature as [0.9×0.5, 0.3×0.2, 4.36×0.3]=[0.45, 0.06, 1.308]; the dynamic bitrate feature can be obtained by further processing the bitrate adaptation deviation value. For example, determine the trend and amplitude of the dynamic bitrate adjustment according to the size and direction of the bitrate adaptation deviation value. Assume the dynamic bitrate feature is [0.6, 0.4], indicating that there is a 60% probability of needing to increase the bitrate and a 40% probability of needing to decrease the bitrate; the decoding delay feature is generated by combining the sensitivity feature and other relevant factors. Assume the decoding delay feature is [0.7, 0.3], indicating that there is a 70% probability of needing to optimize the decoding delay and a 30% probability of maintaining the current state.

[0051] Step S130: Invoke the pre - trained reinforcement learning optimization model to perform dynamic aggregation processing on the audio - video synchronization feature, the dynamic bitrate feature, and the decoding delay feature, and generate the transmission quality assessment result of the transmission session sequence.

[0052] In this embodiment, step S130 may include: Step S131: Input the audio - video synchronization feature into the first policy sub - network to generate a first candidate policy set for audio - video synchronization optimization.

[0053] For example, the previously generated audio-visual synchronization feature [0.45, 0.06, 1.308] can be input into the first policy sub-network. The first policy sub-network is specifically designed and optimized for audio-visual synchronization. For example, the first policy sub-network includes multiple neuron layers. After inputting the audio-visual synchronization feature, the first layer of neurons performs weighted summation on the feature. Assuming the weights are 0.2, 0.3, and 0.5 respectively, the calculation is 0.45×0.2 + 0.06×0.3 + 1.308×0.5 = 0.09 + 0.018 + 0.654 = 0.762. The second layer of neurons transforms this result through a non-linear function, such as using the ReLU function, to obtain max(0, 0.762) = 0.762. The subsequent layers continue to process the data, and finally generate a series of policies optimized for audio-visual synchronization, forming the first candidate policy set. Assume that the first candidate policy set includes Policy P1: Adjust the video frame rate to optimize audio-visual synchronization; Policy P2: Increase the audio delay to match the video.

[0054] Step S132: Input the dynamic bitrate feature into the second policy sub-network to generate a second candidate policy set optimized for bitrate adaptation.

[0055] For example, the dynamic bitrate feature [0.6, 0.4] can be input into the second policy sub-network. The second policy sub-network is optimized for bitrate adaptation. Similarly, it also has a multi-layer structure. After inputting the feature, the first layer of neurons performs weighted processing on the feature. Assuming the weights are 0.4 and 0.6 respectively, the calculation is 0.6×0.4 + 0.4×0.6 = 0.24 + 0.24 = 0.48. After non-linear transformation and processing by subsequent layers, policies optimized for bitrate adaptation are generated. For example, the second candidate policy set includes Policy Q1: Increase the bitrate to 2200 kbps; Policy Q2: Decrease the bitrate to 1800 kbps.

[0056] Step S133: Input the decoding delay feature into the third policy sub-network to generate a third candidate policy set optimized for delay.

[0057] For example, the decoding delay feature [0.7, 0.3] can be input into the third policy sub-network. The third policy sub-network focuses on delay optimization. After inputting the decoding delay feature, the first layer of the third policy sub-network calculates the feature. Assuming the weights are 0.5 and 0.5 respectively, the calculation is 0.7×0.5 + 0.3×0.5 = 0.35 + 0.15 = 0.5. After multi-layer processing, the third candidate policy set is generated. For example, the third candidate policy set includes Policy R1: Increase the number of decoding threads; Policy R2: Optimize the decoding algorithm to reduce the delay.

[0058] Step S134: Perform policy conflict detection processing on the first candidate policy set, the second candidate policy set, and the third candidate policy set, eliminate conflicting candidate policies, and obtain the remaining candidate policies.

[0059] In this embodiment, step S134 may include: Step S1341: Extract the set of transmission parameter adjustment instructions included in each candidate policy in the first candidate policy set, the second candidate policy set, and the third candidate policy set.

[0060] For example, for policy Q1 in the second candidate policy set, the set of transmission parameter adjustment instructions is {Increase the bitrate to 2200 kbps}; the instruction set for policy Q2 is {Decrease the bitrate to 1800 kbps}. For policy R1 in the third candidate policy set, the set of transmission parameter adjustment instructions is {Increase the number of decoding threads}; the instruction set for policy R2 is {Optimize the decoding algorithm to reduce latency}.

[0061] Step S1342: Detect whether there is a conflict combination in which a bitrate adjustment instruction and a buffer expansion instruction exist simultaneously in the set of transmission parameter adjustment instructions.

[0062] For example, in the above candidate policies, there is no explicit buffer expansion instruction currently, so there is no such conflict combination for the time being. However, in other actual scenarios or more policy sets, if there is a similar situation, such as a policy requiring an increase in bitrate while also expanding the buffer capacity, further analysis is needed. Suppose in a certain scenario, a policy set contains two instructions: "Increase the bitrate to 2500 kbps" and "Expand the buffer capacity from 500 KB to 800 KB", which belongs to this kind of conflict combination. Because increasing the bitrate may increase the data transmission volume, while expanding the buffer simultaneously may lead to conflicts in resource allocation, requiring more complex resource management and scheduling.

[0063] Step S1343: If so, calculate the failure probability of the conflict combination in the historical transmission data. If the failure probability exceeds the preset threshold, eliminate the corresponding candidate policy.

[0064] Suppose in the historical transmission data, it is statistically found that when the bitrate increase and buffer expansion instructions are executed simultaneously, transmission instability, stuttering, or even interruption and other failure situations occur 60 times out of 100 transmissions. Then the failure probability is 60%. If the preset threshold is 50%, since 60% exceeds the preset threshold, the candidate policy containing this conflict combination needs to be eliminated. This is because a high failure probability means that this policy combination has poor performance in actual applications and may seriously affect the user experience.

[0065] Step S1344: Detect whether there is an instruction combination that simultaneously increases the number of decoding threads and reduces the bit rate. If it exists and the current terminal hardware performance parameters do not meet the parallel processing requirements, the corresponding candidate strategy is excluded.

[0066] For example, in the current candidate strategy set, there is a potential conflict combination of strategy Q2 (reducing the bit rate to 1800 kbps) and strategy R1 (increasing the number of decoding threads). Assume that the current terminal hardware performance parameters are: the number of CPU cores is 4, and the memory bandwidth is 8 GB / s. According to experience, when increasing the number of decoding threads and reducing the bit rate simultaneously, at least 6 CPU cores and at least 10 GB / s of memory bandwidth are required to meet the parallel processing requirements. Since the current terminal hardware performance does not meet the requirements, if such an instruction combination appears, it may lead to insufficient system resources and unable to effectively process audio and video data, thus affecting the transmission quality and user experience. Therefore, in this case, the candidate strategy containing this instruction combination needs to be excluded. After the above strategy conflict detection and processing, the remaining candidate strategy set is obtained.

[0067] Step S135: Based on the strategy priority scores of the remaining candidate strategies, perform weighted fusion to generate a comprehensive transmission optimization strategy, and generate the transmission quality evaluation result according to the predicted value of the effect of the comprehensive transmission optimization strategy in the simulated transmission environment.

[0068] In this embodiment, step S135 may include: Step S1351: Obtain the execution success rate, execution efficiency score, and user satisfaction score of each remaining candidate strategy in the historical transmission data.

[0069] Assume that after conflict detection, the remaining candidate strategies are strategy P1 (adjusting the video frame rate to optimize audio-visual synchronization), strategy Q1 (increasing the bit rate to 2200 kbps), and strategy R2 (optimizing the decoding algorithm to reduce latency). Through the analysis of historical transmission data, the execution success rate of strategy P1 is 70%, which means that in the past 100 transmissions applying this strategy, 70 times successfully optimized the audio-visual synchronization; the execution efficiency score is 6 points (out of 10), and this execution efficiency score is comprehensively obtained based on factors such as the time and resource consumption required for the strategy execution; the user satisfaction score is 7 points, reflecting the user experience after applying this strategy. The execution success rate of strategy Q1 is 60%, the execution efficiency score is 5 points, and the user satisfaction score is 6 points. The execution success rate of strategy R2 is 75%, the execution efficiency score is 7 points, and the user satisfaction score is 8 points.

[0070] Step S1352: Determine the strategy effectiveness coefficient according to the product of the execution success rate and the execution efficiency score.

[0071] For example, for strategy P1, the strategy effectiveness coefficient = 70% × 6 = 4.2. The strategy effectiveness coefficient of strategy Q1 = 60% × 5 = 3. The strategy effectiveness coefficient of strategy R2 = 75% × 7 = 5.25. The above-mentioned strategy effectiveness coefficients reflect the comprehensive performance of each strategy in terms of execution success and execution efficiency.

[0072] Step S1353: Determine the strategy priority weight according to the ratio of the user satisfaction score to the preset satisfaction threshold.

[0073] For example, assume that the preset satisfaction threshold is 5 points. For strategy P1, the strategy priority weight = 7 ÷ 5 = 1.4. The strategy priority weight of strategy Q1 = 6 ÷ 5 = 1.2. The strategy priority weight of strategy R2 = 8 ÷ 5 = 1.6. This strategy priority weight reflects the importance of the user's satisfaction with each strategy in the comprehensive evaluation.

[0074] Step S1354: Perform weighted summation of the strategy effectiveness coefficient and the strategy priority weight to obtain the comprehensive priority score of each candidate strategy.

[0075] For example, assume that the weight of the strategy effectiveness coefficient is 0.6 and the weight of the strategy priority weight is 0.4. For strategy P1, the comprehensive priority score = 4.2 × 0.6 + 1.4 × 0.4 = 2.52 + 0.56 = 3.08. The comprehensive priority score of strategy Q1 = 3 × 0.6 + 1.2 × 0.4 = 1.8 + 0.48 = 2.28. The comprehensive priority score of strategy R2 = 5.25 × 0.6 + 1.6 × 0.4 = 3.15 + 0.64 = 3.79.

[0076] Step S1355: Select the top N candidate strategies with the highest comprehensive priority scores, and de-duplicate and merge the transmission parameter adjustment instructions in the top N candidate strategies to generate the comprehensive transmission optimization strategy.

[0077] For example, assume N = 2. The two strategies with the highest comprehensive priority scores are strategy R2 and strategy P1. The transmission parameter adjustment instruction of strategy R2 is {Optimize the decoding algorithm to reduce latency}, and the transmission parameter adjustment instruction of strategy P1 is {Adjust the video frame rate to optimize audio-visual synchronization}. After de-duplicating and merging the above instructions, the comprehensive transmission optimization strategy obtained is: Simultaneously execute optimizing the decoding algorithm to reduce latency and adjusting the video frame rate to optimize audio-visual synchronization.

[0078] Step S1356: Generate the transmission quality assessment result according to the predicted value of the effect of the comprehensive transmission optimization strategy in the simulated transmission environment.

[0079] In this embodiment, in a simulated transmission environment, the comprehensive transmission optimization strategy is tested. It is assumed that the settings of the simulated transmission environment are similar to those of the actual transmission environment, including parameters such as network bandwidth, packet loss rate, and terminal hardware performance. After multiple simulation tests, it is found that after applying this comprehensive transmission optimization strategy, the audio-visual synchronization problem has been significantly improved, the video fluency has been increased, and the satisfaction of user feedback has also been improved. Based on the above simulation test results, it is predicted that good results can also be achieved in actual transmission. Based on the above predicted effect values, a transmission quality assessment result is generated. For example, if it is found in the simulation test that the audio-visual synchronization error is reduced from the original average of 50 milliseconds to 20 milliseconds, and the number of video freezes is reduced from 10 times per hour to 3 times, the transmission quality assessment result can be rated as "good" according to the above specific data and the preset assessment criteria. This transmission quality assessment result reflects the expected transmission quality situation of the current transmission session sequence after applying the comprehensive transmission optimization strategy.

[0080] Step S140: Determine a dynamic transmission optimization strategy based on the transmission quality assessment result, and feedback the dynamic transmission optimization strategy to the audio-visual transmission system to trigger a real-time transmission parameter adjustment operation.

[0081] In this embodiment, step S140 may include: Step S141: Divide the transmission quality assessment result into multiple quality levels, where each quality level corresponds to a different optimization strategy template.

[0082] Suppose the transmission quality assessment result is divided into four levels: "excellent", "good", "average", and "poor". The optimization strategy template corresponding to the "excellent" level focuses on maintaining the current transmission state and making minor optimization adjustments to maintain high-quality transmission. For example, the bitrate may be slightly adjusted to adapt to minor network fluctuations, while maintaining the efficient scheduling of decoding threads. The optimization strategy template corresponding to the "good" level aims to further improve the transmission quality, and may optimize some encoding parameters of the video or add some additional buffering mechanisms. The optimization strategy template corresponding to the "average" level requires a relatively large adjustment of transmission parameters, such as adjusting the bitrate range and optimizing the decoding algorithm. The optimization strategy template corresponding to the "poor" level may involve more radical measures such as reselecting the transmission path and replacing the codec scheme.

[0083] Step S142: Detect the target quality level to which the transmission quality assessment result belongs, and match the optimization strategy template corresponding to the target quality level from the predefined policy library.

[0084] For example, since the previously generated transmission quality assessment result is at the "good" level, the corresponding optimization policy template can be searched for in the predefined policy library. The policy library stores various optimization policy templates for different quality levels, and the above optimization policy templates are obtained through a large number of experiments and data analyses. After finding the optimization policy template corresponding to the "good" level, the optimization policy template may include some basic adjustment instructions, such as appropriately increasing the bit rate to enhance video quality, optimizing the buffer management policy to reduce data loss, etc.

[0085] Step S143: Dynamically adjust the parameter thresholds in the optimization policy template according to the network fluctuation characteristics of the current transmission session sequence to generate a dynamic transmission optimization policy adapted to the current network state.

[0086] In this embodiment, step S143 may include: Step S1431: Extract the bandwidth change rate gradient sequence, packet loss rate distribution trend, and jitter peak-valley interval duration in the network fluctuation characteristics of the current transmission session sequence.

[0087] For example, reviewing the previously obtained network fluctuation characteristics, the bandwidth change rate gradient sequence is assumed to be [0.05, 0.1, -0.05, 0.03], indicating the bandwidth change rate gradient conditions at different time periods in the transmission session sequence; the packet loss rate distribution trend shows a gradually increasing trend, from 0.01 at the beginning to 0.03 later; the average jitter peak-valley interval duration is 8 milliseconds.

[0088] Step S1432: Calculate the bandwidth stability index based on the bandwidth change rate gradient sequence, and determine the packet loss risk level according to the coincidence degree between the packet loss rate distribution trend and the preset packet loss tolerance curve.

[0089] In this embodiment, to calculate the bandwidth stability index, assume the calculation method is to sum the absolute values of the bandwidth change rate gradient sequence and then divide by the sequence length. For the bandwidth change rate gradient sequence [0.05, 0.1, -0.05, 0.03], the absolute value sequence is [0.05, 0.1, 0.05, 0.03], the sum is 0.05 + 0.1 + 0.05 + 0.03 = 0.23, and the sequence length is 4. Then the bandwidth stability index = 0.23 ÷ 4 = 0.0575. The smaller the bandwidth stability index, the more stable the bandwidth. The preset packet loss tolerance curve is set according to historical data and experience, and the packet loss rate distribution trend is compared with the preset packet loss tolerance curve. Assume that the packet loss rate distribution trend is relatively close to the preset packet loss tolerance curve in most time periods, but exceeds the tolerance curve in some time periods. After comprehensive judgment, the packet loss risk level is "medium".

[0090] Step S1433: Analyze the correlation between the duration of the jitter peak-valley interval and the buffer capacity parameter in the codec parameter characteristics to generate the cumulative impact coefficient of jitter on decoding delay.

[0091] For example, given that the average duration of the jitter peak-valley interval is 8 milliseconds and the buffer capacity parameter is 500 KB, by analyzing historical data and experimental results, it is found that there is a certain relationship between the duration of the jitter peak-valley interval and decoding delay. Suppose it is concluded through analysis that for every 1-millisecond increase in the duration of the jitter peak-valley interval, under the current buffer capacity, the decoding delay increases by 0.1 millisecond. Then the cumulative impact coefficient of an 8-millisecond jitter peak-valley interval duration on decoding delay is 8 × 0.1 = 0.8.

[0092] Step S1434: Adjust the dynamic bitrate adjustment threshold in the optimization strategy template according to the bandwidth stability index, lower the buffer overflow protection threshold according to the packet loss risk level, and synchronously correct the decoding thread scheduling interval parameter in combination with the cumulative impact coefficient to generate a dynamic parameter adjustment rule.

[0093] For example, according to the bandwidth stability index of 0.0575, since the index is small, it indicates that the bandwidth is relatively stable, and the range of the dynamic bitrate adjustment threshold can be appropriately narrowed. Suppose the original dynamic bitrate adjustment threshold was ±200 kbps, and after adjustment, it is ±150 kbps. According to the packet loss risk level of "medium", lower the buffer overflow protection threshold. The original buffer overflow protection threshold was 450 KB, and after lowering, it is 400 KB. In combination with the cumulative impact coefficient of 0.8, correct the decoding thread scheduling interval parameter. Suppose the original scheduling interval was 10 milliseconds, and now it is adjusted to 10 + 0.8 = 10.8 milliseconds. The above adjusted parameters form a dynamic parameter adjustment rule.

[0094] Step S1435: Inject the dynamic parameter adjustment rule into the bitrate adaptation module, buffer management module, and decoding control module of the optimization strategy template to generate a dynamic transmission optimization strategy adapted to the current network fluctuation characteristics.

[0095] In this embodiment, in the bitrate adaptation module, a bitrate increase / decrease step instruction is generated based on the difference between the dynamic bitrate adjustment threshold and the current actual bandwidth occupancy rate, and the bitrate increase / decrease step instruction and the buffer capacity parameter are dynamically weighted to generate an adaptive bitrate adjustment instruction that takes into account both bandwidth fluctuations and buffering capabilities. Assume that the current actual bandwidth occupancy rate is 1800 kbps and the dynamic bitrate adjustment threshold is ±150 kbps. After calculating the difference, a bitrate increase / decrease step instruction is generated according to the magnitude and direction of the difference. For example, if the difference is positive and within a certain range, an instruction with a bitrate increase step of 50 kbps is generated. This instruction and the buffer capacity parameter of 500 KB are dynamically weighted. Assume that the weight of the buffer capacity is 0.4 and the weight of the bitrate difference is 0.6. After calculation and processing, the generated adaptive bitrate adjustment instruction is: when the bandwidth is stable and there is room for improvement, gradually increase the bitrate to 2000 kbps, but make appropriate adjustments according to the usage of the buffer.

[0096] In the buffer management module, the memory allocation ratio between the receive buffer and the decoding buffer is re-divided according to the adjusted buffer overflow protection threshold, and the trigger frequency of the redundant data retransmission request is dynamically set in combination with the packet loss risk level to generate a buffer expansion and data recovery linkage instruction. The adjusted buffer overflow protection threshold is 400 KB, and the memory allocation ratio between the receive buffer and the decoding buffer is re-divided. Assume that it is adjusted from the original 6:4 to 5:5. In combination with the packet loss risk level of "medium", the trigger frequency of the redundant data retransmission request is dynamically set. When the packet loss rate reaches 0.03, the redundant data retransmission request is triggered, and the number and frequency of retransmissions are dynamically adjusted according to factors such as the continuous number and interval time of packet losses.

[0097] In the decoding control module, the priority queue of the audio-visual decoding task is adjusted according to the decoding thread scheduling interval parameter, and the resource occupancy rate of the hardware decoder is dynamically quota-allocated based on the cumulative impact factor to generate a decoding delay equalization control instruction. According to the adjusted decoding thread scheduling interval parameter of 10.8 milliseconds, the priority queue of the audio-visual decoding task is adjusted. For video decoding tasks, if their data volume is large and may cause an increase in delay, they are given priority for processing. Based on the cumulative impact factor of 0.8, the resource occupancy rate of the hardware decoder is dynamically quota-allocated. Assume that the total resources of the hardware decoder are 100%. According to the cumulative impact factor, 58% of the resources are allocated for video decoding and 42% for audio decoding to achieve the equalization control of decoding delay.

[0098] Perform cross - dependency detection on the adaptive bitrate adjustment instruction, buffer expansion and data recovery linkage instruction, and decoding delay equalization control instruction, eliminate resource competition conflicts caused by threshold adjustment between instructions, and generate a dynamic transmission optimization strategy under mutually exclusive operation constraints. For example, it is detected that the adaptive bitrate adjustment instruction may affect the usage of the buffer, and the buffer expansion and data recovery linkage instruction is also related to the buffer. Through analysis and adjustment, it is determined that when adjusting the bitrate, the buffer expansion operation is first suspended, and relevant buffer operations are carried out after the bitrate adjustment is stable to avoid resource competition conflicts. After such processing, the final dynamic transmission optimization strategy is generated.

[0099] Step S144: Parse the bitrate adjustment instruction, buffer management instruction, and decoding priority instruction in the dynamic transmission optimization strategy.

[0100] For example, the bitrate adjustment instruction in the dynamic transmission optimization strategy is: when the bandwidth is stable and there is room for improvement, gradually increase the bitrate to 2000 kbps, but make appropriate adjustments according to the usage of the buffer. The buffer management instruction is: adjust the memory allocation ratio of the receiving buffer to the decoding buffer to 5:5. When the packet loss rate reaches 0.03, trigger a redundant data re - transmission request and dynamically adjust the re - transmission frequency according to the packet loss situation. The decoding priority instruction is: according to the adjusted decoding thread scheduling interval parameter of 10.8 milliseconds, adjust the priority queue of the audio - video decoding task, allocate 58% of the hardware decoder resources for video decoding, and 42% of the resources for audio decoding.

[0101] Step S145: Adjust the encoding bitrate of the current video stream according to the bitrate adjustment instruction, and synchronously update the sensitivity parameter of the adaptive bitrate algorithm.

[0102] In this embodiment, according to the bitrate adjustment instruction, start to adjust the encoding bitrate of the current video stream. If the current bitrate is lower than 2000 kbps and the bandwidth is stable, gradually increase the bitrate according to the instruction. Suppose the current bitrate is 1800 kbps, first increase the bitrate by 50 kbps to 1850 kbps. At the same time, synchronously update the sensitivity parameter of the adaptive bitrate algorithm. The sensitivity parameter of the adaptive bitrate algorithm determines the response speed of the algorithm to network fluctuations. According to the current network fluctuation situation and bitrate adjustment requirements, adjust the sensitivity parameter from the original 0.5 to 0.6, so that the algorithm can more quickly sense network changes and make bitrate adjustments.

[0103] Step S146: Dynamically allocate the memory ratio of the receiving buffer to the decoding buffer according to the buffer management instruction, and set the buffer overflow protection threshold.

[0104] In this embodiment, according to the buffer management instruction, the memory ratio of the receive buffer to the decoding buffer is adjusted from the original 6:4 to 5:5. Assuming the total memory is 1000 KB, then 500 KB is allocated to the receive buffer and 500 KB is also allocated to the decoding buffer. At the same time, the buffer overflow protection threshold is set to 400 KB. This means that when the data volume in the receive buffer or the decoding buffer reaches 400 KB, the system will take corresponding measures, such as pausing data reception or accelerating the decoding speed, to prevent buffer overflow.

[0105] Step S147: Adjust the scheduling order of the audio - video decoding thread according to the decoding priority instruction, and allocate the usage priority of the hardware acceleration resources.

[0106] In this embodiment, according to the decoding priority instruction, the priority queue of the audio - video decoding task can be adjusted according to the adjusted decoding thread scheduling interval parameter of 10.8 milliseconds. During video playback, for video decoding tasks with a large amount of data that may cause an increase in latency, they are advanced in the priority queue. For example, when complex scenes appear in the video frame, such as a large number of dynamic special effects or high - resolution detail displays, the corresponding video decoding task will be preferentially arranged for the decoding thread to process.

[0107] At the same time, based on the instruction, 58% of the hardware decoder resources are allocated to video decoding, and 42% of the hardware decoder resources are allocated to audio decoding. The hardware acceleration resources include the computing power of the GPU, the processing power of the dedicated decoding chip, etc. Assuming the hardware decoder is a video decoding functional module integrated with the GPU, the total resources are measured by the number of computing cores and bandwidth of the GPU. For example, 58% of the GPU computing cores and the corresponding bandwidth can be allocated to the video decoding task, and 42% to the audio decoding task. In this way, during decoding, video decoding can preferentially obtain more hardware acceleration resources, thus processing video data more efficiently, reducing decoding latency, achieving balanced control of decoding latency, and improving the overall audio - video transmission quality and playback effect.

[0108] In this embodiment, the training process of the reinforcement learning optimization model includes: Step S210: Collect the audio - visual synchronization feature set, dynamic bitrate feature set, and decoding latency feature set of the sample transmission session sequence as the model input sample set, and obtain the user feedback score, transmission interruption rate, and bitrate smoothness index of the corresponding transmission session as the quality evaluation true value label set.

[0109] In this embodiment, it still focuses on the scenario where user A watches a popular TV drama on a certain video - playing platform. To improve the user viewing experience, the platform deeply analyzes the audio - video transmission situations of many users, and user A's viewing behavior this time is one of the important samples.

[0110] For each transmission session sequence in which user A watches a TV drama, various required feature sets can be collected. Regarding the audio-visual synchronization feature set, the subjective scoring feature sequence among them is obtained relying on the real-time feedback of the user during the viewing process. The platform has specifically set up a convenient and fast feedback channel for users. For example, a scoring button is set on the playback interface, and user A can rate the current audio-visual synchronization status at any time. The scoring range is from 1 to 10 points. In a 30-minute transmission session sequence, each minute is used as a feedback interval. For example, at the 1st minute, user A gives a score of 7 to the audio-visual synchronization situation based on his intuitive feeling; at the 2nd minute, feeling that the audio-visual synchronization effect is slightly better, gives 8 points; at the 3rd minute, feeling that it is about the same as the 1st minute, gives 7 points... and so on. Finally, the subjective scoring feature sequence is obtained as [7, 8, 7, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7].

[0111] The acquisition of the bitrate adaptation deviation value sequence is relatively complex. It needs to comprehensively consider various factors such as network fluctuation characteristics and codec parameters. During this transmission session process, the network fluctuation characteristics cover data such as packet loss rate, jitter rate, and bandwidth change rate. The above network data is collected in real time, and combined with information such as the initial bitrate setting value, dynamic bitrate adjustment threshold, and buffer capacity parameter in the codec parameters, through a series of calculation processes. For example, first evaluate the stability of the network based on data such as packet loss rate, jitter rate, and bandwidth change rate, and then combine the initial bitrate setting value and the dynamic bitrate adjustment threshold to calculate the bitrate that should be adapted theoretically at different times to achieve the best transmission effect, and then compare it with the actual bitrate to obtain the bitrate adaptation deviation value. After this series of operations, the final bitrate adaptation deviation value sequence is obtained as [40, -20, 30, 50, -10, 20, -5, 40, 60, 30, -20, 40, 50, 30, -10, 20, 30, -5, 40, 50, 30, -10, 20, 40, 50, 30, -10, 20, 30, -5]. The subjective scoring feature sequence and the bitrate adaptation deviation value sequence are corresponded one by one according to the transmission time period order to construct a two-dimensional input vector. Specifically, the two-dimensional input vector corresponding to the 1st minute is composed of the subjective score of 7 points at the 1st minute and the bitrate adaptation deviation value of 40, that is, [7, 40]; the two-dimensional input vector corresponding to the 2nd minute is [8, -20]; the 3rd minute is [7, 30]... and so on. In this way, a complete audio-visual synchronization feature set is formed.

[0112] The collection of the dynamic bitrate feature set cannot be ignored either. The sequence of bitrate adaptation deviation values is the one obtained through complex calculations before, which is [40, -20, 30, 50, -10, 20, -5, 40, 60, 30, -20, 40, 50, 30, -10, 20, 30, -5, 40, 50, 30, -10, 20, 40, 50, 30, -10, 20, 30, -5]. And the sequence of bandwidth change trends is obtained by monitoring the network bandwidth in real time. During the entire 30-minute transmission session, the system continuously tracks the changes in network bandwidth, with each minute as the recording unit. For example, at the 1st minute, the bandwidth change rate is 0.08; at the 2nd minute, the bandwidth change rate becomes 0.06; at the 3rd minute, it becomes 0.09... Recorded in this way in sequence, the sequence of bandwidth change trends is [0.08, 0.06, 0.09, 0.1, 0.07, 0.08, 0.06, 0.09, 0.11, 0.09, 0.06, 0.08, 0.1, 0.09, 0.07, 0.08, 0.09, 0.06, 0.09, 0.1, 0.09, 0.07, 0.08, 0.09, 0.1, 0.09, 0.07, 0.08, 0.09, 0.06]. This sequence of bandwidth change trends clearly reflects the change trend of the network bandwidth per minute. Combining the sequence of bitrate adaptation deviation values and the sequence of bandwidth change trends, a dynamic bitrate fluctuation input matrix is constructed. Described in words, it is to put the bitrate adaptation deviation value at each moment and the corresponding bandwidth change rate together to form a data set with a matrix-like structure, which is convenient for subsequent analysis and processing of dynamic bitrate features.

[0113] The acquisition of the decoding delay feature set depends on the accurate monitoring of the delay situation during the decoding process. In this transmission session, the system sets up an accurate timestamp recording mechanism in the decoding link, and records the decoding delay data in detail at intervals of one minute. For example, the decoding delay time at the 1st minute is 12 milliseconds, at the 2nd minute is 10 milliseconds, at the 3rd minute is 11 milliseconds... And so on. Finally, the decoding delay feature set obtained is [12, 10, 11, 13, 10, 12, 10, 11, 14, 11, 10, 12, 13, 11, 10, 12, 11, 10, 11, 13, 11, 10, 12, 11, 13, 11, 10, 12, 11, 10], with the unit of milliseconds. This decoding delay feature set intuitively shows the decoding delay situation at each moment during the entire transmission session.

[0114] Meanwhile, for this transmission session, the system also obtains the user feedback score, transmission interruption rate, and bitrate smoothness index of the corresponding transmission session as the true quality evaluation label set. The user feedback score is a comprehensive score given by User A based on their overall viewing experience after watching the entire TV drama clip. After this viewing, User A felt that the overall experience was good and gave an 8-point evaluation. The transmission interruption rate is determined by calculating the ratio of the number of interruptions during the transmission process to the total transmission duration. In this 30-minute transmission session, there were 2 short transmission interruptions, so the transmission interruption rate is 2 divided by 30, approximately equal to 0.067. The bitrate smoothness index is an index used to measure the fluctuation of the bitrate during the entire transmission process. It is obtained by analyzing and calculating the law of the bitrate changing over time and reflects whether the bitrate is stable. In this transmission session, after a series of complex calculations and analyses (the specific calculation process involves considering various factors such as the amplitude and frequency of the bitrate change), the bitrate smoothness index is a specific value (assumed to be a value that conforms to the actual situation), and this index will be used as an important reference for measuring the transmission quality in subsequent model training.

[0115] Step S220: Input the audio-visual synchronization feature set into the first policy sub-network for audio-visual synchronization policy inference training to generate the first quality influence coefficient sequence. At the same time, input the dynamic bitrate feature set into the second policy sub-network for bitrate adaptation policy inference training to generate the second quality influence coefficient sequence, and input the decoding delay feature set into the third policy sub-network for delay control policy inference training to generate the third quality influence coefficient sequence.

[0116] Step S221: Input the audio-visual synchronization feature set into the first policy sub-network for audio-visual synchronization policy inference training to generate the first quality influence coefficient sequence.

[0117] Step S2211: Extract the subjective score feature sequence and bitrate adaptation deviation value sequence from the audio-visual synchronization features, and construct a two-dimensional input vector in the order of transmission time periods.

[0118] For example, in the instance where User A watches a TV drama, from the already collected audio-visual synchronization feature set, accurately extract the subjective score feature sequence [7, 8, 7, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7, 8, 9, 8, 7] and the bitrate adaptation deviation value sequence [40, -20, 30, 50, -10, 20, -5, 40, 60, 30, -20, 40, 50, 30, -10, 20, 30, -5, 40, 50, 30, -10, 20, 40, 50, 30, -10, 20, 30, -5].

[0119] Then, strictly in accordance with the order of transmission time periods, the corresponding elements in these two sequences are combined one by one to construct a two-dimensional input vector. Specifically, in the transmission time period of the 1st minute, the subjective score of 7 points and the bitrate adaptation deviation value of 40 are combined to form the first two-dimensional input vector [7, 40]; at the 2nd minute, the subjective score of 8 points and the bitrate adaptation deviation value of -20 are combined to obtain the two-dimensional input vector [8, -20]; at the 3rd minute, the two-dimensional input vector [7, 30] is formed... and so on, until the 30th minute, a complete set of two-dimensional input vectors is formed. Each two-dimensional input vector in this set of two-dimensional input vectors represents the audio-visual synchronization related feature information of a specific transmission time period.

[0120] Step S2212: Perform a non-linear transformation on the two-dimensional input vector through the three-layer fully connected layer of the first policy sub-network, and output the quality attenuation coefficient of the audio-visual synchronization anomaly on the user perception within each transmission time period.

[0121] The three-layer fully connected layer of the first policy sub-network processes the input two-dimensional input vector step by step and meticulously. Taking the first two-dimensional input vector [7, 40] as an example, when it enters the first layer of the fully connected layer, weighted summation and non-linear transformation operations will be performed.

[0122] Specific weight and bias parameters are set inside the first layer of the fully connected layer. Assume the weight parameters are: the weight corresponding to the subjective score of 7 is 0.3, and the weight corresponding to the bitrate adaptation deviation value of 40 is 0.7; the bias parameter is 0.1. Then, through the weighted summation calculation, the result is: 0.3×7 + 0.7×40 + 0.1 = 2.1 + 28 + 0.1 = 30.2. Among them, the calculation process here is to multiply each element in the two-dimensional input vector by the corresponding weight, then add the products, and then add the bias value.

[0123] Next, this result is transformed through a non-linear function. Assume the non-linear function used here is the ReLU function (Rectified Linear Unit), and its definition is: when the input value is greater than 0, the function output is equal to the input value; when the input value is less than or equal to 0, the function output is 0. For the calculation result of 30.2, since it is greater than 0, after being transformed by the ReLU function, the output value is still 30.2.

[0124] The result after the first - layer processing will be used as the input of the second - layer fully - connected layer. The second - layer fully - connected layer also has its own independent weight and bias parameters. Suppose the weights corresponding to the output values of the first layer are 0.5 and 0.5 respectively (the weights here are assumed for simplicity of explanation and will be adjusted according to model training in practice), and the bias parameter is 0.3. Perform the weighted - sum calculation again: 0.5×30.2 + 0.5×30.2+0.3 = 15.1+15.1 + 0.3 = 30.5. Then, this result passes through the ReLU function for non - linear transformation. Since 30.5 is greater than 0, the output value after transformation is still 30.5.

[0125] Finally, the output of the second layer will enter the third - layer fully - connected layer. The third - layer fully - connected layer also has corresponding weights and biases. Suppose the weights are 0.8 and 0.2 respectively, and the bias is 0.5. Perform the weighted - sum: 0.8×30.5+0.2×30.5 + 0.5 = 24.4+6.1 + 0.5 = 31. After non - linear transformation through the ReLU function, the final output value is 31, and this final output value represents the quality decay coefficient of the audio - visual synchronization anomaly on the user perception during the first transmission period.

[0126] In the same way, for each two - dimensional input vector corresponding to each transmission period, such a three - layer fully - connected layer processing is carried out, and the quality decay coefficient of the audio - visual synchronization anomaly on the user perception during each transmission period can be obtained. For example, after the same three - layer fully - connected layer processing process for the second two - dimensional input vector [8, - 20], a corresponding quality decay coefficient will also be obtained (the specific calculation process is similar to the above, and finally a value that conforms to the model operation result is obtained). In this way, the corresponding quality decay coefficients are generated for each transmission period.

[0127] Step S2213: Align and calibrate the quality decay coefficient with the occurrence time of the synchronization - related complaint event in the user feedback score, and generate the first quality impact coefficient sequence reflecting the correlation strength between the audio - visual synchronization feature and the quality evaluation result.

[0128] During the process of user A watching a TV drama, the occurrence time of the synchronization - related complaint event in the user feedback score is recorded simultaneously. Suppose at the 10th minute, user A submitted a complaint about audio - visual synchronization through the feedback channel, and the feedback content indicated that there was an obvious out - of - sync situation of the audio and video at that time.

[0129] Previously, through the processing of the first - strategy sub - network, the quality decay coefficients of each transmission period have been obtained. For example, the quality decay coefficient at the 1st minute is 31 (assumed value), the coefficient at the 2nd minute is another calculated value... The quality decay coefficient at the 10th minute is 40 (assumed value).

[0130] Now, align and calibrate the mass attenuation coefficient with the occurrence time of synchronization-related complaint events. Specifically, take the 10th minute after the occurrence of the complaint event as the key node to adjust and correlate the entire mass attenuation coefficient sequence.

[0131] The system will re-examine the previously calculated mass attenuation coefficient based on the occurrence of the complaint event. For the mass attenuation coefficient before the occurrence of the complaint event, appropriate weighted adjustments may be made according to the time distance from the complaint moment and other relevant factors. For example, the mass attenuation coefficient at the 9th minute, due to its proximity to the complaint moment, may be given a relatively high weight, making it play a more important role in the final first quality impact coefficient sequence. Suppose after a series of complex calibration algorithms (involving comprehensive considerations of various factors such as time distance and the severity of the complaint event), the mass attenuation coefficient at the 9th minute is adjusted to obtain a new value (assumed to be 38).

[0132] For the mass attenuation coefficient after the occurrence of the complaint event, corresponding processing will also be carried out. For example, the mass attenuation coefficient at the 11th minute may be processed differently during the adjustment due to the influence of the complaint event compared to when there is no complaint event. After calibration, a new value (assumed to be 35) is obtained.

[0133] By comprehensively and meticulously calibrating the mass attenuation coefficient of each transmission period according to the occurrence time of synchronization-related complaint events, a first quality impact coefficient sequence reflecting the correlation strength between the audio-visual synchronization characteristics and the quality assessment results is finally generated. Each element in this first quality impact coefficient sequence is no longer just a simple mass attenuation coefficient, but more accurately reflects the correlation degree between the audio-visual synchronization characteristics and the overall quality assessment results after comprehensively considering the synchronization-related complaint events in the user feedback. For example, after calibration, the first quality impact coefficient sequence may be [30, 32, 31, 33, 32, 31, 32, 33, 34, 40, 38, 36, 35, 34, 33, 32, 33, 34, 35, 36, 35, 34, 33, 32, 33, 34, 35, 36, 35, 34], and this first quality impact coefficient sequence will be used as an important basis for subsequent model training and evaluation.

[0134] Step S222: Input the dynamic bitrate feature set into the second policy sub-network for bitrate adaptation policy inference training to generate a second quality impact coefficient sequence.

[0135] Step S2221: Extract the bitrate adaptation deviation value sequence and the bandwidth change trend sequence from the dynamic bitrate features to construct a dynamic bitrate fluctuation input matrix.

[0136] In the transmission session where user A watches a TV drama, a sequence of bitrate adaptation deviation values [40, -20, 30, 50, -10, 20, -5, 40, 60, 30, -20, 40, 50, 30, -10, 20, 30, -5, 40, 50, 30, -10, 20, 40, 50, 30, -10, 20, 30, -5] and a sequence of bandwidth change trends [0.08, 0.06, 0.09, 0.1, 0.07, 0.08, 0.06, 0.09, 0.11, 0.09, 0.06, 0.08, 0.1, 0.09, 0.07, 0.08, 0.09, 0.06, 0.09, 0.1, 0.09, 0.07] are extracted from the dynamic bitrate feature set. These two sequences are constructed into a dynamic bitrate fluctuation input matrix. Taking the transmission period as the dimension, the bitrate adaptation deviation value and the bandwidth change trend value corresponding to each period are combined together. In the first transmission period, the bitrate adaptation deviation value is 40 and the bandwidth change trend value is 0.08, forming a combined unit; in the second transmission period, the bitrate adaptation deviation value is -20 and the bandwidth change trend value is 0.06, forming another combined unit. And so on, the corresponding values of all 30 transmission periods are combined in this way.

[0137] Step S2222: Capture the association pattern between bitrate mutation events and transmission interruption events through the temporal convolutional layer of the second policy sub-network, and output the negative impact weight of bitrate instability on quality assessment for each period.

[0138] The temporal convolutional layer of the second policy sub-network begins to deeply analyze the input dynamic bitrate fluctuation input matrix. This temporal convolutional layer has a specially designed convolutional kernel, and the above-mentioned convolutional kernel can scan and process data in the time dimension to capture the potential association pattern between bitrate mutation events and transmission interruption events.

[0139] In the first transmission period, the convolutional kernel first acts on the first unit of the dynamic bitrate fluctuation input matrix (composed of the bitrate adaptation deviation value 40 and the bandwidth change trend value 0.08) and several adjacent units (here assume the first three units, which will actually be determined according to the size and design of the convolutional kernel). The convolutional kernel performs specific weighted calculations and combinations on the bitrate adaptation deviation value and the bandwidth change trend value in the above units. For example, for the bitrate adaptation deviation value 40, a weight of 0.6 is assigned, and for the bandwidth change trend value 0.08, a weight of 0.4 is assigned. After a series of calculations (such as multiplication, addition, etc.), an intermediate result is obtained, and this intermediate result will undergo a non-linear transformation (such as ReLU function transformation) to highlight the important features in the data.

[0140] Over time, the convolutional kernel gradually moves along the time dimension, performing the same operation on the unit data for each transmission period. During the movement, it pays special attention to the change of the coding rate adaptation deviation value. When it is found that the coding rate adaptation deviation value changes significantly in a certain period, for example, suddenly changes from a small value to a large value (such as suddenly changing from -20 to 60), this may indicate the occurrence of a coding rate mutation event. At the same time, combining relevant data such as the transmission interruption rate (in this scenario, it is assumed that the system will incorporate the occurrence information of the transmission interruption event into this analysis process in some way, such as marking a special identifier in the period when the transmission interruption event occurs), analyze whether there is an association between this coding rate mutation event and the transmission interruption event.

[0141] After traversing and processing the entire dynamic coding rate fluctuation input matrix, the temporal convolutional layer will calculate the negative impact weight of coding rate instability on quality assessment for each period according to the captured association pattern. For example, for the first transmission period, after a series of complex calculations and analyses, the negative impact weight is obtained as 0.3; for the second period, due to the specific combination of the coding rate adaptation deviation value and the bandwidth change trend and the association analysis with the possible transmission interruption event, the negative impact weight is obtained as 0.2. Each period will obtain a corresponding negative impact weight value according to its own data characteristics and the association with the transmission interruption event, and the above negative impact weight value reflects the degree of negative impact that the coding rate instability may have on the overall quality assessment in each period.

[0142] Step S2223: Perform a moving window average process on the negative impact weight in combination with the coding rate smoothness true label to generate a second quality impact coefficient sequence.

[0143] After obtaining the negative impact weight of coding rate instability on quality assessment for each period, it is necessary to further process the above weight in combination with the coding rate smoothness true label. The coding rate smoothness true label is an important index obtained when collecting the sample transmission session sequence before, and it reflects the actual smoothness of the coding rate in this transmission session.

[0144] Suppose we use a moving window of size 3 to average the negative impact weight. Taking the first period as an example, the moving window includes the negative impact weights of the first period and its two adjacent subsequent periods (i.e., the first, second, and third periods). The negative impact weight of the first period is 0.3, the second period is 0.2, and the third period is assumed to be 0.4 after calculation.

[0145] First, calculate the sum of the negative impact weights within the sliding window: 0.3 + 0.2 + 0.4 = 0.9. Then, divide the sum by the size of the sliding window, which is 3, to obtain the average negative impact weight within the window as 0.9 ÷ 3 = 0.3. This average value is the result after the average processing of the first time period through the sliding window and serves as the first element of the second quality impact coefficient sequence.

[0146] Next, move the sliding window backward by one time period, which includes the negative impact weights of the second, third, and fourth time periods. Assume the negative impact weight of the fourth time period is 0.5. Then, the sum within the window is 0.2 + 0.4 + 0.5 = 1.1, and the average negative impact weight is 1.1 ÷ 3 ≈ 0.37, which is the second element of the second quality impact coefficient sequence.

[0147] In this way, the sliding window continuously moves backward, calculating a new average value each time until all time periods are processed. For example, when the sliding window moves to include the twenty-eighth, twenty-ninth, and thirtieth time periods, assume the negative impact weight of the twenty-eighth time period is 0.25, the twenty-ninth time period is 0.3, and the thirtieth time period is 0.2. Then, the sum is 0.25 + 0.3 + 0.2 = 0.75, and the average negative impact weight is 0.75 ÷ 3 = 0.25, which is the thirtieth element of the second quality impact coefficient sequence.

[0148] Through such average processing of the sliding window, the second quality impact coefficient sequence is generated. This second quality impact coefficient sequence comprehensively considers the negative impact weight of the unstable bitrate in each time period on the quality assessment and the overall situation of bitrate smoothness, and more accurately reflects the degree of influence of the dynamic bitrate characteristics on the quality assessment. For example, the generated second quality impact coefficient sequence may be [0.3, 0.37, 0.33, 0.4, 0.35, 0.32, 0.3, 0.33, 0.38, 0.35, 0.32, 0.3, 0.33, 0.36, 0.34, 0.31, 0.3, 0.32, 0.35, 0.34, 0.32, 0.3, 0.31, 0.33, 0.35, 0.33, 0.31, 0.3, 0.32, 0.25].

[0149] Step S223: Input the decoding delay feature set into the third policy sub-network for delay control policy inference training to generate the third quality impact coefficient sequence.

[0150] Step S2231: Extract the decoding delay value sequence from the decoding delay feature set and construct a one-dimensional input vector in the order of transmission time periods.

[0151] In the scenario where user A is watching a TV drama, from the set of decoding delay features [12, 10, 11, 13, 10, 12, 10, 11, 14, 11, 10, 12, 13, 11, 10, 12, 11, 10, 11, 13, 11, 10, 12, 11, 13, 11, 10, 12, 11, 10], a sequence of decoding delay values is directly extracted. According to the order of transmission time periods, the above values are constructed into a one-dimensional input vector. The decoding delay value in the first transmission time period is 12, forming the first element of the one-dimensional input vector; the decoding delay value in the second transmission time period is 10, serving as the second element; the decoding delay value in the third transmission time period is 11, becoming the third element... and so on, until the decoding delay value 10 in the thirtieth transmission time period, finally constructing the one-dimensional input vector [12, 10, 11, 13, 10, 12, 10, 11, 14, 11, 10, 12, 13, 11, 10, 12, 11, 10, 11, 13, 11, 10, 12, 11, 13, 11, 10, 12, 11, 10]. This one-dimensional input vector will be used as the input data for the third policy sub-network for subsequent analysis and processing.

[0152] Step S2232: Extract and transform features from the one-dimensional input vector through a specific neuron layer of the third policy sub-network, and output the influence coefficients of decoding delay in each time period on quality assessment.

[0153] A specific neuron layer of the third policy sub-network starts to process the input one-dimensional input vector. This specific neuron layer has a specially designed neuron structure and connection weights, aiming to extract key features related to decoding delay and convert the above features into influence coefficients meaningful for quality assessment.

[0154] When the one-dimensional input vector enters the specific neuron layer, the neurons in the first layer first perform a weighted sum on each element of the input vector. Assume that the weight corresponding to the first element (decoding delay value 12) is 0.4, the weight corresponding to the second element (decoding delay value 10) is 0.3, and the weight corresponding to the third element (decoding delay value 11) is 0.3 (the weights here are assumed for simplicity of explanation and will be adjusted according to model training in practice). After the weighted sum calculation, an intermediate result is obtained: 0.4×12 + 0.3×10 + 0.3×11 = 4.8 + 3 + 3.3 = 11.1.

[0155] This intermediate result will be transformed through a non-linear function, such as using the Sigmoid function ((f(x)=\frac{1}{1 + e^{-x}})). Substituting 11.1 into the Sigmoid function, a transformed value (assumed to be 0.95, and the actual calculation result is obtained according to the function calculation) is obtained, and then it is passed to the next layer of neurons.

[0156] The second - layer neurons will further process the output of the first layer. It may perform weighted summation and non - linear transformation operations again. Suppose the weight of the second layer for the output value of the first layer is 0.6. After weighted calculation and non - linear transformation (such as another non - linear function), a new result is obtained. This process will continue in multiple neuron layers, and each neuron layer processes the data according to its own design to extract and decode deeper features related to the delay.

[0157] After being processed by multiple neuron layers, finally, an influence coefficient of the decoding delay on the quality assessment is output for each transmission period. For example, for the first transmission period, after a series of complex processes, the influence coefficient is 0.8; for the second transmission period, the influence coefficient is 0.7. Each period will obtain a coefficient value reflecting the degree of influence of the decoding delay in that period on the quality assessment according to its decoding delay value and the processing process in the neuron layer.

[0158] Step S2233: Adjust the influence coefficient by combining the true - value label of the transmission interruption rate to generate a third sequence of quality influence coefficients.

[0159] After obtaining the influence coefficients of the decoding delay in each period on the quality assessment, it is necessary to adjust the above - mentioned influence coefficients by combining the true - value label of the transmission interruption rate. The true - value label of the transmission interruption rate reflects the actual occurrence of transmission interruptions in the entire transmission session.

[0160] Suppose the true - value label of the transmission interruption rate is 0.067. For the influence coefficient of each period, it can be adjusted according to the correlation analysis of the influence of the transmission interruption event on the decoding delay. If a transmission interruption event occurs near a certain transmission period, then the influence coefficient of the decoding delay in that period on the quality assessment may be appropriately amplified.

[0161] For example, suppose a transmission interruption event occurs at the 15th minute. For several periods before and after this transmission interruption event, such as the 14th, 15th, and 16th periods, the influence coefficients of the decoding delay on the quality assessment are 0.7, 0.8, and 0.7 respectively. Due to the occurrence of the transmission interruption event, the system will analyze the potential relationship between the transmission interruption and the decoding delay. Considering that the transmission interruption may lead to instability in the decoding process, thereby affecting the quality assessment, the influence coefficients of these periods are adjusted. Suppose the adjustment rule is to multiply the influence coefficients of the period when the transmission interruption event occurs and the one - period before and after it by 1.2 (this adjustment rule is for illustrative purposes, and the actual rule will be determined based on a large number of experiments and data analyses).

[0162] Then, the adjusted influence coefficient for the 14th period is 0.7×1.2 = 0.84; the adjusted influence coefficient for the 15th period is 0.8×1.2 = 0.96; the adjusted influence coefficient for the 16th period is 0.7×1.2 = 0.84.

[0163] For other periods not directly affected by the transmission interruption event, appropriate fine-tuning may also be carried out according to the transmission interruption rate and the overall evaluation logic. For example, for periods far from the transmission interruption event, the influence coefficient may be adjusted by a small proportion according to the magnitude of the transmission interruption rate. Suppose the adjustment rule for periods far from the transmission interruption event is to multiply the influence coefficient by (1 + transmission interruption rate × 0.1). For the 5th period, the original influence coefficient is 0.6, and after adjustment, it is 0.6×(1 + 0.067×0.1) = 0.6×(1 + 0.0067) = 0.6×1.0067 = 0.60402.

[0164] By comprehensively adjusting the influence coefficient of each period in this way in combination with the true value label of the transmission interruption rate, a third quality influence coefficient sequence is finally generated. This third quality influence coefficient sequence comprehensively considers the impact of decoding delay itself on quality assessment and the correction of the impact of transmission interruption events on decoding delay, and more accurately reflects the role of decoding delay characteristics in the entire quality assessment. For example, the generated third quality influence coefficient sequence may be [0.60402, 0.72, 0.75, 0.8, 0.7, 0.78, 0.7, 0.75, 0.84, 0.8, 0.7, 0.78, 0.84, 0.8, 0.84, 0.96, 0.84, 0.8, 0.75, 0.8, 0.8, 0.7, 0.78, 0.8, 0.84, 0.8, 0.7, 0.78, 0.8, 0.75], and this sequence will be used for subsequent fusion with other quality influence coefficient sequences and model training.

[0165] Step S230: Perform time window alignment processing on the first quality influence coefficient sequence, the second quality influence coefficient sequence, and the third quality influence coefficient sequence, calculate the predicted value of the comprehensive quality score for each transmission period through a weighted fusion network, and perform error backpropagation on the predicted value of the comprehensive quality score and the quality assessment true value label set, and iteratively update the parameters of the first policy sub-network, the second policy sub-network, and the weighted fusion network until convergence.

[0166] Step S231: Perform time window alignment processing on the first quality influence coefficient sequence, the second quality influence coefficient sequence, and the third quality influence coefficient sequence.

[0167] In this embodiment, the first quality influence coefficient sequence, the second quality influence coefficient sequence, and the third quality influence coefficient sequence are all generated based on the same transmission period. However, in actual processing, due to minor differences in the calculation process or data acquisition, there may be slight misalignments in the sequences in the time dimension. Therefore, time window alignment processing is required.

[0168] For example, the first quality influence coefficient sequence is [0.3, 0.4, 0.35,...], the second quality influence coefficient sequence is [0.25, 0.3, 0.28,...], and the third quality influence coefficient sequence is [0.4, 0.45, 0.42,...]. Taking the first transmission period as the reference, the three sequences are strictly aligned in time to ensure that the coefficients corresponding to the same transmission period in each sequence can be accurately matched. After the alignment processing, the three sequences are completely corresponding in the time dimension.

[0169] Step S232: Calculate the predicted value of the comprehensive quality score for each transmission period through a weighted fusion network.

[0170] Step S2321: Concatenate the aligned first quality influence coefficient, second quality influence coefficient, and third quality influence coefficient into a three-dimensional input tensor according to the transmission period.

[0171] The three sequences after the time window alignment processing are concatenated according to the transmission period. In the first transmission period, the first quality influence coefficient 0.3, the second quality influence coefficient 0.25, and the third quality influence coefficient 0.4 are concatenated together to form a three-dimensional vector [0.3, 0.25, 0.4]. In the second transmission period, the corresponding coefficients are similarly concatenated into a three-dimensional vector [0.4, 0.3, 0.45]. And so on, the coefficients of all transmission periods are concatenated in this way, and finally a three-dimensional input tensor is formed. Each dimension of this tensor corresponds to the three quality influence coefficient sequences respectively, and in the time dimension, it corresponds to each transmission period, comprehensively integrating the quality influence information of the three aspects.

[0172] Step S2322: Capture the temporal dependence relationship across feature dimensions through a two-layer gated recurrent unit and output the joint influence intensity of multiple factors for each period.

[0173] The two-layer gated recurrent unit (GRU) processes the three-dimensional input tensor. When the three-dimensional vector [0.3, 0.25, 0.4] in the first transmission period enters the first layer of GRU, the GRU will comprehensively analyze the three coefficients according to its internal gating mechanism and weight parameters. For example, through the control of the update gate and the reset gate, it determines the importance of each coefficient at the current moment and captures the mutual relationship between them. Suppose that after being processed by the first layer of GRU, an intermediate result vector [0.32, 0.27, 0.41] is obtained.

[0174] This intermediate result vector enters the second-layer GRU, and the second-layer GRU further processes it to deeply explore the temporal dependence relationships across feature dimensions. For example, it will consider the coefficient changes between the current time period and the previous time period, as well as the long-term and short-term dependencies between different feature coefficients. After being processed by the second-layer GRU, the multi-factor joint influence intensity for each time period is output. For the first transmission time period, assume the output joint influence intensity is 0.35, which comprehensively reflects the joint influence degree of factors such as audio-visual synchronization, bitrate adaptation, and decoding delay on the quality during this time period.

[0175] Step S2323: Map the joint influence intensity to the quality score interval, and superimpose the baseline quality parameter at the transmission session level to generate an end-to-end comprehensive quality score prediction value.

[0176] Assume the quality score interval is [0, 10]. Map the joint influence intensity output for each time period to this quality score interval through a mapping function. For example, for the joint influence intensity of 0.35 output in the first transmission time period, through the mapping function, the mapped score is (10 * 0.35 = 3.5).

[0177] Meanwhile, consider the baseline quality parameter at the transmission session level. Assume the baseline quality parameter is 5, which is a fixed parameter comprehensively considering factors such as the quality of the video content itself and the platform's basic service level. Superimpose the mapped score and the baseline quality parameter to obtain the comprehensive quality score prediction value for the first transmission time period as (3.5 + 5 = 8.5). Calculate in the same way for each transmission time period, and finally generate a sequence of end-to-end comprehensive quality score prediction values.

[0178] Step S233: Perform error backpropagation on the comprehensive quality score prediction value and the quality assessment true label set, and iteratively update the parameters of the first policy sub-network, the second policy sub-network, and the weighted fusion network until convergence.

[0179] Step S2331: Calculate the mean square error between the comprehensive quality score prediction value and the true value of the user feedback score as the main loss function.

[0180] Suppose the true value of the user feedback score is 8 points, and the predicted value of the comprehensive quality score in the first transmission period is 8.5 points. The mean squared error (MSE) is calculated as the average of the squares of the differences between the predicted value and the true value. For the first transmission period, the error is (8.5 - 8 = 0.5), and after squaring it is (0.5^2 = 0.25). Such calculations are performed for the errors of all transmission periods, and then the average is taken to obtain the value of the main loss function. Suppose after calculation, the average value of the mean squared error for all 30 transmission periods is 0.3. This value of the main loss function reflects the overall error degree between the predicted value of the comprehensive quality score and the true value of the user feedback score.

[0181] Step S2332: Increase the cross - entropy loss between the output of the first policy sub - network and the true value of the coding rate adaptation deviation value as an auxiliary constraint term.

[0182] Suppose the results related to the coding rate adaptation deviation value output by the first policy sub - network are [predicted deviation value 1, predicted deviation value 2,...], and the true value of the coding rate adaptation deviation value is [true deviation value 1, true deviation value 2,...]. Cross - entropy loss is used to measure the difference between two probability distributions, and here it is used to measure the difference between the output of the first policy sub - network and the true value of the coding rate adaptation deviation value. By calculating the cross - entropy loss, a value of the auxiliary constraint term is obtained. Suppose after calculation, the value of this auxiliary constraint term is 0.2.

[0183] Step S2333: Use the dynamic weight allocation algorithm to balance the ratio of the main loss function and the auxiliary constraint term, and control the parameter update amplitude.

[0184] In this embodiment, the dynamic weight allocation algorithm can automatically adjust the weights of the main loss function and the auxiliary constraint term according to the situation during the training process. For example, at the beginning of training, the weight of the main loss function may be set to 0.8, and the weight of the auxiliary constraint term is 0.2; as the training progresses, the weights are dynamically adjusted according to the convergence situation and performance of the model. Suppose after several iterations, the weight of the main loss function is adjusted to 0.7, and the weight of the auxiliary constraint term is adjusted to 0.3.

[0185] According to the adjusted weights, calculate the total loss value. Suppose the total loss value is (0.8 * 0.3+0.2 * 0.2 = 0.24 + 0.04 = 0.28). Then, based on this total loss value, perform error backpropagation, calculate the gradient through the backpropagation algorithm, and update the parameters of the first policy sub - network, the second policy sub - network, and the weighted fusion network.

[0186] In each iteration, the above process of calculating the loss value, adjusting the weights, backpropagating the error, and updating the parameters is continuously repeated until the loss value of the model no longer decreases, that is, it reaches the convergence state. At this time, the parameters of the first policy sub-network, the second policy sub-network, and the weighted fusion network reach the optimal state, and can more accurately evaluate and optimize the audio-visual transmission quality.

[0187] In practical application scenarios, the above user experience optimization method can be applied to various online video platforms, live broadcast platforms, and video conferencing systems, etc. Taking an online video platform as an example, there are a large number of users with different network environments and different terminal devices accessing the platform to watch videos every day. Through the above optimization method, it is possible to monitor the transmission quality data of each user terminal in real time, and perform targeted optimization processing according to the network fluctuations, codec parameters, and user feedback of different users.

[0188] It should be noted that in terms of data security and privacy protection, in this embodiment, when obtaining and processing the transmission quality data, relevant laws, regulations, and privacy policies will be strictly followed. The user data is encrypted to ensure that the user's personal information and transmission data are not leaked. During the data storage and transmission process, a secure storage mechanism and encryption protocol are adopted to prevent the data from being illegally obtained and tampered with.

[0189] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a user experience optimization system 100 for an audio-visual transmission system according to some embodiments of the present invention that can implement the idea of the present invention. For example, the processor 120 can be used on the user experience optimization system 100 for the audio-visual transmission system and is used to execute the functions in the present invention.

[0190] The user experience optimization system 100 for the audio-visual transmission system can be a general-purpose server or a special-purpose server, both of which can be used to implement the user experience optimization method for the audio-visual transmission system of the present invention. Although only one server is shown in the present invention, for convenience, the functions described in the present invention can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0191] For example, the user experience optimization system 100 for an audio-video transmission system may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the user experience optimization system 100 for an audio-video transmission system may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present invention can be implemented according to the above program instructions. The user experience optimization system 100 for an audio-video transmission system further includes an input / output (I / O) interface 150 between the computer and other input / output devices.

[0192] For ease of explanation, only one processor is described in the user experience optimization system 100 for an audio-video transmission system. However, it should be noted that the user experience optimization system 100 for an audio-video transmission system in the present invention may also include multiple processors. Therefore, the steps executed by one processor described in the present invention may also be jointly executed or separately executed by multiple processors. For example, if the processor of the user experience optimization system 100 for an audio-video transmission system executes step A and step B, it should be understood that step A and step B may also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0193] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the user experience optimization method for an audio-video transmission system as described above is implemented.

[0194] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A method for optimizing the user experience of an audio - video transmission system, characterized in that, The method includes: Obtaining a set of transmission quality data generated during the audio - video transmission of a target terminal, where the set of transmission quality data includes multiple transmission session sequences, and each transmission session sequence consists of network fluctuation characteristics, codec parameter characteristics, and user feedback characteristics of at least one transmission period; Performing transmission feature extraction processing on the set of transmission quality data to obtain the audio - video synchronization feature, dynamic bitrate feature, and decoding delay feature of each transmission session sequence; Invoking a pre - trained reinforcement learning optimization model to perform dynamic aggregation processing on the audio - video synchronization feature, dynamic bitrate feature, and decoding delay feature to generate a transmission quality evaluation result of the transmission session sequence; Determining a dynamic transmission optimization strategy based on the transmission quality evaluation result and feeding back the dynamic transmission optimization strategy to the audio - video transmission system to trigger real - time transmission parameter adjustment operations.

2. The user experience optimization method for an audio and video transmission system according to claim 1, wherein The performing transmission feature extraction processing on the set of transmission quality data to obtain the audio - video synchronization feature, dynamic bitrate feature, and decoding delay feature of each transmission session sequence includes: Performing time - series segmentation processing on the network fluctuation characteristics in the transmission session sequence to obtain multiple network fluctuation sub - segments; Invoking a pre - trained encoder model to encode each network fluctuation sub - segment to generate a fluctuation intensity vector of each network fluctuation sub - segment; Determining the bitrate adaptation deviation value of each transmission period based on the correlation calculation between the fluctuation intensity vector and the codec parameter characteristics; Performing semantic parsing processing on the user feedback characteristics to extract the subjective scoring feature of the user's perception of audio - video synchronization and the sensitivity feature to delay; Fusing the bitrate adaptation deviation value, the subjective scoring feature, and the sensitivity feature to generate the audio - video synchronization feature, dynamic bitrate feature, and decoding delay feature.

3. The user experience optimization method for an audio and video transmission system according to claim 2, wherein The invoking a pre - trained encoder model to encode each network fluctuation sub - segment to generate a fluctuation intensity vector of each network fluctuation sub - segment includes: Obtaining the packet loss rate sequence, jitter rate sequence, and bandwidth change rate sequence included in each network fluctuation sub - segment; Performing normalization processing on the packet loss rate sequence to obtain the packet loss rate distribution feature; Performing sliding window statistical processing on the jitter rate sequence to generate the jitter peak - valley difference feature and the jitter duration feature; Performing gradient calculation on the bandwidth change rate sequence to obtain the bandwidth change trend feature; Inputting the packet loss rate distribution feature, jitter peak - valley difference feature, jitter duration feature, and bandwidth change trend feature into the encoder model to generate the fluctuation intensity vector through multi - layer non - linear transformation.

4. The user experience optimization method for an audio-video transmission system according to claim 2, characterized in that The determining the bitrate adaptation deviation value of each transmission period based on the correlation calculation between the fluctuation intensity vector and the codec parameter characteristics includes: Extracting the initial bitrate setting value, dynamic bitrate adjustment threshold, and buffer capacity parameter in the codec parameter characteristics; Calculating the cosine similarity between the fluctuation intensity vector and the dynamic bitrate adjustment threshold to obtain the influence coefficient of network fluctuation on the bitrate; Determining the delay tolerance of bitrate adjustment according to the buffer capacity parameter and generating a bitrate adaptation weight in combination with the influence coefficient; Determine the theoretical bitrate adaptation value based on the product of the initial bitrate setting value and the bitrate adaptation weight; Calculate the bitrate adaptation deviation value according to the difference between the theoretical bitrate adaptation value and the actual bitrate change value.

5. The method for optimizing the user experience for an audio and video transmission system according to claim 2, wherein The semantic parsing process for the user feedback features to extract the subjective scoring features of the user's perception of audio - video synchronization and the sensitivity features to latency includes: Obtain the keyword set included in the user feedback text, where the keyword set includes the first - type keywords related to audio - video synchronization and the second - type keywords related to latency; Conduct sentiment polarity analysis on the first - type keywords to determine the user's satisfaction score for audio - video synchronization, and map the satisfaction score to a numerical subjective scoring feature; Perform frequency statistics on the second - type keywords to determine the occurrence frequency and context association strength of the latency - related keywords; Generate the sensitivity feature based on the weighted sum of the occurrence frequency and the context association strength.

6. The user experience optimization method for an audio-video transmission system according to claim 1, characterized in that The reinforcement learning optimization model includes a first policy sub - network, a second policy sub - network, and a third policy sub - network. Invoke the pre - trained reinforcement learning optimization model to perform dynamic aggregation processing on the audio - video synchronization feature, dynamic bitrate feature, and decoding latency feature to generate the transmission quality evaluation result of the transmission session sequence, including: Input the audio - video synchronization feature into the first policy sub - network to generate a first candidate policy set for audio - video synchronization optimization; Input the dynamic bitrate feature into the second policy sub - network to generate a second candidate policy set for bitrate adaptation optimization; Input the decoding latency feature into the third policy sub - network to generate a third candidate policy set for latency optimization; Perform policy conflict detection processing on the first candidate policy set, the second candidate policy set, and the third candidate policy set, and eliminate the conflicting candidate policies to obtain the remaining candidate policies; Generate a comprehensive transmission optimization policy through weighted fusion based on the policy priority scores of the remaining candidate policies, and generate the transmission quality evaluation result according to the effect prediction value of the comprehensive transmission optimization policy in the simulated transmission environment.

7. The user experience optimization method for an audio and video transmission system according to claim 6, wherein The policy conflict detection processing on the first candidate policy set, the second candidate policy set, and the third candidate policy set to eliminate the conflicting candidate policies includes: Extract the set of transmission parameter adjustment instructions included in each candidate policy in the first candidate policy set, the second candidate policy set, and the third candidate policy set; Detect whether there is a conflict combination where a bitrate adjustment instruction and a buffer expansion instruction exist simultaneously in the set of transmission parameter adjustment instructions; If it exists, calculate the failure probability of the conflict combination in the historical transmission data. If the failure probability exceeds the preset threshold, eliminate the corresponding candidate policy; Detect whether there is an instruction combination of increasing the number of decoding threads and decreasing the bitrate at the same time. If it exists and the current terminal hardware performance parameters do not meet the parallel processing requirements, eliminate the corresponding candidate policy.

8. The method for optimizing user experience for an audio and video transmission system according to claim 6, characterized in that, The weighted fusion based on the policy priority scores of the remaining candidate policies to generate a comprehensive transmission optimization policy includes: Obtain the execution success rate, execution efficiency score, and user satisfaction score of each remaining candidate strategy in the historical transmission data; Determine the strategy effectiveness coefficient according to the product of the execution success rate and the execution efficiency score; Determine the strategy priority weight according to the ratio of the user satisfaction score to the preset satisfaction threshold; Perform weighted summation of the strategy effectiveness coefficient and the strategy priority weight to obtain the comprehensive priority score of each candidate strategy; Select the top N candidate strategies with the highest comprehensive priority scores, and de-duplicate and merge the transmission parameter adjustment instructions in the top N candidate strategies to generate the comprehensive transmission optimization strategy.

9. The method for optimizing user experience for an audio and video transmission system according to claim 1, wherein The determining of the dynamic transmission optimization strategy based on the transmission quality evaluation result includes: Divide the transmission quality evaluation result into multiple quality levels, where each quality level corresponds to a different optimization strategy template; Detect the target quality level to which the transmission quality evaluation result belongs, and match the optimization strategy template corresponding to the target quality level from the predefined strategy library; Dynamically adjust the parameter thresholds in the optimization strategy template according to the network fluctuation characteristics of the current transmission session sequence to generate a dynamic transmission optimization strategy adapted to the current network state; And, the feedback of the dynamic transmission optimization strategy to the audio-video transmission system to trigger real-time transmission parameter adjustment operations includes: Parse the bitrate adjustment instruction, buffer management instruction, and decoding priority instruction in the dynamic transmission optimization strategy; Adjust the encoding bitrate of the current video stream according to the bitrate adjustment instruction, and synchronously update the sensitivity parameter of the adaptive bitrate algorithm; Dynamically allocate the memory ratio of the receive buffer and the decoding buffer according to the buffer management instruction, and set the buffer overflow protection threshold; Adjust the scheduling order of the audio-video decoding threads according to the decoding priority instruction, and allocate the usage priority of the hardware acceleration resources.

10. A user experience optimization system for an audio-video transmission system, characterized in that, The user experience optimization system for the audio-video transmission system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the user experience optimization method for the audio-video transmission system according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • A wireless DASH streaming media code rate smoothing adaptive transmission method

    CN109040855A

  • A method and system for resisting weak network for audio and video transmission

    CN118101941A

  • Real-time streaming media transmission method and system based on cloud game and storage medium

    CN119232713A

  • Digital audio signal transmission verification method and system based on dynamic feedback enhancement

    CN120015061A

  • System, streaming media optimizer and methods for use therewith

    US20140181266A1

Cited By

  • Alarm condition video communication system applied to police management

    CN120935321A

  • Data transmission method and device based on low-power-consumption Bluetooth device and medium

    CN121078412A

  • Video quality and time delay joint optimization coding method and system under limited bandwidth

    CN122137963A