AI-based audio and video data processing method and system

Through the AI-based audio and video data processing method, the scheduling priority weight of audio and video clips is dynamically adjusted, which solves the problem of failing to fully consider the link end warning status in the prior art, and improves the efficiency and quality of audio and video transmission.

CN120223683AActive Publication Date: 2025-06-27BEIJING JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510671966.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-27
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art audio and video data transmission strategies fail to fully consider the real-time warning status of link end devices, resulting in the inability to adjust the scheduling strategy in time when the link quality decreases, affecting the transmission quality.

Method used

Using AI-based audio and video data processing method, by obtaining the initial scheduling priority weight and historical transmission data records of the target audio and video clips, analyzing and filtering out historical fragments consistent with the environmental characteristics of the target clips, counting the probability of link transmission performance degradation amplitude, and dynamically adjusting the scheduling priority weight based on the change trend of reconstruction cost value.

Benefits of technology

It significantly improves the efficiency and stability of link transmission, and can dynamically adjust the scheduling priority weight of audio and video clips according to real-time link performance and historical data trends, ensuring that highly dynamic voice clips are given priority in an unstable network environment, and improving the quality and reliability of audio and video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223683A_ABST
    Figure CN120223683A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video data processing method and system based on AI. The AI-based audio and video data processing method comprises the following steps: after determining that a target audio and video clip belongs to a high-dynamic voice characteristic type, acquiring an initial scheduling priority value which is generated for the target audio and video clip and is in a link end early warning state, and acquiring a historical transmission data record of associated equipment bearing a transmission task of the target audio and video clip; according to the method, the scheduling priority values of the audio and video clips can be dynamically adjusted according to the real-time link performance and the historical data trend in the link end early warning state, and the link transmission efficiency and stability are remarkably improved. The core innovation point is that a first correction parameter and a second correction parameter are combined, the first correction parameter is dynamically corrected based on real-time link performance, and the second correction parameter is optimized and adjusted according to the trend change of historical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video data transmission optimization, and in particular, to an AI - based audio - video data processing method and system. Background Art

[0002] With the wide application of audio - video communication and real - time data transmission technologies, the quality of the network link has become a key factor affecting transmission stability and user experience. Especially when facing high - dynamic audio - video content, performance fluctuations such as bandwidth, latency, and packet loss in the link often directly affect the transmission quality. The audio - video data transmission strategies in the prior art usually rely on static priority scheduling algorithms, which are scheduled based on indicators such as the bandwidth, latency, and packet loss rate of the link. However, these strategies do not fully consider the real - time warning state of the end - device of the link. In practical applications, the end - link warning state refers to the warning state triggered by the system when the end - device of the link (such as a user terminal or an access device) monitors that the link performance is close to or exceeds the lower limit of the quality of service (QoS). At this time, the system should timely adjust the scheduling strategy to avoid the deterioration of the audio - video transmission quality caused by the decline of the link quality.

[0003] Although the scheduling priority allocation methods in the prior art schedule based on the warning state of the end - device of the link, there are still problems of static setting and inability to flexibly respond to the dynamic changes of the link. In a network environment with multi - hop relay, bandwidth fluctuations, or large link delays, traditional methods cannot dynamically adjust according to the actual link conditions, which may lead to the failure to timely guarantee the transmission priority of key content such as high - dynamic voice segments. Therefore, the existing scheduling mechanisms fail to effectively combine the dynamic changes of the link state with the characteristics of audio - video data, cannot accurately identify and adapt to real - time transmission requirements, and affect the overall transmission efficiency and quality. Summary of the Invention

[0004] The purpose of the present invention is to provide an AI - based audio - video data processing method and system to solve the technical problem of insufficient dynamic scheduling of the link in the prior art.

[0005] The present invention is implemented as follows. An AI - based audio - video data processing method, the method includes: After determining that the target audio - video segment belongs to the high - dynamic voice characteristic type, obtain the initial scheduling priority value generated for the target audio - video segment in the end - link warning state, and obtain the historical transmission data record of the associated device carrying the transmission task of the target audio - video segment; Analyze the historical transmission data record, and screen out several first - type historical segments that are consistent with the environmental characteristics of the target audio - video segment and belong to the high - dynamic voice characteristic type, and several second - type historical segments that do not belong to the high - dynamic voice characteristic type; Statistically analyze the occurrence probabilities of the link transmission performance degradation exceeding the preset threshold standard in a number of first - type historical segments and a number of second - type historical segments respectively, and determine a first correction parameter based on the deviation magnitude between the two; Calculate the reconstruction cost value corresponding to each first - type historical segment, and determine a second correction parameter according to the change trend characteristics of the reconstruction cost value; Combine the first correction parameter and the second correction parameter to correct the initial scheduling priority value.

[0006] As a further limitation of the technical solution of the embodiment of the present invention, the steps of statistically analyzing the occurrence probabilities of the link transmission performance degradation exceeding the preset threshold standard in a number of first - type historical segments and a number of second - type historical segments respectively, and determining a first correction parameter based on the deviation magnitude between the two include: Parse each first - type historical segment and second - type historical segment in sequence, extract the link transmission degradation event data recorded in each historical segment to determine the corresponding link transmission performance degradation magnitude; Count the number of segments in a number of first - type historical segments and a number of second - type historical segments where the link transmission performance degradation magnitude exceeds the preset threshold standard respectively, and calculate their corresponding occurrence probabilities respectively; Determine the first correction parameter based on the deviation magnitude of the occurrence probabilities of the link transmission performance degradation exceeding the preset threshold standard in the first - type historical segments and the second - type historical segments.

[0007] As a further limitation of the technical solution of the embodiment of the present invention, the link transmission degradation event data includes packet loss rate, delay jitter magnitude, and instantaneous bandwidth fluctuation magnitude, and the link transmission performance degradation magnitude is calculated by weighted combination of the packet loss rate, delay jitter magnitude, and instantaneous bandwidth fluctuation magnitude.

[0008] As a further limitation of the technical solution of the embodiment of the present invention, the steps of calculating the reconstruction cost value corresponding to each first - type historical segment and determining a second correction parameter according to the change trend characteristics of the reconstruction cost value include: Parse the link transmission data recorded in each first - type historical segment, extract the actual re - transmission times and transmission delay corresponding to each first - type historical segment, and calculate the reconstruction cost value corresponding to each first - type historical segment based on the actual re - transmission times and transmission delay; Statistically analyze the reconstruction cost values corresponding to all first - type historical segments, and construct a reconstruction cost value change curve according to the segment time sequence; Calculate the average slope of the reconstruction cost value change curve, and use this average slope as the second correction parameter.

[0009] As a further limitation of the technical solution of the embodiment of the present invention, the actual retransmission times are obtained from the number of data retransmissions occurring during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and the receiving time of each first-type historical segment; The reconstruction cost value is calculated by a weighted linear combination of the actual retransmission times and the transmission delay.

[0010] As a further limitation of the technical solution of the embodiment of the present invention, the steps of correcting the initial scheduling priority value by combining the first correction parameter and the second correction parameter include: Retrieve the preset scheduling priority value correction formula, substitute the first correction parameter and the second correction parameter into the scheduling priority value correction formula, and correct the initial scheduling priority value to obtain the corrected scheduling priority value; Apply the corrected scheduling priority value to the link resource scheduling strategy of the target video segment to dynamically adjust the scheduling priority level of the target video segment in the transmission link.

[0011] As a further limitation of the technical solution of the embodiment of the present invention, the scheduling priority value correction formula is: ; where refers to the corrected scheduling priority value; refers to the initial scheduling priority value, refers to the first correction parameter, that is, the deviation amplitude of the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first-type historical segment and the second-type historical segment, refers to the adjustment coefficient corresponding to the first correction parameter, refers to the second correction parameter, that is, the average slope of the reconstruction cost value change curve, refers to the adjustment coefficient corresponding to the second correction parameter; In the scheduling priority value correction formula, ; where refers to the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first-type historical segment, refers to the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the second-type historical segment, is a preset minimum protection value, which is used to avoid calculation anomalies when the denominator is zero or close to zero.

[0012] An AI-based audio and video data processing system, the system includes: a data acquisition module, a data screening module, a first correction parameter determination module, a second correction parameter determination module, and a correction module, where: A data acquisition module, configured to obtain an initial scheduling priority value in a warning state at the end of the link generated for the target audio-visual segment after determining that the target audio-visual segment belongs to the high-dynamic voice feature type, and obtain the historical transmission data record of the associated device carrying the transmission task of the target audio-visual segment; A data screening module, configured to parse the historical transmission data record, and screen out a number of first-class historical segments that are consistent with the environmental characteristics of the target audio-visual segment and belong to the high-dynamic voice feature type, and a number of second-class historical segments that do not belong to the high-dynamic voice feature type; A first correction parameter determination module, configured to respectively count the occurrence probabilities of the link transmission performance deterioration amplitude exceeding a preset threshold standard in a number of first-class historical segments and a number of second-class historical segments, and determine a first correction parameter based on the deviation amplitude between the two; A second correction parameter determination module, configured to calculate the reconstruction cost value corresponding to each first-class historical segment, and determine a second correction parameter according to the change trend characteristics of the reconstruction cost value; A correction module, configured to correct the initial scheduling priority value by combining the first correction parameter and the second correction parameter.

[0013] As a further limitation of the technical solution of the embodiment of the present invention, the first correction parameter determination module specifically includes: A historical segment analysis unit, configured to sequentially analyze each first-class historical segment and second-class historical segment, and extract the link transmission deterioration event data recorded in each historical segment to determine the corresponding link transmission performance deterioration amplitude; the link transmission deterioration event data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude, and the link transmission performance deterioration amplitude is calculated by weighted combination of the packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude; An occurrence probability calculation unit, configured to respectively count the number of segments in a number of first-class historical segments and a number of second-class historical segments where the link transmission performance deterioration amplitude exceeds a preset threshold standard, and calculate their corresponding occurrence probabilities respectively; A deviation amplitude determination unit, configured to determine a first correction parameter based on the deviation amplitude of the occurrence probabilities of the link transmission performance deterioration amplitude exceeding a preset threshold standard in the first-class historical segments and the second-class historical segments.

[0014] As a further limitation of the technical solution of the embodiment of the present invention, the second correction parameter determination module specifically includes: A transmission data parsing unit, which is used to parse the link transmission data recorded in each first - type historical segment, extract the actual re - transmission times and transmission delays corresponding to each first - type historical segment, and calculate the reconstruction cost value corresponding to each first - type historical segment based on the actual re - transmission times and transmission delays; the actual re - transmission times are obtained from the number of data re - transmissions that occur during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and the receiving time of each first - type historical segment; the reconstruction cost value is calculated by a weighted linear combination of the actual re - transmission times and the transmission delays; A change curve construction unit, which is used to count the reconstruction cost values corresponding to all first - type historical segments and construct a change curve of the reconstruction cost values in the order of segment time; An average slope calculation unit, which is used to calculate the average slope of the change curve of the reconstruction cost values and use this average slope as the second correction parameter.

[0015] Adopting the above technical solution, the present invention has the following beneficial effects: The present invention proposes an AI - based audio - visual data processing method, which can dynamically adjust the scheduling priority values of audio - visual segments according to real - time link performance and historical data trends under the warning state at the link end, and significantly improve the efficiency and stability of link transmission. The core innovation lies in the combination of the first correction parameter and the second correction parameter, where the first correction parameter is dynamically corrected based on real - time link performance (such as packet loss rate, delay, etc.), and the second correction parameter is optimized and adjusted according to the trend change of historical data (such as the change trend of reconstruction cost).

[0016] Through this comprehensive correction mechanism, the present invention can effectively solve the problem that the prior art fails to fully consider link fluctuations and long - term trend changes. Especially for high - dynamic voice segments, in an unstable network environment, the present invention can give priority to ensuring their transmission according to the warning signal at the link end, thereby improving the quality and reliability of audio - visual transmission. Through this innovative mechanism, the present invention provides more stable and reliable support for real - time communication and remote video conferencing in complex network environments. Description of the Drawings

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention; Figure 2Flowchart for determining the first correction parameter in the method provided by the embodiments of the present invention; Figure 3 Flowchart for determining the second correction parameter in the method provided by the embodiments of the present invention; Figure 4 Flowchart for correcting the initial scheduling priority value in the method provided by the embodiments of the present invention; Figure 5 Application architecture diagram of the system provided by the embodiments of the present invention; Figure 6 Structural block diagram of the first correction parameter determination module in the system provided by the embodiments of the present invention; Figure 7 Structural block diagram of the second correction parameter determination module in the system provided by the embodiments of the present invention. Detailed implementation manners

[0019] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The present invention will be further explained below in conjunction with specific implementation manners.

[0021] Figure 1 The flowchart of the method provided by the embodiments of the present invention is shown.

[0022] Specifically, an AI-based audio and video data processing method, the method specifically includes the following steps: Step S100, after determining that the target audio and video segment belongs to the high-dynamic voice characteristic type, obtain the initial scheduling priority value generated for the target audio and video segment in the warning state at the end of the link, and obtain the historical transmission data record of the associated device carrying the transmission task of the target audio and video segment.

[0023] In the embodiments of the present invention, with the wide application of audio and video transmission technologies, the quality and stability of the link have become key factors affecting the user experience. Especially in the process of real-time audio and video data transmission, fluctuations in link performance (such as packet loss, latency, bandwidth fluctuations, etc.) directly affect the transmission quality of audio and video content. Therefore, in the case of unstable links, it is particularly important to adopt a method of dynamically adjusting the scheduling priority value. The AI-based audio and video data processing method of the present invention can dynamically adjust the scheduling priority value by intelligently evaluating the characteristics of the target audio and video segment and the link state, so as to ensure the priority transmission of high-dynamic audio and video segments and optimize the overall transmission quality.

[0024] The link end warning state refers to a warning state triggered by the system when the link performance indicators (such as bandwidth, delay, packet loss rate, etc.) of the end device of the link in the network (such as user terminal or access device) approach or exceed the preset warning threshold during the transmission process. Specifically, when the transmission performance monitored by the device reaches or approaches the predetermined service quality (QoS) lower limit, the system will trigger a warning signal, indicating that the transmission quality of the current link may not meet the requirements of real-time audio and video content transmission, resulting in quality degradation or transmission delay.

[0025] High-dynamic speech characteristics refer to the fact that the speech content in the audio signal has large fluctuations and changes, such as changes in speech speed, volume fluctuations, and ups and downs in tone. Compared with other types of audio and video clips, high-dynamic speech clips have more stringent requirements on link performance, because these clips are particularly sensitive to factors such as delay, packet loss, and bandwidth fluctuations. Any slight transmission problem will have a significant impact on the continuity and clarity of the speech. Therefore, for such clips, it is urgent to ensure their priority in the transmission process by dynamically scheduling priority values ​​to ensure their transmission quality.

[0026] The initial scheduling priority value is derived based on the characteristics of the target audio and video clip and the current status of the link. This value is usually set according to the type of audio and video data, transmission conditions, and device performance. In the prior art, the acquisition of scheduling priority values ​​generally relies on static settings or link status evaluations, such as assigning an initial priority value to each clip through a fixed algorithm or rule based on standards such as network bandwidth, latency, and packet loss.

[0027] Associated devices refer to devices that participate in the transmission, processing, and scheduling of data streams during the transmission of target audio and video clips. Specifically, associated devices can be user terminal devices (such as smartphones, computers, tablets, etc.), network access devices (such as routers, switches, base stations, etc.), and transmission devices (such as video encoders, transmission servers, CDN nodes, etc.). These devices are responsible for the transmission of audio and video data streams, and their link status directly affects the transmission quality of audio and video data.

[0028] Historical transmission data records are usually collected through network monitoring systems, link management platforms, or the monitoring modules built into the device. These records contain information about the transmission performance of the link, including but not limited to the following types of data: packet loss rate, i.e., the proportion of packets lost per unit time; latency, i.e., the transmission time of packets from the source device to the destination device; bandwidth fluctuation, i.e., the fluctuation range of network bandwidth per unit time; number of retransmissions, i.e., the number of times a packet is retransmitted; throughput, i.e., the amount of data successfully transmitted per unit time; and link load, i.e., the utilization rate of the link.

[0029] Further, the AI-based audio and video data processing method further includes the following steps: Step S200: Analyze the historical transmission data record, and screen out a number of first-class historical segments that are consistent with the environmental characteristics of the target audio and video segment and belong to the high-dynamic voice characteristic type, and a number of second-class historical segments that do not belong to the high-dynamic voice characteristic type.

[0030] In the embodiments of the present invention, the historical transmission data record not only includes conventional link performance information, but also contains environmental characteristic data for determining whether the target audio and video segment belongs to the high-dynamic voice characteristic type. These environmental characteristic data can be obtained by detailed analysis of the voice signal of the target audio and video segment. Specifically, it includes spectrum analysis data for evaluating the frequency distribution and frequency change trend of the audio signal; volume fluctuation data for detecting the volume change of the audio signal in the time dimension, especially when the volume fluctuates greatly in speech; and speech rate change data for measuring the change rate of the speech rate in the speech segment, and a rapidly changing speech rate is usually associated with high-dynamic voice characteristic segments.

[0031] Through comprehensive analysis of these environmental characteristic data, the system can determine whether the target audio and video segment belongs to the high-dynamic voice characteristic type. Next, the system screens out a number of historical segments that are consistent with the environmental characteristics of the target segment from the historical transmission data record. The screening process uses a similarity calculation method. For example, by calculating the similarity between the target segment and the historical segment in terms of spectrum characteristics, volume fluctuation, and speech rate change, the eligible historical segments are identified. The system further classifies these segments. First, it screens out the first-class historical segments that meet the characteristics of the target segment and belong to the high-dynamic voice characteristic type, and then screens out the second-class historical segments that do not have the high-dynamic voice characteristic type with the target segment.

[0032] This process realizes the precise screening of historical segments through dynamic feature matching, similarity calculation, and classification marking, enabling the system to provide a more accurate transmission scheduling optimization scheme based on the environmental characteristics and dynamic voice characteristic type of the target audio and video segment.

[0033] Further, the AI-based audio and video data processing method further includes the following steps: Step S300: Respectively count the occurrence probabilities of the link transmission performance degradation exceeding the preset threshold standard in a number of first-class historical segments and a number of second-class historical segments, and determine the first correction parameter based on the deviation amplitude between the two.

[0034] Specifically, Figure 2 shows a flowchart for determining the first correction parameter.

[0035] Among them, specifically including the following steps to respectively count the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in several first-type historical segments and several second-type historical segments, and determine the first correction parameter based on the deviation amplitude between the two: Step S301: Analyze each first-type historical segment and second-type historical segment in sequence, extract the link transmission degradation event data recorded in each historical segment, so as to determine the corresponding link transmission performance degradation amplitude; Step S302: Respectively count the number of segments in several first-type historical segments and several second-type historical segments where the link transmission performance degradation amplitude exceeds the preset threshold standard, and calculate their corresponding occurrence probabilities respectively; Step S303: Determine the first correction parameter based on the deviation amplitude of the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first-type historical segments and the second-type historical segments.

[0036] The link transmission degradation event data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude, and the link transmission performance degradation amplitude is calculated by weighted combination of the packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude.

[0037] In the embodiment of the present invention, the specific implementation process of step S302 is as follows: First, the system analyzes each first-type historical segment and second-type historical segment in sequence, and extracts the link transmission performance degradation event data in each segment. These data include packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude. By performing weighted combination calculation on these data, the system obtains the link transmission performance degradation amplitude of each segment.

[0038] Then, the system statistically analyzes all first-type historical segments and second-type historical segments, and calculates the number of segments in which the link transmission performance degradation amplitude exceeds the preset threshold standard. Next, the system calculates the probability of the segments exceeding the threshold standard appearing in each type of segment. Specifically, the system divides the number of segments in each type of segment where the link transmission performance degradation amplitude exceeds the threshold by the total number of segments in that type, so as to obtain its occurrence probability.

[0039] Next, the system determines the first correction parameter based on the difference between the occurrence probabilities of the link transmission performance degradation amplitude exceeding the threshold standard in the first type of historical segments and the second type of historical segments. This difference can be obtained by calculating the deviation of the occurrence probabilities of the two. If the occurrence probability of the link transmission performance degradation amplitude exceeding the threshold standard in the first type of historical segments is greater than that in the second type of historical segments, then the deviation is larger, indicating that the first type of segments (i.e., high-dynamic voice segments) are more likely to experience performance degradation during link transmission, and the system needs to adjust the priority accordingly. If the probability difference between the two is small, it means that the difference in link performance between high-dynamic voice segments and ordinary segments is not significant, and the corresponding priority adjustment will be smaller.

[0040] The significance and technical effect of this process are that by comparing the difference in the occurrence probabilities of link transmission performance degradation between the first type and the second type of historical segments, the system can more accurately measure the difference in link transmission performance between high-dynamic voice segments and other segments. This enables the system to dynamically adjust the scheduling priorities of these segments according to the actual situation, ensuring that high-dynamic voice segments are preferentially processed when the link conditions are poor, thereby improving the audio-visual transmission quality and user experience.

[0041] Furthermore, the AI-based audio-visual data processing method further includes the following steps: Step S400, calculate the reconstruction cost value corresponding to each first type of historical segment, and determine the second correction parameter according to the change trend characteristics of the reconstruction cost value.

[0042] Specifically, Figure 3 shows the flowchart for determining the second correction parameter.

[0043] Among them, calculating the reconstruction cost value corresponding to each first type of historical segment and determining the second correction parameter according to the change trend characteristics of the reconstruction cost value specifically includes the following steps: Step S401, parse the link transmission data recorded in each first type of historical segment, extract the actual retransmission times and transmission delay corresponding to each first type of historical segment, and calculate the reconstruction cost value corresponding to each first type of historical segment based on the actual retransmission times and transmission delay; Step S402, count the reconstruction cost values corresponding to all first type of historical segments, and construct a change curve of the reconstruction cost value in the order of segment time; Step S403, calculate the average slope of the change curve of the reconstruction cost value, and use this average slope as the second correction parameter.

[0044] The actual retransmission times are obtained from the number of data retransmissions that occur during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and the receiving time of each first-type historical segment; the reconstruction cost value is calculated by a weighted linear combination of the actual retransmission times and the transmission delay.

[0045] In the embodiment of the present invention, the specific implementation process of step S401 is as follows: First, the system parses the link transmission data recorded in each first-type historical segment and extracts the actual retransmission times and the transmission delay corresponding to each segment. The actual retransmission times are obtained by monitoring the number of packet retransmissions during the link transmission process, that is, by calculating the retransmission event frequency of each segment. The transmission delay is obtained by calculating the time difference between the data sending time and the receiving time of each historical segment, which usually represents the time required for a packet to be transmitted in the network.

[0046] Next, the system calculates the reconstruction cost value of each first-type historical segment by performing a weighted linear combination of the actual retransmission times and the transmission delay. Before calculating the reconstruction cost value of each first-type historical segment, the system first needs to perform a normalization process on the retransmission times and the transmission delay. Since the retransmission times and the transmission delay are two data with different dimensions and different magnitudes, directly performing a weighted linear combination may cause one item of data to have too much influence on the final result. Therefore, a normalization process is required to ensure that these two items of data are weighted on the same scale.

[0047] After completing the normalization process, the system weights the retransmission times and the transmission delay according to a preset weight factor and sums them to obtain the reconstruction cost value of each segment. Through this method of weighted linear combination, the system can comprehensively consider the influence of the link stability (reflected by the retransmission times) and the transmission delay (reflected by the transmission delay) on the reconstruction cost, so as to calculate the reconstruction difficulty and the required resource consumption of each segment.

[0048] In step S402, the system counts the reconstruction cost values of all first-type historical segments and constructs a change curve of the reconstruction cost values according to the time sequence of the segments. Specifically, the system arranges the reconstruction cost values of each historical segment in chronological order, thereby obtaining a curve reflecting the change of the reconstruction cost over time. This curve can show the fluctuation of the reconstruction demand in each time period during the transmission process.

[0049] In step S403, the system calculates the average slope of the reconstruction cost value change curve and uses this average slope as the second correction parameter. By calculating the average slope of the curve, the system can obtain the overall trend of the reconstruction cost changing over time. If the slope of the curve is positive, it indicates that as time goes by, the reconstruction cost value gradually increases, suggesting that the quality of the transmission link may deteriorate, resulting in a higher reconstruction cost; if the slope of the curve is negative, it means that the reconstruction cost gradually decreases, possibly indicating an improvement in the link quality and a reduction in the reconstruction cost. The value of this average slope, as the second correction parameter, can reflect the volatility of the link during the transmission process and its impact on the reconstruction cost.

[0050] The significance of calculating the average slope of the reconstruction cost value change curve and using it as the second correction parameter is that the average slope can effectively reflect the changing trend of the link transmission quality. Through this method, the system can not only understand the current state of the link in real time but also predict future link changes based on historical data. This trend-based adjustment parameter can help the system adjust priorities in real time in a dynamic network environment, ensuring that audio and video segments that may require more resources for reconstruction during transmission can be processed preferentially, thereby optimizing the efficiency and quality of the entire transmission process.

[0051] Furthermore, the AI-based audio and video data processing method further includes the following steps: Step S500, modifying the initial scheduling priority value by combining the first correction parameter and the second correction parameter.

[0052] Specifically, Figure 4 shows a flowchart for modifying the initial scheduling priority value.

[0053] Among them, modifying the initial scheduling priority value by combining the first correction parameter and the second correction parameter specifically includes the following steps: Step S501, retrieving the preset scheduling priority value correction formula, substituting the first correction parameter and the second correction parameter into the scheduling priority value correction formula, and modifying the initial scheduling priority value to obtain the modified scheduling priority value; Step S502, applying the modified scheduling priority value to the link resource scheduling strategy of the target video segment to dynamically adjust the scheduling priority level of the target video segment in the transmission link.

[0054] The scheduling priority value correction formula is: ; where refers to the modified scheduling priority value, refers to the initial scheduling priority value, Refers to the first correction parameter, that is, the deviation amplitude of the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first type of historical segments and the second type of historical segments. Refers to the adjustment coefficient corresponding to the first correction parameter. Refers to the second correction parameter, that is, the average slope of the reconstructed cost value change curve. Refers to the adjustment coefficient corresponding to the second correction parameter. In the scheduling priority value correction formula, , where Refers to the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first type of historical segments. Refers to the occurrence probability of the link transmission performance degradation amplitude exceeding the preset threshold standard in the second type of historical segments. Is a preset minimum protection value, used to avoid calculation anomalies when the denominator is zero or close to zero.

[0055] In the embodiments of the present invention, the purpose of comprehensively correcting by combining the first correction parameter and the second correction parameter is to be able to comprehensively adjust the scheduling priority value of the target audio-visual segment to cope with the changes in the network link and the transmission requirements of different types of audio-visual segments. The first correction parameter is mainly based on the real-time status of the link transmission performance, such as factors like packet loss rate and latency. It reflects the quality of the current link. Especially when there are large fluctuations in the network during the transmission process, it can timely adjust the priority of the segment. The second correction parameter is based on the trend in historical data, such as the change trend of the reconstructed cost value. It captures the fluctuations of the network link in the time dimension.

[0056] Through this comprehensive correction, the system can more flexibly cope with sudden link fluctuations and long-term trend changes based on historical data, thereby more accurately adjusting the transmission priority to ensure that high-dynamic voice segments are preferentially guaranteed in an unstable network environment. The combination of the first correction parameter and the second correction parameter provides a solution that balances real-time and historical factors, avoiding the deviation that may be caused by a single factor. In this way, the system can more accurately judge the transmission priority of each audio-visual segment, improving the overall transmission efficiency and stability.

[0057] For example, assume that the system is transmitting a high-dynamic voice segment. First, the system calculates the first correction parameter according to the real-time link status. If the link has large bandwidth fluctuations, a high packet loss rate, and an increase in latency, the system will increase the priority of this segment to ensure its preferential transmission. Then, the system calculates the second correction parameter by analyzing historical data, especially the change trend of the reconstructed cost. If the historical data shows that the reconstructed cost shows an upward trend over time, then the system will also increase the priority of this segment to ensure its smooth transmission.

[0058] Finally, the corrected scheduling priority value is obtained by combining these two correction parameters into the calculation of the initial priority value to derive a new scheduling priority. In this way, the system can dynamically adjust the scheduling priority according to the real-time link status and historical transmission data, ensuring that high-dynamic voice segments can still be prioritized even when network conditions are poor, thereby improving the overall audio-visual transmission quality.

[0059] The scheduling priority value correction formula proposed in the present invention is only an intuitive and effective solution, which can be further optimized through other different calculation methods. For example, in some systems, the correction coefficient can be adjusted according to different scenarios or real-time link conditions to make the correction process more flexible and refined. By introducing technologies such as machine learning, the system can also adaptively adjust these correction coefficients based on historical data, thereby achieving a higher optimization effect.

[0060] Furthermore, Figure 5 The application architecture diagram of the system provided by the embodiments of the present invention is shown.

[0061] Among them, in another preferred embodiment provided by the present invention, an AI-based audio-visual data processing system includes: A data acquisition module 100, configured to, after determining that the target audio-visual segment belongs to the high-dynamic voice characteristic type, acquire the initial scheduling priority value generated for the target audio-visual segment in the early warning state at the end of the link, and acquire the historical transmission data record of the associated device carrying the transmission task of the target audio-visual segment.

[0062] In the embodiments of the present invention, with the wide application of audio-visual transmission technology, the quality and stability of the link have become key factors affecting the user experience. Especially during the transmission of real-time audio-visual data, fluctuations in link performance (such as packet loss, latency, bandwidth fluctuations, etc.) directly affect the transmission quality of audio-visual content. Therefore, in the face of unstable links, the method of dynamically adjusting the scheduling priority value is particularly important. The AI-based audio-visual data processing method of the present invention can dynamically adjust the scheduling priority value by intelligently evaluating the characteristics of the target audio-visual segment and the link status, thereby ensuring the priority transmission of high-dynamic audio-visual segments and optimizing the overall transmission quality.

[0063] The link end warning state refers to a warning state triggered by the system when the link performance indicators (such as bandwidth, delay, packet loss rate, etc.) of the end device of the link in the network (such as user terminal or access device) approach or exceed the preset warning threshold during the transmission process. Specifically, when the transmission performance monitored by the device reaches or approaches the predetermined service quality (QoS) lower limit, the system will trigger a warning signal, indicating that the transmission quality of the current link may not meet the requirements of real-time audio and video content transmission, resulting in quality degradation or transmission delay.

[0064] High-dynamic speech characteristics refer to the fact that the speech content in the audio signal has large fluctuations and changes, such as changes in speech speed, volume fluctuations, and ups and downs in tone. Compared with other types of audio and video clips, high-dynamic speech clips have more stringent requirements on link performance, because these clips are particularly sensitive to factors such as delay, packet loss, and bandwidth fluctuations. Any slight transmission problem will have a significant impact on the continuity and clarity of the speech. Therefore, for such clips, it is urgent to ensure their priority in the transmission process by dynamically scheduling priority values ​​to ensure their transmission quality.

[0065] The initial scheduling priority value is derived based on the characteristics of the target audio and video clip and the current status of the link. This value is usually set according to the type of audio and video data, transmission conditions, and device performance. In the prior art, the acquisition of scheduling priority values ​​generally relies on static settings or link status evaluations, such as assigning an initial priority value to each clip through a fixed algorithm or rule based on standards such as network bandwidth, latency, and packet loss.

[0066] Associated devices refer to devices that participate in the transmission, processing, and scheduling of data streams during the transmission of target audio and video clips. Specifically, associated devices can be user terminal devices (such as smartphones, computers, tablets, etc.), network access devices (such as routers, switches, base stations, etc.), and transmission devices (such as video encoders, transmission servers, CDN nodes, etc.). These devices are responsible for the transmission of audio and video data streams, and their link status directly affects the transmission quality of audio and video data.

[0067] Furthermore, the AI-based audio and video data processing system also includes: The data screening module 200 is used to parse the historical transmission data records, and screen out a number of first-category historical segments that are consistent with the target audio and video segment environmental characteristics and belong to the high dynamic voice characteristic type, and a number of second-category historical segments that do not belong to the high dynamic voice characteristic type.

[0068] In the embodiments of the present invention, the historical transmission data record not only includes conventional link performance information, but also contains environmental feature data used to determine whether the target audio-visual segment belongs to the high-dynamic voice characteristic type. These environmental feature data can be obtained by detailed analysis of the voice signal of the target audio-visual segment. Specifically, it includes spectral analysis data for evaluating the frequency distribution and frequency change trend of the audio signal; volume fluctuation data for detecting the volume change of the audio signal in the time dimension, especially when the volume fluctuates greatly in the voice; and speech rate change data for measuring the change rate of the speech rate in the speech segment, and a rapidly changing speech rate is usually associated with high-dynamic voice characteristic segments.

[0069] Through comprehensive analysis of these environmental feature data, the system can determine whether the target audio-visual segment belongs to the high-dynamic voice characteristic type. Next, the system filters out several historical segments from the historical transmission data record that are consistent with the environmental features of the target segment. The filtering process uses a similarity calculation method. For example, by calculating the similarity between the target segment and the historical segment in terms of spectral features, volume fluctuations, and speech rate changes, the historical segments that meet the conditions are identified. The system further classifies these segments. First, it filters out the first type of historical segments that meet the characteristics of the target segment and belong to the high-dynamic voice characteristic type, and then filters out the second type of historical segments that do not have the high-dynamic voice characteristic type as the target segment.

[0070] This process realizes the precise filtering of historical segments through dynamic feature matching, similarity calculation, and classification marking, enabling the system to provide a more accurate transmission scheduling optimization scheme based on the environmental features and dynamic voice characteristic type of the target audio-visual segment.

[0071] Furthermore, the AI-based audio-visual data processing system further includes: A first correction parameter determination module 300, which is used to respectively count the occurrence probabilities of the link transmission performance deterioration amplitude exceeding the preset threshold standard in several first type of historical segments and several second type of historical segments, and determine the first correction parameter based on the deviation amplitude between the two.

[0072] Specifically, Figure 6 The structural block diagram of the first correction parameter determination module 300 in the system provided by the embodiments of the present invention is shown.

[0073] Among them, in the preferred embodiment provided by the present invention, the first correction parameter determination module 300 specifically includes: A historical segment analysis unit 301 is configured to sequentially analyze each first - type historical segment and second - type historical segment, extract the link transmission degradation event data recorded in each historical segment, and determine the corresponding link transmission performance degradation amplitude; the link transmission degradation event data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude, and the link transmission performance degradation amplitude is obtained by weighted combination calculation of the packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude; An occurrence probability calculation unit 302 is configured to respectively count the number of segments in a number of first - type historical segments and a number of second - type historical segments where the link transmission performance degradation amplitude exceeds a preset threshold standard, and calculate their corresponding occurrence probabilities respectively; A deviation amplitude determination unit 303 is configured to determine a first correction parameter based on the deviation amplitude of the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first - type historical segments and the second - type historical segments.

[0074] In an embodiment of the present invention, the specific working process of the historical segment analysis unit 301 is as follows: First, the system sequentially analyzes each first - type historical segment and second - type historical segment, and extracts the link transmission performance degradation event data in each segment. This data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude. By performing weighted combination calculation on this data, the system obtains the link transmission performance degradation amplitude of each segment.

[0075] Then, the system statistically analyzes all first - type historical segments and second - type historical segments, calculates the number of segments where the link transmission performance degradation amplitude exceeds the preset threshold standard. Next, the system calculates the probability of occurrence of the segments exceeding the threshold standard in each type of segment. Specifically, the system divides the number of segments in each type of segment where the link transmission performance degradation amplitude exceeds the threshold by the total number of segments in that type, so as to obtain its occurrence probability.

[0076] Next, the system determines the first correction parameter according to the difference between the occurrence probabilities of the link transmission performance degradation amplitude exceeding the threshold standard in the first - type historical segments and the second - type historical segments. This difference can be obtained by calculating the deviation of the two occurrence probabilities. If the occurrence probability of the link transmission performance degradation amplitude exceeding the threshold standard in the first - type historical segments is greater than that in the second - type historical segments, then the deviation is larger, indicating that the first - type segments (i.e., high - dynamic voice segments) are more likely to have performance degradation during link transmission, and the system needs to adjust the priority accordingly. If the probability difference between the two is small, it means that the difference in link performance between high - dynamic voice segments and ordinary segments is small, and the corresponding priority adjustment will be small.

[0077] The significance and technical effects of this process are that by comparing the difference in the occurrence probability of transmission performance degradation between the first type and the second type of historical segments, the system can more accurately measure the performance difference between high-dynamic voice segments and other segments during link transmission. This enables the system to dynamically adjust the scheduling priorities of these segments according to the actual situation, ensuring that high-dynamic voice segments are preferentially processed when the link conditions are poor, thereby improving the audio-visual transmission quality and user experience.

[0078] Furthermore, the AI-based audio-visual data processing system further includes: A second correction parameter determination module 400, configured to calculate the reconstruction cost value corresponding to each first type of historical segment, and determine the second correction parameter according to the change trend feature of the reconstruction cost value.

[0079] Specifically, Figure 7 FIG. shows the structural block diagram of the second correction parameter determination module 400 in the system provided by the embodiment of the present invention.

[0080] Among them, in the preferred embodiment provided by the present invention, the second correction parameter determination module 400 specifically includes: A transmission data analysis unit 401, configured to analyze the link transmission data recorded in each first type of historical segment, extract the actual retransmission times and transmission delays corresponding to each first type of historical segment, and calculate the reconstruction cost value corresponding to each first type of historical segment based on the actual retransmission times and transmission delays; the actual retransmission times are obtained from the number of data retransmissions occurring during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and the receiving time of each first type of historical segment; the reconstruction cost value is calculated by a weighted linear combination of the actual retransmission times and the transmission delays; A change curve construction unit 402, configured to count the reconstruction cost values corresponding to all first type of historical segments, and construct a change curve of the reconstruction cost values in the order of segment time; An average slope calculation unit 403, configured to calculate the average slope of the change curve of the reconstruction cost values, and use this average slope as the second correction parameter.

[0081] In the embodiment of the present invention, the specific working process of the transmission data analysis unit 401 is as follows: First, the system analyzes the link transmission data recorded in each first type of historical segment, and extracts the actual retransmission times and transmission delays corresponding to each segment. The actual retransmission times are obtained by monitoring the number of packet retransmissions during the link transmission process, that is, calculating the retransmission event frequency of each segment. The transmission delay is obtained by calculating the time difference between the data sending time and the receiving time of each historical segment, and usually represents the time required for the packet to be transmitted in the network.

[0082] Next, the system calculates the reconstruction cost value of each first - type historical segment by performing a weighted linear combination of the actual re - transmission times and the transmission delay. Before calculating the reconstruction cost value of each first - type historical segment, the system first needs to normalize the re - transmission times and the transmission delay. Since the re - transmission times and the transmission delay are two data with different dimensions and magnitudes, directly performing a weighted linear combination may cause one item of data to have too much influence on the final result. Therefore, normalization is required to ensure that these two items of data are weighted on the same scale.

[0083] After completing the normalization process, the system weights the re - transmission times and the transmission delay according to the preset weight factors and sums them to obtain the reconstruction cost value of each segment. Through this method of weighted linear combination, the system can comprehensively consider the influence of the link stability (reflected by the re - transmission times) and the transmission delay (reflected by the transmission delay) on the reconstruction cost, thereby calculating the reconstruction difficulty of each segment and the required resource consumption.

[0084] In the change - curve construction unit 402, the system counts the reconstruction cost values of all first - type historical segments and constructs a change curve of the reconstruction cost values according to the time sequence of the segments. Specifically, the system arranges the reconstruction cost values of each historical segment in chronological order, thereby obtaining a curve that reflects the change of the reconstruction cost over time. This curve can show the fluctuation of the reconstruction demand in each time period during the transmission process.

[0085] In the average - slope calculation unit 403, the system calculates the average slope of the change curve of the reconstruction cost value and uses this average slope as the second correction parameter. By calculating the average slope of the curve, the system can obtain the overall trend of the change of the reconstruction cost over time. If the slope of the curve is positive, it means that as time goes by, the reconstruction cost value gradually increases, indicating that the quality of the transmission link may deteriorate, resulting in a higher reconstruction cost; if the slope of the curve is negative, it means that the reconstruction cost gradually decreases, which may mean that the link quality improves and the reconstruction cost decreases. This value of the average slope, as the second correction parameter, can reflect the volatility of the link during the transmission process and its influence on the reconstruction cost.

[0086] Furthermore, the AI - based audio - video data processing system further includes: A correction module 500, configured to correct the initial scheduling priority value by combining the first correction parameter and the second correction parameter.

[0087] In the embodiments of the present invention, the purpose of comprehensively correcting by combining the first correction parameter and the second correction parameter is to be able to comprehensively adjust the scheduling priority value of the target audio-visual segment to cope with the changes in the network link and the transmission requirements of different types of audio-visual segments. The first correction parameter is mainly based on the real-time status of the link transmission performance, such as factors like packet loss rate and latency. It reflects the quality of the current link. Especially when there are large fluctuations in the network during the transmission process, it can timely adjust the priority of the segment. The second correction parameter is based on the trend in historical data, such as the change trend of the reconstruction cost value. It captures the fluctuations of the network link in the time dimension.

[0088] Through this comprehensive correction, the system can more flexibly cope with sudden link fluctuations and long-term trend changes based on historical data, thereby more accurately adjusting the transmission priority to ensure that high-dynamic voice segments are preferentially guaranteed when the network environment is unstable. The combination of the first correction parameter and the second correction parameter provides a solution that balances real-time and historical factors, avoiding the deviation that may be caused by a single factor. In this way, the system can more accurately judge the transmission priority of each audio-visual segment, improving the overall transmission efficiency and stability.

[0089] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI-based audio and video data processing method, characterized in that, The method includes: After determining that the target audio - video segment belongs to the high - dynamic voice characteristic type, obtaining the initial scheduling priority value in the early - warning state at the end of the link generated for the target audio - video segment, and obtaining the historical transmission data record of the associated device carrying the transmission task of the target audio - video segment; Analyzing the historical transmission data record, screening out several first - type historical segments that are consistent with the environmental characteristics of the target audio - video segment and belong to the high - dynamic voice characteristic type, and several second - type historical segments that do not belong to the high - dynamic voice characteristic type; Respectively counting the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in several first - type historical segments and several second - type historical segments, and determining the first correction parameter based on the deviation amplitude between the two; Calculating the reconstruction cost value corresponding to each first - type historical segment, and determining the second correction parameter according to the change trend characteristics of the reconstruction cost value; Combining the first correction parameter and the second correction parameter to correct the initial scheduling priority value.

2. The AI-based audio and video data processing method according to claim 1, wherein The steps of respectively counting the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in several first - type historical segments and several second - type historical segments, and determining the first correction parameter based on the deviation amplitude between the two include: Sequentially analyzing each first - type historical segment and second - type historical segment, extracting the link transmission degradation event data recorded in each historical segment to determine the corresponding link transmission performance degradation amplitude; Counting the number of segments with the link transmission performance degradation amplitude exceeding the preset threshold standard in several first - type historical segments and several second - type historical segments respectively, and calculating their corresponding occurrence probabilities respectively; Determining the first correction parameter based on the deviation amplitude of the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first - type historical segments and the second - type historical segments.

3. The AI-based audio and video data processing method according to claim 2, wherein The link transmission degradation event data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude, and the link transmission performance degradation amplitude is calculated by weighted combination of the packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude.

4. The AI-based audio and video data processing method according to claim 1, wherein The steps of calculating the reconstruction cost value corresponding to each first - type historical segment, and determining the second correction parameter according to the change trend characteristics of the reconstruction cost value include: Analyzing the link transmission data recorded in each first - type historical segment, extracting the actual re - transmission times and transmission delay corresponding to each first - type historical segment, and calculating the reconstruction cost value corresponding to each first - type historical segment based on the actual re - transmission times and transmission delay; Counting the reconstruction cost values corresponding to all first - type historical segments, and constructing a reconstruction cost value change curve according to the segment time sequence; Calculating the average slope of the reconstruction cost value change curve, and using this average slope as the second correction parameter.

5. The AI-based audio and video data processing method according to claim 4, wherein The actual re - transmission times are obtained from the number of data re - transmissions occurring during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and receiving time of each first - type historical segment; The reconstruction cost value is calculated by weighted linear combination of the actual re - transmission times and the transmission delay.

6. The AI-based audio and video data processing method according to claim 4, wherein The steps of combining the first correction parameter and the second correction parameter to correct the initial scheduling priority value include: Retrieve the preset scheduling priority value correction formula, substitute the first correction parameter and the second correction parameter into the scheduling priority value correction formula, and correct the initial scheduling priority value to obtain the corrected scheduling priority value; Apply the corrected scheduling priority value to the link resource scheduling strategy of the target video segment to dynamically adjust the scheduling priority level of the target video segment in the transmission link.

7. The AI-based audio and video data processing method according to claim 6, wherein The scheduling priority value correction formula is: ; Among them refers to the corrected scheduling priority value; refers to the initial scheduling priority value, refers to the first correction parameter, that is, the deviation amplitude of the occurrence probability that the link transmission performance degradation amplitude in the first type of historical segment and the second type of historical segment exceeds the preset threshold standard, refers to the adjustment coefficient corresponding to the first correction parameter, refers to the second correction parameter, that is, the average slope of the reconstruction cost value change curve, refers to the adjustment coefficient corresponding to the second correction parameter; In the scheduling priority value correction formula, ; wherein refers to the occurrence probability that the link transmission performance degradation amplitude in the first type of historical segments exceeds the preset threshold standard, refers to the occurrence probability that the link transmission performance degradation amplitude in the second type of historical segments exceeds the preset threshold standard, is a preset minimum protection value, used to avoid abnormal calculation when the denominator is zero or close to zero.

8. An AI-based audio and video data processing system, characterized in that, The system includes: a data acquisition module, a data screening module, a first correction parameter determination module, a second correction parameter determination module, and a correction module, where: The data acquisition module is used to obtain the initial scheduling priority value in the link end warning state generated for the target audio-visual segment and the historical transmission data record of the associated device carrying the transmission task of the target audio-visual segment after determining that the target audio-visual segment belongs to the high-dynamic voice characteristic type; The data screening module is used to analyze the historical transmission data record, screen out several first-class historical segments that are consistent with the environmental characteristics of the target audio-visual segment and belong to the high-dynamic voice characteristic type, and several second-class historical segments that do not belong to the high-dynamic voice characteristic type; The first correction parameter determination module is used to respectively count the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in several first-class historical segments and several second-class historical segments, and determine the first correction parameter based on the deviation amplitude between the two; The second correction parameter determination module is used to calculate the reconstruction cost value corresponding to each first-class historical segment and determine the second correction parameter according to the change trend characteristics of the reconstruction cost value; The correction module is used to correct the initial scheduling priority value by combining the first correction parameter and the second correction parameter.

9. The AI-based audio and video data processing system according to claim 8, wherein The first correction parameter determination module specifically includes: The historical segment analysis unit is used to sequentially analyze each first-class historical segment and second-class historical segment, extract the link transmission degradation event data recorded in each historical segment to determine the corresponding link transmission performance degradation amplitude; the link transmission degradation event data includes packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude, and the link transmission performance degradation amplitude is calculated by weighted combination of the packet loss rate, delay jitter amplitude, and instantaneous bandwidth fluctuation amplitude; The occurrence probability calculation unit is used to respectively count the number of segments with the link transmission performance degradation amplitude exceeding the preset threshold standard in several first-class historical segments and several second-class historical segments, and calculate their corresponding occurrence probabilities respectively; The deviation amplitude determination unit is used to determine the first correction parameter based on the deviation amplitude of the occurrence probabilities of the link transmission performance degradation amplitude exceeding the preset threshold standard in the first-class historical segments and the second-class historical segments.

10. The AI-based audio and video data processing system according to claim 9, wherein, The second correction parameter determination module specifically includes: A transmission data parsing unit, configured to parse the link transmission data recorded in each first - type historical segment, extract the actual re - transmission times and transmission delays corresponding to each first - type historical segment, and calculate the reconstruction cost value corresponding to each first - type historical segment based on the actual re - transmission times and transmission delays; the actual re - transmission times are obtained from the number of data re - transmissions that occur during the link transmission process, and the transmission delay is calculated from the time difference between the data sending time and the receiving time of each first - type historical segment; the reconstruction cost value is calculated by a weighted linear combination of the actual re - transmission times and the transmission delays; A change curve construction unit, configured to count the reconstruction cost values corresponding to all first - type historical segments and construct a change curve of the reconstruction cost values in the order of segment time; An average slope calculation unit, configured to calculate the average slope of the change curve of the reconstruction cost values and use this average slope as the second correction parameter.

Citation Information

Patent Citations

  • Low-delay audio and video real-time transmission system based on 5G

    CN118301139A

  • Link fault analysis method for digital audio and video signal transmission processing equipment

    CN118646860A

  • Data flow priority management method and system of optical communication device

    CN118764438A

  • Intelligent optimization system for video monitoring traffic

    CN119788617A

  • Scheduling method applied in industrial heterogeneous network in which TSN and non-TSN are interconnected

    US20220353195A1