Audio and video transmission method and system for online conference, electronic equipment and storage medium

By configuring the transmission priority of audio and video data on the main speaker and determining the transmission channel and encoding strategy based on the network state prediction model, the key information loss problem of traditional audio and video transmission strategies in a weak network environment is solved, and the reliable and real-time transmission of audio and video data is achieved.

CN120075387AActive Publication Date: 2025-05-30BEIJING ZHENSHITONG DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510505006.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-30
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the case of unstable network environment, traditional audio and video transmission strategies have problems such as indiscriminate degradation, lack of priority classification and insufficient redundancy mechanism, resulting in the loss of key information.

Method used

By configuring the transmission priority of audio and video data on the main speaker, and predicting future network status based on the network state prediction model, determining the transmission channel and dynamic encoding strategy that adapt to transmission priority can realize multi-path transmission of audio and video data.

Benefits of technology

In a weak network state, it ensures the accurate transmission of key information in audio and video data, solves the problem of loss of key information, and ensures the reliability and real-timeness of audio and video data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075387A_ABST
    Figure CN120075387A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video transmission method and system for an online conference, electronic equipment and a storage medium, and relates to the technical field of audio and video data transmission. The audio and video transmission method comprises the following steps: configuring the transmission priority of audio and video data of a main speaking end according to a preset priority rule; predicting a network prediction state of the main speaking end in a preset future time period based on the network state prediction model; when the network prediction state is a weak network state, determining a transmission channel corresponding to the transmission priority and a dynamic coding strategy corresponding to the transmission priority; and synchronously transmitting the audio and video data of the corresponding transmission priority to the conference participating end through the corresponding transmission channel based on the dynamic coding strategy. According to the invention, accurate transmission of key information in the audio and video data can be ensured, and the problem of key information loss in a weak network state is solved. And through multi-path transmission, the reliability and real-time performance of audio and video data transmission can be ensured at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video data transmission. Specifically, it relates to an audio - video transmission method and system for online meetings, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of Internet technology and communication technology, application scenarios such as real - time audio - video communication, online education, and remote meetings have become increasingly popular.

[0003] However, the inventors of the present invention have found that in an unstable network environment (such as a weak network environment), traditional audio - video transmission strategies often cannot effectively guarantee the transmission quality of key audio - video information, resulting in the loss of key information.

[0004] For example, the inventors have found that in an unstable network environment, traditional audio - video transmission strategies (such as weak network degradation strategies) have problems such as undifferentiated degradation, lack of priority division, and insufficient redundancy mechanisms.

[0005] The content of the background art section is only the technology known to the applicant and does not necessarily represent the prior art in this field. Summary of the Invention

[0006] The present invention provides an audio - video transmission method and system for online meetings, an electronic device, and a storage medium, aiming to solve the problems of undifferentiated degradation, lack of priority division, and insufficient redundancy mechanisms existing in traditional audio - video transmission strategies in an unstable network environment.

[0007] According to one aspect of the present invention, there is provided an audio - video transmission method for an online meeting, including: configuring the transmission priority of the audio - video data of the host end according to a preset priority rule; predicting the network prediction state of the host end in a preset future time period based on a network state prediction model; in the case where the network prediction state is a weak network state, determining a transmission channel corresponding to the transmission priority and a dynamic encoding strategy corresponding to the transmission priority; and synchronously transmitting the audio - video data of the corresponding transmission priority to the participant end through the corresponding transmission channel based on the dynamic encoding strategy.

[0008] According to some embodiments of the present invention, predicting the network prediction state of the host end in a preset future time period based on a network state prediction model includes: determining the historical network data information of a preset historical time period; predicting the predicted bandwidth information of the preset future time period based on the network state prediction model according to the historical network data information; and in the case where the predicted bandwidth information is less than a preset threshold, determining the network prediction state as a weak network state.

[0009] According to some embodiments of the present invention, synchronously transmitting audio and video data with corresponding transmission priorities to the participating end through a corresponding transmission channel based on a dynamic coding strategy includes: dynamically adjusting the forward error correction redundancy of the audio and video data based on predicted bandwidth information.

[0010] According to some embodiments of the present invention, the audio and video transmission method further includes: dynamically determining the visual fixation area of the user of the participating end based on visual tracking; in the case where the network prediction state is a weak network state, increasing the display resolution of the visual fixation area to a preset resolution.

[0011] According to some embodiments of the present invention, the audio and video transmission method further includes: in the case where frame loss of the audio and video data received by the participating end is detected, determining the lost frame data based on the previous frame data and the subsequent frame data of the audio and video data received by the participating end.

[0012] According to some embodiments of the present invention, the audio and video transmission method further includes: in the case where the delay between the audio data and the video data in the audio and video data is greater than a preset delay threshold, dynamically stretching the audio data based on the audio waveform or compressing or doubling the speed of the video data so that the audio data and the video data are synchronously transmitted.

[0013] According to another aspect of the present invention, the present invention further provides an audio and video transmission method for an online meeting, including: configuring the transmission priorities of the audio and video data of the participating end according to a preset priority rule; predicting the network prediction state of the participating end in a preset future time period based on a network state prediction model; in the case where the network prediction state is a weak network state, determining a transmission channel corresponding to the transmission priority and a dynamic coding strategy corresponding to the transmission priority; synchronously transmitting the audio and video data with the corresponding transmission priority to the host end through the corresponding transmission channel based on the dynamic coding strategy.

[0014] According to another aspect of the present invention, the present invention further provides an audio and video transmission system for an online meeting, including a content grading determination module, a network state prediction module, a channel and strategy control module, and an audio and video data transmission module. The content grading determination module configures the transmission priorities of the audio and video data of the host end according to a preset priority rule; the network state prediction module predicts the network prediction state of the host end in a preset future time period based on a network state prediction model; the channel and strategy control module determines a transmission channel corresponding to the transmission priority and a dynamic coding strategy corresponding to the transmission priority in the case where the network prediction state is a weak network state; the audio and video data transmission module synchronously transmits the audio and video data with the corresponding transmission priority to the participating end through the corresponding transmission channel based on the dynamic coding strategy.

[0015] According to another aspect of the present invention, the present invention further provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, enable the one or more processors to implement the audio and video transmission method as described above.

[0016] According to another aspect of the present invention, the present invention further provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the audio and video transmission method as described above.

[0017] According to another aspect of the present invention, the present invention further provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, which when executed by a computer, cause the computer to execute the audio and video transmission method as described above.

[0018] Beneficial effects By configuring the transmission priorities of the audio and video data at the host end, in the case where the current network environment is predicted to be a weak network state, the transmission channels and dynamic encoding strategies corresponding to the transmission priorities can be determined, so that the synchronous transmission of audio and video data with different transmission priorities can be achieved based on the corresponding transmission channels and corresponding dynamic encoding strategies.

[0019] The present invention realizes the content classification of audio and video data according to the content importance of the audio and video data, and can achieve multi-path transmission of audio and video data with different transmission priorities based on the dynamic encoding strategy in a weak network state. Thus, the accurate transmission of key information in the audio and video data can be ensured, and the problem of loss of key information in a weak network state is solved. And through multi-path transmission, the reliability and real-time performance of audio and video data transmission can be ensured simultaneously. Description of the drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0021] Figure 1 A flowchart showing the audio and video transmission method according to an embodiment of the present invention; Figure 2 Another flowchart showing the audio and video transmission method according to an embodiment of the present invention; Figure 3 Another flowchart showing the audio and video transmission method according to an embodiment of the present invention; Figure 4 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention; Figure 5 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention; Figure 6 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention; Figure 7 A schematic structural diagram of the audio - video transmission system according to an embodiment of the present invention.

[0022] Explanation of reference numerals: Audio - video transmission system 1; Content classification determination module 10; Network status prediction module 20; Channel and policy control module 30; Audio - video data transmission module 40. Detailed implementation manners

[0023] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted.

[0024] The features, structures, or characteristics described may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of these specific details, or can be implemented in other ways, components, materials, devices, etc. In these cases, well - known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0025] In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0026] The terms "first", "second", etc. in the description and claims of the present invention and in the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order.

[0027] The following will clearly and completely describe the technical solutions of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0028] The inventors found that in the case of unstable network environments, the traditional audio and video transmission strategies have at least the following problems: 1. Indiscriminate degradation: Traditional weak network degradation strategies usually adopt an indiscriminate approach, that is, when the bandwidth is insufficient, the bit rate or resolution of all content is uniformly reduced. Since this strategy fails to distinguish the importance of content, key information (such as text, charts, etc.) and secondary content (such as background videos, decorative elements, etc.) are degraded equally.

[0029] For example, in scenarios such as online meetings or distance education, the key text information in a PPT (PowerPoint, presentation) may become blurred due to degradation, seriously affecting the information transmission effect.

[0030] 2. Lack of priority division: Since traditional audio and video transmission strategies do not effectively divide the priority of the transmitted content, high-importance content (such as voice, text, human pictures, etc.) cannot obtain sufficient transmission guarantee.

[0031] For example, in scenarios such as online meetings or distance education, the importance of voice and key frames of PPT is much higher than that of background videos, but traditional solutions fail to dynamically adjust the transmission strategy according to the importance of content, resulting in the loss or serious degradation of key information in a weak network environment.

[0032] 3. Insufficient redundancy mechanism: Usually relying on a single transmission path, such as only using the TCP (Transmission Control Protocol) or UDP (User Datagram Protocol) protocol during data transmission, lacking a multi-path redundant transmission and intelligent retransmission mechanism. When the network fluctuates or data packets are lost, key data cannot be quickly restored, resulting in information transmission interruption or quality degradation.

[0033] For example, in a weak network environment, although the TCP protocol can ensure the reliability of data, its retransmission mechanism may lead to increased latency; while the UDP protocol has low latency but lacks reliability guarantee, easily resulting in the loss of key data packets.

[0034] Therefore, the inventors believe that the existing audio and video transmission strategies in a weak network environment have limitations and cannot effectively guarantee the transmission quality of key information.

[0035] According to one aspect of the present invention, the present invention provides an audio and video transmission method for online meetings. Figure 1 A flowchart showing the audio and video transmission method according to an embodiment of the present invention is as follows. Figure 1 As shown, the audio and video transmission method may include steps S100 - S400.

[0036] Exemplarily, the audio and video transmission method may be executed by an audio and video transmission system (such as a server or a host, etc.) with computing capabilities.

[0037] An online meeting is a virtual meeting form that connects participants located in different geographical locations through Internet technology, enabling functions such as real - time audio and video communication, screen sharing, and file transfer.

[0038] According to an exemplary embodiment, an online meeting may include at least one host end and multiple participant ends. For example, the host end may be the terminal device where the host of the meeting is located; the participant ends may be the terminal devices where other participants in the meeting are located. In the present invention, the terminal device includes, but is not limited to, devices such as computers, mobile phones, and tablets, and the present invention does not limit this.

[0039] Exemplarily, the audio and video transmission method may be used for the host end. Hereinafter, the present invention will be described and introduced taking the host end as an example.

[0040] According to an exemplary embodiment, in step S100, the audio and video transmission system configures the transmission priority of the audio and video data of the host end according to a preset priority rule.

[0041] For example, the audio and video data of the host end may include audio data and video data. The audio data at least includes the sound data generated by the host end during the meeting, such as the voice signal of the host and other background sound signals, etc.; the video data at least includes the video data generated by the host end during the meeting, such as the real - time video image of the host, the shared screen page (such as display pages of PPT, Word documents, software operations, etc.), and other background images, etc.

[0042] The preset priority rule may be custom - set by the user according to actual needs. For example, the audio and video transmission system may respond to the user's configuration instruction, divide the audio and video data into priorities according to the importance of the data to achieve content classification. It can be understood here that the preset priority rule can be adjusted in real - time according to actual needs, rather than relying on a fixed rule. Such a setting can ensure the configuration flexibility of audio and video data transmission.

[0043] As an example, the audio - video transmission system can configure the voice signal of the main speaker and the shared screen page (such as display pages of PPT, Word documents, software operations, etc.) as the first transmission priority according to the preset priority rule, configure the real - time video image of the main speaker as the second transmission priority, and configure other background images as the third transmission priority.

[0044] For example, the audio - video transmission system can extract preset meeting content keywords (such as "budget", "deadline", etc.) through natural language processing technology and mark the keywords as the first transmission priority, etc. Another example is that the audio - video transmission system can automatically identify complex content such as charts and formulas in PPT based on convolutional neural networks, thereby ensuring the accuracy of content classification.

[0045] In step S200, the audio - video transmission system predicts the network prediction state of the main speaker end in a preset future time period based on the network state prediction model.

[0046] For example, the network prediction state can include a strong network state, a medium network state, and a weak network state. The network state prediction model can be an LSTM (Long Short - Term Memory) prediction model. In the prediction of the network state, LSTM can predict future bandwidth fluctuations by learning the change rules of historical data, thereby predicting the network state within a specific future time period.

[0047] Exemplarily, the audio - video transmission system can also optimize the input features of the LSTM prediction model by combining the real - time routing node state (such as congestion situation), thereby improving the prediction accuracy of the LSTM prediction model.

[0048] Figure 2 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention.

[0049] Optionally, as Figure 2 shown, step S200 may further include S210 - S30.

[0050] In step S210, the audio - video transmission system determines the historical network data information of a preset historical time period.

[0051] In step S220, the audio - video transmission system predicts the predicted bandwidth information of a preset future time period based on the network state prediction model according to the historical network data information.

[0052] In step S230, when the predicted bandwidth information is less than the preset threshold, the audio - video transmission system determines the network prediction state as the weak network state.

[0053] For example, the historical network data information includes at least bandwidth information, latency information, packet loss rate information, etc. Exemplarily, the preset historical time period can be the previous 30s of the current moment, and the preset future time period can be the next 10s of the current moment.

[0054] The audio - video transmission system trains the LSTM prediction model based on the historical network data information so that the LSTM prediction model can predict the predicted bandwidth information for the next 10s. When the predicted bandwidth information is less than a preset threshold (such as 1Mbps), the audio - video transmission system determines that the network prediction state is a weak network state.

[0055] Through the above - mentioned embodiments, the present invention can accurately predict the network state based on the LSTM prediction model and the historical network data information.

[0056] In step S300, when the network prediction state of the audio - video transmission system is a weak network state, the audio - video transmission system determines the transmission channel corresponding to the transmission priority and the dynamic encoding strategy corresponding to the transmission priority.

[0057] For example, the transmission channels can include a high - reliability transmission channel (such as TCP) and a low - latency transmission channel (such as UDP). The high - reliability transmission channel can provide reliable data transmission, ensuring that data packets are not lost, not repeated, and arrive in order; the low - latency transmission channel can provide low - latency data transmission.

[0058] The audio - video transmission system can configure different transmission channels according to the transmission priority of the audio - video data, and transmit the audio - video data based on different transmission channels, so as to achieve multi - path transmission in a weak network state.

[0059] As an embodiment, in a weak network state, the first transmission priority can correspond to a high - reliability transmission channel (such as TCP), the second transmission priority can correspond to both a high - reliability transmission channel and a low - latency transmission channel (such as TCP + UDP), and the third transmission priority can correspond to a low - latency transmission channel (such as UDP).

[0060] The dynamic encoding strategy can include a dynamic bitrate adjustment strategy. For example, in a weak network state, the audio - video transmission system can determine the dynamic bitrate adjustment strategy corresponding to the transmission priority, and each piece of audio - video data with a transmission priority can correspond to a transmission bitrate, so that the audio - video data can be transmitted based on the transmission bitrate.

[0061] As an embodiment, in a weak network state, the first transmission priority can correspond to a high compression bitrate (such as 1080p high compression), the second transmission priority can correspond to a medium compression bitrate (such as 720p medium compression), and the third transmission priority can correspond to a low compression bitrate (such as 480p low compression).

[0062] In step S400, the audio - video transmission system synchronously transmits the audio - video data with the corresponding transmission priority to the participating end through the corresponding transmission channel based on the dynamic coding strategy.

[0063] For example, in a weak network state, the audio - video transmission system can synchronously transmit the audio - video data corresponding to the first transmission priority to the participating end through the high - reliability transmission channel based on a high compression code rate; the audio - video transmission system synchronously transmits the audio - video data corresponding to the second transmission priority to the participating end through the high - reliability transmission channel and the low - latency transmission channel based on a medium compression code rate; and the audio - video transmission system synchronously transmits the audio - video data corresponding to the third transmission priority to the participating end through the low - latency transmission channel based on a low compression code rate.

[0064] As an embodiment, in a weak network state, the audio - video transmission system can also perform multi - path transmission in combination with other network transmission paths (including but not limited to network transmission paths such as Wi - Fi6 and 5G).

[0065] Through the above - mentioned embodiments, the present invention can configure the transmission priority of the audio - video data of the host end. In the case where the current network environment is predicted to be in a weak network state, the transmission channel and the dynamic coding strategy corresponding to the transmission priority can be determined. Thus, the present invention can realize the synchronous transmission of the audio - video data with different transmission priorities based on the corresponding transmission channel and the corresponding dynamic coding strategy.

[0066] The present invention realizes the content classification of the audio - video data according to the content importance of the audio - video data, and can perform multi - path transmission of the audio - video data with different transmission priorities based on the dynamic coding strategy in a weak network state. Thus, the accurate transmission of the key information in the audio - video data can be ensured, and the problem of key information loss in a weak network state is solved. And through multi - path transmission, the present invention can ensure the reliability and real - time performance of the audio - video data transmission at the same time.

[0067] Optionally, in step S400, the audio - video transmission system can also dynamically adjust the forward error correction redundancy of the audio - video data based on the predicted bandwidth information.

[0068] For example, the dynamic coding strategy can also include a dynamic forward error correction redundancy adjustment strategy. This dynamic forward error correction redundancy adjustment strategy can be to dynamically adjust the forward error correction redundancy of the audio - video data according to the predicted bandwidth information.

[0069] Exemplarily, the audio - video transmission system can determine whether the bandwidth of the current network is predicted to increase or decrease according to the predicted bandwidth information and the current bandwidth information.

[0070] When the bandwidth of the current network is predicted to increase, the audio - video transmission system can reduce the forward error correction redundancy, thereby reducing the occupied bandwidth resources and improving the transmission quality of audio - video. When the bandwidth of the current network is predicted to decrease, the audio - video transmission system can increase the forward error correction redundancy, thereby increasing the error - correction ability of data transmission and avoiding problems such as video stuttering or image quality degradation caused by packet loss.

[0071] As an embodiment, when the LSTM prediction model predicts that the bandwidth of the current network is predicted to decrease by 30%, the audio - video transmission system increases the forward error correction redundancy of the audio - video data corresponding to the first transmission priority by 20%.

[0072] Through the above - mentioned embodiments, the present invention can, through the dynamic forward error correction redundancy adjustment strategy, perform corresponding optimization adjustments on the forward error correction redundancy when the network condition changes, thereby balancing the error - correction ability and bandwidth occupancy, and thus improving the quality and stability of audio - video data transmission.

[0073] Figure 3 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention.

[0074] Optionally, as Figure 3 shown, the audio - video transmission method may further include steps S510 - S520.

[0075] In step S510, the audio - video transmission system dynamically determines the visual fixation area of the user at the participating end based on visual tracking.

[0076] In step S520, when the network prediction state is a weak network state, the audio - video transmission system increases the display resolution of the visual fixation area to a preset resolution.

[0077] For example, the audio - video transmission system can collect the user's image based on the front - facing camera device at the participating end and use visual tracking technology (such as a deep - learning model) to track the user's eye movement in real - time, thereby determining the user's fixation area (i.e., the viewing area). Then, the audio - video transmission system can increase the display resolution of this fixation area (such as increasing the resolution from 1080p to 4k).

[0078] Through the above - mentioned embodiments, the present invention can, through visual tracking, determine the visual fixation area of the user at the participating end, thereby giving priority to ensuring the packet transmission of the key fixation area, reducing the re - transmission and redundancy of non - critical areas, and improving the effective information volume generated per Mbps of bandwidth.

[0079] Figure 4 Another flowchart showing the audio - video transmission method according to an embodiment of the present invention.

[0080] Optionally, as Figure 4 shown, the audio - video transmission method may further include step S600.

[0081] In step S600, when the audio - video transmission system detects that the audio - video data received by the participating end is dropped, based on the previous frame data and the subsequent frame data of the audio - video data received by the participating end, the dropped frame data is determined.

[0082] For example, when the audio - video transmission system detects that some frames are missing in the audio - video data of the participating end (such as directly jumping from frame 10 to frame 12), the audio - video transmission system automatically determines the missing intermediate data (such as frame 11) based on the received previous and subsequent frame data (frame 10 and frame 12). This setting can ensure the smoothness of audio - video data playback.

[0083] Optionally, the audio - video transmission system may also trigger data re - transmission only for the dropped frame data corresponding to the first transmission priority, without re - transmitting all the audio - video data corresponding to the first transmission priority.

[0084] This setting can reduce the traffic resources required for redundant data re - transmission. The audio - video transmission system can combine NACK (Negative Acknowledgment) message fast feedback and high - reliability channel transmission, which can shorten the recovery time of critical data.

[0085] According to the exemplary embodiment, the audio - video transmission system triggering data re - transmission for the dropped frame data corresponding to the first transmission priority may include S1 - S10.

[0086] In S1, the audio - video transmission system embeds its corresponding priority identification field in each audio - video data to mark its corresponding transmission priority (for example, the first transmission priority is marked as 0x01).

[0087] In S2, the audio - video transmission system assigns a globally increasing sequence number (such as 32 - bit) to each audio - video data for tracking the order and integrity of the audio - video data.

[0088] In S3, the audio - video transmission system divides the audio - video data with the first transmission priority (such as the key frames of the PPT) into data blocks of a fixed size (such as 512 bytes per block), and each data block can be independently numbered. This setting can accurately identify the lost units.

[0089] In S4, the audio - video transmission system can detect whether there are missing data blocks at the participating end through the continuity of the sequence numbers.

[0090] For example, the participating end can maintain a dynamic receiving buffer to record the sequence numbers of the received data blocks. The audio - video transmission system determines the lost data blocks by comparing the continuity of adjacent sequence numbers (for example, if the sequence numbers received by the participating end are 100 and 102, then 101 can be determined as the lost data block). With such a setting, the audio - video transmission system can real - time identify the loss situation of the audio - video data with the first transmission priority at the participating end. Exemplarily, the audio - video transmission system can also identify the continuous packet loss range through a sliding window mechanism, and the window size of the sliding window can be dynamically adjusted according to network latency.

[0091] In S5, when the audio - video transmission system detects a lost data block, the participating end generates a NACK packet. The NACK packet can include information such as the sequence number range of the lost data block and the corresponding priority identification field.

[0092] In S6, the audio - video transmission system sends the NACK packet generated by the participating end to the host end through a low - latency channel (such as UDP).

[0093] In S7, the host end can maintain a circular sending buffer for the audio - video data with the first transmission priority to save the recently sent data blocks (such as retaining the audio - video data within the most recent 5 s).

[0094] Exemplarily, the capacity of the circular sending buffer can be dynamically adjusted according to network latency to ensure coverage of the possible packet loss time window.

[0095] In S8, after the audio - video transmission system receives and parses the NACK packet at the host end, it extracts only the corresponding lost data blocks from the circular sending buffer. The audio - video transmission system adds the lost data blocks to the high - priority retransmission queue and then immediately sends the lost data blocks through a highly reliable channel (such as TCP).

[0096] Exemplarily, if the lost data block has expired (i.e., beyond the time window of the circular sending buffer), the audio - video transmission system can ignore the request and feedback an error code.

[0097] Exemplarily, the audio - video transmission system can combine the network state prediction result for data retransmission. For example, when the current bandwidth is sufficient, the audio - video transmission system directly triggers the retransmission of the lost data blocks; when the current bandwidth is tight, it pauses the non - critical data transmission to prioritize the retransmission of the lost data blocks. With such a setting, the audio - video transmission system can combine dynamic bandwidth allocation and forward error correction - assisted redundancy to balance the network load, thereby improving the data transmission efficiency.

[0098] In S9, after the participating end obtains the retransmitted data, the audio-visual transmission system inserts it into the target position of the original data stream according to the sequence number and updates the play buffer to ensure continuous playback of the audio-visual stream.

[0099] Exemplarily, the audio-visual transmission system can also dynamically add redundant error correction packets to the audio-visual data with the first transmission priority (for example, 1 redundant block can be attached to every 5 data blocks). In the case of a small number of lost data blocks, the audio-visual transmission system can directly recover through the redundant error correction packets, thereby reducing the retransmission requirement.

[0100] In S10, the audio-visual transmission system can also ensure the synchronization of the retransmitted data through a synchronization correction mechanism.

[0101] Exemplarily, in the case where the audio-visual data is out of sync due to the delay caused by the retransmission of lost data blocks, the audio-visual transmission system can ensure the synchronization of the audio-visual data by means of dynamic stretching of the audio waveform (such as fine-tuning the duration of the audio data through the dynamic time warping algorithm) or fine-tuning of the video frame rate (such as discarding redundant video frames or inserting interpolation frames to keep the synchronization error of the audio-visual data within an acceptable threshold range).

[0102] Optionally, the audio-visual transmission system can also set a timeout timer (such as 200 ms) for each NACK packet. In the case where the participating end does not receive the retransmitted data after the timeout, a secondary NACK packet request is triggered, and at the same time, the transmission priority of the lost data block is reduced to avoid blocking subsequent data transmissions.

[0103] Figure 5 Another schematic flowchart showing the audio-visual transmission method according to an embodiment of the present invention.

[0104] Optionally, as Figure 5 shown, the audio-visual transmission method may further include step S700.

[0105] In step S700, when the delay of the audio data and the video data in the audio-visual data is greater than a preset delay threshold, the audio-visual transmission system dynamically stretches the audio data based on the audio waveform or compresses the video data so that the audio data and the video data are transmitted synchronously.

[0106] For example, in the audio-visual transmission system, when the delay difference between the audio data and the video data exceeds a preset threshold (such as ±80 ms), the audio-visual transmission system will achieve precise synchronization of the audio data and the video data through dynamic stretching / compression of the audio waveform or adjustment of the video frame rate by multiples.

[0107] According to another aspect of the present invention, the present invention also provides an audio-visual transmission method for an online meeting. Figure 6Another flowchart diagram showing the audio-video transmission method according to an embodiment of the present invention. As Figure 6 shown, the audio-video transmission method may include steps S100a - S400a.

[0108] Exemplarily, the audio-video transmission method may be used for the participating end. Hereinafter, taking the participating end as an example, the present invention will be described and introduced.

[0109] According to an example embodiment, in step S100a, the audio-video transmission system configures the transmission priority of the audio-video data of the participating end according to a preset priority rule.

[0110] In step S200a, the audio-video transmission system predicts the network prediction state of the preset future time period of the participating end based on a network state prediction model.

[0111] In step S300a, when the network prediction state is a weak network state, the audio-video transmission system determines the transmission channel corresponding to the transmission priority and the dynamic encoding strategy corresponding to the transmission priority.

[0112] In step S400a, the audio-video transmission system synchronously transmits the audio-video data of the corresponding transmission priority to the host end through the corresponding transmission channel based on the dynamic encoding strategy.

[0113] It can be understood here that the technical solution adopted by the audio-video transmission method for the participating end is exactly the same as that of the audio-video transmission method for the host end described above, except for the execution end. The technical solution adopted by the audio-video transmission method for the host end has been described in detail above, so it will not be repeated here.

[0114] According to another aspect of the present invention, the present invention also provides an audio-video transmission system for an online meeting. Figure 7 A structural diagram showing the audio-video transmission system according to an embodiment of the present invention. As Figure 7 shown, the audio-video transmission system 1 may include a content grading determination module 10, a network state prediction module 20, a channel and policy control module 30, and an audio-video data transmission module 40.

[0115] Exemplarily, the audio-video transmission system may be used for the host end. Hereinafter, taking the host end as an example, the present invention will be described and introduced.

[0116] According to an example embodiment, the content grading determination module 10 configures the transmission priority of the audio-video data of the host end according to a preset priority rule.

[0117] For example, the audio-visual data at the presenting end may include audio data and video data. The audio data at least includes the sound data generated by the presenting end during the meeting, such as the voice signal of the presenter and other background sound signals, etc.; the video data at least includes the video data generated by the presenting end during the meeting, such as the real-time video image of the presenter, the shared screen page (such as the display pages of PPT, Word documents, software operations, etc.), and other background images, etc.

[0118] The preset priority rule can be custom-set by the user according to actual needs. For example, the content classification determination module 10 can respond to the user's configuration instruction, divide the audio-visual data into priorities according to the importance of the data, so as to achieve content classification. It can be understood here that the preset priority rule can be adjusted in real time according to actual needs, rather than relying on fixed rules. Such a setting can ensure the configuration flexibility of the audio-visual data transmission.

[0119] As an embodiment, the content classification determination module 10 can configure the voice signal of the presenter and the shared screen page (such as the display pages of PPT, Word documents, software operations, etc.) as the first transmission priority according to the preset priority rule, configure the real-time video image of the presenter as the second transmission priority, and configure other background images as the third transmission priority.

[0120] For example, the content classification determination module 10 can extract preset meeting content keywords (such as "budget", "deadline", etc.) through natural language processing technology, and mark the keyword as the first transmission priority, etc.). For another example, the content classification determination module 10 can automatically identify complex content such as charts and formulas in the PPT based on a convolutional neural network, so as to ensure the accuracy of content classification.

[0121] According to the exemplary embodiment, the network state prediction module 20 predicts the network prediction state of the presenting end in a preset future time period based on the network state prediction model.

[0122] For example, the network prediction state may include a strong network state, a medium network state, and a weak network state. The network state prediction model can be an LSTM (Long Short-Term Memory) prediction model. In the prediction of the network state, LSTM can predict future bandwidth fluctuations by learning the change rules of historical data, so as to predict the network state within a specific future time period.

[0123] Exemplarily, the network state prediction module 20 can also optimize the input features of the LSTM prediction model by combining the real-time routing node state (such as congestion situation), so as to improve the prediction accuracy of the LSTM prediction model.

[0124] Optionally, the network status prediction module 20 determines the historical network data information for a preset historical time period.

[0125] Based on the network status prediction model, the network status prediction module 20 predicts the predicted bandwidth information for a preset future time period according to the historical network data information.

[0126] When the predicted bandwidth information is less than the preset threshold, the network status prediction module 20 determines that the network prediction status is a weak network status.

[0127] For example, the historical network data information at least includes bandwidth information, latency information, packet loss rate information, etc. Exemplarily, the preset historical time period can be the previous 30s of the current moment, and the preset future time period can be the next 10s of the current moment.

[0128] The network status prediction module 20 trains the LSTM prediction model based on the historical network data information so that the LSTM prediction model can predict the predicted bandwidth information for the next 10s. When the predicted bandwidth information is less than the preset threshold (such as 1Mbps), the network status prediction module 20 determines that the network prediction status is a weak network status.

[0129] Through the above embodiments, the present invention can accurately predict the network status based on the LSTM prediction model and the historical network data information.

[0130] According to the exemplary embodiment, when the network prediction status is a weak network status, the channel and policy control module 30 determines the transmission channel corresponding to the transmission priority and the dynamic coding policy corresponding to the transmission priority.

[0131] For example, the transmission channels can include a high-reliability transmission channel (such as TCP) and a low-latency transmission channel (such as UDP). The high-reliability transmission channel can provide reliable data transmission to ensure that data packets are not lost, not repeated, and arrive in order; the low-latency transmission channel can provide low-latency data transmission.

[0132] The channel and policy control module 30 can configure different transmission channels according to the transmission priority of the audio and video data to transmit the audio and video data based on different transmission channels, thereby realizing multi-path transmission in a weak network state.

[0133] As an embodiment, in a weak network state, the first transmission priority can correspond to a high-reliability transmission channel (such as TCP), the second transmission priority can correspond to a high-reliability transmission channel and a low-latency transmission channel (such as TCP+UDP), and the third transmission priority can correspond to a low-latency transmission channel (such as UDP).

[0134] The dynamic encoding strategy may include a dynamic bitrate adjustment strategy. For example, in a weak network state, the channel and policy control module 30 may determine a dynamic bitrate adjustment strategy corresponding to the transmission priority, and each piece of audio-visual data with a transmission priority may correspond to a transmission bitrate, so that the audio-visual data can be transmitted based on the transmission bitrate.

[0135] As an embodiment, in a weak network state, the first transmission priority may correspond to a high compression bitrate (such as 1080p high compression), the second transmission priority may correspond to a medium compression bitrate (720p medium compression), and the third transmission priority may correspond to a low compression bitrate (such as 480p low compression).

[0136] According to the exemplary embodiment, the audio-visual data transmission module 40 synchronously transmits the audio-visual data with the corresponding transmission priority to the participating end through the corresponding transmission channel based on the dynamic encoding strategy.

[0137] For example, in a weak network state, the audio-visual data transmission module 40 may synchronously transmit the audio-visual data corresponding to the first transmission priority to the participating end through a highly reliable transmission channel based on the high compression bitrate; the audio-visual data transmission module 40 synchronously transmits the audio-visual data corresponding to the second transmission priority to the participating end through a highly reliable transmission channel and a low-latency transmission channel based on the medium compression bitrate; and the audio-visual data transmission module 40 synchronously transmits the audio-visual data corresponding to the third transmission priority to the participating end through a low-latency transmission channel based on the low compression bitrate.

[0138] As an embodiment, in a weak network state, the audio-visual data transmission module 40 may also perform multipath transmission in combination with other network transmission paths (including but not limited to network transmission paths such as Wi-Fi6 and 5G).

[0139] Through the above embodiments, the present invention can configure the transmission priority of the audio-visual data of the host end. In the case where the current network environment is predicted to be in a weak network state, the transmission channel and the dynamic encoding strategy corresponding to the transmission priority can be determined. Thus, the present invention can realize the synchronous transmission of audio-visual data with different transmission priorities based on the corresponding transmission channel and the corresponding dynamic encoding strategy.

[0140] The present invention realizes the content grading of audio-visual data according to the content importance of the audio-visual data, and can perform multipath transmission of audio-visual data with different transmission priorities based on the dynamic encoding strategy in a weak network state. Thus, the accurate transmission of key information in the audio-visual data can be ensured, and the problem of key information loss in a weak network state can be solved. And through multipath transmission, the present invention can simultaneously ensure the reliability and real-time performance of the audio-visual data transmission.

[0141] Optionally, the channel and policy control module 30 can also dynamically adjust the forward error correction redundancy of the audio and video data based on the predicted bandwidth information.

[0142] For example, the dynamic encoding policy can also include a dynamic forward error correction redundancy adjustment policy. The dynamic forward error correction redundancy adjustment policy can be to dynamically adjust the forward error correction redundancy of the audio and video data according to the predicted bandwidth information.

[0143] Exemplarily, the channel and policy control module 30 can determine whether the bandwidth of the current network is predicted to increase or decrease based on the predicted bandwidth information and the current bandwidth information.

[0144] In the case where the bandwidth of the current network is predicted to increase, the channel and policy control module 30 can reduce the forward error correction redundancy, which can reduce the occupied bandwidth resources and improve the transmission quality of the audio and video. In the case where the bandwidth of the current network is predicted to decrease, the channel and policy control module 30 can increase the forward error correction redundancy, which can increase the error correction ability of data transmission and avoid problems such as video stuttering or image quality degradation caused by packet loss.

[0145] As an embodiment, in the case where the LSTM prediction model predicts that the bandwidth of the current network is predicted to decrease by 30%, the channel and policy control module 30 increases the forward error correction redundancy of the audio and video data corresponding to the first transmission priority by 20%.

[0146] Through the above embodiments, the present invention can perform corresponding optimization adjustments on the forward error correction redundancy in the case of changes in the network condition through the dynamic forward error correction redundancy adjustment policy, which can balance the error correction ability and bandwidth occupancy, thereby improving the quality and stability of the audio and video data transmission.

[0147] Optionally, the channel and policy control module 30 dynamically determines the visual fixation area of the user at the participating end based on visual tracking.

[0148] When the network prediction state of the channel and policy control module 30 is a weak network state, the display resolution of the visual fixation area is increased to a preset resolution.

[0149] For example, the channel and policy control module 30 can collect an image of the user based on the front camera device at the participating end and use visual tracking technology (such as a deep learning model) to track the user's eye movement in real time, so as to determine the user's fixation area (i.e., the viewing area). Then, the channel and policy control module 30 can increase the display resolution of this fixation area (such as increasing the resolution from 1080p to 4k).

[0150] Through the above embodiments, the present invention can determine the visual fixation area of the user at the participating end through visual tracking, so that it is possible to preferentially ensure the data packet transmission in the key fixation area, reduce the retransmission and redundancy in the non-critical area, and improve the effective information volume generated per Mbps of bandwidth.

[0151] Optionally, when the channel and policy control module 30 detects that the audio-visual data received at the participating end is dropped, it determines the dropped data based on the previous frame data and the subsequent frame data of the audio-visual data received at the participating end.

[0152] For example, when the channel and policy control module 30 detects that some frames are missing in the audio-visual data at the participating end (such as directly jumping from the 10th frame to the 12th frame), the channel and policy control module 30 automatically determines the missing data in the middle (such as the 11th frame) based on the received previous and subsequent frame data (the 10th frame and the 12th frame). This setting can ensure the smooth playback of audio-visual data.

[0153] Optionally, the audio-visual data transmission module 40 can also trigger data retransmission only for the lost data of the audio-visual data corresponding to the first transmission priority, without retransmitting all the audio-visual data corresponding to the first transmission priority.

[0154] This setting is to reduce the traffic resources required for redundant data retransmission. The audio-visual data transmission module 40 can combine NACK packet fast feedback and high-reliability channel transmission to shorten the recovery time of key data.

[0155] According to the exemplary embodiment, the audio-visual transmission system triggers data retransmission for the dropped data corresponding to the first transmission priority.

[0156] For example, the audio-visual data transmission module 40 embeds its corresponding priority identification field in each audio-visual data to mark its corresponding transmission priority (such as the first transmission priority marked as 0x01).

[0157] The audio-visual data transmission module 40 assigns a globally increasing serial number (such as 32 bits) to each audio-visual data for tracking the order and integrity of the audio-visual data.

[0158] The audio-visual data transmission module 40 can divide the audio-visual data with the first transmission priority (such as the key frame of the PPT) into data blocks of a fixed size (such as 512 bytes per block), and each data block can be independently numbered. This setting can accurately identify the lost unit.

[0159] The audio-visual data transmission module 40 can detect whether there are lost data blocks at the participating end through the continuity of the serial numbers.

[0160] For example, the participating end can maintain a dynamic reception buffer for recording the sequence numbers of the received data blocks. The audio-video data transmission module 40 determines the lost data blocks by comparing the continuity of adjacent sequence numbers (for example, if the sequence numbers received by the participating end are 100 and 102, then 101 can be determined as the lost data block). With such a setting, the audio-video data transmission module 40 can identify the loss situation of the audio-video data with the first transmission priority at the participating end in real time. Exemplarily, the audio-video data transmission module 40 can also identify the continuous packet loss range through a sliding window mechanism, and the window size of the sliding window can be dynamically adjusted according to the network latency.

[0161] When the audio-video data transmission module 40 detects lost data blocks, the participating end generates a NACK (Negative Acknowledgment) message. The NACK message can include information such as the sequence number range of the lost data blocks and the corresponding priority identifier.

[0162] The audio-video data transmission module 40 sends the NACK message generated by the participating end to the presenter end through a low-latency channel (such as UDP).

[0163] The presenter end can maintain a circular send buffer for the audio-video data with the first transmission priority to store the recently sent data blocks (for example, retain the audio-video data within the most recent 5 seconds).

[0164] Exemplarily, the capacity of the circular send buffer can be dynamically adjusted according to the network latency to ensure coverage of the possible packet loss time window.

[0165] After the audio-video data transmission module 40 receives and parses the NACK message at the presenter end, it extracts only the corresponding lost data blocks from the circular send buffer. The audio-video data transmission module 40 adds the lost data blocks to the high-priority retransmission queue and then immediately sends the lost data blocks through a high-reliability channel (such as TCP).

[0166] Exemplarily, if the lost data block has expired (that is, it exceeds the time window of the circular send buffer), the audio-video data transmission module 40 can ignore the request and feedback an error code.

[0167] Exemplarily, the audio-video data transmission module 40 can combine the network status prediction result for data retransmission. For example, when the current bandwidth is sufficient, the audio-video data transmission module 40 directly triggers the retransmission of the lost data blocks; when the current bandwidth is tight, it pauses the non-critical data transmission to prioritize the retransmission of the lost data blocks. With such a setting, the audio-video data transmission module 40 can combine dynamic bandwidth allocation and forward error correction-assisted redundancy to balance the network load, thereby improving the data transmission efficiency.

[0168] After the retransmitted data is obtained at the participating end, the audio and video data transmission module 40 inserts it into the target position of the original data stream according to the sequence number and updates the play buffer to ensure continuous playback of the audio and video stream.

[0169] Exemplarily, the audio and video data transmission module 40 can also dynamically add redundant error correction packets to the audio and video data of the first transmission priority (for example, 1 redundant block can be attached to every 5 data blocks). In the case of a small number of lost data blocks, the audio and video data transmission module 40 can directly recover through the redundant error correction packets, thereby reducing the retransmission requirement.

[0170] The audio and video data transmission module 40 can also ensure the synchronization of the retransmitted data through a synchronization correction mechanism.

[0171] Exemplarily, in the case where the audio and video data is out of sync due to the delay caused by the retransmission of lost data blocks, the audio and video data transmission module 40 can ensure the synchronization of the audio and video data by dynamically stretching the audio waveform (such as fine-tuning the duration of the audio data through the dynamic time warping algorithm) or fine-tuning the video frame rate (such as discarding redundant video frames or inserting interpolation frames to keep the synchronization error of the audio and video data within an acceptable threshold).

[0172] Optionally, the audio and video data transmission module 40 can also set a timeout timer (such as 200 ms) for each NACK message. In the case where the participating end does not receive the retransmitted data after the timeout, a secondary NACK message request is triggered, and at the same time, the transmission priority of the lost data block is reduced to avoid blocking subsequent data transmission.

[0173] Optionally, when the delay of the audio data and the video data in the audio and video data is greater than a preset delay threshold, the audio and video data transmission module 40 dynamically stretches the audio data based on the audio waveform or compresses the video data to enable synchronous transmission of the audio data and the video data.

[0174] For example, when the delay difference between the audio data and the video data exceeds a preset threshold (such as ±80 ms), the audio and video data transmission module 40 will achieve precise synchronization of the audio data and the video data through dynamic stretching / compression of the audio waveform or adjustment of the video frame rate by multiples.

[0175] It can be understood here that, similarly, this audio and video transmission system can also be used at the participating end, and details will not be elaborated here.

[0176] According to another aspect of the present invention, the present invention also provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the audio and video transmission method as described above.

[0177] According to another aspect of the present invention, the present invention further provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the audio-video transmission method as described above.

[0178] According to another aspect of the present invention, the present invention further provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the audio-video transmission method as described above.

[0179] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions of the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for transmitting audio and video in an online conference, characterized in that: The online conference includes a speaker terminal and a participant terminal, and the audio and video transmission method includes: Configure the transmission priority of the audio and video data of the main speaker terminal according to a preset priority rule; Predicting the network prediction state of the main speaker terminal in a preset future time period based on the network state prediction model; When the predicted network state is a weak network state, determining a transmission channel corresponding to the transmission priority and a dynamic encoding strategy corresponding to the transmission priority; Based on the dynamic encoding strategy, the audio and video data of the corresponding transmission priority are synchronously transmitted to the participating end through the corresponding transmission channel.

2. The audio and video transmission method according to claim 1, characterized in that: The predicting of the network prediction state of the speaker terminal in a preset future time period based on the network state prediction model includes: Determine historical network data information for a preset historical time period; Based on the network status prediction model, predicting the predicted bandwidth information for the preset future time period according to the historical network data information; When the predicted bandwidth information is less than a preset threshold, it is determined that the predicted network state is the weak network state.

3. The audio and video transmission method according to claim 2, characterized in that: The synchronously transmitting the audio and video data of the corresponding transmission priority to the conference participant through the corresponding transmission channel based on the dynamic encoding strategy includes: The forward error correction redundancy of the audio and video data is dynamically adjusted based on the predicted bandwidth information.

4. The audio and video transmission method according to claim 1, characterized in that: The audio and video transmission method also includes: Dynamically determine the visual gaze area of ​​the user at the conference participant based on visual tracking; When the predicted network state is the weak network state, the display resolution of the visual attention area is increased to a preset resolution.

5. The audio and video transmission method according to claim 1, characterized in that: The audio and video transmission method also includes: When it is detected that the audio and video data received by the conference participant has lost frames, the lost frame data is determined based on the previous frame data and the next frame data of the audio and video data received by the conference participant.

6. The audio and video transmission method according to claim 1, characterized in that: The audio and video transmission method also includes: When the delay between the audio data and the video data in the audio and video data is greater than a preset delay threshold, the audio data is dynamically stretched or the video data is compressed or doubled in speed based on the audio waveform so that the audio data and the video data are transmitted synchronously.

7. A method for transmitting audio and video in an online conference, characterized in that: The online conference includes a speaker terminal and a participant terminal, and the audio and video transmission method includes: Configure the transmission priority of the audio and video data of the participant terminal according to the preset priority rule; Predicting the network status of the participant in a preset future time period based on the network status prediction model; When the predicted network state is a weak network state, determining a transmission channel corresponding to the transmission priority and a dynamic encoding strategy corresponding to the transmission priority; Based on the dynamic encoding strategy, the audio and video data of the corresponding transmission priority are synchronously transmitted to the main speaker through the corresponding transmission channel.

8. An audio and video transmission system for online conferences, characterized in that: The online conference includes a speaker terminal and a participant terminal, and the audio and video transmission system is used to execute the audio and video transmission method according to any one of claims 1 to 6, and the audio and video transmission system includes: A content classification determination module, which configures the transmission priority of the audio and video data of the main speaker terminal according to a preset priority rule; A network status prediction module, which predicts the network status of the main speaker terminal in a preset future time period based on a network status prediction model; A channel and strategy control module, when the network prediction state is a weak network state, determines a transmission channel corresponding to the transmission priority and a dynamic encoding strategy corresponding to the transmission priority; The audio and video data transmission module synchronously transmits the audio and video data of the corresponding transmission priority to the conference participant through the corresponding transmission channel based on the dynamic encoding strategy.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the audio and video transmission method as described in any one of claims 1-7.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio and video transmission method as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Systems and methods for adaptive throughput management for event-driven message-based data

    CN101491036A

  • Media low-delay communication method and system for network video live

    CN109194974A

  • Data processing method and device

    CN110299963A

  • Communication data processing method and device thereof, storage medium and processor

    CN113660175A

  • Audio and video transmission guarantee method and system based on channel system

    CN117857519A

Cited By

  • Remote guidance method and system based on multi-dimensional perception

    CN121262249A