Video and audio data transmission system and method

By employing multi-channel packet splitting and unpacking processes, combined with network detection and dynamically adjusted compression strategies, the problem of single transmission channel limitation was solved, enabling stable and high-quality transmission of ultra-high-definition real-time media data and improving network resource utilization and media quality.

CN121644822APending Publication Date: 2026-03-10BEIJING DIGITAL VIDEO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies limit media quality and application scenarios when transmitting ultra-high-definition real-time media data through a single network channel. In particular, it is difficult to achieve stable and high-quality transmission in low-bandwidth environments, and problems such as pixelation, stuttering, delay, and distortion are prone to occur during transmission.

Method used

Multiple first transmission channels and multiple second transmission channels are used to process audio and video data into packets and unpack them. Combining network detection and dynamic adjustment compression strategies, data packets are transmitted through multiple channels, and stable and high-quality transmission is achieved by using layered packing and decompression technology.

Benefits of technology

It effectively improves data transmission capabilities in scenarios with dispersed network bandwidth resources, enhances network resource utilization, achieves stable and high-quality transmission of ultra-high-definition real-time media data, and solves the media quality limitations caused by the single transmission channel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644822A_ABST
    Figure CN121644822A_ABST
Patent Text Reader

Abstract

The invention provides a video and audio data transmission system and method, a first transmission and management unit comprises a plurality of first transmission channels, and a second transmission and management unit comprises a plurality of second transmission channels. The compression processing module can send the obtained multiple data sub-packets to the decompression processing module through the multiple first transmission channels and the multiple second transmission channels in sequence, and after the decompression processing module decompresses all the received data sub-packets, target video and audio data can be obtained. According to the system, the data sub-packets are transmitted through the plurality of first transmission channels and the plurality of second transmission channels, so that the limitation of the singleness of the transmission channels on the media quality can be avoided, the data transmission capability in a network bandwidth resource dispersion scene is effectively improved, and the stable and high-quality transmission of the ultra-high-definition real-time media data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data transmission, in particular to a kind of audio and video data transmission system and method. BACKGROUND

[0002] Real-time audio and video media data needs large network bandwidth guarantee due to its large data volume characteristics in transmission, existing transmission network types include Ethernet, optical fiber, satellite, 4G (The 4th generation mobile communication technology, the fourth generation mobile communication and its technology), 5G (The 5th generation mobile communication technology, the fifth generation mobile communication technology) and the like, the usual transmission mode is to encode real-time media data, according to the different network conditions, different encoding strategies and packet loss processing strategies are adopted, such as encoding strategy CBR (Constant Bit Rate, constant bit rate), ABR (Average Bit Rate, average bit rate) and the like, packet loss processing strategy is packet loss retransmission, FEC (Forward Error Correction, forward error correction) and the like, after encoding, it is sent through a selected transmission channel for transmission. However, in this way, the singleness of the transmission network channel limits the media quality and application scenarios, especially the current video application demand has developed to ultra-high-definition quality, it is difficult to realize stable high-quality transmission of ultra-high-definition real-time media data through a single low-bandwidth network. SUMMARY

[0003] The purpose of the present application is to provide an audio and video data transmission system and method to realize stable high-quality transmission of ultra-high-definition real-time media data.

[0004] The audio and video data transmission system provided by the present application, the system comprises: media source component and media destination component;Wherein, media source component includes: compression processing module and first transmission unit;Media destination component includes: second transmission unit and decompression processing module;First transmission unit includes a plurality of first transmission channels, second transmission unit includes a plurality of second transmission channels, a plurality of first transmission channels and a plurality of second transmission channels are connected one by one corresponding;Compression processing module is used for receiving target audio and video data, target audio and video data is compressed and processed, a plurality of data packets are obtained, each data packet is sent to the corresponding first transmission channel respectively;Wherein, the number of a plurality of data packets is same as the number of a plurality of first transmission channels, each data packet includes at least one data packet; The first transmission unit is configured to send the received data packets to the corresponding connected second transmission channels respectively through each first transmission channel; the second transmission unit is configured to send all the data packets to the decompression processing module through each second transmission channel; and the decompression processing module is configured to perform decompression processing on all the received data packets to obtain target audio and video data.

[0005] Further, the compression processing module comprises a pre-compression unit and a hierarchical packing unit; the pre-compression unit is configured to receive the target audio and video data, input each frame of image in the target audio and video data into a pre-trained detection model, identify a key content area in each frame of image through the detection model, perform block processing on each frame of image according to a preset block processing mode to obtain a plurality of block processing results, and perform compression processing on each block processing result to obtain a compression result; wherein, if the key content area is included in a specified block processing result, more computing resources are allocated to the compression of the specified block processing result than to the compression of other block processing results; and the hierarchical packing unit is configured to perform packet splitting on the compression result to obtain a plurality of data packets, and perform packet processing on the plurality of data packets to obtain a plurality of data packets.

[0006] Further, the preset block processing mode comprises at least one of the following: a binary tree block processing mode, a ternary tree block processing mode, and a quad tree block processing mode.

[0007] Further, the first transmission unit is further configured to send a network probe signal to the second transmission unit; the second transmission unit is configured to send a feedback result to the first transmission unit according to the network probe signal; wherein, the feedback result carries network state information of each second transmission channel; the first transmission unit is further configured to feed back the feedback result to the pre-compression unit; and the pre-compression unit is configured to determine a network overall state according to the feedback result, determine a compression strategy matched with the network overall state according to the network overall state, and perform compression processing on each block processing result according to the compression strategy to obtain a compression result.

[0008] Further, the hierarchical packing unit is further configured to estimate a transmission bandwidth of each first transmission channel according to the feedback result, and perform packet processing on the plurality of data packets according to the transmission bandwidth of each first transmission channel to obtain a plurality of data packets.

[0009] Further, the first transmission unit is further configured to obtain channel data of each first transmission channel, determine an encapsulation mode and a transmission strategy corresponding to the first transmission channel according to the channel data, encapsulate the data packets corresponding to the first transmission channel according to the encapsulation mode, and send the encapsulated data packets to the corresponding connected second transmission channel according to the transmission strategy.

[0010] Further, the decompression processing module comprises a media aggregation unit and a decompression unit; the media aggregation unit is configured to parse all the received data packets to obtain the sequence code of each data packet in all the data packets; sort and aggregate all the data packets according to the sequence code to obtain an aggregation result; and the decompression unit is configured to perform decompression processing on the aggregation result to obtain the target audio and video data.

[0011] Further, the decompression unit is further configured to perform decompression processing on the aggregation result to obtain a decompression result; perform processing on the decompression result according to a preset processing mode to obtain the target audio and video data; and the preset processing mode comprises image sharpening processing, image enhancement processing, image reconstruction processing and image frame insertion processing.

[0012] Further, the second transmission unit is further configured to feed back the data packet receiving result of each second transmission channel to the first transmission unit after receiving the corresponding data packet in each second transmission channel.

[0013] The application provides an audio and video data transmission method, wherein the first transmission unit comprises a plurality of first transmission channels, the second transmission unit comprises a plurality of second transmission channels, and the plurality of first transmission channels are connected with the plurality of second transmission channels in one-to-one correspondence; the method comprises the following steps: a compression processing module receives target audio and video data, performs compression processing on the target audio and video data to obtain a plurality of data packets, and sends each data packet to a corresponding first transmission channel; the number of the plurality of data packets is the same as the number of the plurality of first transmission channels, and each data packet comprises at least one data packet; the first transmission unit is configured to send the received data packets to the corresponding second transmission channels through each first transmission channel; the second transmission unit is configured to send all the data packets to a decompression processing module through each second transmission channel; and the decompression processing module is configured to perform decompression processing on all the received data packets to obtain the target audio and video data.

[0014] The audio and video data transmission system and method provided by the application can avoid the limitation of media quality caused by the singleness of the transmission channel, effectively improve the data transmission capacity in the network bandwidth resource dispersion scene, and realize stable and high-quality transmission of ultra-high-definition real-time media data. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the specific embodiments or the prior art of the present application, the drawings needed to be used in the specific embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 A schematic diagram of a video and audio data transmission system provided by an embodiment of the present application is shown in FIG. 1. Figure 2 A schematic diagram of another video and audio data transmission system provided by an embodiment of the present application is shown in FIG. 2. Figure 3 A schematic diagram of a dynamic adjustment mode provided by an embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION

[0017] The technical solutions of the present application will be described clearly and completely in combination with embodiments. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0018] The existing video and audio data transmission technology has several defects in actual use: first, in a weak network transmission environment, excessive compression is performed to adapt to a low-bandwidth environment, which brings a large loss of image details. In order to solve the problem of data loss caused by network packet loss, retransmission, error correction and other technical means are usually used, which will increase the bandwidth and cause real-time video transmission in a low-bandwidth environment to easily appear problems such as mosaic, lag, delay, distortion and the like. Second, the singleness of the transmission network channel limits the media quality and application scenarios. In particular, the current video application demand has developed to ultra-high-definition quality, and it is difficult to realize stable and high-quality transmission of ultra-high-definition real-time media data through a single low-bandwidth network. Based on this, the embodiments of the present application provide a video and audio data transmission system and method, which can be applied to scenarios requiring stable and high-quality transmission of ultra-high-definition real-time media data.

[0019] In order to facilitate the understanding of the present embodiment, first, a video and audio data transmission system disclosed by the present embodiment will be introduced, as shown in FIG. 1. Figure 1As shown, the system comprises a media source component 10 and a media destination component 20; wherein the media source component 10 comprises a compression processing module 100 and a first transmission unit 101; the media destination component 20 comprises a second transmission unit 200 and a decompression processing module 201; the first transmission unit 101 comprises a plurality of first transmission channels, the second transmission unit 200 comprises a plurality of second transmission channels, the plurality of first transmission channels and the plurality of second transmission channels are connected one by one in a one-to-one correspondence; the plurality of first transmission channels and the plurality of second transmission channels can be set according to actual needs, and the number of the plurality of first transmission channels is the same as the number of the plurality of second transmission channels.

[0020] The compression processing module 100 is used for receiving target audio and video data, compressing the target audio and video data, obtaining a plurality of data packets, and sending each data packet to a corresponding first transmission channel; wherein the number of the plurality of data packets is the same as the number of the plurality of first transmission channels, and each data packet comprises at least one data packet; the above-mentioned target audio and video data is usually a combination of video data and audio data, used to describe information containing dynamic images and sound; in actual implementation, after the compression processing module 100 receives the target audio and video data, it can be compressed, processed, etc. to obtain a plurality of data packets, each data packet is composed of one or more data packets, and the number of data packets contained in different data packets can be the same or different; since the number of data packets is the same as the number of first transmission channels, the plurality of data packets obtained can be sent to the corresponding first transmission channel respectively.

[0021] The first transmission unit 101 is used for sending the received data packets to the corresponding connected second transmission channel through each first transmission channel; after each first transmission channel in the first transmission unit 101 receives the corresponding data packet, it can continue to send to the second transmission unit 200, and specifically can be sent to the corresponding connected second transmission channel.

[0022] The second transmission unit 200 is used for sending all data packets to the decompression processing module 201 through each second transmission channel; after the second transmission unit 200 receives the corresponding data packet through each second transmission channel, it can transmit all the received data packets to the decompression processing module 201.

[0023] The decompression processing module 201 is used for decompressing all the received data packets to obtain target audio and video data. After the decompression processing module 201 receives all the data packets, it can decompress, process, etc. all the data packets, and finally restore the target audio and video data.

[0024] The aforementioned audio-visual data transmission system includes a first transmission unit comprising multiple first transmission channels and a second transmission unit comprising multiple second transmission channels. The compression processing module sequentially sends multiple data packets to the decompression processing module via the multiple first and second transmission channels. The decompression processing module decompresses all received data packets to obtain the target audio-visual data. This system, by transmitting data packets through multiple first and second transmission channels, avoids the limitations on media quality caused by a single transmission channel, effectively improves data transmission capabilities in scenarios with dispersed network bandwidth resources, and achieves stable and high-quality transmission of ultra-high-definition real-time media data.

[0025] like Figure 2 The diagram shows another audio / video data transmission system. The compression processing module includes a pre-compression unit and a layered packaging unit. The pre-compression unit receives the target audio and video data, inputs each frame of the target audio and video data into a pre-trained detection model, and identifies key content regions in each frame of the image through the detection model; it divides each frame of the image into blocks according to a preset block division method to obtain multiple block results; it compresses each block result to obtain a compressed result; wherein, if a specified block result contains key content regions, the computing resources allocated to compressing the specified block result are more than those allocated to other block results.

[0026] The aforementioned detection model can be implemented using various network forms; the aforementioned key content regions can be understood as the parts of the image that are of high importance and contain key information, usually related to the main visual focus or core content of the image; the aforementioned computing resources can include processor resources, memory resources, storage resources, etc.; in actual implementation, after receiving the target audio and video data, the pre-compression unit can perform intelligent content understanding and analysis on the target audio and video data. Specifically, each frame of the target audio and video data can be input into a pre-trained detection model, which identifies the key content regions in each frame of the image; to improve the subsequent compression efficiency, each frame of the image is usually divided into blocks, resulting in multiple block results for each frame of the image. Each block result can be compressed independently to obtain the aforementioned compression result, which can then be sent to the layered packaging unit; during the compression process, a higher quality quantization parameter (QP) is usually set for key content regions. If a specified block result contains key content regions, more computing resources are usually allocated to it when compressing the specified block result, in order to improve the subjective quality of the image under the same bitrate bandwidth conditions.

[0027] The hierarchical packing unit is configured to perform packet splitting on the compression result to obtain a plurality of data packets, perform packet processing on the plurality of data packets, and obtain a plurality of data packets.

[0028] Further, the preset block mode includes at least one of the following: binary tree block mode, ternary tree block mode, and quadtree block mode. The binary tree is a tree structure in which each node has at most two child nodes. In image processing, the binary tree block mode indicates that the image can be further divided into two sub-blocks in the vertical or horizontal direction. The ternary tree is a tree structure in which each node has at most three child nodes. In image processing, the ternary tree block mode indicates that the image block can be divided into three sub-blocks in the vertical or horizontal direction, one of which is in the center and the other two are on the sides. The quadtree is a data structure that recursively divides a two-dimensional space into four quadrants. In image processing, the quadtree block mode indicates that the image can be divided into four sub-blocks of equal size. That is, in this embodiment, after identifying the key content area in each frame of image, the pre-compression unit can perform block processing on each frame of image according to at least one of the binary tree block mode, the ternary tree block mode, and the quadtree block mode to obtain a plurality of block results.

[0029] Further, the first transmission unit is further configured to send a network probe signal to the second transmission unit; the second transmission unit is configured to send a feedback result to the first transmission unit according to the network probe signal; wherein the feedback result carries network state information of each second transmission channel; the first transmission unit is further configured to feed back the feedback result to the pre-compression unit; and the pre-compression unit is configured to determine the overall network state according to the feedback result, determine a compression strategy matching the overall network state according to the overall network state, and perform compression processing on each block result according to the compression strategy to obtain a compression result.

[0030] The network detection signal can be used to indicate a request for obtaining network state information of each second transmission channel in the second transmission unit; the network state information can include network type, bandwidth size, packet loss rate, time delay, channel jitter, and the like; the network overall state refers to the overall performance and state of the current network environment, including overall network bandwidth, packet loss rate, and the like; the compression strategy can include compression bandwidth, resolution, frame rate, transmission service mode, and the like; in actual implementation, in the process of audio and video data transmission, the first transmission unit can send a network detection signal to the second transmission unit, the second transmission unit can return a feedback result carrying network state information of each second transmission channel to the first transmission unit after receiving the network detection signal, and the first transmission unit can further feed back the received feedback result to the front compression unit, so that the front compression unit can determine the network overall state of the current network environment according to the feedback result, and can select a suitable compression strategy according to the network overall state; the process is real-time detection and dynamic updating, and each block result is compressed according to the new compression strategy to obtain a compression result.

[0031] Further, the hierarchical packaging unit is further configured to estimate the transmission bandwidth of each first transmission channel according to the feedback result, and perform packet processing on the plurality of data packets according to the transmission bandwidth of each first transmission channel to obtain a plurality of data packets. The hierarchical packaging unit can dynamically perform packet processing on the plurality of data packets according to the network state information of each second transmission channel carried by the feedback result. Specifically, in the process of audio and video data transmission, the delay variation of data packet arrival and the channel jitter condition can be recorded and analyzed in real time, and the optimal transmission bandwidth of each first transmission channel is estimated accordingly. For details, refer to related technologies; the data packet proportion of each first transmission channel is adjusted in real time according to the current total number of channels and the proportion of the estimated transmission bandwidth of each first transmission channel, and then the plurality of data packets are dynamically packet processed according to the proportion to obtain a plurality of data packets.

[0032] Further, the first transmission unit is further configured to obtain channel data of each first transmission channel, determine an encapsulation mode and a transmission strategy corresponding to the first transmission channel according to the channel data, encapsulate the data packets corresponding to the first transmission channel by using the encapsulation mode, and send the encapsulated data packets to the corresponding connected second transmission channel according to the transmission strategy.

[0033] The channel data generally includes: channel type, state data, etc. In actual implementation, different channel types can be pre-set to correspond to different packaging modes, and different network state information can be pre-set to correspond to different transmission strategies. The first transmission unit can dynamically adjust the packaging mode and the transmission strategy corresponding to each first transmission channel according to the channel type (such as 4G / 5G / IP / V.35, etc.) and the network state information of each first transmission channel, so as to adapt to the channel state of different first transmission channels. For example, in the application of mutual priority, RTP (Real-time Transport Protocol), RTCP (Real-time Transport Control Protocol), and SDP (Session Description Protocol) packaging modes are used, in the application of bandwidth priority, TS (Transport Stream) and HLS (HTTP Live Streaming) packaging modes are used, and whether to start the anti-packet loss and FEC processing mechanism is evaluated according to the network state data.

[0034] For example, as shown in a schematic diagram of a dynamic adjustment mode, the front compression unit can continuously determine the overall network state according to the feedback result, judge whether the current compression strategy can meet the network transmission requirement, if yes, the current compression strategy can be maintained to continue compression, if not, the compression strategy can be dynamically adjusted. The layered packaging unit can continuously judge whether the packaging mode and the transmission strategy of each first transmission channel meet the corresponding transmission channel requirement according to the feedback result, if yes, the current packaging and packaging mode can be maintained, if not, the data packaging and packaging mode can be dynamically adjusted. Figure 3 Further, the decompression processing module includes: a media aggregation unit and a decompression unit; the media aggregation unit is used for analyzing all received data packages to obtain the sequence coding of each data package in the data packages; all data packages are sorted and aggregated according to the sequence coding to obtain an aggregation result; and the decompression unit is used for decompressing the aggregation result to obtain target audio and video data.

[0035]

[0036] ​In actual implementation, when the layered packing unit performs data packet splitting on the compression result to obtain multiple data packets, the layered packing unit usually sets a corresponding sequence number for each data packet in sequence, such as 1, 2, 3, and the like. The media aggregation unit can aggregate the data packets, dynamically analyze the data packets, obtain the sequence number corresponding to each data packet in each data packet, sort all the data packets in ascending order of the sequence number, and aggregate the data packets into a complete real-time media stream, that is, the aggregation result. The aggregation result can be transmitted to the decompression unit, and the decompression unit can decode the received aggregation result to generate the target audio and video data.

[0037] Further, the decompression unit is further configured to perform decompression processing on the aggregation result to obtain a decompression result, perform processing on the decompression result in a preset processing manner to obtain the target audio and video data, and the preset processing manner includes image sharpening processing, image enhancement processing, image reconstruction processing, and image frame interpolation processing.

[0038] Image sharpening processing is a technique that compensates and increases the high-frequency components of an image to make the boundaries of ground objects, region edges, lines, texture features, and fine structure features in the image clearer and more distinct. The purpose is to enhance the image outline and details to make the image look clearer. Image enhancement processing is a technique that improves image quality or highlights key information in the image through a series of algorithms to meet the needs of subsequent tasks. The purpose is to improve the visual effect of the image to make the image more suitable for human eye observation or computer analysis. Image reconstruction processing is a technique that restores the original image from partial or degraded image data through mathematical models and algorithms. The purpose is to restore the true information of the image and eliminate noise and distortion introduced in the transmission, storage, or acquisition process. Image frame interpolation processing is a technique that generates new frames between adjacent image frames through algorithms to improve the frame rate and smoothness of the video. The purpose is to reduce the stuttering and jitter in the video and improve the viewing experience. In actual implementation, after the decompression unit performs decompression processing on the aggregation result, the decompression unit can obtain the decompression result, identify the image outline of the decompression result using an image intelligent edge detection algorithm, and perform sharpening and detail enhancement processing. At the same time, the decompression unit can reconstruct and interpolate the missing image details based on a deep learning training model to optimize the image quality and smoothness. The specific processing process can be referred to related technologies, which will not be described here. Finally, the target audio and video data is obtained.

[0039] Further, the second transmission unit is further configured to feed back the data packet reception result of each second transmission channel to the first transmission unit after each second transmission channel receives the corresponding data packet.

[0040] In actual implementation, each second transmission channel in the second transmission unit respectively receives a corresponding data packet, and can feed back a data packet receiving result to the first transmission unit, such as an ACK (Acknowledge character), indicating successful data receiving.

[0041] The above-mentioned audio and video data transmission system, the media source component completes the receiving and processing of the real-time audio and video original baseband data, compresses it into a code stream suitable for bandwidth, and decomposes it into multiple first transmission channels according to a certain strategy; the media destination component receives the data packets through multiple second transmission channels respectively, and performs analysis and aggregation processing on each data packet, restores it into a complete real-time media stream, and then decompresses and outputs it.

[0042] The system has the following beneficial effects: 1. The single network relied on by the traditional real-time media stream transmission is improved, multiple transmission channels can be bound and transmitted simultaneously, the data transmission capability in the network bandwidth resource dispersion scene is effectively improved, the network resource utilization rate is improved, and the use scene of the real-time media transmission and media service organization application is greatly widened.

[0043] 2. The two-level dynamic adaptive adjustment mechanism sets different data packet transmission amounts and channel adaptability strategies according to the characteristics of different network channels, and improves the stability and fault tolerance of the real-time media stream transmission.

[0044] 3. The image quality loss caused by excessive compression in the extremely narrow bandwidth transmission network environment is solved, and the image quality of the real-time media in the weak network environment is improved.

[0045] The embodiment of the application also provides an audio and video data transmission method, a first transmission unit includes multiple first transmission channels, a second transmission unit includes multiple second transmission channels, and the multiple first transmission channels are connected in one-to-one correspondence with the multiple second transmission channels; the method comprises the following steps: Step one, the compression processing module receives target audio and video data, performs compression processing on the target audio and video data, obtains multiple data packets, and sends each data packet to a corresponding first transmission channel; wherein the number of the multiple data packets is the same as the number of the multiple first transmission channels, and each data packet includes at least one data packet. Step two, the first transmission unit is used for sending the received data packets to the corresponding connected second transmission channels through each first transmission channel. Step three, the second transmission unit is used for sending all data packets to the decompression processing module through each second transmission channel. Step four, the decompression processing module is used for decompressing all received data packets to obtain target audio and video data.

[0046] The above-mentioned audio and video data transmission method can avoid the limitation of media quality caused by the singleness of the transmission channel, effectively improve the data transmission capability in the network bandwidth resource dispersion scene, and realize stable and high-quality transmission of the ultra-high-definition real-time media data.

[0047] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A video and audio data transmission system, characterized by The system comprises a media source component and a media destination component; wherein the media source component comprises a compression processing module and a first transmission unit; the media destination component comprises a second transmission unit and a decompression processing module; the first transmission unit comprises a plurality of first transmission channels, the second transmission unit comprises a plurality of second transmission channels, and the plurality of first transmission channels are connected to the plurality of second transmission channels one by one; The compression processing module is configured to receive target audio-visual data, perform compression processing on the target audio-visual data, obtain a plurality of data packets, and send each data packet to a corresponding first transmission channel; wherein the number of data packets is the same as the number of first transmission channels, and each data packet comprises at least one data packet; The first transmission unit is configured to send the received data packets to the corresponding second transmission channels through each first transmission channel; The second transmission unit is configured to send all data packets to the decompression processing module through each second transmission channel; The decompression processing module is configured to perform decompression processing on all received data packets to obtain the target audio-visual data.

2. The system of claim 1, wherein, The compression processing module comprises a pre-compression unit and a hierarchical packing unit; The pre-compression unit is configured to receive target audio-visual data, input each frame of image in the target audio-visual data into a pre-trained detection model, identify key content areas in each frame of image through the detection model, perform block processing on each frame of image according to a preset block mode to obtain a plurality of block results, and perform compression processing on each block result to obtain a compression result; wherein if a specified block result contains the key content area, more computing resources are allocated to the specified block result than to other block results during compression processing. The hierarchical packing unit is configured to split the compression result into a plurality of data packets, and perform packet processing on the plurality of data packets to obtain a plurality of data packets.

3. The system of claim 2, wherein The preset block mode comprises at least one of a binary tree block mode, a ternary tree block mode, and a quad tree block mode.

4. The system of claim 2, wherein The first transmission unit is further configured to send a network probe signal to the second transmission unit; The second transmission unit is configured to send a feedback result to the first transmission unit according to the network probe signal; wherein the feedback result carries network state information of each second transmission channel; The first transmission unit is further configured to feed back the feedback result to the pre-compression unit; The pre-compression unit is configured to determine a network overall state according to the feedback result, determine a compression strategy matched with the network overall state according to the network overall state, and perform compression processing on each block result according to the compression strategy to obtain a compression result.

5. The system of claim 4, wherein The layered packing unit is further configured to estimate a transmission bandwidth of each of the first transmission channels according to the feedback result, and perform packet processing on the plurality of data packets according to the transmission bandwidth of each of the first transmission channels to obtain a plurality of data packets.

6. The system of claim 1, wherein, The first transmission unit is further configured to, for each of the first transmission channels, acquire channel data of the first transmission channel, determine an encapsulation mode and a transmission strategy corresponding to the first transmission channel according to the channel data, encapsulate data packets corresponding to the first transmission channel in the encapsulation mode, and send the encapsulated data packets to a corresponding second transmission channel according to the transmission strategy.

7. The system of claim 1, wherein, The decompression processing module comprises a media aggregation unit and a decompression unit. The media aggregation unit is configured to parse all the received data packets to obtain an order code of each of the data packets in all the data packets, sort and aggregate all the data packets according to the order code to obtain an aggregation result. The decompression unit is configured to perform decompression processing on the aggregation result to obtain the target audio and video data.

8. The system of claim 7, wherein, The decompression unit is further configured to perform decompression processing on the aggregation result to obtain a decompression result. The decompression result is processed according to a preset processing mode to obtain the target audio and video data, wherein the preset processing mode comprises image sharpening processing, image enhancement processing, image reconstruction processing, and image frame insertion processing.

9. The system of claim 1, wherein, The second transmission unit is further configured to feed back a data packet reception result of each of the second transmission channels to the first transmission unit after each of the second transmission channels receives a corresponding data packet.

10. An audiovisual data transmission method, characterized in that, The first transmission unit comprises a plurality of first transmission channels, and the second transmission unit comprises a plurality of second transmission channels, wherein the plurality of first transmission channels are connected to the plurality of second transmission channels one by one; and the method comprises: A compression processing module receives target audio and video data, performs compression processing on the target audio and video data to obtain a plurality of data packets, and sends each of the data packets to a corresponding first transmission channel; wherein the number of the plurality of data packets is the same as the number of the plurality of first transmission channels, and each of the data packets comprises at least one data packet; The first transmission unit is configured to send the received data packets to corresponding second transmission channels through each of the first transmission channels; The second transmission unit is configured to send all the data packets to a decompression processing module through each of the second transmission channels; The decompression processing module is configured to perform decompression processing on all the received data packets to obtain the target audio and video data.