Audio and video gateway dynamic coding cooperative transmission method and device, equipment and storage medium

By preprocessing audio and video data and sensing network status, dynamically adjusting encoding parameters and bandwidth allocation, and adopting adaptive encoding and cooperative transmission strategies, the problem of low transmission quality in traditional audio and video communication is solved, and more stable audio and video data transmission is achieved.

CN121940565APending Publication Date: 2026-04-28SHENZHEN DINSTAR TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN DINSTAR TECH
Filing Date
2026-01-14
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In traditional audio and video communication, fixed bitrate encoding standards cannot adapt to dynamically changing network bandwidth in real time, resulting in poor audio and video data transmission quality.

Method used

The system collects raw audio and video data for preprocessing, obtains network transmission status information, dynamically determines audio and video encoding parameters and bandwidth allocation ratios, adopts adaptive encoding algorithms and cooperative transmission strategies, optimizes timestamp alignment and buffer adjustment of the encoded stream, and monitors packet loss rate for adaptive processing.

Benefits of technology

It improves the overall transmission quality of audio and video data, avoids resource competition and quality imbalance, and achieves matching between encoding strategies and network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940565A_ABST
    Figure CN121940565A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video gateway dynamic coding cooperative transmission method, device and equipment and a storage medium, and relates to the technical field of communication transmission, and the audio and video gateway dynamic coding cooperative transmission method comprises the steps: collecting original audio data and original video data, and carrying out the preprocessing of the original audio data and the original video data, obtaining target audio data and target video data; acquiring network transmission state information, and determining an audio coding parameter, a video coding parameter and a bandwidth allocation proportion based on the network transmission state information; encoding target audio data according to the audio encoding parameter to obtain an encoded audio stream; encoding target video data according to the video encoding parameter to obtain an encoded video stream; and transmitting the coded audio stream and the coded video stream according to the bandwidth allocation proportion. The transmission quality of the audio and video data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication transmission technology, and in particular to a dynamic encoding collaborative transmission method, apparatus, device, and storage medium for audio and video gateways. Background Technology

[0002] In audio and video communication technologies based on Voice over Internet Protocol (VoIP) gateways, traditional methods typically employ fixed-bitrate encoding standards and independent transmission strategies. These methods often fail to adapt in real-time to dynamically changing network bandwidth, resulting in poor audio and video data transmission quality. Therefore, improving the transmission quality of audio and video data remains a problem that needs to be solved.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a dynamic encoding collaborative transmission method, apparatus, device, and storage medium for audio and video gateways, aiming to solve the technical problem of how to improve the transmission quality of audio and video data.

[0005] To achieve the above objectives, this application proposes a dynamic coding collaborative transmission method for audio and video gateways, the method comprising: Collect raw audio data and raw video data, and preprocess the raw audio data and raw video data to obtain target audio data and target video data; Obtain network transmission status information, and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information; The target audio data is encoded according to the audio encoding parameters to obtain an encoded audio stream; The target video data is encoded according to the video encoding parameters to obtain an encoded video stream; The encoded audio stream and the encoded video stream are transmitted according to the bandwidth allocation ratio.

[0006] In one embodiment, the step of preprocessing the original audio data and original video data to obtain target audio data and target video data includes: The original audio data is subjected to noise suppression and speech signal-to-noise ratio enhancement to obtain the target audio data; The original video data is then optimized for image contrast and improved for video signal-to-noise ratio to obtain the target video data.

[0007] In one embodiment, the step of determining the audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information includes: Based on the network transmission status information, bandwidth fluctuation data is determined, and audio encoding parameters are determined based on the bandwidth fluctuation data; Based on the network transmission status information, future bandwidth data is predicted, and video encoding parameters are determined according to the future bandwidth data. The width allocation ratio is determined based on the network transmission status information, the target audio data, and the target video data.

[0008] In one embodiment, the step of encoding the target audio data according to the audio encoding parameters to obtain an encoded audio stream includes: The target bitrate is determined based on the audio encoding parameters; Obtain the speech activity features of the target audio data; Based on the target bitrate and the speech activity characteristics, the target audio data is encoded using an adaptive frame length linear predictive coding algorithm to obtain a quantized signal; The quantized signal is then subjected to high-frequency component compensation and sound quality enhancement to generate an encoded audio stream.

[0009] In one embodiment, the step of encoding the target video data according to the video encoding parameters to obtain an encoded video stream includes: The target compression algorithm and compression ratio are determined based on the video encoding parameters. The target video data is encoded according to the target compression algorithm and the compression ratio to obtain an encoded video stream.

[0010] In one embodiment, before the step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio, the method further includes: The encoded audio stream and the encoded video stream are timestamped to obtain the target encoded data; Acquire real-time network jitter data, adjust the size of the transmission buffer based on the real-time network jitter data, and buffer the target encoded data.

[0011] In one embodiment, the step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio includes: According to the bandwidth allocation ratio, the encoded audio stream and the encoded video stream are encapsulated into audio data packets and video data packets, and then transmitted. During transmission, the packet loss rate is monitored. When the packet loss rate exceeds a preset threshold, the audio data packets are muted by interpolation, the video data packets are copied from adjacent frames, and the transmission is completed.

[0012] Furthermore, to achieve the above objectives, this application also proposes an audio / video gateway dynamic coding collaborative transmission device, which includes: The acquisition module is used to acquire raw audio data and raw video data, and to preprocess the raw audio data and raw video data to obtain target audio data and target video data. The determination module is used to acquire network transmission status information and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information. The audio module is used to encode the target audio data according to the audio encoding parameters to obtain an encoded audio stream; The video module is used to encode the target video data according to the video encoding parameters to obtain an encoded video stream; The transmission module is used to transmit the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

[0013] In addition, to achieve the above objectives, this application also proposes an audio and video gateway dynamic coding collaborative transmission device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio and video gateway dynamic coding collaborative transmission method described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the audio and video gateway dynamic encoding collaborative transmission method described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the audio and video gateway dynamic encoding collaborative transmission method described above.

[0016] This application provides a dynamic encoding and collaborative transmission method for audio and video gateways. The method involves collecting raw audio and video data, preprocessing the raw audio and video data to obtain target audio and video data, acquiring network transmission status information, and determining audio encoding parameters, video encoding parameters, and bandwidth allocation ratios based on the network transmission status information. The target audio data is then encoded according to the audio encoding parameters to obtain an encoded audio stream. Similarly, the target video data is encoded according to the video encoding parameters to obtain an encoded video stream. Finally, the encoded audio and video streams are transmitted according to the bandwidth allocation ratios. This application, by collecting and preprocessing audio and video data, intelligently determining encoding parameters and bandwidth allocation ratios based on network transmission status information, and performing encoding and collaborative transmission accordingly, ensures that the encoding strategy and transmission resources match network conditions, avoids resource competition and quality imbalance between audio and video streams, and improves the overall transmission quality of audio and video data. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the audio / video gateway dynamic encoding collaborative transmission method of this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the audio / video gateway dynamic encoding collaborative transmission method of this application; Figure 3 A simplified flowchart illustrating the audio / video gateway dynamic encoding collaborative transmission method provided in Embodiment 1 of this application; Figure 4 This is a schematic diagram of the module structure of the audio / video gateway dynamic encoding collaborative transmission device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio and video gateway dynamic encoding collaborative transmission method in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] This application collects raw audio data and raw video data, preprocesses the raw audio data and raw video data to obtain target audio data and target video data; acquires network transmission status information, and determines audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information; encodes the target audio data according to the audio encoding parameters to obtain an encoded audio stream; encodes the target video data according to the video encoding parameters to obtain an encoded video stream; and transmits the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

[0024] In audio and video communication technologies based on Voice over Internet Protocol (VoIP) gateways, traditional methods typically employ fixed-bitrate encoding standards and independent transmission strategies. These methods often fail to adapt in real-time to dynamically changing network bandwidth, resulting in poor audio and video data transmission quality. Therefore, improving the transmission quality of audio and video data remains a problem that needs to be solved.

[0025] This application collects and preprocesses audio and video data, intelligently determines encoding parameters and bandwidth allocation ratios based on network transmission status information, and performs encoding and coordinated transmission accordingly. This ensures that the encoding strategy and transmission resources match the network conditions and avoids resource competition and quality imbalance between audio and video streams, thereby improving the overall transmission quality of audio and video data.

[0026] Based on this, embodiments of this application provide a dynamic encoding and collaborative transmission method for audio and video gateways, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the audio / video gateway dynamic encoding collaborative transmission method of this application.

[0027] In this embodiment, the audio / video gateway dynamic encoding collaborative transmission method includes steps S10~S40: Step S10: Collect raw audio data and raw video data, and preprocess the raw audio data and raw video data to obtain target audio data and target video data; It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as an audio / video gateway dynamic encoding collaborative transmission device. The following description uses an audio / video gateway dynamic encoding collaborative transmission device as an example to illustrate this embodiment and the subsequent embodiments.

[0028] It should be noted that raw audio and raw video data can be simultaneously captured through terminal devices (such as cameras and microphones). The audio sampling rate is 16kHz, the initial video resolution supports 4K and below, and the frame rate is 30fps by default.

[0029] In one feasible approach, the step of preprocessing the original audio data and original video data to obtain target audio data and target video data includes: performing noise suppression and speech signal-to-noise ratio enhancement on the original audio data to obtain target audio data; and optimizing image contrast and improving video signal-to-noise ratio on the original video data to obtain target video data.

[0030] It should be noted that the raw audio data can be processed using a combination algorithm of spectral subtraction and Least Mean Squares (LMS) adaptive filtering to suppress environmental noise above -15dB (such as background noise and electromagnetic interference), improve the speech signal-to-noise ratio by 15dB, output a noise-free broadband speech stream, and simultaneously acquire Voice Activity Detection (VAD) features, marking active and silent segments to obtain the target audio data. For the raw video data, image enhancement algorithms can be used to optimize image contrast in low-light and backlight scenes, preserve details in dynamic areas (such as human movement and object shaking), improve the video signal-to-noise ratio by 10dB, output enhanced video frames, and extract video motion intensity features to obtain the target video data.

[0031] Step S20: Obtain network transmission status information, and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information; It should be noted that the network status monitoring module collects network transmission status information of the transmission link in real time, including current bandwidth, bandwidth fluctuation trends, and packet loss rate. The data sampling period is 100ms to ensure rapid response to network changes. Audio encoding parameters include audio bitrate, and video encoding parameters include video compression mode.

[0032] Step S30: Encode the target audio data according to the audio encoding parameters to obtain an encoded audio stream; It should be noted that the target bitrate can be determined based on the audio encoding parameters, and the encoding process can be performed using an encoder and a residual compensator to obtain an encoded audio stream that compensates for the loss of sound quality.

[0033] In one feasible approach, the step of encoding the target audio data according to the audio encoding parameters to obtain an encoded audio stream includes: determining a target bitrate according to the audio encoding parameters; obtaining speech activity features of the target audio data; encoding the target audio data using an adaptive frame length linear predictive coding algorithm according to the target bitrate and the speech activity features to obtain a quantized signal; and performing high-frequency component compensation and sound quality enhancement on the quantized signal to generate an encoded audio stream.

[0034] It should be noted that the target bitrate can be either 16Kbps or 32Kbps. During encoding, a hybrid architecture combining an improved Code Excited Linear Prediction (CELP) encoder and lightweight Generative Adversarial Network (GAN) residual compensation can be employed. The improved CELP encoder, based on the G.729.1 standard, can adaptively adjust the frame length according to VAD features. For example, 5ms short frames are used for active speech segments to ensure real-time performance, while 20ms long frames are used for silent segments to reduce bandwidth consumption. The linear prediction parameters (LSP) are optimized through split vector quantization to reduce encoding errors. The lightweight GAN residual compensator receives the quantized signal after CELP encoding and, based on trained 48kHz ultrawideband speech data, uses perceptual weighting error and Mel-spectral distortion as loss functions to repair the 2-4kHz high-frequency speech components, compensating for audio quality loss at low bitrates. The final output is an encoded audio stream.

[0035] Step S40: Encode the target video data according to the video encoding parameters to obtain an encoded video stream; It should be noted that the video encoding parameters include the video compression mode. Based on the determined compression mode, the target video data is encoded using the corresponding algorithm to obtain the encoded video stream.

[0036] In one feasible approach, the step of encoding the target video data according to the video encoding parameters to obtain an encoded video stream includes: determining a target compression algorithm and compression ratio according to the video encoding parameters; and encoding the target video data according to the target compression algorithm and the compression ratio to obtain an encoded video stream.

[0037] It should be noted that when the video encoding parameters are in shallow compression mode, the Low Latency Lightweight Image Coding (JPEG XS) algorithm is used. This algorithm processes the target video data through 6-level wavelet decomposition and adaptive quantization, achieving a compression ratio of 1:2-1:4 and an encoding latency of <17ms, suitable for high-bandwidth local area network scenarios. When the video encoding parameters are in shallow compression mode, the Versatile Video Coding (VVC) algorithm is used to process the target video data, and an attention mechanism is introduced to optimize motion vector search, reducing computation by 30%, achieving a compression ratio of 100:1-300:1. The final result is a coded video stream that meets the requirements.

[0038] Step S50: Transmit the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

[0039] It should be noted that the transmission rates of the encoded audio and video streams are controlled based on the determined bandwidth allocation ratio (e.g., 3:7). Subsequently, both are encapsulated and transmitted through the network interface. During transmission, if the total bandwidth is 1Mbps, the system will ensure that the audio stream occupies approximately 300Kbps and the video stream occupies approximately 700Kbps, thereby achieving coordinated transmission and avoiding inter-stream contention.

[0040] In one feasible approach, before the step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio, the method further includes: aligning the timestamps of the encoded audio stream and the encoded video stream to obtain target encoded data; acquiring real-time network jitter data; adjusting the size of the transmission buffer according to the real-time network jitter data; and buffering the target encoded data.

[0041] It should be noted that the encoded audio and video streams can be timestamped using Real-time Transport Protocol (RTP) to keep the audio and video synchronization error within 20ms. The system has a built-in adaptive buffer pool with a capacity of 50-200ms, which can adjust the buffer capacity according to the degree of network jitter. When the network is stable, the buffer capacity is set to 50ms to reduce latency, and when the network jitter is severe, it is expanded to 200ms to offset the impact of fluctuations.

[0042] In one feasible approach, the step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio includes: encapsulating the encoded audio stream and the encoded video stream into audio data packets and video data packets according to the bandwidth allocation ratio, and transmitting them; monitoring the packet loss rate during transmission; when the packet loss rate is greater than a preset threshold, performing silence interpolation on the audio data packets, performing adjacent frame duplication on the video data packets, and completing the transmission.

[0043] It should be noted that when the packet loss rate is greater than the threshold, such as 3%, the packet loss hiding strategy is enabled. This involves inserting a silent segment into the audio or repairing it based on interpolation between previous and subsequent frames, and copying adjacent frames to supplement the video, thus avoiding a precipitous drop in image / sound quality.

[0044] This embodiment acquires raw audio and video data, preprocesses them to obtain target audio and video data, acquires network transmission status information, and determines audio encoding parameters, video encoding parameters, and bandwidth allocation ratios based on the network transmission status information. The target audio data is encoded according to the audio encoding parameters to obtain an encoded audio stream; the target video data is encoded according to the video encoding parameters to obtain an encoded video stream; and the encoded audio and video streams are transmitted according to the bandwidth allocation ratios. This embodiment, by acquiring and preprocessing audio and video data, intelligently determining encoding parameters and bandwidth allocation ratios based on network transmission status information, and performing encoding and coordinated transmission accordingly, ensures that the encoding strategy and transmission resources match network conditions, avoids resource competition and quality imbalance between audio and video streams, and improves the overall transmission quality of audio and video data.

[0045] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S20 also includes steps S201 to S202: Step S201: Determine bandwidth fluctuation data based on the network transmission status information, and determine audio encoding parameters based on the bandwidth fluctuation data; It should be noted that bandwidth fluctuation data represents the degree of bandwidth fluctuation, such as the rate of bandwidth decrease over a period of time. For example, if the bandwidth suddenly drops by ≥20%, the audio bitrate will be reduced from 32Kbps to 16Kbps, and the audio encoding parameters at this time will be the target bitrate of 16Kbps. In addition, if high-priority scenarios such as "command speech" are detected, the audio encoding parameters will also include the priority of audio bandwidth allocation.

[0046] Step S202: Predict future bandwidth data based on the network transmission status information, and determine video encoding parameters based on the future bandwidth data; It should be noted that, based on the network transmission status information, a Long Short-Term Memory (LSTM) network can be used to predict the bandwidth for the next second, thus obtaining future bandwidth data. If the predicted bandwidth is ≥2Mbps and the packet loss rate is <3%, the video encoding parameters are determined to be in shallow compression mode; if the predicted bandwidth is <2Mbps or the packet loss rate is ≥3%, the video encoding parameters are determined to be in deep compression mode.

[0047] Step S203: Determine the width allocation ratio based on the network transmission status information, the target audio data, and the target video data.

[0048] It should be noted that the Proximal Policy Optimization (PPO) reinforcement learning algorithm can be used to dynamically adjust the audio and video bandwidth allocation ratio using the perceptual evaluation of speech quality (PESQ) for audio, the peak signal-to-noise ratio (PSNR) for video, and the total bandwidth usage as a ternary reward function. For example, in normal scenarios, the allocation ratio adaptively fluctuates between 1:9 (audio 10%, video 90%) and 3:7 (audio 30%, video 70%). For high-priority scenarios (such as when command speech is detected or the static area of ​​video accounts for ≥80%), the audio bandwidth allocation is temporarily increased to 40% to avoid the loss of critical information due to audio packet loss.

[0049] This embodiment determines bandwidth fluctuation data based on the network transmission status information and determines audio encoding parameters based on the bandwidth fluctuation data; it predicts future bandwidth data based on the network transmission status information and determines video encoding parameters based on the future bandwidth data; and it determines the bandwidth allocation ratio based on the network transmission status information, the target audio data, and the target video data. This embodiment achieves joint dynamic optimization of audio encoding parameters, video encoding parameters, and bandwidth allocation ratio, which can effectively coordinate the real-time adaptation of audio streams and bandwidth allocation, thereby improving the stability and quality of audio and video transmission.

[0050] For example, to help understand the implementation process of the audio / video gateway dynamic encoding cooperative transmission method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3A simplified flowchart of a dynamic encoding and collaborative transmission method for audio and video gateways is provided. Specifically: audio and video data are acquired, preprocessed, and bandwidth is monitored. Then, audio and video encoding are performed, and the encoded audio and video data are collaboratively optimized, bandwidth is allocated, and synchronous encapsulation is conducted. Finally, the data is dynamically buffered and transmitted, outputting the audio and video streams.

[0051] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the audio and video gateway dynamic coding collaborative transmission method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0052] This application also provides an audio / video gateway dynamic encoding and collaborative transmission device, please refer to... Figure 4 The audio / video gateway dynamic encoding collaborative transmission device includes: The acquisition module 10 is used to acquire raw audio data and raw video data, and to preprocess the raw audio data and raw video data to obtain target audio data and target video data. The determination module 20 is used to acquire network transmission status information and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information. Audio module 30 is used to encode target audio data according to the audio encoding parameters to obtain an encoded audio stream; Video module 40 is used to encode the target video data according to the video encoding parameters to obtain an encoded video stream; The transmission module 50 is used to transmit the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

[0053] This embodiment acquires raw audio and video data, preprocesses them to obtain target audio and video data, acquires network transmission status information, and determines audio encoding parameters, video encoding parameters, and bandwidth allocation ratios based on the network transmission status information. The target audio data is encoded according to the audio encoding parameters to obtain an encoded audio stream; the target video data is encoded according to the video encoding parameters to obtain an encoded video stream; and the encoded audio and video streams are transmitted according to the bandwidth allocation ratios. This embodiment, by acquiring and preprocessing audio and video data, intelligently determining encoding parameters and bandwidth allocation ratios based on network transmission status information, and performing encoding and coordinated transmission accordingly, ensures that the encoding strategy and transmission resources match network conditions, avoids resource competition and quality imbalance between audio and video streams, and improves the overall transmission quality of audio and video data.

[0054] In one embodiment, the acquisition module 10 is further configured to perform noise suppression and speech signal-to-noise ratio enhancement on the original audio data to obtain target audio data; and to optimize image contrast and improve video signal-to-noise ratio on the original video data to obtain target video data.

[0055] In one embodiment, the determining module 20 is further configured to determine bandwidth fluctuation data based on the network transmission status information, and determine audio encoding parameters based on the bandwidth fluctuation data; predict future bandwidth data based on the network transmission status information, and determine video encoding parameters based on the future bandwidth data; and determine a bandwidth allocation ratio based on the network transmission status information, the target audio data, and the target video data.

[0056] In one embodiment, the audio module 30 is further configured to: determine a target bitrate based on the audio encoding parameters; acquire speech activity features of the target audio data; encode the target audio data using an adaptive frame length linear predictive coding algorithm based on the target bitrate and the speech activity features to obtain a quantized signal; and perform high-frequency component compensation and sound quality enhancement on the quantized signal to generate an encoded audio stream.

[0057] In one embodiment, the video module 40 is further configured to determine a target compression algorithm and compression ratio based on the video encoding parameters; and to encode the target video data according to the target compression algorithm and compression ratio to obtain an encoded video stream.

[0058] In one embodiment, the transmission module 50 is further configured to timestamp-align the encoded audio stream and the encoded video stream to obtain target encoded data; acquire real-time network jitter data; adjust the size of the transmission buffer based on the real-time network jitter data; and buffer the target encoded data.

[0059] In one embodiment, the transmission module 50 is further configured to encapsulate the encoded audio stream and the encoded video stream into audio data packets and video data packets according to the bandwidth allocation ratio, and transmit them; monitor the packet loss rate during transmission, and when the packet loss rate is greater than a preset threshold, perform mute interpolation processing on the audio data packets, perform adjacent frame duplication processing on the video data packets, and complete the transmission.

[0060] The audio / video gateway dynamic coding collaborative transmission device provided in this application, employing the audio / video gateway dynamic coding collaborative transmission method in the above embodiments, can solve the technical problem of how to improve the transmission quality of audio and video data. Compared with the prior art, the beneficial effects of the audio / video gateway dynamic coding collaborative transmission device provided in this application are the same as those of the audio / video gateway dynamic coding collaborative transmission method provided in the above embodiments, and other technical features in the audio / video gateway dynamic coding collaborative transmission device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0061] This application provides an audio / video gateway dynamic encoding collaborative transmission device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the audio / video gateway dynamic encoding collaborative transmission method in the above embodiment 1.

[0062] The following is for reference. Figure 5 This document illustrates a structural schematic diagram of an audio / video gateway dynamic coding collaborative transmission device suitable for implementing embodiments of this application. The audio / video gateway dynamic coding collaborative transmission device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The audio and video gateway dynamic encoding cooperative transmission device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0063] like Figure 5As shown, the audio / video gateway dynamic encoding cooperative transmission device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the audio / video gateway dynamic encoding cooperative transmission device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the audio / video gateway dynamic coding cooperative transmission device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an audio / video gateway dynamic coding cooperative transmission device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0064] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0065] The audio / video gateway dynamic coding collaborative transmission device provided in this application, employing the audio / video gateway dynamic coding collaborative transmission method in the above embodiments, can solve the technical problem of how to improve the transmission quality of audio and video data. Compared with the prior art, the beneficial effects of the audio / video gateway dynamic coding collaborative transmission device provided in this application are the same as those of the audio / video gateway dynamic coding collaborative transmission method provided in the above embodiments, and other technical features in this audio / video gateway dynamic coding collaborative transmission device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0066] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0067] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0068] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the audio and video gateway dynamic encoding cooperative transmission method in the above embodiments.

[0069] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0070] The aforementioned computer-readable storage medium may be included in the audio / video gateway dynamic coding cooperative transmission device; or it may exist independently and not be assembled into the audio / video gateway dynamic coding cooperative transmission device.

[0071] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the audio / video gateway dynamic encoding collaborative transmission device, the audio / video gateway dynamic encoding collaborative transmission device performs the following actions: acquires raw audio data and raw video data, and preprocesses the raw audio data and raw video data to obtain target audio data and target video data; acquires network transmission status information, and determines audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information; encodes the target audio data according to the audio encoding parameters to obtain an encoded audio stream; encodes the target video data according to the video encoding parameters to obtain an encoded video stream; and transmits the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

[0072] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0074] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0075] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described audio / video gateway dynamic coding cooperative transmission method, which can solve the technical problem of how to improve the transmission quality of audio and video data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the audio / video gateway dynamic coding cooperative transmission method provided in the above embodiments, and will not be repeated here.

[0076] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio / video gateway dynamic encoding collaborative transmission method described above.

[0077] The computer program product provided in this application can solve the technical problem of how to improve the transmission quality of audio and video data. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the audio and video gateway dynamic coding cooperative transmission method provided in the above embodiments, and will not be repeated here.

[0078] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A dynamic encoding and collaborative transmission method for audio and video gateways, characterized in that, The method includes: Collect raw audio data and raw video data, and preprocess the raw audio data and raw video data to obtain target audio data and target video data; Obtain network transmission status information, and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information; The target audio data is encoded according to the audio encoding parameters to obtain an encoded audio stream; The target video data is encoded according to the video encoding parameters to obtain an encoded video stream; The encoded audio stream and the encoded video stream are transmitted according to the bandwidth allocation ratio.

2. The method as described in claim 1, characterized in that, The step of preprocessing the original audio data and original video data to obtain the target audio data and target video data includes: The original audio data is subjected to noise suppression and speech signal-to-noise ratio enhancement to obtain the target audio data; The original video data is then optimized for image contrast and improved for video signal-to-noise ratio to obtain the target video data.

3. The method as described in claim 1, characterized in that, The step of determining audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information includes: Based on the network transmission status information, bandwidth fluctuation data is determined, and audio encoding parameters are determined based on the bandwidth fluctuation data; Based on the network transmission status information, future bandwidth data is predicted, and video encoding parameters are determined according to the future bandwidth data. The width allocation ratio is determined based on the network transmission status information, the target audio data, and the target video data.

4. The method as described in claim 1, characterized in that, The step of encoding the target audio data according to the audio encoding parameters to obtain the encoded audio stream includes: The target bitrate is determined based on the audio encoding parameters; Obtain the speech activity features of the target audio data; Based on the target bitrate and the speech activity characteristics, the target audio data is encoded using an adaptive frame length linear predictive coding algorithm to obtain a quantized signal; The quantized signal is then subjected to high-frequency component compensation and sound quality enhancement to generate an encoded audio stream.

5. The method as described in claim 1, characterized in that, The step of encoding the target video data according to the video encoding parameters to obtain the encoded video stream includes: The target compression algorithm and compression ratio are determined based on the video encoding parameters. The target video data is encoded according to the target compression algorithm and the compression ratio to obtain an encoded video stream.

6. The method as described in claim 1, characterized in that, Before the step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio, the method further includes: The encoded audio stream and the encoded video stream are timestamped to obtain the target encoded data; Acquire real-time network jitter data, adjust the size of the transmission buffer based on the real-time network jitter data, and buffer the target encoded data.

7. The method as described in claim 1, characterized in that, The step of transmitting the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio includes: According to the bandwidth allocation ratio, the encoded audio stream and the encoded video stream are encapsulated into audio data packets and video data packets, and then transmitted. During transmission, the packet loss rate is monitored. When the packet loss rate exceeds a preset threshold, the audio data packets are muted by interpolation, the video data packets are copied from adjacent frames, and the transmission is completed.

8. An audio / video gateway dynamic encoding and transmission device, characterized in that, The device includes: The acquisition module is used to acquire raw audio data and raw video data, and to preprocess the raw audio data and raw video data to obtain target audio data and target video data. The determination module is used to acquire network transmission status information and determine audio encoding parameters, video encoding parameters, and bandwidth allocation ratio based on the network transmission status information. The audio module is used to encode the target audio data according to the audio encoding parameters to obtain an encoded audio stream; The video module is used to encode the target video data according to the video encoding parameters to obtain an encoded video stream; The transmission module is used to transmit the encoded audio stream and the encoded video stream according to the bandwidth allocation ratio.

9. An audio / video gateway dynamic encoding and transmission device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio / video gateway dynamic encoding and transmission method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the audio and video gateway dynamic encoding and transmission method as described in any one of claims 1 to 7.