Audio and video low-delay backhaul method and system in extreme environments
By monitoring the transmission status of multiple links and generating adaptive encoding parameters, the stability and latency issues of audio and video data transmission in extreme environments are resolved. This achieves precise matching between the encoding strategy and network transmission, improving the transmission stability and equipment resource utilization efficiency of the system in extreme environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-10
AI Technical Summary
In extreme environments, existing audio and video data transmission technologies struggle to achieve stable, low-latency backhauls. Especially when network bandwidth fluctuates and links are interrupted, video stuttering, image quality degradation, and insufficient device energy management lead to communication interruptions and wasted device resources.
By monitoring the transmission status of multiple links, analyzing real-time data status, and generating adaptive encoding parameters, the encoding strategy and equipment management are dynamically adjusted to form encoded audio and video streams suitable for redundant transmission. Through multi-link collaborative distribution and transmission, precise management of equipment and energy is achieved.
It achieves precise matching between encoded output and network transmission in extreme environments, reduces end-to-end latency, improves transmission stability and equipment resource utilization efficiency, and enhances the system's task adaptability and continuous survivability in extreme environments.
Smart Images

Figure CN121486642B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio and video emergency transmission, in particular to an audio and video low-delay back transmission method and system in extreme environment. BACKGROUND
[0002] In extreme environments such as the field and disaster sites, it is a major technical challenge to realize stable and low-delay back transmission of audio and video data. The existing conventional scheme relies on fixed coding parameters and a single transmission link. When the network bandwidth fluctuates, the delay increases, or the link is interrupted, it is easy to cause video freezing, serious degradation of picture quality, and even communication interruption, making it difficult to guarantee the real-time and continuity of critical information.
[0003] The existing adaptive technology optimization direction is relatively single. The common method only adjusts the coding rate according to the bandwidth, packet loss rate and other link quality indicators measured by the network side, without considering the real-time complexity change of the video content itself, resulting in that when the picture motion is violent and the details are rich, the instantaneous code rate demand may exceed the actual carrying capacity of the link. Another idea focuses on intelligent analysis of video content to improve coding efficiency, but lacks real-time linkage with the current state of multiple available links, and the generated code stream may not adapt to the differentiated transmission capabilities of different links, causing mismatch between quality and resources.
[0004] In resource-limited extreme scenarios, the energy of the front-end acquisition and relay device is usually limited. The current technical system usually separates data transmission and device management modules. The working mode of the device is mostly statically configured or simply controlled by time sequence, and cannot be closed-loop and finely adjusted according to the actual resource consumption of each data transmission task and the transmission efficiency of the selected link. This management mode is disconnected from the actual transmission efficiency, which may cause energy waste or affect the overall continuous working ability and task reliability of the system at critical moments due to the non-optimal state of the device. SUMMARY
[0005] The purpose of the present application is to provide an audio and video low-delay back transmission method and system in extreme environment to solve the problems raised in the background art.
[0006] To achieve the above purpose, the present application provides an audio and video low-delay back transmission method in extreme environment, which comprises:
[0007] Obtaining an original multi-modal data set from an audio and video acquisition device deployed in a target area;
[0008] Performing multi-link transmission state monitoring processing on the original multi-modal data set to generate a set of currently available network transmission links and their link quality evaluation information;
[0009] Performing state analysis processing on the original multi-modal data set before encoding to obtain a real-time data state set;
[0010] generate an adaptive coding parameter set matching the current environment in combination with the link quality assessment information and the real-time data state set;
[0011] perform dynamic coding processing on the original multi-modal data set according to the adaptive coding parameter set, to form an encoded audio-video stream suitable for redundant transmission;
[0012] perform collaborative distribution transmission processing based on the network transmission link set and the encoded audio-video stream, to generate a backhaul audio-video stream data;
[0013] form a device and energy management instruction set according to the generation process of the backhaul audio-video stream data;
[0014] issue the device and energy management instruction set to the audio-video collection device and related network relay devices.
[0015] Preferably, the multi-link transmission state monitoring processing on the original multi-modal data set generates a set of currently available network transmission links and their link quality assessment information, including:
[0016] periodically detect the availability and basic connection parameters of satellite communication channels, wireless ad hoc network channels, and mobile communication network slice channels;
[0017] perform stability and bandwidth capacity tests on each detected channel to obtain the instantaneous performance indicators of each channel;
[0018] based on the instantaneous performance indicators of all channels, construct a dynamic network map reflecting the topological relationship and quality difference between channels;
[0019] calculate the weight factor of each channel as a transmission path within a preset time window according to the dynamic network map;
[0020] integrate the weight factor and a preset link switching threshold to filter out channels meeting the low delay requirement to form the set of currently available network transmission links, and label each link in the set of network transmission links with the link quality assessment information.
[0021] Preferably, the state analysis processing on the original multi-modal data set before encoding obtains a real-time data state set, including:
[0022] separate audio data streams, video image frame sequences, and three-axis acceleration data from devices from the original multi-modal data set;
[0023] perform decibel level and spectral feature analysis on the audio data stream to obtain an audio state vector;
[0024] performing content variation rate and scene complexity analysis on the video image frame sequence to obtain a video state vector;
[0025] performing posture stability and motion mode analysis on the three-axis acceleration data to obtain a device motion state vector;
[0026] aligning and merging the audio state vector, the video state vector and the device motion state vector in time stamp to generate the real-time data state set.
[0027] Preferably, the adaptive encoding parameter set matched with the current environment is generated by combining the link quality assessment information and the real-time data state set, comprising:
[0028] establishing a mapping relationship table between the bandwidth fluctuation and delay jitter characteristics in the link quality assessment information and the scene complexity and motion intensity in the real-time data state set;
[0029] predefining a set of encoding parameter templates for different types of combinations according to the mapping relationship table, the encoding parameter templates including target code rate, key frame interval and error recovery strength;
[0030] inputting the current link quality assessment information and the real-time data state set into the mapping relationship table for matching to determine the encoding parameter template most suitable for the current environmental characteristics;
[0031] finely adjusting according to the encoding parameter template combined with the specific numerical values in the real-time data state set to finally output the adaptive encoding parameter set.
[0032] Preferably, the original multi-modal data set is dynamically encoded according to the adaptive encoding parameter set to form an encoded audio and video stream suitable for redundant transmission, comprising:
[0033] compressively encoding the video image frame sequence using the target code rate in the adaptive encoding parameter set to generate a main video code stream;
[0034] performing secondary encoding on the same video image frame sequence using parameters lower than the target code rate to generate a redundant video code stream;
[0035] independently encoding the audio data stream using a clock reference matched with video encoding to generate a synchronous audio code stream;
[0036] packing the main video code stream, the redundant video code stream and the synchronous audio code stream according to a predefined packaging format to form a multi-level encoded audio and video stream.
[0037] Preferably, the cooperative distribution transmission processing is performed based on the set of network transmission links and the encoded audio-video stream to generate the returned audio-video stream data, including:
[0038] The encoded audio-video stream is segmented into a plurality of data blocks, and a sequence identifier and a timestamp are attached to each data block;
[0039] The parallel distribution proportion of each data block on multiple links is calculated according to the current bandwidth and delay parameters of each link in the set of network transmission links;
[0040] Each data block is copied into multiple copies according to the distribution proportion, and is sent simultaneously through multiple selected network transmission links;
[0041] At the receiving end, the same data block arriving through different links is de-duplicated and sorted according to the sequence identifier and the timestamp, and the first correctly arrived data block is reassembled into a continuous code stream;
[0042] For data blocks that fail to arrive within a specified time through any link, a quick recovery mechanism based on redundant video code streams in the encoded audio-video stream is started to generate complete returned audio-video stream data.
[0043] Preferably, the generation process of the returned audio-video stream data forms a set of device and energy management instructions, including:
[0044] The actual transmission success rate, average delay, and energy consumption data of each network transmission link in the cooperative distribution transmission processing are monitored in real time;
[0045] The final end-to-end delay and data integrity of the returned audio-video stream data are analyzed and compared with a preset transmission quality threshold to generate a quality evaluation result;
[0046] According to the quality evaluation result and the energy consumption data, a preset device operation mode strategy library is queried to determine corresponding device operation parameter adjustment strategies and energy scheduling strategies;
[0047] The device operation parameter adjustment strategies and energy scheduling strategies are materialized into the set of device and energy management instructions including device operation frequency, transmission power, sleep cycle, and relay node wake-up rules.
[0048] Preferably, the set of device and energy management instructions is sent to the audio-video collection device and related network relay devices, including:
[0049] The set of device and energy management instructions is encapsulated into a specific control protocol data unit;
[0050] Forward error correction encoding is performed on the control protocol data unit to generate a control signaling data packet with strong fault tolerance capability;
[0051] The control signaling data packet is sent to the target device through a currently available and lowest delay network transmission link;
[0052] If the audio and video collection device and the related network relay device are in a dormant or low power consumption state, the control signaling data packet is sent after the wake-up signaling with the highest priority is sent and the audio and video collection device is activated;
[0053] A signaling confirmation response returned by the target device is received to complete the device and the energy management instruction set issuing process.
[0054] Preferably, the instant performance indicators of all channels are used to construct a dynamic network map reflecting the topology relationship and quality difference between channels, including:
[0055] Each detected channel is used as a node, and the round-trip delay and hop count information in the instant performance indicators are used to calculate the logical distance between any two nodes;
[0056] The current available bandwidth and packet loss rate in the instant performance indicators are quantified as the current load quality score of the corresponding node;
[0057] The node, the logical distance and the current load quality score are stored in association using a graph data structure to form an initial network topology map;
[0058] The instant performance indicators are periodically updated, and the logical distance and the current load quality score are recalculated according to the updated indicators to refresh the graph data structure and generate the dynamic network map.
[0059] Preferably, the application further includes an audio and video low delay back transmission system in an extreme environment, which includes a memory, a processor and a computer program stored in the memory and running on the processor, and the processor implements the steps of the above-mentioned audio and video low delay back transmission method in the extreme environment when executing the computer program.
[0060] Compared with the prior art, the application has the following beneficial effects:
[0061] By simultaneously analyzing the link quality evaluation information of multiple network links monitored in real time and the real-time state of the original audio and video data itself, and dynamically generating a set of adaptive encoding parameters based on the joint analysis results, the real-time collaborative optimization of the encoding strategy under the dual factors of transmission constraints and content characteristics is realized. This makes the audio and video stream output accurately match the current unstable network transmission capacity in terms of code rate, picture quality and fault tolerance, avoiding the problem of high complexity scene sudden stall caused by ignoring content characteristics when only considering network conditions, or the problem of encoding output and link carrying capacity mismatch caused by ignoring network when only considering content, so as to obtain more stable visual quality and lower end-to-end delay that better matches the channel capacity in a fluctuating network environment.
[0062] After the audio and video stream is transmitted through multi-link collaborative distribution, control instructions are generated in reverse according to the whole process information of this return task, and are issued to the audio and video acquisition devices and network relay devices at the front end to directly adjust their acquisition, communication and energy consumption related working states. This scheme establishes a direct feedback loop from transmission results to device control, so that the device management strategy can be dynamically and accurately adjusted based on the actual transmission efficiency. When the transmission efficiency is high, the device can be instructed to enter a lower power consumption state to save energy; when the transmission encounters difficulties, the relevant devices can be instructed to improve performance to ensure key data acquisition, realizing the dynamic balance between data transmission tasks and device resource consumption, and enhancing the overall task adaptability and continuous survival ability of the system in extreme environments. BRIEF DESCRIPTION OF DRAWINGS
[0063] Figure 1 The working principle diagram of the audio and video low-delay return method in extreme environments described in the present application;
[0064] Figure 2 The flowchart of multi-link transmission state monitoring processing;
[0065] Figure 3 The flowchart of generating an adaptive encoding parameter set;
[0066] Figure 4 The hierarchical proportion column chart of the encoded audio and video stream;
[0067] Figure 5 The double-index column chart of multi-link transmission performance evaluation. DETAILED DESCRIPTION
[0068] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0069] Referring to Figure 1 The application provides an audio and video low-delay feedback method in an extreme environment, which comprises the following steps: acquiring an original multi-modal data set from an audio and video collection device deployed in a target area, performing multi-link transmission state monitoring processing on the original multi-modal data set to generate a set of currently available network transmission links and link quality evaluation information thereof, performing state analysis processing on the original multi-modal data set before encoding to obtain a real-time data state set, combining the link quality evaluation information and the real-time data state set to generate a set of adaptive encoding parameters matched with the current environment, performing dynamic encoding processing on the original multi-modal data set according to the set of adaptive encoding parameters to form an encoded audio and video stream suitable for redundant transmission, performing cooperative distribution transmission processing based on the set of network transmission links and the encoded audio and video stream to generate a feedback audio and video stream data, and forming a set of device and energy management instructions according to the generation process of the feedback audio and video stream data and issuing the set of device and energy management instructions to the audio and video collection device and related network relay devices.
[0070] In an embodiment of the application, referring to Figure 2 The availability and basic connection parameters of a satellite communication channel, a wireless ad hoc network channel and a mobile communication network slice channel are periodically detected, the stability and bandwidth capacity of each detected channel are tested to obtain instant performance indicators of each channel, a dynamic network map reflecting the topological relationship and quality difference between channels is constructed based on the instant performance indicators of all channels, each detected channel is used as a node, the round-trip delay and hop count information in the instant performance indicators are used to calculate the logical distance between any two nodes, the current available bandwidth and packet loss rate in the instant performance indicators are quantified into the current load quality score of the corresponding node, the node, the logical distance and the current load quality score are associated and stored using a graph data structure to form an initial network topology map, the instant performance indicators are periodically updated, the logical distance and the current load quality score are recalculated based on the updated indicators to refresh the graph data structure and generate the dynamic network map, the weight factor of each channel as a transmission path within a preset time window is calculated based on the dynamic network map, the channels meeting the low-delay requirement are selected based on the weight factor and a preset link switching threshold to form a set of currently available network transmission links, and the link quality evaluation information is marked for each link in the set of network transmission links.
[0071] In a specific implementation, taking a polar scientific expedition scenario as an example, the expedition team members wear devices integrated with audio and video collection and communication functions and are active in the glacier area. In this scenario, the periodic detection of satellite communication channels, wireless ad hoc network channels, and mobile communication network slice channels is triggered at fixed time intervals. The detection of satellite communication channels obtains the current signal-to-noise ratio, Doppler shift, and link budget parameters. The detection of wireless ad hoc network channels measures the received signal strength indication, link quality indication, and adjacent node list between nodes. The detection of mobile communication network slice channels confirms the slice identifier, network attachment state, and radio resource control connection state. The stability and bandwidth capacity test of each detected channel is completed by sending a series of probe packets of different lengths and counting the response time and the number of successful receptions, and then the real-time performance indicators of each channel are obtained, including the current available bandwidth, average round-trip delay, instantaneous packet loss rate, and signal jitter variance.
[0072] In some embodiments, the process of constructing a dynamic network map based on the real-time performance indicators of all channels abstracts each available communication channel entity as a network node. The average round-trip delay and network hop information in the real-time performance indicators are used to solve the logical distance between any two nodes through a distance calculation function, where the logical distance combines the effects of transmission delay and path hop count. The current available bandwidth and instantaneous packet loss rate in the real-time performance indicators are quantified as the current load quality score of the corresponding node through a scoring function. The design of the scoring function makes nodes with high bandwidth and low packet loss rate obtain higher scores. A graph data structure is used to store the nodes, the logical distance between the nodes, and the current load quality score of each node. This structure constitutes the initial network topology graph. The system periodically initiates a new round of channel detection and performance testing, recalculates the logical distance between all nodes and the current load quality score of each node based on the updated real-time performance indicators, and refreshes the entire graph data structure to generate a dynamic network map that reflects the real-time state of the network.
[0073] In a specific implementation, the calculation of the logical distance can use the following formula:
[0074]
[0075] Wherein: represents the logical distance between node and node , represents the average round-trip delay of the link between node and node , represents the transmission of a data packet from node to node The minimum number of hops required. and These are preset normalized weighting coefficients used to balance the contributions of latency and hop count to the logical distance. When calculating the weighting factor for each channel within a preset time window as a transmission path based on the dynamic network map, the weighting factor... The calculation will integrate nodes Current load quality score ,node Logical distance to the aggregation gateway node and nodes In the time window Historical fractional variance An example computational relationship is , where the function This allows channel nodes with high load quality scores, short logical distances, and low historical variance to receive higher weighting factors. When selecting channels that meet low latency requirements by combining the weighting factors with a preset link switching threshold (set to a value between 0 and 1), only channel nodes with weighting factors greater than this threshold are included in the currently available network transmission link set. When labeling each link in the network transmission link set with link quality assessment information, the labeled information includes at least the link's corresponding weighting factor, current load quality score, latest detected available bandwidth, average round-trip delay, and instantaneous packet loss rate.
[0076] It is understandable that the periodic time interval of channel probing can be dynamically adjusted according to the stability of the network environment, and the probing period can be automatically shortened when drastic fluctuations in link quality are detected. Optionally, the probing of wireless ad hoc network channels can be based on the Hello message mechanism of an optimized link-state routing protocol or an on-demand distance vector routing protocol to obtain some real-time performance indicators, thereby reducing the overhead of dedicated probe data packets.
[0077] In one embodiment of the present invention, see [reference] Figure 3The audio data stream, the video image frame sequence and the three-axis acceleration data from the device are separated from the original multi-modal data set, the audio data stream is subjected to decibel level and spectral feature analysis to obtain an audio state vector, the video image frame sequence is subjected to content change rate and scene complexity analysis to obtain a video state vector, the three-axis acceleration data is subjected to posture stability and motion mode analysis to obtain a device motion state vector, the audio state vector, the video state vector and the device motion state vector are time-stamped aligned and merged to generate a real-time data state set, a mapping relationship table between the bandwidth fluctuation and delay jitter features in the link quality assessment information and the scene complexity and motion intensity in the real-time data state set is established, a group of coding parameter templates are predefined for different types of combinations according to the mapping relationship table, the coding parameter templates include a target code rate, a key frame interval and an error recovery strength, the current link quality assessment information and the real-time data state set are input into the mapping relationship table for matching to determine a coding parameter template most suitable for the current environmental features, the coding parameter template is fine-tuned in combination with specific values in the real-time data state set to finally output an adaptive coding parameter set.
[0078] In a specific implementation, the audio and video data state analysis and adaptive coding parameter generation method described in the embodiment can be applied in a rescue scene, and an integrated device on a helmet of a rescuer continuously collects audio and video and motion information of a scene. The operation of separating the audio data stream, the video image frame sequence and the three-axis acceleration data from the device from the original multi-modal data set is performed by a data demultiplexing module, the audio data stream is a real-time audio sampling sequence in a pulse code modulation format, the video image frame sequence is a set of image frames arranged in a time sequence according to capture time in an RGB or YUV format, and the three-axis acceleration data is a sampling sequence of three-axis acceleration values from an inertial measurement unit.
[0079] When performing the decibel level and spectral feature analysis on the audio data stream, the root mean square value of the audio sampling value in a time window is calculated to obtain the average decibel level, and the fast Fourier transform is performed on the audio data in the window to obtain the frequency spectrum, and the energy proportion of the preset frequency band is extracted from the frequency spectrum as the spectral feature, and the decibel level value and the spectral feature value jointly constitute the audio state vector. When performing the content change rate and scene complexity analysis on the video image frame sequence, the content change rate is measured by calculating the sum of the absolute values of the pixel differences between the continuous video image frames, and the scene complexity is quantified by performing edge detection and texture analysis on a single video image frame to count the density of edge points and the entropy value of texture direction, and the quantization results of the content change rate and the scene complexity constitute the video state vector. When performing the posture stability and motion mode analysis on the three-axis acceleration data, the posture stability is evaluated by calculating the variance of the vector sum of the three-axis acceleration data within a period of time, and the motion mode analysis is performed by applying a pattern classification algorithm to the time series of the three-axis acceleration data to determine which state the device is in, such as stationary, walking, running or violent shaking, and the posture stability index and the motion mode classification identifier jointly constitute the device motion state vector. The audio state vector, the video state vector and the device motion state vector are timestamp aligned and merged, and the system timestamp carried by each data vector is used to synchronize all vectors to a unified time reference, and then the timestamp order is spliced into a multi-dimensional real-time data state set.
[0080] In some embodiments, a mapping relationship table is established between the bandwidth fluctuation, delay jitter characteristics in the link quality evaluation information and the scene complexity, motion intensity in the real-time data state set, the mapping relationship table is a predefined query table, the row index is composed of the combination of the bandwidth fluctuation level and the delay jitter level, and the column index is composed of the combination of the scene complexity level and the motion intensity level. The bandwidth fluctuation level is divided according to the standard deviation of the bandwidth history data, the delay jitter level is divided according to the standard deviation of the delay history data, the scene complexity level is divided according to the quantization value of the scene complexity, and the motion intensity level is divided according to the posture stability index and the motion mode classification in the device motion state vector. A set of predefined encoding parameter templates is stored in each cell of the mapping relationship table, and the encoding parameter template explicitly includes three parameters of target code rate, key frame interval and error recovery strength. When a set of encoding parameter templates is predefined for different types of combinations according to the mapping relationship table, the encoding parameter template corresponding to the combination of high bandwidth fluctuation, high delay jitter, high scene complexity and high motion intensity will be configured with lower target code rate, shorter key frame interval and higher error recovery strength; and the encoding parameter template corresponding to the combination of low bandwidth fluctuation, low delay jitter, low scene complexity and low motion intensity will be configured with higher target code rate, longer key frame interval and lower error recovery strength.
[0081] In a specific implementation, the current link quality assessment information is matched with the real-time data state set input mapping relationship table to determine the encoding parameter template that best fits the current environmental characteristics. This process involves a query operation: first, determine the row index according to the latest bandwidth fluctuation value and delay jitter value; second, determine the column index according to the scene complexity value and motion intensity value at the latest time in the real-time data state set; finally, locate the unique cell in the mapping relationship table through the row index and column index, and obtain the encoding parameter template stored in the cell as the basic template. According to the encoding parameter template, fine-tuning operations are performed in combination with the specific values in the real-time data state set, for example, using a linear interpolation function, adjusting the specific value of scene complexity in a small range around the target code rate reference value set in the basic template. The relationship between the specific value of scene complexity and the adjusted target code rate can be expressed as:
[0082]
[0083] wherein: is the target code rate reference value set in the basic template, is the scene complexity level reference value based on which the mapping relationship table column index is divided, is a preset adjustment coefficient. After similar specific value-based fine-tuning of all encoding parameters, the complete adaptive encoding parameter set is finally output.
[0084] It can be understood that the generation frequency of the audio state vector, the video state vector, and the device motion state vector can be configured according to the computing resources and real-time requirements. Optionally, the motion pattern analysis of the three-axis acceleration data can be completed using a lightweight algorithm based on threshold judgment or a pre-trained lightweight neural network model. In some embodiments, the construction of the mapping relationship table can be based on statistical analysis of a large amount of historical transmission data and encoding effects, or based on expert experience knowledge. The specific implementation of error recovery strength can refer to the redundancy ratio of forward error correction coding or the number of retransmissions of the automatic retransmission request mechanism. It can be understood that the adjustment coefficient used in the fine-tuning process can be preset according to the preference of code rate and quality in different application scenarios. Optionally, the scene complexity analysis of the video image frame sequence can be performed on the reduced resolution image to reduce the computational load. The level division of motion intensity can mainly be based on the value range of the posture stability index in the device motion state vector, and the result of motion pattern classification can be referred to.
[0085] In one embodiment of the application, a sequence of video image frames is compressed and encoded using a target bit rate from a set of adaptive encoding parameters to generate a primary video bitstream, the same sequence of video image frames is re-encoded using a lower bit rate to generate a redundant video bitstream, an audio data stream is independently encoded using a clock reference matching the video encoding to generate a synchronized audio bitstream, and the primary video bitstream, the redundant video bitstream and the synchronized audio bitstream are packaged according to a predefined encapsulation format to form a multi-layered encoded audio-visual stream.
[0086] In a specific implementation, a scenario of a high mountain search and rescue is taken as an example, and a helmet device of a leader of a mountaineering team works as an audio-visual acquisition device in a complex mountain environment. The operation of compressing and encoding a sequence of video image frames using a target bit rate from a set of adaptive encoding parameters is performed by a primary video encoder, which is configured as an H.264 or H.265 encoder, and a bit rate control module of the primary video encoder is set as a constant bit rate mode or a variable bit rate mode, a target bit rate value is directly from a target bit rate parameter in the set of adaptive encoding parameters, and the primary video encoder performs motion estimation, transformation, quantization and entropy encoding on an input sequence of YUV format video image frames to output a primary video bitstream conforming to a predetermined format. The operation of re-encoding the same sequence of video image frames using a lower bit rate is performed by a redundant video encoder, which can be another instance of the same type of encoder as the primary video encoder, and an input of the redundant video encoder is the same sequence of video image frames, but a target bit rate value configured for the redundant video encoder is a fixed proportion of a target bit rate of the primary video encoder or a lower value obtained through calculation, a secondary encoding process is independently performed and a redundant video bitstream is generated, and the redundant video bitstream corresponds to the same original image sequence as the primary video bitstream in content but has differences in encoding precision and detail retention.
[0087] In a specific implementation, the operation of encoding the audio data stream with a clock reference matching the video encoding is performed by an audio encoder, which receives the audio data stream in pulse code modulation format, and the system clock reference used inside the encoder is strictly synchronized with the system clock reference of the main video encoder, usually based on the same high-precision clock source. The audio encoder selects an appropriate encoding algorithm and bit rate according to the audio quality requirement implied by the adaptive encoding parameter set, for example, using adaptive differential pulse code modulation or advanced audio coding algorithm, and embeds timestamp information consistent with the video stream timestamp system in the output synchronous audio code stream. The operation of packaging the main video stream, redundant video stream and synchronous audio stream according to the predefined packaging format is performed by a multiplexer, and the predefined packaging format can be MPEG-2 transport stream format or other custom container format. The multiplexer attaches a packet header to each segment of main video stream data, redundant video stream data and synchronous audio stream data, which contains stream type identifier, sequence number and timestamp based on the common clock reference, and then interleaves the data packets of the three streams to form a logically unified but physically multi-level data stream, i.e. a multi-level encoded audio-video stream.
[0088] In some embodiments, the calculation of the target code rate of the redundant video stream can be determined proportionally based on the target code rate of the main video stream, for example, setting the target code rate of the redundant video stream to 30% to 50% of the target code rate of the main video stream. The key frame interval of the encoding parameter set used in the generation of the redundant video stream can be consistent with the main video stream to ensure the alignment of random access points, but the precision-affected settings such as quantization parameter can be adjusted to achieve a lower code rate. It can be understood that the redundant video encoding process can use a simpler set of encoding tools or reduce the image resolution to further reduce the computational complexity and output code rate.
[0089] In some embodiments, the encoding of the main video stream and the redundant video stream can be performed in parallel on different processor cores or hardware encoding units to reduce the overall processing delay. The encoding of the synchronous audio stream must ensure that its timestamp strictly corresponds to the timestamp of the video frame, and the multiplexer will align the audio data packets with the video data packets in the time axis during the packaging process. Alternatively, the redundant video stream can only encode key frames or some important predicted frames in the sequence of video image frames, rather than the complete sequence, to balance the redundancy protection and bandwidth occupation.
[0090] In a specific implementation, the specific determination of the parameters below the target code rate can be calculated according to a formula, which is:
[0091]
[0092] wherein: Rtarget represents the target bit rate of the redundant video bitstream, Rtarget represents the target bit rate of the main video bitstream, a is a base scaling factor between 0 and 1, a is a decay factor, Qavg is the average of the main quantization parameter used by the main video encoder, Qmin is the minimum quantization parameter allowed by the system. This formula makes the target bit rate of the redundant video bitstream decrease more when the main video bitstream uses a larger quantization parameter (i.e. lower quality) due to the content complexity, because the main video bitstream itself can have a lower tolerance to errors or the value of the redundancy protection is relatively lower at this time.
[0093] Optionally, the pre-defined encapsulation format can set an identification bit in the packet header to distinguish the main video bitstream data packet, the redundant video bitstream data packet and the synchronous audio bitstream data packet. It can be understood that the generated multi-level encoded audio and video stream maintains decoupling in structure, allowing stripping or separate processing of a certain level of bitstream as needed during transmission or processing.
[0094] Referring to Figure 4 , which is a bar chart of the proportion of each level of the encoded audio and video stream, a professional visualization chart in the "bitstream packaging" stage of the extreme environment audio and video low-delay backhaul. The proportion of the main video bitstream is always the highest (about 70%), which is consistent with the design logic of "main bitstream carrying core content", ensuring the main quality of the audio and video backhaul. The proportion of the redundant video bitstream is stable at about 20%, which provides redundancy for data recovery, and avoids excessive bandwidth occupation. The proportion of the audio bitstream is low and stable, which embodies the characteristics of "low bandwidth overhead of audio", while ensuring the synchronization with the video. This chart is used to verify the rationality of the bitstream packaging strategy, to ensure the balanced allocation of bandwidth among the main bitstream, the redundant bitstream and the audio bitstream. The stable proportion distribution shows that the dynamic adjustment of the encoding parameters (main / redundant bit rate) is controllable and adapts to the bandwidth fluctuation in extreme environments.
[0095] In an embodiment of the present application, the encoded audio and video stream is divided into multiple data blocks, and a sequence identifier and a timestamp are attached to each data block. According to the current bandwidth and delay parameters of each link in the network transmission link set, the parallel distribution proportion of each data block on multiple links is calculated. According to the distribution proportion, each data block is copied into multiple copies and sent simultaneously through multiple selected network transmission links. At the receiving end, the same data blocks arriving through different links are de-duplicated and sorted according to the sequence identifier and the timestamp, and the first correctly arrived data block is reassembled into a continuous bitstream. For data blocks that fail to arrive within a specified time through any link, a fast recovery mechanism based on the redundant video bitstream in the encoded audio and video stream is started to generate a complete backhaul audio and video stream data.
[0096] In a specific implementation, taking a forest fire fighting scenario as an example, the communication device carried by a firefighter as a sender operates at the edge of the fire. The operation of splitting the encoded audio and video stream into multiple data blocks and attaching a sequence identifier and a timestamp to each data block is performed by the splitting and packaging module of the sender. The encoded audio and video stream is input as a continuous binary stream, and the splitting module cuts the binary stream into multiple data blocks according to a preset fixed size or a dynamically determined size according to the maximum transmission unit of the network. Each data block is assigned a strictly increasing sequence identifier and a current timestamp taken from the high-precision clock of the sender, and the sequence identifier and the timestamp are packaged together into the header of the data block. According to the current bandwidth and delay parameters of each link in the network transmission link set, the process of calculating the parallel distribution proportion of each data block on multiple links is completed by a dynamic scheduler. The dynamic scheduler maintains a table containing the real-time parameters of each link in the network transmission link set. For each data block to be sent, referring to Table 1, the dynamic scheduler calculates the distribution proportion based on the latest link parameters.
[0097] Table 1: Network transmission link parameters and data block distribution proportion calculation table
[0098]
[0099] According to the distribution proportion, each data block is copied into multiple copies and sent through the selected multiple network transmission links simultaneously. The operation of copying and routing each data block into multiple copies and sending them through the selected multiple network transmission links simultaneously is performed by the copying and routing module of the sender. The distribution proportion determines the number of copies of each data block sent on different links. For example, a data block is copied into 100 logical subunits. According to the proportion in Table 1, 30 subunits are sent through the satellite link, 45 subunits are sent through the ad hoc network link A, and 25 subunits are sent through the cellular network link. These subunits are placed in the corresponding link's sending queue and sent out simultaneously through the respective network interfaces.
[0100] In practice, at the receiving end, the deduplication and sorting of the same data block arriving from different links based on the sequence identifier and timestamp are handled by the reassembly buffer and processing logic. The receiving end maintains a state for each expected sequence identifier. When a data block is received from any link, its sequence identifier and timestamp are extracted. If the data block with that sequence identifier is arriving for the first time, it is stored in the reassembly buffer, sorted by sequence identifier, and marked as "received". Subsequent data blocks with the same sequence identifier are considered redundant copies and discarded. The operation of reassembling the first correctly arriving data block into a continuous bitstream continues at the receiving end. The reassembly buffer outputs data blocks in sequence by sequence identifier, removes the sequence identifier and timestamp information appended to their headers, and splices the payload portion of the data blocks to restore a continuous bitstream. For data blocks that fail to arrive within the specified time via any link, a fast recovery mechanism based on redundant video bitstreams in the encoded audio and video stream is activated to generate complete return audio and video stream data. The fast recovery mechanism is triggered when a data block with a certain sequence identifier is still missing after the timeout. The receiver sends a negative acknowledgment request for the missing sequence identifier to the sender. Alternatively, the receiver uses redundant video bitstreams in the successfully received multi-level encoded audio and video stream, combined with correctly received data blocks before and after, to perform error concealment or approximate reconstruction to fill the data gaps and ultimately form complete return audio and video stream data.
[0101] In some embodiments, the distribution ratio of data blocks can be calculated using a weighted function based on link bandwidth and latency. For example, for a single link... Its distribution ratio weight It can be calculated as:
[0102]
[0103] in: Indicates link Current bandwidth, Indicates link The current delay, Indicates the number of selected network transmission links and the distribution ratio. That is It is understandable that the size of the data block needs to be balanced between header overhead, transmission efficiency, and error recovery granularity. Data blocks that are too large will reduce the flexibility of distribution, while data blocks that are too small will increase header overhead.
[0104] Optionally, the deduplication operation at the receiving end can be performed at the data link layer or the application layer. It can be understood that requesting retransmission from the sending end in the fast recovery mechanism is a reliable but possibly increases the delay, while local recovery using redundant video code streams is a low-delay but possibly loses part of the quality, the system can dynamically select according to the strategy. In some embodiments, the timestamp attached by the sending end to the data block can be used by the receiving end to evaluate the current delay of each link and perform audio and video synchronization. Optionally, when distributing in parallel, multiple copies of the same data block can be sent through the same path or through different paths to maximize path diversity. For data blocks that fail to arrive within a specified time through any link, the specific process of starting the fast recovery mechanism based on the redundant video code stream in the encoded audio and video stream can include decoding the data at the corresponding position in the redundant video code stream and up-converting it to replace the lost main video code stream data.
[0105] In one embodiment of the application, the actual transmission success rate, average delay and energy consumption data of each network transmission link in the cooperative distribution transmission process are monitored in real time, the end-to-end delay and data integrity of the backhaul audio and video stream data are analyzed, and compared with the preset transmission quality threshold to generate a quality evaluation result. According to the quality evaluation result and the energy consumption data, the preset device operation mode strategy library is queried to determine the corresponding device operation parameter adjustment strategy and energy scheduling strategy. The device operation parameter adjustment strategy and energy scheduling strategy are materialized into a device and energy management instruction set containing device operation frequency, transmission power, sleep cycle and relay node wake-up rule. The device and energy management instruction set is encapsulated into a control protocol data unit, the control protocol data unit is forward error correction encoded to generate a control signaling data packet with strong fault tolerance capability, the control signaling data packet is sent to the target device through the currently available and lowest delay network transmission link. If the audio and video acquisition device and the related network relay device are in a sleep or low-power state, the wake-up signaling with the highest priority is sent first, and then the control signaling data packet is sent after it is activated. The signaling acknowledgment response returned by the target device is received to complete the device and energy management instruction set delivery process.
[0106] In a specific implementation, the operation of monitoring the actual transmission success rate, average delay and energy consumption data of each network transmission link in the cooperative distribution transmission process is performed by a monitoring agent module, which resides on a key node on the transmission path. The ratio of the number of data packets successfully transmitted to the total number of transmitted data packets on each active link is collected as the actual transmission success rate. The time delay samples of each data block from transmission to reception confirmation are collected and the sliding window average is calculated as the average delay. The integral value of the product of current and voltage or the value of the power consumption count register is read from the device power management unit and network interface chip as the energy consumption data.
[0107] The process of analyzing the end-to-end latency and data integrity of the returned audio and video stream data and comparing them with preset transmission quality thresholds to generate a quality assessment result is completed by the central control unit. The final end-to-end latency is calculated by comparing the source timestamp on the audio and video acquisition device with the timestamp when the application layer at the receiving end receives the complete frame. Data integrity is assessed by checking the sequence identifier continuity of the reassembled audio and video stream at the receiving end and the error report of the decoder. The preset transmission quality thresholds include the maximum allowable end-to-end latency limit and the minimum allowable data integrity percentage. The quality assessment result is output as a status code or vector to indicate whether the current transmission quality is "good", "compliant", or "non-compliant". Based on the quality assessment results and energy consumption data, the system queries the pre-defined equipment operating mode strategy library to determine the corresponding equipment operating parameter adjustment strategy and energy scheduling strategy. The equipment operating mode strategy library is a predefined decision table or rule set. Its inputs are the quality assessment result status code and the energy consumption data of each link, and its output is the specific adjustment strategy. For example, when the quality assessment result is "excellent" and the energy consumption per unit bit of a certain link exceeds the threshold, the strategy library may output a strategy to reduce the transmission power of that link; when the quality assessment result is "unsatisfactory" and the energy consumption data still has a margin, the strategy library may output a strategy to increase the operating frequency of critical links.
[0108] The process of concretizing equipment operating parameter adjustment strategies and energy scheduling strategies into a set of device and energy management instructions, including equipment operating frequency, transmit power, sleep cycle, and relay node wake-up rules, is executed by an instruction generator. The instruction generator converts the abstract strategies output from the strategy library into specific parameter values that can be executed by the device. For example, the "reduce transmit power" strategy is converted into the instruction "set transmit power to 15dBm," and the "adjust sleep cycle" strategy is converted into the instruction "set sleep duration to 200 milliseconds." All these specific instructions are aggregated and encapsulated into a set of device and energy management instructions. The encapsulation of this set of instructions into control protocol data units follows a lightweight binary protocol that defines the field structure and encoding method for instruction types, target device identifiers, parameter lists, and checksums.
[0109] In specific implementations, the operation of forward error correction encoding on control protocol data units to generate control signaling data packets with strong fault tolerance is performed using Reed-Solomon encoding or convolutional encoding, and redundant check bytes are added to the control protocol data units so that the control signaling data packets can tolerate a certain degree of bit errors during transmission without the need for retransmission. The routing of the control signaling data packets to the target device through the currently available network transmission links with the lowest delay is based on real-time updates of the network state, and the sending module selects the link with the smallest average delay value from the set of network transmission links as the preferred path for carrying the control signaling. If the audio and video collection device and the related network relay device are in a sleep or low-power state, the control signaling data packet is first sent after the wake-up signaling with the highest priority is sent to activate it. The wake-up signaling is a special, energy-concentrated short pulse signal or a broadcast packet containing a minimal protocol, which is designed to reliably trigger the radio frequency front end of the device in a deep sleep state to start. The signaling acknowledgment response returned by the receiving target device completes the device and energy management instruction set delivery process. The sending end starts a timer after sending the control signaling data packet. If the acknowledgment response packet returned by the target device is received before the timeout, the delivery is considered successful; otherwise, retransmission or delivery through a backup link may be triggered.
[0110] In some embodiments, the rules of the device operation mode strategy library can be formally represented as a set of "IF (condition) THEN (action)", such as "IF the quality assessment result is up to standard AND the remaining battery power is less than 20% THEN execute the energy scheduling strategy: turn off the redundant relay node". The instruction generator may, in particularizing the device operating frequency parameter, balance performance and energy consumption according to a formula:
[0111]
[0112] wherein: represents the newly set device operating frequency, represents the basic operating frequency of the device, is the basic adjustment coefficient mapped from the strategy library according to the quality assessment result, represents the difference between the current measured end-to-end delay and the target delay, is the maximum allowed delay deviation constant. It can be understood that the strength of forward error correction encoding on control protocol data units can be dynamically adjusted according to the average bit error rate of the current link.
[0113] Reference Figure 5This is a double-index column chart for multi-link transmission performance evaluation, which is a professional chart in the "transmission link quality evaluation" stage of the audio and video low-delay backhaul in extreme environments. This chart is used to guide the dynamic selection of links in extreme environments. If low delay is prioritized, "wireless ad hoc network channel" can be selected first. If long-distance transmission is prioritized, "satellite communication channel" can be selected. The "data integrity is 100%" of the three types of links reflects the effectiveness of the "redundant transmission + error recovery" strategy in the backhaul system. Even if the single-link delay is high, the data will not be lost through multi-link cooperation.
[0114] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one entity or action from another, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0115] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, alternatives, and variations can be made in the embodiments without departing from the spirit and scope of the present application as defined by the appended claims and their equivalents.
Claims
1. A method for low-latency audio and video backhaul under extreme environments, characterized in that, The method includes: Obtain the raw multimodal data set from the audio and video acquisition devices deployed in the target area; Perform multi-link transmission status monitoring processing on the original multimodal data set to generate a set of currently available network transmission links and their link quality assessment information; The original multimodal data set is subjected to state analysis processing before encoding to obtain a real-time data state set; By combining the link quality assessment information and the real-time data status set, an adaptive coding parameter set that matches the current environment is generated; The original multimodal data set is dynamically encoded according to the adaptive encoding parameter set to form an encoded audio and video stream suitable for redundant transmission; Based on the network transmission link set and the encoded audio and video stream, a collaborative distribution transmission process is performed to generate return audio and video stream data. Based on the generation process of the returned audio and video stream data, a set of device and energy management instructions is formed; The device and energy management command set are sent to the audio and video acquisition device and related network relay devices; The step of performing multi-link transmission status monitoring processing on the original multimodal data set to generate a set of currently available network transmission links and their link quality assessment information includes: Periodically probe the availability and basic connectivity parameters of satellite communication channels, wireless ad hoc network channels, and mobile communication network slice channels; Stability and bandwidth capacity tests were performed on each detected channel to obtain real-time performance metrics for each channel. Based on the real-time performance metrics of all channels, a dynamic network map reflecting the topological relationships and quality differences between channels is constructed. Based on the dynamic network map, calculate the weight factor for each channel as a transmission path within a preset time window; By combining the weighting factors with the preset link switching threshold, channels that meet the low latency requirements are selected to form the currently available network transmission link set, and the link quality assessment information is labeled for each link in the network transmission link set. The process of generating the returned audio and video stream data forms a set of device and energy management instructions, including: Real-time monitoring of the actual transmission success rate, average latency, and energy consumption data of each network transmission link in the collaborative distribution transmission process; The end-to-end latency and data integrity of the returned audio and video stream data are analyzed and compared with a preset transmission quality threshold to generate a quality assessment result. Based on the quality assessment results and the energy consumption data, query the preset equipment working mode strategy library to determine the corresponding equipment working parameter adjustment strategy and energy dispatch strategy. The device operating parameter adjustment strategy and energy scheduling strategy are specified as a set of device and energy management instructions, including device operating frequency, transmission power, sleep cycle, and relay node wake-up rules.
2. The method for low-latency audio and video backhaul under extreme environments according to claim 1, characterized in that, The pre-encoding state analysis processing of the original multimodal data set yields a real-time data state set, including: The audio data stream, video image frame sequence, and three-axis acceleration data from the device are separated from the original multimodal data set; The audio data stream is analyzed for decibel level and spectral characteristics to obtain an audio state vector; The video image frame sequence is analyzed for content change rate and scene complexity to obtain a video state vector; The attitude stability and motion pattern analysis of the triaxial acceleration data are performed to obtain the device motion state vector; The audio state vector, the video state vector, and the device motion state vector are timestamped and merged to generate the real-time data state set.
3. The method for low-latency audio and video backhaul under extreme environments according to claim 1, characterized in that, The step of combining the link quality assessment information and the real-time data status set to generate an adaptive coding parameter set that matches the current environment includes: Establish a mapping table between bandwidth fluctuation and latency jitter characteristics in the link quality assessment information and scene complexity and motion intensity in the real-time data status set; Based on the mapping table, a set of encoding parameter templates are predefined for different types of combinations. The encoding parameter templates include target bitrate, keyframe interval and error recovery strength. The current link quality assessment information is matched with the real-time data status set in the mapping table to determine the coding parameter template that best matches the current environment characteristics; Based on the encoding parameter template, fine-tuning is performed using the specific values in the real-time data state set, and finally the adaptive encoding parameter set is output.
4. The method for low-latency audio and video backhaul under extreme environments according to claim 1, characterized in that, The step of dynamically encoding the original multimodal data set according to the adaptive encoding parameter set to form an encoded audio and video stream suitable for redundant transmission includes: The video image frame sequence is compressed and encoded using the target bitrate in the adaptive coding parameter set to generate the main video stream. The same video image frame sequence is re-encoded using parameters lower than the target bitrate to generate a redundant video bitstream; The audio data stream is independently encoded using a clock reference that matches the video encoding to generate a synchronized audio bitstream; The main video stream, the redundant video stream, and the synchronous audio stream are packaged according to a predefined encapsulation format to form a multi-layered encoded audio and video stream.
5. The method for low-latency audio and video backhaul under extreme environments according to claim 1, characterized in that, The step of performing coordinated distribution and transmission processing based on the network transmission link set and the encoded audio and video stream to generate return audio and video stream data includes: The encoded audio and video stream is divided into multiple data blocks, and a sequence identifier and timestamp are attached to each data block; Based on the current bandwidth and delay parameters of each link in the network transmission link set, calculate the parallel distribution ratio of each data block on multiple links; Based on the distribution ratio, each data block is copied into multiple copies and sent simultaneously through multiple selected network transmission links; At the receiving end, based on the sequence identifier and timestamp, the same data block arriving through different links is deduplicated and sorted, and the first correctly arriving data block is reassembled into a continuous bitstream; For data blocks that fail to arrive within the specified time via any link, a fast recovery mechanism based on redundant video bitstreams in the encoded audio and video stream is initiated to generate complete return audio and video stream data.
6. The method for low-latency audio and video backhaul under extreme environments according to claim 1, characterized in that, The step of sending the device and energy management command set to the audio and video acquisition device and related network relay devices includes: The device and energy management instruction set are encapsulated into a specific control protocol data unit; The control protocol data unit is forward-corrected and encoded to generate a control signaling data packet with strong fault tolerance. The control signaling data packet is sent to the target device via the currently available network transmission link with the lowest latency. If the audio / video acquisition device and the related network relay device are in a sleep or low-power state, a wake-up signaling message with the highest priority will be sent first. After they are activated, the control signaling data packet will be sent. Receive the signaling confirmation response returned by the target device to complete the process of issuing the device and the set of energy management instructions.
7. The method for low-latency audio and video backhaul under extreme environments according to claim 2, characterized in that, The method for constructing a dynamic network map reflecting the topological relationships and quality differences between channels based on the real-time performance metrics of all channels includes: Using each detected channel as a node, the logical distance between any two nodes is calculated using the round-trip delay and hop count information in the real-time performance metrics. The current available bandwidth and packet loss rate in the real-time performance metrics are quantified into the current load quality score of the corresponding node. Using a graph data structure, the nodes, the logical distances, and the current load quality scores are associated and stored to form an initial network topology graph. The real-time performance metrics are periodically updated, and the logical distance and the current load quality score are recalculated based on the updated metrics to refresh the graph data structure and generate the dynamic network map.
8. A low-latency audio and video transmission system for extreme environments, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the low-latency audio and video backhaul method under extreme environments as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Adaptive, scalable packet loss recovery
CN101779377A
Video encoder parameter dynamic adjustment method
CN121037570A
Multidirectional frame audio stream transmission method, device, equipment and medium
CN121151327A