Security monitoring interactive video and audio real-time synchronization transmission system

By implementing physical binding of audio and video payloads at the bitstream encapsulation level of the audio and video synchronous transmission system, the audio and video synchronization problem in traditional systems under sudden large-volume congestion scenarios is solved, and the stability of audio and video synchronization and decoding end in complex networks is achieved.

CN122138000APending Publication Date: 2026-06-02CHONGQING JIZHONG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING JIZHONG TECH CO LTD
Filing Date
2026-04-14
Publication Date
2026-06-02

Smart Images

  • Figure CN122138000A_ABST
    Figure CN122138000A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio-visual synchronization transmission technology and discloses an interactive real-time synchronous audio-visual transmission system for security monitoring. The system includes: video encoding, audio feature extraction, video encoding rate adjustment, interactive control, and a composite bitstream encapsulation module. The video encoding module extracts the motion vector magnitude of the core motion region, and the video encoding rate adjustment module converts the magnitude into audio frequency band energy allocation weights. The composite bitstream encapsulation module embeds audio slices and previous audio historical state parameters into video frames to supplement and enhance information, and fills bits according to a fixed load span to maintain a constant physical load length. This invention integrates audio data into the video frame syntax structure, eliminates synchronization deviations caused by conflicts between different sampling mechanisms, ensures the determinism of decoding offset calculation delay, and ensures semantic continuity during command and interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an interactive real-time synchronous transmission system for security monitoring, belonging to the field of audio-visual synchronous transmission technology. Background Technology

[0002] Current audio-visual synchronous transmission systems typically employ a dual-track independent stream architecture. In this architecture, the acquisition end processes the video and audio streams separately and sets timestamps based on real-time transmission protocols. The receiving end constructs a buffer queue, calculates the arrival time difference of each data packet, and performs audio-visual alignment based on timestamp mapping logic. This mode uses increased receiver buffer depth to achieve consistency at the logic layer, maintaining basic synchronization in typical streaming media playback scenarios. Analysis reveals that when the system experiences sudden high-volume congestion, there is a physical mismatch between the instantaneous network pressure generated by video keyframes and the constant bitrate characteristics of the audio stream. Due to the lack of strong coupling between the audio and video payloads at the bitstream encapsulation level, random jitter in the transmission channel causes nonlinear arrival time difference offsets between the two types of payloads. If the receiver buffer is compressed to pursue real-time interaction, it will inevitably induce audio-visual timing deviations at the decoding end.

[0003] To address the aforementioned technical constraints, existing improvement approaches typically focus on enhancing the accuracy of the global clock protocol or increasing network channel redundancy. However, these linear optimization methods fail to address the underlying characteristics of the independent syntactic structure of audio and video payloads. As long as the audio and video information lacks a physically bound deterministic structure within the video network abstraction layer unit, the receiver cannot achieve time-dependent parsing of composite bitstreams without establishing a jitter smoothing mechanism. This synchronization method based on out-of-band control information exhibits significant systemic vulnerability to the dynamic evolution of network fluctuations. As mentioned above, existing systems suffer from limitations not only in the underlying transmission architecture hardware but also in the software control methods. For example, Chinese invention patent application CN121664937A discloses a system medium for synchronizing and aligning various types of video signals by receiving the transmission time from a PTP clock source. The PTP timestamp is embedded in the signal, and the delay frame number is calculated based on the timestamp difference at the receiving end, thereby implementing buffer alignment. However, penetrating the surface alignment logic, it can be seen that the underlying objective attribute is still the traditional control method of out-of-band clock reference and receiver buffer alignment. The software control method implicitly relies on the ideal prerequisite that the global PTP clock network always maintains an absolute steady state under complex heterogeneous routing. In non-ideal dynamic evolution scenarios such as security monitoring and other sudden large-volume congestion channels with extreme degradation, the external clock distribution link itself is very susceptible to nonlinear jitter, causing the PTP time base used as the calculation benchmark to drift. The approach of forcibly constructing the virtual delay at the receiving end by relying on difference conversion cannot reduce the absolute system latency. On the contrary, the dynamic adjustment of the buffer depth causes the decoding end to be disordered. The fundamental mismatch between the core preset premise and the actual harsh network boundary conditions determines that it is impossible to guarantee the rigid alignment of audio and video timing under extremely low latency requirements.

[0004] Therefore, how to achieve physical binding of audio and video payloads at the bitstream encapsulation level, thereby avoiding transmission delay caused by synchronous buffering at the receiving end, is the technical problem to be solved by this invention. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A security monitoring interactive audio-visual real-time synchronous transmission system, comprising a video encoding module, an audio feature extraction module, a video encoding rate adjustment module, an interactive control module, and a composite bitstream encapsulation module:

[0006] The video encoding module is used to acquire the monitoring video stream and extract the motion vector magnitude values ​​of video macroblocks within the core motion region;

[0007] The audio feature extraction module is used to synchronously acquire audio streams and extract the current audio slice and previous audio historical state parameters based on a time axis sliding window.

[0008] The video encoding rate adjustment module is connected to both the video encoding module and the audio feature extraction module. It is used to establish a mapping relationship between motion vector magnitude and audio bit allocation weights, and to convert the distribution characteristics of motion vector magnitude into frequency band energy allocation weights for audio slices.

[0009] The composite bitstream encapsulation module is used to encapsulate the audio slices adjusted by frequency band energy allocation weights and the composite audio descriptor formed by nesting the preceding audio historical state parameters into the supplementary enhancement information (SEI) information of the video network abstraction layer unit. The SEI information has a fixed byte payload span, and the composite bitstream encapsulation module performs bit padding on the SEI information according to the fixed byte payload span to keep the output SEI information constant in physical length.

[0010] Preferably, the composite bitstream encapsulation module includes: a data anchoring unit and a redundancy padding unit; the data anchoring unit is used to rigidly anchor the audio slice data to the starting offset address of the SEI information when the bitrate of the audio slice fluctuates; the redundancy padding unit is used to calculate the free byte span outside the audio slice data in the SEI information and fill the free byte span with zero-value bitstream to isolate the impact of the characteristic changes of the audio slice on the syntax parsing pointer at the decoding end.

[0011] Preferably, the audio feature extraction module acquires the original audio PCM pulse code modulation signal of the current sampling period, and obtains the line spectrum pair parameters reflecting the characteristics of the vocal tract through linear prediction analysis; the line spectrum pair parameters are nested with the recoded preceding audio historical state parameters to form a composite audio descriptor; the composite audio descriptor encapsulated by the composite bitstream encapsulation module includes audio slices and preceding audio historical state parameters.

[0012] Preferably, the video encoding rate adjustment module statistically analyzes the distribution gradient of macroblock coordinates and motion vector magnitude values ​​of the core motion region in the monitored video stream, and maps the spatial motion features of the video image to the energy allocation weights of the corresponding frequency bands in the audio slices based on the macroblock coordinates; under the condition of limited total bandwidth, it reduces the bit allocation resources of the background environmental noise frequency band to improve the quantization accuracy of the video macroblocks corresponding to the core motion region.

[0013] Preferably, the composite bitstream encapsulation module inserts SEI information with a specific identifier into the video network abstraction layer unit, making the audio payload an endogenous property of the video frame syntax structure.

[0014] Preferably, the video encoding module supports the H.265 or VVC video encoding standard, and the SEI information is inserted into the header of the video network abstraction layer unit, so that the decoder can obtain the corresponding audio synchronization parameters before parsing the image macroblock data.

[0015] Preferably, the length of the fixed byte payload span is set according to the maximum preset audio bitrate, and the zero-value bitstream filled by the redundant padding units conforms to the byte alignment constraint.

[0016] Preferably, the audio feature extraction module dynamically adjusts the step size of the sliding window based on the packet loss rate index fed back by the channel; when the packet loss rate exceeds a preset threshold, the redundant nesting depth of the preceding audio historical state parameters is increased.

[0017] Preferably, when the interactive control module triggers an audio compression request at the receiving end, it sends a dimension reduction instruction to the audio feature extraction module and simultaneously increases the padding span of the zero-value bit stream in the redundant padding unit to maintain the consistency of the SEI information output length.

[0018] Compared with the prior art, the beneficial effects of the present invention are:

[0019] 1. In the real-time synchronous transmission of interactive audio and video in security monitoring, by directly configuring discrete audio sampling slices as custom supplementary and enhanced information payloads within the video network abstraction layer unit, the traditional dual-track parallel transmission method that relies on timestamps for logical alignment is changed. This transforms the audio bitstream into the structured accompanying attributes of video frames. When the image payload of the video network abstraction layer unit is stripped at the receiving end, audio slices within the same syntax node can be extracted synchronously and the playback loop can be driven. This forces the relative time drift of audio and video to converge within a single video frame refresh cycle, eliminating the need for the receiving end to set up a deep buffer queue to bridge the time difference between audio and video arrival, and reducing the inherent latency in the interaction process.

[0020] 2. By pre-setting a target bearer window with a constant byte length in the isomorphic encapsulation control unit and locking the physical size of the bearer window within the fragmentation-free safety boundary of the maximum transmission unit of the network layer, and in conjunction with the redundant array injection mechanism of the secondary padding field, it is ensured that the output supplementary enhancement information sequence remains absolutely constant in physical byte size regardless of whether the audio bitstream dimensionality reduction algorithm is triggered. This mechanism effectively isolates the impact of source feature dimensionality reduction operation on the syntax parsing pointer of the decoding end, avoids macroblock parsing pipeline clock jitter caused by the sudden change of custom payload length, and provides the decoding end with a constant and fluctuation-free syntax tree parsing cycle, thereby ensuring the determinism of audio and video presentation under low bitrate conditions.

[0021] 3. By constructing a nested mechanism of pre-sequence audio history state parameters based on a time-axis sliding window, the network abstraction layer unit of the current video frame can simultaneously carry the current slice and the historical redundant slice recoded at a low bit rate. Combined with dynamic redundancy depth adjustment instructions, a single composite bitstream frame has the ability to fill the gaps in adjacent frames. This mechanism breaks the strong dependence of audio-visual synchronization on the reliability of the network channel. Even in extreme fading channels with high one-way packet loss rates, as long as the receiver obtains isolated video frames, it can complete audio-visual analysis and state self-recovery, ensuring the seamless continuity of voice and semantics during security command interaction. Attached Figure Description

[0022] Figure 1 This is a diagram illustrating the system composition and data flow of the present invention;

[0023] Figure 2 This is a logic diagram of the joint reorganization and allocation of audio and video resources in this invention.

[0024] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0026] An interactive real-time synchronous audio and video transmission system for security monitoring includes a video encoding module, an audio feature extraction module, a video encoding rate adjustment module, an interactive control module, and a composite bitstream encapsulation module.

[0027] The video encoding module is used to acquire the monitoring video stream and extract the motion vector magnitude values ​​of video macroblocks within the core motion region;

[0028] The audio feature extraction module is used to synchronously acquire audio streams and extract the current audio slice and previous audio historical state parameters based on a time axis sliding window.

[0029] The video encoding rate adjustment module is connected to both the video encoding module and the audio feature extraction module. It is used to establish a mapping relationship between motion vector magnitude and audio bit allocation weights, and to convert the distribution characteristics of motion vector magnitude into frequency band energy allocation weights for audio slices.

[0030] The composite bitstream encapsulation module is used to encapsulate the audio slices adjusted by frequency band energy allocation weights and the composite audio descriptor formed by nesting the preceding audio historical state parameters into the supplementary enhancement information (SEI) information of the video network abstraction layer unit. The SEI information has a fixed byte payload span, and the composite bitstream encapsulation module performs bit padding on the SEI information according to the fixed byte payload span to keep the output SEI information constant in physical length.

[0031] Preferably, the composite bitstream encapsulation module includes: a data anchoring unit and a redundancy padding unit; the data anchoring unit is used to rigidly anchor the audio slice data to the starting offset address of the SEI information when the bitrate of the audio slice fluctuates; the redundancy padding unit is used to calculate the free byte span outside the audio slice data in the SEI information and fill the free byte span with zero-value bitstream to isolate the impact of the characteristic changes of the audio slice on the syntax parsing pointer at the decoding end.

[0032] Preferably, the audio feature extraction module acquires the original audio PCM pulse code modulation signal of the current sampling period, and obtains the line spectrum pair parameters reflecting the characteristics of the vocal tract through linear prediction analysis; the line spectrum pair parameters are nested with the recoded preceding audio historical state parameters to form a composite audio descriptor; the composite audio descriptor encapsulated by the composite bitstream encapsulation module includes audio slices and preceding audio historical state parameters.

[0033] Preferably, the video encoding rate adjustment module statistically analyzes the distribution gradient of macroblock coordinates and motion vector magnitude values ​​of the core motion region in the monitored video stream, and maps the spatial motion features of the video image to the energy allocation weights of the corresponding frequency bands in the audio slices based on the macroblock coordinates; under the condition of limited total bandwidth, it reduces the bit allocation resources of the background environmental noise frequency band to improve the quantization accuracy of the video macroblocks corresponding to the core motion region.

[0034] Preferably, the composite bitstream encapsulation module inserts SEI information with a specific identifier into the video network abstraction layer unit, making the audio payload an endogenous property of the video frame syntax structure.

[0035] Preferably, the video encoding rate adjustment module calculates the frequency band energy allocation weighting coefficient G based on the following formula: Where G is the frequency band energy allocation weighting coefficient; α is the motion vector magnitude of the i-th macroblock within the core motion region; M is the total number of macroblocks within the core motion region; α is the preset sensitivity gain factor; β is the preset base offset.

[0036] Preferably, the video encoding module supports the H.265 or VVC video encoding standard, and the SEI information is inserted into the header of the video network abstraction layer unit, so that the decoder can obtain the corresponding audio synchronization parameters before parsing the image macroblock data.

[0037] Preferably, the length of the fixed byte payload span is set according to the maximum preset audio bitrate, and the zero-value bitstream filled by the redundant padding units conforms to the byte alignment constraint.

[0038] Preferably, the audio feature extraction module dynamically adjusts the step size of the sliding window based on the packet loss rate index fed back by the channel; when the packet loss rate exceeds a preset threshold, the redundant nesting depth of the preceding audio historical state parameters is increased.

[0039] Preferably, when the interactive control module triggers an audio compression request at the receiving end, it sends a dimension reduction instruction to the audio feature extraction module and simultaneously increases the padding span of the zero-value bit stream in the redundant padding unit to maintain the consistency of the SEI information output length.

[0040] Example 1: When the system faces remote interactive security command and dispatch under high-stress emergencies, network bandwidth fluctuations and low-latency interaction requirements create structural conflicts. The traditional dual-track independent transmission architecture relies on the receiver to build a deep buffer queue to bridge the audio and video arrival time difference. However, under the condition that network backpressure leads to a high compression ratio of video frames, changing the payload length of the accompanying audio causes the syntax tree parsing pointer at the decoding end to jump between different memory steps, causing clock tick disorder in the macroblock parsing pipeline. This security monitoring interactive audio and video real-time synchronous transmission system eliminates the interference of variable-length custom payloads on the video decoding pipeline by performing structural reconstruction at the multimedia reuse level.

[0041] The video encoding module acquires the monitored video stream and extracts the motion vector magnitude of video macroblocks within the core motion region. Simultaneously, the audio feature extraction module extracts the current audio slice and the historical state parameters of the preceding audio based on a time-axis sliding window. The historical state parameters of the preceding audio refer to the feature set representing the audio channel characteristics of at least one historical sampling period prior to the current sampling period, extracted based on the time-axis sliding window. Specifically, this parameter is extracted from the line spectrum pair (LSP) parameter matrix of the original audio from the preceding sampling period, generated by scalar quantization compression of the main diagonal feature elements of this matrix. Its physical significance lies in providing an independently decodeable audio background reference for the current video frame, enabling the receiver to complete audio-visual reconstruction and state self-recovery based on this redundant historical state information when data is lost in the current sampling period. The video encoding rate adjustment module establishes a mapping relationship between the motion vector magnitude and the audio bit allocation weights, converting the distribution characteristics of the motion vector magnitude into the frequency band energy allocation weights of the audio slices, according to the formula... Calculate the frequency band energy allocation weighting coefficient, where G is the frequency band energy allocation weighting coefficient. Let M be the motion vector magnitude of the i-th macroblock within the core motion region, M be the total number of macroblocks within the core motion region, α be the preset sensitivity gain factor, and β be the preset base offset. After calculating the frequency band energy allocation weights, the system proceeds to the composite audio payload encapsulation stage. At the data organization level, audio slices represent the audio feature data of the current sampling period, and the preceding audio historical state parameters represent the historical state data extracted to provide anti-packet loss redundancy. The composite bitstream encapsulation module nests and combines the audio slices adjusted by the frequency band energy allocation weights and the preceding audio historical state parameters to form a composite audio description. Therefore, the composite audio descriptor includes audio slices and previous audio history parameters. Based on this, the composite bitstream encapsulation module encapsulates the composite audio descriptor as a single object into the supplemental enhancement information (SEI) of the video network abstraction layer unit. The SEI has a fixed byte payload span set according to the maximum preset audio bitrate. When network bandwidth fluctuations cause changes in the audio slice bitrate, the data anchoring unit of the composite bitstream encapsulation module anchors the audio slice data to the starting offset address of the SEI. The redundancy padding unit calculates the free byte span outside the audio slice data within the SEI and fills this free byte span with zero-value bitstreams conforming to byte alignment constraints, maintaining a constant total output payload length through physical byte padding. The fixed-length SEI structure provides a transmission architecture unaffected by sudden changes in payload volume for frequency band energy allocation weight adjustment, allowing limited channel capacity to be distributed towards high-dynamic visual regions and key audio segments, resolving the technical contradiction between compression efficiency and parsing stability. When extracting the motion vector magnitude of video macroblocks within the core motion region, the video encoding module calculates the motion vector magnitude of all macroblocks within the current video frame. The global arithmetic mean of the motion vector magnitude is used to identify the set of continuous macroblocks whose motion vector magnitude is greater than 1.5 times the global arithmetic mean as the core motion region. The audio feature extraction module divides the original audio pulse code modulation signal into low-frequency sub-bands with frequencies below 2kHz and high-frequency sub-bands with frequencies between 2kHz and 4kHz through a sub-band filter bank. The audio feature extraction module extracts line spectrum pair parameters that reflect the characteristics of the vocal tract, extracts and quantizes the baseband prediction residual sequence of the corresponding sub-band, and uses the line spectrum pair parameters and the baseband prediction residual sequence together as the current audio slice load to provide complete excitation parameters for waveform reconstruction at the decoding end.

[0042] After calculating the frequency band energy allocation weight coefficient G, the video coding rate adjustment module allocates G×100% of the total available audio bit quota to the high-frequency subband, and the remaining quota to the low-frequency subband. The composite bitstream encapsulation module writes a payload length identifier with a constant length of 2 bytes at the starting offset address of the supplementary enhancement information (SEI) information, records the total number of bytes of the current audio slice data and the previous audio historical state parameters, and then writes audio data sequentially following the length identifier. It calculates the difference between the fixed-byte payload span and the total length of the written data, and fills the memory space corresponding to the difference with hexadecimal zero-value bitstream byte by byte. The accompanying parsing unit at the receiving end reads the starting payload length identifier, extracts the audio data according to the value, discards the remaining zero-value bitstream, and establishes the variable-length data syntax parsing boundary within the fixed-length physical encapsulation. When the accompanying parsing unit at the receiving end strips the image payload of the video network abstraction layer unit and sends it to the decoding pipeline, it extracts the internally nested fixed-length supplementary enhancement information (SEI) payload to drive the audio reconstruction loop. The constant physical byte size of the supplementary enhancement information (SEI) payload isolates the source feature dimensionality reduction operation from the source feature dimensionality reduction operation. The impact of the syntax parsing pointer at the decoding end causes the video decoding pipeline to have a definite syntax tree parsing time when processing accompanying audio information. This deep nesting at the bitstream level not only smooths out the surface clock beat on the decoding side, but also forces the originally independently linked audio and video data packets to merge into a single network layer maximum transmission unit (MTU) at the network transmission physical layer. When experiencing overall routing congestion and random jitter in the wide area network, the audio slices are reduced in dimension and solidified into the internal structural attributes of video keyframes, and together with the video macroblocks as an indivisible physical whole. Having experienced the same network latency, the relative arrival time difference caused by the channel fading difference between the two independent data streams is eliminated. After entering the receiving end, due to the co-arrival of the audio and video payloads at the network layer, coupled with the absolute stability of the parsing time of the syntax tree at the surface layer, the system effectively bridges the nonlinear offset of the transmission path under the dual synergy of overall same-frequency delay and surface parsing lock. The relative time drift of audio and video converges within the refresh cycle of a single video frame, eliminating the delay of timing alignment relying on the buffer queue, and forming a frame-level synchronous data transmission stream that is not affected by network bandwidth fluctuations.

[0043] Example 2: Network bandwidth degradation causes timing tearing in audio and video synchronization during high-stress security command and dispatch operations. A test environment was constructed using a hardware-level network impairment simulator, capable of injecting random packet loss rates from 0% to 50% and adjusting latency jitter from 1ms to 500ms. The video source was a standard 4K resolution security monitoring dataset, and the audio source was a live command voice stream with a sampling frequency of 48kHz and superimposed with 65dB of background power frequency interference noise. The setting of the fixed byte payload span involved considerations of data acquisition real-time performance and channel bandwidth redundancy consumption. The technical trade-offs are as follows: When the channel is limited and accompanied by audio bitrate fluctuations, too small a payload span leads to audio truncation distortion, while too large a payload span increases the risk of network transport layer fragmentation; the logic is set based on the maximum transmission unit model. When the underlying network basic transmission bandwidth is lower than the congestion threshold, the fixed byte payload span is determined to be less than the maximum transmission unit byte number after subtracting the network layer header overhead, and covers the slice length corresponding to the highest audio coding bitrate; the maximum transmission unit is selected as 1500 bytes, and after subtracting 28 bytes of header overhead, the fixed byte payload span is established as 512 bytes.

[0044] The verification process sets three bandwidth limiting gradients: 10Mbps, 2Mbps, and 0.5Mbps, and globally injects a 15% random packet loss rate. Four sets of comparative test entities are constructed. Comparative sample group one uses an independent real-time transmission protocol with timestamp alignment; comparative sample group two uses variable-length supplementary enhanced information payload to encapsulate audio slices; comparative sample group three selects an out-of-range sensitivity gain factor α and sets its value to 5.0; the present invention sample group uses a 512-byte constant-length supplementary enhanced information payload encapsulation, selects a sensitivity gain factor α of 0.8, and base... With an offset β of 0.2, under a 10Mbps bandwidth baseline, the relative time drift of the audio and video of the four test entities was distributed in the range of 12.5ms to 14.2ms. When the bandwidth limit was tightened to 2Mbps, due to the influence of network back pressure, the relative time drift of the first comparison sample group, which relied on the receiver buffer queue, increased to 145.3ms. The variable-length load of the second comparison sample group caused the video decoder syntax tree parsing pointer to jump frequently, and the measured standard deviation of the macroblock parsing pipeline clock beat reached 2.3ms, and the relative time drift deteriorated to 45.8ms.

[0045] The present invention's sample group initiates feature linkage allocation and fixed-length encapsulation mechanisms under 2Mbps conditions; the video encoding module extracts the normalized average motion vector magnitude of the core motion region, which is 0.6; according to the formula The calculation outputs a frequency band energy allocation weighting coefficient G of 0.68, where G is the frequency band energy allocation weighting coefficient. Let M be the motion vector magnitude of the i-th macroblock within the core motion region, M be the total number of macroblocks within the core motion region, α be a preset sensitivity gain factor, and β be a preset base offset. The audio feature extraction module allocates 68% of the available bits to the high-frequency command audio segment based on this weighting coefficient. The audio slice data after this bit allocation actually occupies 128 bytes. The composite bitstream encapsulation module anchors this 128 bytes of data to the starting offset address of the 512-byte supplementary enhancement information payload, calculates the free byte span to be 384 bytes, and fills it with zero-value bitstreams. Actual test data shows that the standard deviation of the macroblock parsing pipeline clock tick in the sample group of this invention is maintained at 0.05ms. The receiving end video decoding pipeline obtains a stable syntax tree parsing time with a relative time drift locked at 16.6ms, corresponding to a single frame refresh cycle of 60fps video. When the bandwidth limit is reached... When the bandwidth was reduced to 0.5 Mbps, the test data showed a non-linear inflection point. Compared with the third sample group, the calculated value of the frequency band energy allocation weight coefficient G exceeded the saturation boundary of 1.0, and all audio bits were allocated to the high frequency band. The loss of the low frequency envelope caused the speech intelligibility score to drop sharply to 32.4%. The speech intelligibility score of the sample group of this invention remained at 85.6%, and the physical byte filling mechanism continuously output a fixed-length payload of 512 bytes. The relative time drift was maintained at 16.6 ms. The test data confirmed that constructing a supplementary enhanced information payload of constant byte size inside the video network abstraction layer unit isolated the impact of the physical length jump brought about by the source feature dimensionality reduction on the underlying syntax parsing pointer of the video decoder. The structural reconstruction of the multimedia reuse layer made the relative time drift of audio and video converge within a single video frame refresh cycle, eliminated the buffer queue delay, and maintained the determinism of the decoding offset calculation delay.

[0046] Example 3: In a heterogeneous network routing environment, the basic transmission bandwidth fluctuates continuously. A preset, static SEI (Supplementary Enhancement Information) payload length leads to bandwidth consumption during idle periods or fragmentation during contraction periods. The system faces physical constraints due to the mismatch between the audio feature parameter encapsulation span and the dynamic network boundary. The security monitoring interactive audio-visual real-time synchronous transmission system introduces an adaptive payload span calibration mechanism and a low-level byte padding mechanism to establish an encapsulation structure that matches the physical channel boundary. Before the audio-visual data stream is transmitted, the interactive control module initiates link detection, sending a sequence of probe data packets with progressively increasing lengths from 64 bytes to 1500 bytes to the receiving end. The length increment between adjacent probe data packets is set to 32 bytes. The receiving end calculates the round-trip delay and packet loss rate of each probe data packet length, identifying the critical data packet length corresponding to the first jump in packet loss rate to a preset packet loss threshold of 1%. The composite bitstream encapsulation module obtains this critical data packet length and, according to the formula... Calculate the basic safety span value, where, Based on the basic safety span value, This is the critical data packet length. The header bytes span of the network layer. The header byte span of the video network abstraction layer unit is set to 28 bytes; the composite bitstream encapsulation module sets the header byte span of the network layer to 28 bytes and the header byte span of the video network abstraction layer unit to 12 bytes, outputs the basic security span value, and the redundancy padding unit allocates contiguous memory space based on the basic security span value.

[0047] The data anchoring unit of the composite bitstream encapsulation module writes a 16-byte audio history state parameter to the starting address of the memory space, and then writes the variable-length audio slice data immediately after the audio history state parameter. The redundancy padding unit obtains the basic safety span value, subtracts the sum of the length of the audio history state parameter and the length of the audio slice data from it, and calculates the number of free bytes. The redundancy padding unit generates a hexadecimal zero-value byte stream corresponding to the number of free bytes, and writes it byte by byte to the end of the memory after the audio slice data until the end boundary of the memory space is reached. The payload span adaptive calibration and zero-value byte padding mechanism align the physical size of the supplementary enhancement information SEI information with the fragmented transmission boundary of the current network. The receiving end decoding pipeline extracts data according to the basic safety span value at a fixed memory step size, eliminating the dynamic length identifier parsing and determination step, and maintaining the parsing stability of the syntax tree of the fixed-length feature stream in the heterogeneous network environment.

[0048] Example 4: When the system faces the initial deployment of a security monitoring interactive audio-visual real-time synchronous transmission system node in a completely new physical location, the audio-visual synchronization parameters establish a baseline mapping with the on-site spatial environment; the interactive control module drives the camera to continuously acquire a static background video stream without moving targets for a preset reference duration, while simultaneously activating the microphone to record the ambient background noise frequency sequence; the audio feature extraction module calculates the average power spectral density of the ambient background noise frequency sequence as the background noise reference value; a standard moving calibration object is introduced into the monitoring field of view to form a uniform lateral shuttle trajectory, and the video encoding module extracts the motion vector of the macroblock sequence corresponding to the standard moving calibration object. Peak value of motion vector magnitude; The video encoding rate adjustment module substitutes the peak value of motion vector magnitude and the reference value of background noise into the initial control matrix, and determines the value of the basic offset β by solving a system of univariate linear equations. This maps the environmental background noise to the lowest frequency band energy allocation benchmark, eliminating the ineffective occupation of audio feature data allocation by static background. In the specific solution process, the initial control matrix is ​​configured as a two-row, two-column diagonal matrix. Its main diagonal elements are the reciprocal of the extracted peak value of motion vector magnitude and the preset system reference audio energy quota constant, while all secondary diagonal elements are set to zero. The specific mathematical expression of the system of univariate linear equations is as follows: Where C is the product of the main diagonal elements of the initial control matrix, and β is the base offset to be solved. The measured noise floor reference value is K, which is the acoustic-to-electrical conversion coefficient pre-calibrated based on the hardware sensitivity of the field microphone. The system directly calculates the basic offset β value corresponding to the acoustic environment of the current physical node through the determined linear equation.

[0049] The interactive control module drives the on-site acoustic equipment to broadcast standard test audio with a known spectral envelope. The composite bitstream encapsulation module encapsulates and generates an initial test bitstream based on a determined base offset β, and sends it to the receiving end. The accompanying parsing unit at the receiving end extracts the audio signal inside the initial test bitstream and calculates the signal correlation coefficient between the reconstructed audio and the standard test audio. The video encoding rate adjustment module receives feedback link data containing the signal correlation coefficient, increments the sensitivity gain factor value by a fixed numerical step, and synchronously triggers the transmission and parsing loops. When the signal correlation coefficient first crosses the preset voice fidelity threshold, the video encoding rate adjustment module locks the current sensitivity gain factor value and writes it to the firmware register. The system completes the numerical binding of spatial physical characteristics and frequency band energy allocation weight coefficient G at the current physical node, constructing an audio-visual data processing framework that outputs a constant-length supplementary enhanced information payload based on specific industrial site acoustic and optical conditions.

[0050] Example 5: High-frequency noise interference and fluctuating audio bitrate during security monitoring cause audio data truncation and decoding pointer offset; the interactive control module reads the current video stream's frame rate. The audio feature extraction module is based on the formula Calculate the span threshold of the time axis sliding window ,in, This is the span threshold of the timeline sliding window. To optimize the video frame rate, the audio feature extraction module locks the physical time span of the internal timeline sliding window to a span threshold. This process aligns the audio sampling interval of a single feature extraction with the refresh cycle of a single video frame, eliminating temporal overlap caused by cross-frame audio feature acquisition. The audio feature extraction module acquires the original audio pulse-code modulation signal of the current sampling cycle, inputs it into the linear predictive coding algorithm model to calculate the prediction residual of the signal, and extracts line spectrum pair parameters reflecting the characteristics of the vocal tract. The composite bitstream encapsulation module acquires the allocated supplementary enhancement information (SEI) information in a contiguous memory space, and delineates a header state area and a tail payload area immediately following it within the contiguous memory space. The composite bitstream encapsulation module writes the preceding audio history state parameters into the header. In the state area, the line spectrum parameters are written to the starting physical address of the tail load area. The composite bitstream encapsulation module calculates the remaining available byte capacity of the tail load area, and when it is determined that the total length of the current audio slice data is greater than the remaining available byte capacity, it discards the edge load bits of the audio slice data step by step from the highest frequency band to the lowest frequency band until the total length of the remaining data is equal to the remaining available byte capacity. Based on the sliding window span calibration and structured memory partition truncation writing mechanism based on the refresh cycle, the dynamic audio parameters obtain a unique corresponding physical storage location, and output a multimedia transmission payload with deterministic read offset.

[0051] When the security monitoring channel faces continuous data packet loss, extracting audio parameters at a single moment causes a break in the timing context at the decoding end. The audio feature extraction module initiates the historical parameter recoding and data nesting calibration process. The audio feature extraction module reads the original audio line spectrum parameter matrix of the preceding continuous sampling period from the underlying buffer register, extracts the main diagonal feature elements of the original audio line spectrum parameter matrix to form the preceding audio historical state parameters, and reads the current channel packet loss rate fed back by the on-site network probe. According to the formula Calculate the scalar quantization step size Q, where, The current channel packet loss rate, The preset baseline quantization constant, The audio feature extraction module uses a scalar quantization algorithm with a scalar quantization step size Q to compress the preceding audio historical state parameters and outputs the recoded preceding audio historical state parameters. Simultaneously, the audio feature extraction module dynamically adjusts the step size of the time axis sliding window based on the current channel packet loss rate fed back by the on-site network probe. When the current channel packet loss rate is less than or equal to a preset packet loss threshold, the step size of the sliding window is set to be equal to the span threshold, implementing non-overlapping continuous audio sampling. When the current channel packet loss rate is greater than the preset packet loss threshold, the step size of the sliding window is forcibly reduced to half of the span threshold, causing adjacent sampling windows to... A 50% physical overlap is generated on the timeline, thus pre-acquiring surplus audio context slices during the feature extraction stage. Because overlapping sampling doubles the number of preceding audio history state parameters covered by a single window, the system synchronously increases the redundancy nesting depth of these parameters. The parameters of two adjacent overlapping windows are jointly encoded and pushed into the accompanying payload of a single video frame. The audio feature extraction module obtains the line spectrum pair parameters of the current sampling period, writes the re-encoded preceding audio history state parameters into the high-order physical address segment of an independent memory structure, and writes the line spectrum pair parameters of the current sampling period into the low-order physical address segment of the same independent memory structure. The audio descriptor is generated by concatenating the data bitstreams within the physical address segments on both sides using a bitwise OR operation. Before performing this bitwise OR operation, the audio feature extraction module performs a low-level left shift operation on the recoded preceding audio history state parameters based on the pre-allocated byte span of the high-order physical address segment, ensuring that its least significant bit is aligned with the most significant bit of the low-order physical address segment with zero error. A logical AND mask operation matching the length of the low-order address segment is applied to the line spectrum parameters of the current sampling period, zeroing out its high-order redundant regions to clear any remaining random memory fragments. The two sets of ratios, physically isolated by left shift and logically cleaned by masking, are then... The special stream forms absolutely mutually exclusive non-overlapping data boundaries on independent memory distribution, thereby ensuring that subsequent bit OR operations only perform pure payload gluing, avoiding bit cross-contamination between feature matrices. The composite bit stream encapsulation module reads the composite audio descriptor and writes the composite audio descriptor byte by byte into the header state area of ​​the video network abstraction layer unit to supplement the enhanced information SEI information. The parametric recoding and feature nesting mechanism enables the current video frame to directly carry the compressed data mapping of the preceding audio context. In the event of a sudden network outage recovery node, the decoder prediction filter is driven to reset the initial physical state, smoothing the underlying waveform timing jitter caused by channel interruption.

[0052] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A security monitoring interactive audio-visual real-time synchronous transmission system, characterized in that, It includes a video encoding module, an audio feature extraction module, a video encoding rate adjustment module, an interactive control module, and a composite bitstream encapsulation module. The video encoding module is used to acquire the monitoring video stream and extract the motion vector magnitude values ​​of video macroblocks within the core motion region; The audio feature extraction module is used to synchronously acquire audio streams and extract the current audio slice and previous audio historical state parameters based on a time axis sliding window. The video encoding rate adjustment module is connected to both the video encoding module and the audio feature extraction module. It is used to establish a mapping relationship between motion vector magnitude and audio bit allocation weights, and to convert the distribution characteristics of motion vector magnitude into frequency band energy allocation weights for audio slices. The composite bitstream encapsulation module is used to encapsulate the audio slices adjusted by frequency band energy allocation weights and the composite audio descriptor formed by nesting the preceding audio historical state parameters into the supplementary enhancement information (SEI) information of the video network abstraction layer unit. The SEI information has a fixed byte payload span, and the composite bitstream encapsulation module performs bit padding on the SEI information according to the fixed byte payload span to keep the output SEI information constant in physical length.

2. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The composite bitstream encapsulation module includes: a data anchoring unit and a redundancy padding unit; the data anchoring unit is used to rigidly anchor the audio slice data to the starting offset address of the SEI information when the bitrate of the audio slice fluctuates; the redundancy padding unit is used to calculate the free byte span outside the audio slice data in the SEI information and fill the free byte span with zero-value bitstream to isolate the impact of changes in the characteristic quantity of the audio slice on the syntax parsing pointer at the decoding end.

3. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The audio feature extraction module acquires the original audio PCM pulse code modulation signal of the current sampling period and obtains the line spectrum pair parameters reflecting the characteristics of the vocal tract through linear prediction analysis. The line spectrum pair parameters are nested with the recoded previous audio history state parameters to form a composite audio descriptor. The composite audio descriptor encapsulated by the composite bitstream encapsulation module contains audio slices and previous audio history state parameters.

4. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The video encoding rate adjustment module statistically analyzes the distribution gradient of macroblock coordinates and motion vector magnitude values ​​of the core motion region in the monitored video stream, and maps the spatial motion features of the video image to the energy allocation weights of the corresponding frequency bands in the audio slices based on the macroblock coordinates. Under conditions of limited total bandwidth, the bit allocation resources of the background ambient noise frequency band are reduced to improve the quantization accuracy of video macroblocks corresponding to the core motion region.

5. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The composite bitstream encapsulation module inserts SEI information with specific identifiers into the video network abstraction layer unit, making the audio payload an endogenous property of the video frame syntax structure.

6. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The video encoding module supports H.265 or VVC video encoding standards. SEI information is inserted into the header of the video network abstraction layer unit, enabling the decoder to obtain the corresponding audio synchronization parameters before parsing the image macroblock data.

7. A security monitoring interactive audio-visual real-time synchronous transmission system according to claim 2, characterized in that, The length of the fixed byte payload span is set according to the maximum preset audio bitrate, and the zero-value bitstream filled by the redundant padding units conforms to the byte alignment constraints.

8. The security monitoring interactive audio-visual real-time synchronous transmission system according to claim 1, characterized in that, The audio feature extraction module dynamically adjusts the step size of the sliding window based on the packet loss rate index fed back from the channel; when the packet loss rate exceeds the preset threshold, the redundant nesting depth of the preceding audio historical state parameters is increased.

9. A security monitoring interactive audio-visual real-time synchronous transmission system according to claim 2, characterized in that, When the interactive control module triggers an audio compression request at the receiving end, it sends a dimension reduction instruction to the audio feature extraction module and simultaneously increases the padding span of the zero-value bit stream in the redundant padding units to maintain the consistency of the SEI information output length.