A low-delay streaming media no-memory allocation adaptive audio frame reorganization push method

CN122824718APending Publication Date: 2026-09-25FANHUA INTELLIGENT TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610971239.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]然而,上述现有技术实时交互对音频延迟要求极高,频繁垃圾回收(GarbageCollection,GC)引发高延迟与抖动,导致推流时序产生微小延迟和时钟抖动,严重影响流媒体质量;多声道扩展、音量增益缩放及类型转型涉及大量连续的物理拷贝操作,在并发流媒体服务中会消耗大量CPU周期,限制了单服务器的并发处理能力,CPU计算资源消耗大;当TTS音频流自然结束(End Of File,EOF)时,末尾不足一帧的数据由于无法拼满一帧常被直接丢弃;同时,若推流轨道骤然关闭,会导致接收客户端(如游戏引擎、网页端解码器)的缓冲区瞬时抽空,引发尾音被生硬截断,并产生明显的电流爆破音,严重影响收听体验

Benefits of technology

本发明采用自适应滑动缓冲区,在音频流处理的稳定运行期实现零内存再分配,消除了垃圾回收带来的时延抖动,延迟控制相比传统方案大幅改善,能够稳定满足高并发、实时流式音频交互的严苛要求,显著降低流媒体传输延迟与抖动;本发明通过内存地址投影技术,直接将字节缓冲区重投影为短整型视图,省去了转换过程中的临时内存申请;结合合并操作使声道映射与增益限幅在单次遍历中完成,减少了数据遍历次数,显著降低了流媒体服务器的CPU开销,提升了单节点并发处理能力;本发明通过引入高精密尾音补齐与静音冲刷补偿机制,通过向物理通道中注入特定数量的静音帧,强行冲刷接收端的音频解码队列,在物理层面实现平滑淡出,有效消除了音频尾音截断与接收端解码噪声,提升了收听体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824718A_ABST
    Figure CN122824718A_ABST
Patent Text Reader

Abstract

The application provides a low-delay streaming media no-memory allocation adaptive audio frame reorganization and pushing method, and belongs to the technical field of real-time audio and video streaming media processing.The application realizes audio frame reorganization and pushing through three collaborative modules of adaptive sliding reorganization buffer, zero-memory allocation data projection and sound channel expansion conversion, high-precision timing pumping and mute flushing compensation: the application realizes stable-state zero-memory allocation of audio data reorganization through pre-allocated continuous memory and read-write double pointers, reduces CPU conversion overhead through memory view re-projection and single-step sound channel mapping, and solves the problems of tail sound truncation and receiver decoding burst sound by combining precise timing streaming and EOF mute flushing mechanism.The application can eliminate the latency jitter caused by GC from the root, improve the concurrent processing efficiency, completely avoid tail sound truncation and burst sound, and is suitable for low-delay streaming interaction scenes of RTC and large language models and voice synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of real-time audio and video streaming media processing technology, and in particular to a low-latency streaming media adaptive audio frame reassembly and push method without memory allocation. Background Technology

[0002] In scenarios involving real-time audio and video (RTC, such as LiveKit and WebRTC) and streaming interaction with large language models (LLM) and text-to-speech (TTS), TTS engines typically send audio streams in the form of variable-length, high-frequency network data packets, while RTC transmission tracks require pumping out and pushing audio streams with extremely high precision and fixed frame lengths (such as audio sampling frames corresponding to a precise duration of 20ms).

[0003] In the process of reassembling and streaming the above-mentioned variable-length audio packets into fixed audio frames, the existing technology usually adopts the following scheme: (1) Dynamic reconstruction buffer: Whenever a new variable-length packet arrives, a new byte array is dynamically allocated in the memory heap for splicing, truncation and copying; (2) Physical copy and format conversion: When converting variable-length mono byte data (byte) into the stereo short integer value (short) required by RTC, a new array is created one by one and physical data copy is performed; (3) Uncontrolled streaming and abrupt interruption: When the network transmission ends (EOF), the streaming is stopped directly or the socket channel is closed.

[0004] However, the aforementioned existing technologies have extremely high requirements for real-time audio latency. Frequent garbage collection (GC) causes high latency and jitter, resulting in slight delays and clock jitter in the streaming timing, which seriously affects the quality of streaming media. Multichannel expansion, volume gain scaling, and type conversion involve a large number of continuous physical copy operations, which consume a lot of CPU cycles in concurrent streaming media services, limiting the concurrent processing capacity of a single server and consuming a lot of CPU computing resources. When the TTS audio stream ends naturally (End Of File, EOF), the data at the end that is less than one frame is often directly discarded because it cannot be pieced together into a complete frame. At the same time, if the streaming track is suddenly closed, the buffer of the receiving client (such as a game engine or web-based decoder) will be instantly emptied, causing the tail sound to be abruptly truncated and producing obvious electrical popping sounds, which seriously affects the listening experience.

[0005] In addition, existing alternatives use a circular bidirectional circular queue storage structure to complete variable-length data concatenation. The read and write pointers move cyclically within a fixed-size array, eliminating the need for physical translation and copying. However, when the RTC interface requires physically contiguous memory segments for streaming, if the unread data in the circular buffer crosses the boundary between the end and the beginning of the array, the system still needs to allocate temporary physical contiguous memory to concatenate and wrap back the data. This introduces additional dynamic copying and memory allocation overhead, which is insufficient when dealing with strictly contiguous memory frame extraction. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a low-latency streaming media adaptive audio frame reassembly and push method without memory allocation. This method eliminates the memory allocation overhead during the streaming audio frame reassembly process. During long streaming media connections, it enables the reassembly of variable-length audio packets into fixed audio frames to achieve steady-state zero memory allocation, fundamentally eliminating streaming media latency caused by frequent garbage collection (GC). It also reduces CPU time consumption for data transformation and channel expansion, avoids multiple physical memory copies during type conversion and multi-channel expansion, and improves concurrent processing capabilities. Furthermore, it solves the problem of tail truncation and popping sounds at the end of streaming media, ensuring the complete output of the tail fragments at the end of the network audio stream and providing a smooth physical transition to the receiving end, eliminating decoding popping sounds.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: This invention provides a low-latency streaming media adaptive audio frame reassembly and push method without memory allocation, applicable to scenarios where variable-length streaming audio packets are converted into fixed-length audio frames and pushed to a real-time audio and video transmission track. The method includes adaptive sliding reassembly buffer processing, zero-memory-allocation data projection and channel expansion conversion processing, and high-precision timed pumping and silent flushing compensation processing. The specific steps are as follows: S1. Adaptive sliding reassembly buffer processing: When a streaming media session is established, a contiguous storage space is pre-allocated as a shared reassembly buffer, and the read pointer and write pointer are initialized; when a variable-length audio data packet is received, it is determined whether there is enough remaining space at the end of the buffer to accommodate the current data packet. If not, and there is already free space at the beginning of the buffer, the unread data is shifted to the beginning address of the buffer as a whole, the read and write pointers are reset, and the current variable-length audio data packet is appended to the buffer. No dynamic memory allocation is generated during the steady-state operation of streaming media. S2. Zero-memory allocation data projection and channel expansion conversion processing: When retrieving fixed-frame-length audio data from the shared reassembly buffer, the corresponding byte memory segment is directly reprojected into a short integer value view without physical data copying. During a single traversal of the short integer value view, the channel mapping from mono to stereo, volume gain calculation, and anti-pop noise overflow limiting processing are completed simultaneously to generate a standard stereo audio frame. S3. High-precision timed pumping and silent flushing compensation processing: The streaming rhythm is controlled by a high-precision periodic timer aligned with the physical duration of the audio frame, and audio frames are pushed according to the frame rate; when the audio stream end signal is received, the audio fragments that are less than one frame remaining in the shared reconstruction buffer are first filled into complete audio frames and pushed, and then a preset number of absolute silence data frames are continuously injected into the streaming channel to flush the physical audio pipe and the decoding queue at the receiving end. After the flushing is completed, the streaming state is switched to the idle state.

[0008] As a further aspect of the present invention, the pre-allocated contiguous storage space is a fixed-size contiguous byte array, allocated in the managed heap or unmanaged physical memory; the read pointer is used to mark the end position of the read data, and the write pointer is used to mark the end position of the written data.

[0009] As a further aspect of the present invention, the specific method for determining whether there is sufficient remaining space at the end of the buffer is as follows: determine whether the difference between the total length of the buffer and the write pointer is greater than or equal to the length of the newly arrived audio data packet; if the difference is greater than or equal to the length of the data packet, then directly append the data packet to the end and update the write pointer.

[0010] As a further aspect of the present invention, when the remaining space at the tail is insufficient and the read pointer is greater than 0, the unread data between the read pointer and the write pointer is shifted to the starting address of the buffer offset of 0 by the block copy method; the read pointer is reset to 0, and the write pointer is updated to the length of the unread data; then the current variable-length audio data packet is appended to the updated write pointer position in the buffer and the write pointer is updated.

[0011] As a further aspect of the present invention, when the remaining space at the tail is insufficient and there is no free space at the head of the buffer, the shared buffer is physically expanded.

[0012] As a further aspect of the present invention, the reprojection of the byte memory segment to the short integer value view is performed at the pointer or lower-level view level, achieving physical zero memory copy and zero memory allocation; the short integer value view is a 16-bit integer view.

[0013] As a further aspect of the present invention, when generating dual-channel audio frames, a target frame storage array is rented from the array object pool; after processing, the temporary frame array is returned to the object pool to ensure that there is no heap memory allocation overhead during the conversion process.

[0014] As a further aspect of the present invention, channel mapping, volume gain calculation and overflow limiting processing are completed synchronously during the same traversal of the mono numerical view, and the calculation results are directly written into the target stereo frame storage array.

[0015] As a further aspect of the present invention, the trigger signal of a high-precision periodic timer is monitored asynchronously in the background, and audio frames are pumped out only when the timer is triggered, thereby eliminating streaming time jitter.

[0016] As a further aspect of the present invention, audio fragments with less than one frame at the end are filled and aligned using zero-value silence data to form a complete audio frame before being pushed.

[0017] As a further aspect of the present invention, the number of injected absolute silence data frames is 50, corresponding to a duration of 1000ms; the silence data frames are used to drive the physical audio pipeline and the audio data packets remaining at the receiving end to be output completely, preventing the decoding queue at the receiving end from being interrupted and causing tail truncation and popping sounds.

[0018] Compared with existing technologies, the low-latency streaming media memory-allocation-free adaptive audio frame reassembly and push method provided in this application has the following advantages: This invention employs an adaptive sliding buffer to achieve zero memory reallocation during the stable operation of audio stream processing, eliminating latency jitter caused by garbage collection. Latency control is significantly improved compared to traditional solutions, stably meeting the stringent requirements of high-concurrency, real-time streaming audio interaction, and significantly reducing streaming media transmission latency and jitter. This invention uses memory address projection technology to directly reproject the byte buffer into a short integer view, eliminating temporary memory allocation during the conversion process. Combined with merging operations, channel mapping and gain limiting are completed in a single traversal, reducing the number of data traversals, significantly reducing the CPU overhead of the streaming media server, and improving the single-node concurrent processing capability. This invention introduces a high-precision tail-end completion and silence flushing compensation mechanism, injecting a specific number of silence frames into the physical channel to forcibly flush the audio decoding queue at the receiving end, achieving smooth fade-out at the physical level, effectively eliminating audio tail-end truncation and receiving end decoding noise, and improving the listening experience. Attached Figure Description

[0019] Figure 1 This is a flowchart of a low-latency streaming media adaptive audio frame reassembly and push method provided by the present invention.

[0020] Figure 2 This is a flowchart of the adaptive sliding buffer space compact translation process in a low-latency streaming media memory-free adaptive audio frame reassembly and push method provided by the present invention.

[0021] Figure 3 This is a flowchart of the zero-memory-allocation audio data projection and channel expansion conversion process in a low-latency streaming media zero-memory-allocation adaptive audio frame reassembly and push method provided by the present invention.

[0022] Figure 4This is a flowchart of the high-precision timed pumping and silent flushing compensation process in a low-latency streaming media memory-free adaptive audio frame reassembly and push method provided by the present invention. Detailed Implementation

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0024] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0025] To eliminate memory allocation overhead during streaming audio frame reassembly, this invention achieves steady-state zero memory allocation in long-lived streaming connections by reassembling variable-length audio packets into fixed-length audio frames, fundamentally eliminating streaming latency caused by frequent garbage collection (GC). It also reduces CPU time consumption for data transformation and channel expansion, avoids multiple physical memory copies during type conversion and multi-channel expansion, and improves concurrent processing capabilities. Furthermore, it addresses the issues of truncated and popping sounds at the end of streaming, ensuring complete output of the remaining phrase at the end of the network audio stream and providing a smooth physical transition to the receiving end, eliminating decoding popping sounds. This invention provides a low-latency, memory-allocation-free adaptive audio frame reassembly and push method for streaming media, achieving low-latency, memory-allocation-free adaptive audio frame reassembly and push.

[0026] See Figures 1 to 4 As shown, this embodiment of the invention provides a low-latency streaming media adaptive audio frame reconstruction and push method without memory allocation, including adaptive sliding reconstruction buffer processing, zero-memory-allocation data projection and channel expansion conversion processing, and high-precision timed pumping and silent flushing compensation processing. The specific steps are as follows: Step 1, Adaptive sliding reassembly buffer processing: When a streaming session is established, a contiguous storage space is pre-allocated as a shared reassembly buffer, and the read pointer and write pointer are initialized; when a variable-length audio data packet is received, it is determined whether there is enough remaining space at the end of the buffer to accommodate the current data packet. If not, and there is already free space at the beginning of the buffer, the unread data is shifted to the beginning address of the buffer as a whole, the read and write pointers are reset, and the current variable-length audio data packet is appended to the buffer. No dynamic memory allocation is generated during the steady-state operation of the streaming media.

[0027] In this step, the pre-allocated contiguous storage space is a fixed-size contiguous byte array, allocated in the managed heap or unmanaged physical memory; the read pointer is used to mark the end position of the read data, and the write pointer is used to mark the end position of the written data. In this embodiment, the pre-allocation mechanism is as follows: when the streaming media session is established, the system pre-allocates a fixed-size contiguous storage space in the managed heap or unmanaged physical memory as a shared reassembly buffer, and initializes two integer pointers: a read pointer `readPos` and a write pointer `writePos`; the read pointer is used to mark the end position of the read data, and the write pointer is used to mark the end position of the written data.

[0028] The specific method for determining whether there is sufficient remaining space at the end of the buffer is as follows: determine whether the difference between the total length of the buffer and the write pointer is greater than or equal to the length of the newly arrived audio data packet; if the difference is greater than or equal to the length of the data packet, then directly append the data packet to the end and update the write pointer.

[0029] When there is insufficient space at the end of the buffer and the read pointer is greater than 0, the unread data between the read and write pointers is shifted to the starting address with an offset of 0 in the buffer using a block copy method; the read pointer is reset to 0, and the write pointer is updated to the length of the unread data; then the current variable-length audio data packet is appended to the updated write pointer position in the buffer and the write pointer is updated. When there is insufficient space at the end of the buffer and there is no free space at the beginning of the buffer, the shared buffer is physically expanded.

[0030] In this step, the lightweight translation compaction logic includes: when the network receives a variable-length audio data packet of length L, the system first evaluates whether there is sufficient remaining physical space at the end of the buffer, that is, whether the total buffer length - writePos ≥ L holds true. If there is sufficient space at the end, the current variable-length audio data packet is directly appended to the end of the buffer and the write pointer is updated. If there is insufficient space at the tail end, and the current read pointer readPos > 0 (indicating that there is free space at the head of the buffer that has been read), the system does not allocate a new array. Instead, it calls an efficient block copy method to compactly shift the unread data between readPos and writePos to the starting address of the buffer (offset 0). Then, it updates the pointer state: resets the read pointer readPos = 0, updates the write pointer to the length of the remaining unread data, and finally appends the newly arrived L-length data packet to the reset write pointer and updates writePos. If there is insufficient space at the tail end and no free space at the head of the buffer, it is an extreme case, and the shared buffer will be physically expanded.

[0031] The pre-allocation mechanism of this invention can avoid any dynamic memory allocation during the long-term operation of streaming media connections, achieving steady-state zero allocation.

[0032] Step 2, Zero-Memory Allocation Data Projection and Channel Expansion Conversion Processing: When retrieving fixed-frame-length audio data from the shared reconstruction buffer, the corresponding byte memory segment is directly reprojected into a short integer value view without physical data copying. During a single traversal of the short integer value view, the channel mapping from mono to stereo, volume gain calculation, and anti-pop overflow limiting processing are completed simultaneously to generate a standard stereo audio frame.

[0033] In this step, the reprojection of the byte memory segment to the short integer value view is performed at the pointer or lower-level view level, achieving zero physical memory copy and zero memory allocation; the short integer value view is a 16-bit integer view; when generating a dual-channel audio frame, the target frame storage array is rented from the array object pool; after processing, the temporary frame array is returned to the object pool to ensure that there is no heap memory allocation overhead during the conversion process.

[0034] In this embodiment, when the system needs to retrieve data of a fixed frame length from the buffer for streaming, it does not perform physical data copying. Instead, it uses memory reprojection technology to directly redirect and project the byte view of the corresponding segment in the buffer into a 16-bit integer view, i.e., a short value type view. This projection process is performed at the pointer or underlying view level, achieving zero physical memory copying and zero allocation.

[0035] When converting a mono numerical view into a stereo standard audio frame, the system directly rents the target frame's storage array from the array object pool, avoiding the direct creation of the array in the heap. While traversing the mono numerical view, the system simultaneously performs channel copy calculation, volume gain factor calculation, and anti-pop limiting logic processing, and the calculated data is directly written into the target stereo frame. After data processing is complete, the temporary frame array is returned to the object pool, ensuring that the entire data conversion pipeline has no memory allocation overhead.

[0036] Step 3: High-precision timed pumping and silent flushing compensation: The streaming rhythm is controlled by a high-precision periodic timer aligned with the physical duration of the audio frame, and audio frames are pushed according to the frame rate. When the audio stream end signal is received, the audio fragments that are less than one frame remaining in the shared reconstruction buffer are first filled into complete audio frames and pushed. Then, a preset number of absolute silence data frames are continuously injected into the streaming channel to flush the physical audio pipeline and the decoding queue at the receiving end. After flushing is completed, the streaming state is switched to idle state.

[0037] In this step, the trigger period of the high-precision periodic timer is precisely aligned with the physical duration of a single audio frame. The timer trigger signal is monitored asynchronously in the background, and the audio is pumped out only when the clock triggers. The tail audio fragments are filled and aligned using zero-value silence data. The injected silence frames are absolute silence data frames, numbering 50 frames, corresponding to a duration of approximately 1000ms.

[0038] In high-precision timed pumping and silent flushing compensation, the physical timer control system is configured with a high-precision periodic timer whose trigger period is highly aligned with the physical duration of the audio frame. The system asynchronously and cyclically listens for the timer trigger signal in the background, and only calls the RTC track push interface to pump out audio when the clock ticks, thereby eliminating the jitter of the push time.

[0039] In this embodiment, in the network end-of-flight (EOF) mute flushing logic, when the network audio channel receives the EOF end signal, the following operations are performed: Tail padding: The system first checks whether there are any audio fragments in the adaptive sliding buffer that are less than one frame; if so, it fills them with zero values ​​(absolute silence data) to align them to a complete frame and pushes them to the send queue. Silent flushing injection: Subsequently, the system does not immediately enter an idle state or close the channel, but actively injects a specific number of absolutely silent data frames into the push channel continuously; the timer metronome continues to pump along with these silent frames, and the high-frequency silent frames, as a "physical flushing medium", forcibly push forward the physical audio pipe and the data packets that may remain in the receiver, preventing the receiver from being truncated due to the sudden interruption of the decoding queue. State switching: After flushing is complete, the system gracefully switches the flow state to the idle state.

[0040] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0041] The embodiments of the present invention are implemented on the .NET 10 / C# platform and applied to the zLiveKit streaming media system. They are designed for RTC and TTS streaming interaction scenarios. The audio sampling is 16-bit mono Pulse Code Modulation (PCM) data, and the RTC streaming requirement is 16-bit stereo with a fixed frame length of 20ms.

[0042] like Figure 1 As shown, the low-latency streaming media memory-free adaptive audio frame reassembly and push method of the present invention includes three core processing layers, and the data flow is as follows: 1. The input is a variable-length streaming TTS network audio packet. It first enters the adaptive sliding reconstruction buffer. This buffer uses pre-allocated storage. Data reconstruction is achieved by detecting available space, adaptive compact translation, and resetting the read and write pointers. Then, the data is retrieved according to a fixed frame length. 2. The retrieved data enters the zero-memory allocation data projection and channel expansion conversion stage. First, the byte to short type conversion is completed through memory view projection, and then gain and limiting calculations are performed to generate standard two-channel audio frames. 3. Finally, the high-precision timed pump and silent flush compensation module pushes the audio stream to the physical RTC audio streaming channel with precise timing through a high-precision periodic timer, and injects a silent flush frame at EOF.

[0043] like Figure 2 As shown, the spatial compact translation process of the adaptive sliding buffer is as follows: Step 1: A new data packet arrives, with a length of L; Step 2: Determine if there is enough space at the end of the buffer, i.e., whether the total length of the buffer - writePos ≥ L is true; if true, directly append to the end, update writePos, and the process ends; if not true, proceed to step 3. Step 3: Check if there is any free space to be read, i.e., whether readPos > 0 is true; if true, perform a compact shift: shift the data between readPos and writePos to the beginning of the buffer; reset the pointer: readPos = 0, writePos = remaining length; append the data packet to the new tail, update writePos, and the process ends; if not true, enter the extreme case and perform physical expansion of the shared buffer.

[0044] In this embodiment, the specific implementation process of the adaptive sliding recombination buffer is as follows: When a streaming session is established, a contiguous byte array pcmBuffer of size 16384 bytes is pre-allocated as a shared reassembly buffer, and readPos=0 and writePos=0 are initialized at the same time.

[0045] When a network audio data packet of length L is received: If pcmBuffer.Length - writePos ≥ L, directly call the block copy method to write the data packet data to the writePos position of pcmBuffer, and then writePos... + =L; If there is insufficient space at the tail and readPos > 0, the block copy method is called to copy the unread data in pcmBuffer starting from readPos and with a length of writePos - readPos to the position at offset 0; then readPos = 0, writePos = writePos - readPos (i.e., the length of unread data); then the new data packet is written to the new writePos position and the pointer is updated; If there is insufficient space at the tail and readPos=0, it means that the buffer is full of unread data. At this time, the buffer is expanded, a larger contiguous memory is allocated and the original data is migrated.

[0046] During normal streaming, the data read and write rates are basically matched, and the buffer head will continuously generate read free space. Therefore, the compact translation mechanism can fully cover normal scenarios. There is no dynamic memory allocation during steady-state operation, which avoids GC triggering.

[0047] In the zero-memory allocation data projection and channel expansion conversion of this embodiment, when it is necessary to retrieve data with a fixed frame length of 20ms, the corresponding number of bytes frameSize is calculated, and a memory segment with a length of frameSize starting from readPos in the buffer is retrieved.

[0048] By using memory views and pointer operations, the byte memory segment is directly reinterpreted as a short integer view, without any physical data copying or memory allocation.

[0049] Rent a stereo target array of length frameSize*2 from the array object pool; iterate through each sample value in the short integer view and execute synchronously: 1. Channel duplication: Simultaneously write the mono sample value to the corresponding positions of the left and right channels; 2. Volume Gain: Multiply the sampled value by a preset volume gain factor; 3. Overflow Limiting: Saturates and limits the calculation results to prevent popping sounds caused by values ​​exceeding the range of the short type.

[0050] After processing, the dual-channel array is delivered to the streaming module and returned to the object pool after use, ensuring that there is no heap memory allocation during the entire conversion process.

[0051] In the high-precision timed pumping and silent flushing compensation of this embodiment, a high-precision periodic timer is configured with a trigger period of 20ms, which is perfectly aligned with the duration of a single frame of audio. An asynchronous loop is started in the background. After each timer is triggered, a complete frame of dual-channel audio data is read from the buffer and the push interface of the RTC track is called to complete the push, ensuring that the push timing is accurate and jitter-free.

[0052] When the TTS stream EOF signal is received: 1. Check the length of unread data in the buffer. If it is less than the number of bytes in a single frame, pad the end of the data with zero bytes to the length of a single frame to form a complete audio frame. Push the audio after completing the format conversion. 2. Continuously generate 50 frames of all-zero silence audio (corresponding to a duration of approximately 1000ms) and push them to the RTC channel frame by frame according to the original timing. These silence frames will drive the complete output of residual audio data in the physical link and the decoding queue at the receiving end, avoiding sudden queue interruption. 3. After all silent frames have been pushed, switch the streaming state to idle and wait for the next audio stream to start, without directly closing the channel.

[0053] The method described in this embodiment has been verified in a production environment: it supports compilation into a self-contained independent binary program, which can be deployed on multiple platforms such as Windows Server (Internet Information Services, IIS host) and Linux (x64 architecture); in the Linux environment, it is packaged as a system-level daemon process with a crash auto-start mechanism configured. Under concurrent stress testing of multiple audio streams, it can achieve 24 / 7 uninterrupted operation without any audio stuttering caused by memory overflow (OOM) or garbage collection (GC) jitter.

[0054] The core design principle of this invention has platform universality. In addition to .NET / C#, it can also be ported to languages ​​such as C++, Rust, Go, and Java. Besides RTC intelligent agent voice interaction, it can also be applied to scenarios with low latency requirements such as cloud gaming audio streaming, large-scale TTS audio streaming gateways, and distributed mixing systems.

[0055] This invention employs an adaptive sliding buffer to achieve zero memory reallocation during the stable operation of audio stream processing, eliminating latency jitter caused by garbage collection. Latency control is significantly improved compared to traditional solutions, stably meeting the stringent requirements of high-concurrency, real-time streaming audio interaction, and significantly reducing streaming media transmission latency and jitter. This invention uses memory address projection technology to directly reproject the byte buffer into a short integer view, eliminating temporary memory allocation during the conversion process. Combined with merging operations, channel mapping and gain limiting are completed in a single traversal, reducing the number of data traversals, significantly reducing the CPU overhead of the streaming media server, and improving the single-node concurrent processing capability. This invention introduces a high-precision tail-end completion and silence flushing compensation mechanism. By injecting a specific number of silence frames into the physical channel, it forcibly flushes the audio decoding queue of the receiving end (such as Unreal Engine 5 / browser), achieving smooth fade-out at the physical level, effectively eliminating audio tail-end truncation and receiving end decoding noise, and improving the listening experience.

[0056] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on these embodiments, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art can still combine, add, delete, or otherwise adjust the features of the various embodiments of the present invention according to the circumstances without conflict or creative effort, thereby obtaining different technical solutions that do not fundamentally depart from the concept of the present invention. These technical solutions also fall within the scope of protection of the present invention.

Claims

1. A low-latency streaming media adaptive audio frame reassembly and push method without memory allocation, applied to scenarios where variable-length streaming audio packets are converted into fixed-frame-length audio and pushed to a real-time audio / video transmission track, characterized in that... This method includes adaptive sliding reconstruction buffer processing, zero-memory allocation data projection and channel expansion conversion processing, and high-precision timed pumping and silent flushing compensation processing. The specific steps are as follows: When a streaming session is established, a contiguous storage space is pre-allocated as a shared reassembly buffer, and the read pointer and write pointer are initialized. When a variable-length audio data packet is received, it is determined whether there is enough remaining space at the end of the buffer to accommodate the current data packet. If not, and there is already free space at the beginning of the buffer, the unread data is shifted to the beginning address of the buffer as a whole. After resetting the read and write pointers, the current variable-length audio data packet is appended to the buffer. No dynamic memory allocation occurs during the steady-state operation of the streaming media. When retrieving fixed-frame-length audio data from the shared reassembly buffer, the corresponding byte memory segment is directly reprojected into a short integer value view without physical data copying. During a single traversal of the short integer value view, the channel mapping from mono to stereo, volume gain calculation, and anti-pop noise overflow limiting processing are completed synchronously to generate a standard stereo audio frame. The streaming rhythm is controlled by a high-precision periodic timer aligned with the physical duration of the audio frames, and audio frames are pushed according to the frame rate. When the audio stream end signal is received, the audio fragments that are less than one frame remaining in the shared reconstruction buffer are first filled into complete audio frames and pushed. Then, a preset number of absolute silence data frames are continuously injected into the streaming channel to flush the physical audio pipeline and the decoding queue at the receiving end. After flushing is completed, the streaming state is switched to the idle state.

2. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, The pre-allocated contiguous storage space is a fixed-size contiguous byte array, allocated in the managed heap or unmanaged physical memory; the read pointer is used to mark the end position of the read data, and the write pointer is used to mark the end position of the written data.

3. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 2, characterized in that, The specific way to determine whether there is enough remaining space at the end of the buffer is as follows: determine whether the difference between the total length of the buffer and the write pointer is greater than or equal to the length of the newly arrived audio data packet; if the difference is greater than or equal to the length of the data packet, then directly append the data packet to the end and update the write pointer.

4. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 3, characterized in that, When there is insufficient remaining space at the end and the read pointer is greater than 0, the unread data between the read pointer and the write pointer is shifted to the starting address of the buffer offset 0 using the block copy method; the read pointer is reset to 0, and the write pointer is updated to the length of the unread data; then the current variable-length audio data packet is appended to the updated write pointer position in the buffer and the write pointer is updated.

5. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 3, characterized in that, When there is insufficient remaining space at the tail and no free space at the head of the buffer, the shared buffer is physically expanded.

6. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, The reprojection of byte memory segments to the short integer value view is performed at the pointer or lower-level view level, achieving physical zero-memory copy and zero-memory allocation; the short integer value view is a 16-bit integer view.

7. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, When generating dual-channel audio frames, a target frame storage array is rented from the array object pool; after processing, the temporary frame array is returned to the object pool to ensure that there is no heap memory allocation overhead during the conversion process.

8. The low-latency streaming media memory-allocation-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, Channel mapping, volume gain calculation, and overflow limiting are completed synchronously during the same traversal of the mono numerical view, and the calculation results are directly written to the target stereo frame storage array.

9. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, The system asynchronously listens for the trigger signal of a high-precision periodic timer in the background, and only calls the real-time audio and video track push interface to pump out audio frames when the timer is triggered, thus eliminating streaming time jitter. For audio fragments at the end that are less than one frame, zero-value mute data is used to fill and align them to form a complete audio frame before pushing.

10. The low-latency streaming media memory-free adaptive audio frame reassembly and push method as described in claim 1, characterized in that, The number of injected absolute silence data frames is 50, with a corresponding duration of 1000ms. The silence data frames are used to drive the complete output of the audio data packets remaining in the physical audio pipeline and the receiving end, preventing the decoding queue of the receiving end from being interrupted, resulting in tail truncation and popping sounds.