3D multichannel digital audio real-time IP network transmission live broadcast system

Through modular design and multi-threaded processing, combined with RTP/UDP protocol and MFC interactive interface, real-time transmission and playback of multi-channel digital audio is realized, solving the stability and flexibility problems of traditional audio live streaming platforms. It supports simultaneous transmission of 24 channels and meets the needs of real-time IP network transmission of 3D audio.

CN121887783APending Publication Date: 2026-04-17谭秋林
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
谭秋林
Filing Date
2022-05-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies cannot achieve real-time transmission and playback of multi-channel digital audio. Especially in complex network environments, traditional audio live streaming platforms cannot support live streaming of 22.2 channels or custom channel numbers of 3D audio data. The system has poor stability and flexibility and cannot meet the needs of real-time IP network transmission of multi-channel audio.

Method used

It adopts a modular design and multi-threaded processing, and builds a 3D audio data transmission channel through RTP/UDP protocol to realize the real-time transmission of multi-channel digital audio. It adopts MFC visual interactive interface, supports arbitrary channel configuration and speaker settings, and combines server/client architecture for task allocation and control to ensure the real-time performance and stability of the system.

Benefits of technology

It enables real-time transmission and playback of multi-channel digital audio, supports simultaneous transmission of 24 channels, improves system stability and flexibility, reduces live streaming latency to less than 2 seconds, meets the real-time live streaming needs of different channel configurations, and promotes the popularization of 3D audio technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887783A_ABST
    Figure CN121887783A_ABST
Patent Text Reader

Abstract

The multi-channel digital audio has rich azimuth information and vivid scene sense, the problem of real-time transmission and playing of the multi-channel digital audio is solved, firstly, a 3D audio data transmission channel is constructed based on an RTP / UDP efficient transmission protocol, and the problem of real-time audio transmission is solved; 2, a multi-channel modular design is adopted, independent IP network analysis, design and implementation are carried out on different functional modules of audio encoding and decoding, data transmission and cache playing, distribution and start-stop control of task threads are completed based on multi-thread live broadcast, and the flexibility of audio live broadcast is improved; thirdly, standard data interfaces are applied to all functional modules, so that modification and transplantation of system functions are facilitated; and 4, a loudspeaker configuration module is designed and realized, a user can set a loudspeaker at a specified position for sounding according to any sound channel configuration, and at most 24 sound channels can be set for simultaneous transmission and playing, so that the flexibility of the live broadcast system is further improved, and the targets of real-time performance, flexibility and stability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a multi-channel audio live streaming system, and more particularly to a 3D multi-channel digital audio real-time IP network transmission live streaming system, belonging to the field of multimedia 3D audio technology. Background Technology

[0002] As digital audio technology continues to develop and innovate, people are no longer satisfied with simple audio information; instead, they seek a more realistic listening experience and the entertainment experience brought by diverse audio application scenarios. In recent years, the emergence of 3D audio has brought audio technology to a wider range of applications. 3D audio solves the shortcomings of traditional two-channel audio or horizontal single-layer surround sound in terms of spatial orientation, layer distribution, clarity, and presence. It can be applied in 22.2-channel systems and WFS concert halls, simulating a realistic listening environment. Furthermore, its channel count far exceeds that of 5.1 surround sound, making the listening experience more realistic and providing listeners with a comprehensive listening experience. However, in multi-channel audio playback environments built using 3D audio technology, there is still no mature real-time audio transmission solution. Currently, traditional audio transmission platforms such as online music radio broadcasts, music TV program broadcasts, and real-time audio conferencing are approaching saturation. Therefore, realizing a real-time live streaming system suitable for multi-channel audio environments is of great significance and will also help further expand the application scenarios of 3D audio.

[0003] However, current mainstream audio live streaming platforms cannot support live playback of 22.2-channel or other user-defined channel counts of 3D audio data. Unlike traditional audio live streaming, multi-channel digital audio live streaming systems place higher demands on the real-time performance of data transmission. Ensuring the normal transmission of large amounts of audio data to the receiving end in complex network environments, and rationally integrating encoding / decoding modules and playback modules into the system to guarantee system stability, are all problems that need to be solved.

[0004] Current multi-channel audio live streaming technologies primarily rely on broadcast television networks and satellite television broadcasting, limiting their application to users connected to these networks. This restricts the widespread adoption of multi-channel real-time transmission systems. Transmitting multi-channel digital audio over a network represents a new demand, and IP networks offer a more open platform. IP-based multi-channel live streaming boasts strong real-time transmission capabilities and a high degree of openness, facilitating its expansion into more application scenarios and opening up broader development opportunities for real-time multi-channel digital audio transmission.

[0005] Currently, all existing multi-channel digital audio live broadcasts are carried out on broadcast television networks, with relatively limited application scenarios. Furthermore, these broadcast television-based live broadcasts generally have relatively long broadcast delays. In contrast, IP networks provide real-time data transmission services. Using IP networks for multi-channel digital audio live broadcasts can reduce live broadcast delays compared to traditional television live broadcasts, and people have more diverse ways to access live broadcasts, which precisely meets the demand for real-time multi-channel audio live broadcasts.

[0006] In summary, existing multi-channel audio live streaming technologies still have several problems and shortcomings. The problems and key technical challenges addressed in this application include:

[0007] (1) Currently, people are pursuing a more realistic listening experience and entertainment experience brought by diverse audio application scenarios. However, traditional two-channel audio or horizontal single-layer surround sound has serious shortcomings in terms of spatial orientation, layer distribution, clarity, and presence. It cannot be applied to 22.2-channel systems and WFS concert halls, cannot simulate a real listening environment, and has a small number of channels and surround sound volume, resulting in an unrealistic listening experience and listeners not being able to obtain a full-range listening effect. In the multi-channel audio playback environment constructed using 3D audio technology, there is currently no mature real-time audio transmission solution. At present, the development of traditional audio transmission platforms such as online music radio broadcasts, music TV program broadcasts, and real-time audio conferencing is approaching saturation. Therefore, there is an urgent need to develop a real-time live broadcast system suitable for multi-channel audio environments to further expand the application scenarios of 3D audio.

[0008] (2) Existing mainstream audio live streaming platforms cannot support live streaming of 22.2 channels or other user-defined channels of 3D audio data. Multi-channel digital audio live streaming systems have higher requirements for the real-time performance of data transmission. Existing technologies cannot transmit large amounts of audio data normally to the receiving end in complex network environments. They cannot reasonably integrate the encoding / decoding module and playback module into the system, resulting in poor system stability. In particular, based on the large number of channels and the large amount of audio data, they cannot achieve real-time transmission and playback of audio with dual channels, 5.1 channels, 7.1 channels, 22.2 channels and other user-defined channels. The real-time live streaming effect of multi-channel audio is poor and cannot meet the needs of real-time IP network transmission and live streaming of 3D multi-channel digital audio. They lack 3D audio encoding / decoding technology and speaker simplification technology, resulting in fewer 3D audio live streaming scenarios, which is not conducive to promoting the high-quality popularization and promotion of 3D audio technology applications.

[0009] (3) Existing multi-channel audio live streaming technology is basically transmitted through broadcast television networks and satellite television broadcasting. Its application scope is limited to user scenarios that access broadcast television networks, which restricts the widespread application of multi-channel real-time transmission systems. Transmitting multi-channel digital audio over the network is a new demand, but the current audio live streaming has limitations and low effectiveness. It lacks a real-time audio data transmission network protocol, cannot parse the live streaming model of IP network structure, lacks a 3D audio data transmission channel based on the RTP / UDP high-efficiency transmission protocol, and cannot solve the problem of real-time audio transmission. It lacks multi-channel modular design, and cannot perform independent IP network analysis, design and implementation of different functional modules such as audio encoding and decoding, data transmission, and buffer playback. It cannot complete the allocation and start / stop control of each task thread based on multi-threaded live streaming, resulting in poor flexibility of audio live streaming. Each functional module lacks standardized data interfaces, making it difficult to modify and port system functions. Users cannot set the speaker in a specified position to produce sound according to any channel configuration, resulting in poor system real-time performance, flexibility and stability.

[0010] (4) In the absence of an audio signal acquisition subsystem, the sending end cannot obtain the audio data to be sent by reading audio files under a specified path; the sending end cannot configure the number of live audio channels, and the system cannot read the corresponding number of mono WAV audio files; the sending end cannot encode and compress the audio data stream, and cannot package and send the encoded data; the sending end cannot set the destination IP address and base transmission port, and cannot establish an end-to-end RTP transmission session; the receiving end cannot buffer the received audio data packets, and cannot solve the jitter problem caused by data transmission over the network; the playback end cannot configure the speakers according to the number of channels, and cannot realize the function of playing audio data sent to a specified speaker channel; it does not support multi-channel digital audio live playback with any number of channels configured within 24 channels. The existing multi-channel digital audio data cannot maintain correct consistency before and after network transmission. Under appropriate pressure and long-term continuous operation, the system has difficulty maintaining the stability of the live service, and cannot guarantee the smoothness and good playback quality of audio playback. The live system has a long playback delay, with a real-time performance index of more than 2 seconds. Summary of the Invention

[0011] This application designs and implements a multi-channel digital audio live streaming system, achieving modular design for each function. Through the explanation of the characteristics of multi-channel digital audio live streaming and the analysis of network transmission methods, the RTP real-time transmission protocol is finally set to realize the function of multi-channel digital audio data transmission over IP networks, establishing a transmission channel for multi-channel digital audio data between different IP addresses. The modular programming approach is used to implement the acquisition analog module, encoding / decoding module, speaker configuration module, and ASIO-based audio driver playback control module. Each module has standardized data interfaces, facilitating subsequent system modification and integration. MFC is used to implement the visual interactive interface of this application's live streaming system, allowing channel number settings, client IP address settings, and speaker configuration calculations to be completed on the interactive interface, improving the system's integrity and ease of use.

[0012] To achieve the above technical effects, the technical solution adopted in this application is as follows:

[0013] This 3D multi-channel digital audio real-time IP network transmission and live streaming system utilizes a real-time audio data transmission network protocol to achieve a real-time multi-channel digital audio transmission and playback system based on an IP network. Firstly, by analyzing the IP network structure live streaming model, a 3D audio data transmission channel is constructed based on the high-efficiency RTP / UDP transmission protocol, solving the real-time audio transmission problem. Secondly, in the system architecture design, a multi-channel modular design is adopted, with independent IP network analysis, design, and implementation for different functional modules such as audio encoding / decoding, data transmission, and buffered playback. Multi-threaded live streaming is used to allocate and control the start / stop of each task thread, improving the flexibility of audio live streaming. Thirdly, each functional module uses standardized data interfaces, facilitating system function modification and portability. Fourthly, based on the different channel configuration characteristics of multi-channel digital audio, a speaker configuration module is designed and implemented, allowing users to set speakers in specified positions according to any channel configuration, and up to 24 channels can be transmitted and played simultaneously, further improving the flexibility of the live streaming system.

[0014] First, the RTP real-time transmission protocol is configured to enable multi-channel digital audio data transmission over an IP network, establishing a transmission channel for multi-channel digital audio data between different IP addresses. Then, multi-channel modular programming is used to implement the acquisition analog module, encoding / decoding module, speaker configuration module, and ASIO-based audio driver-based playback control module, with standardized data interfaces set for each module. Next, MFC is used to implement the visual interaction of the live streaming system, allowing channel number settings, client IP address settings, and speaker configuration calculations to be completed within the interactive interface. Finally, multi-threading allocation and control for each module are configured, adopting a server / client-based system architecture. Multi-threading is used to allocate and control the start and stop of each task thread, designing corresponding trigger conditions and termination flags for different tasks on the server and client sides to avoid conflicts between different task threads.

[0015] Furthermore, the overall architecture of the multi-channel audio system: Multi-channel digital audio live streaming is a unidirectional data stream transmission. The C / S architecture is adopted to complete the data acquisition and compression encoding work with large computational loads on the server side, while the decoding and playback tasks are completed on the client side. This makes full use of the hardware on both ends and distributes the tasks to the client and server sides, reducing the system's resource consumption.

[0016] The overall framework of the multi-channel digital audio live streaming system and the deployment of server and client functional modules are as follows: The system is divided into two layers: server and client. The server side includes an audio source acquisition module, an encoding module, and a transmission control module. The client side includes a receive buffer parsing module, a decoding module, a speaker configuration module, and a playback module. During the operation of the live streaming system, once the client starts the live streaming service, it begins monitoring the transmission port and receiving data. There is no signaling control, and the connectionless UDP protocol eliminates the need for a handshake connection establishment process, making it more flexible and convenient to use.

[0017] After the live streaming service is started on the server side, the associated audio channel configuration initialization calculation is performed first. Then, the acquired audio data is frame-encoded and then enters the sending control module, which includes establishing an RTP session, setting the port number, and defining and calculating the header of each audio frame. The data packets are then sent to the lower-layer IP network for transmission. On the client side, the receiving module traverses all data sources, unpacks the received data packets, stores them in the dejitter buffer for calculation, and finally decodes them and outputs them to the playback device through the speaker configuration module to achieve live streaming of multi-channel digital audio.

[0018] Furthermore, the server thread handles real-time processing: reading in multi-channel digital audio data streams and simultaneously encoding, compressing, and sending them; the server's main thread performs system configuration and initialization work based on MFC calculations, and this thread completes system configuration such as setting the destination IP address and setting the RTP session base transmission port.

[0019] Once the initialization is complete, the live stream begins. The server system creates an encoding and sending thread. Under this thread, the server system completes the tasks of reading data, encoding and compressing, establishing RTP sessions, and sending data to the destination address. After the server system's main thread performs the calculation to shut down the service, the system stops encoding and sending audio data and releases resources for the data encoding and sending thread, thus enabling control over the start and stop of the encoding and sending thread.

[0020] Furthermore, the client-side thread handles real-time processing: The client system completes the audio driver settings and speaker configuration initialization in the main thread. When the live streaming service is started, the client system starts the data receiving thread, creates an RTP receiving session, and monitors the session transmission port to receive and buffer data packets in sequence. When the number of data packets in the receiving buffer exceeds the set value, the client starts the decoding buffer thread, retrieves the data from the receiving buffer, and performs decoding calculations.

[0021] After the main thread begins audio playback calculations, the client system opens a new playback thread. This playback thread writes the decoded data into the playback buffer and sends it to the designated speaker for playback. After the main thread performs the stop playback calculation, the client system closes the playback thread. Finally, when the live streaming service is closed in the main thread, the system releases the data receiving and decoding thread, ending the live streaming service.

[0022] Furthermore, the 3D audio data acquisition module: The acquisition module uses WAV file streams for simulation. During system configuration, a corresponding number of WAV files matching the current channel configuration are prepared and placed in the project's default path. When the live streaming service starts, these WAV audio files are read in, and then a multi-channel digital audio data encoding and sending thread is created, thereby enabling the simultaneous parallel processing of audio data from each channel.

[0023] Audio data streaming processing: The file pointer _audio_file is set to point to the beginning of the audio data block, and then the data is retrieved frame by frame to achieve streaming input of audio data. channel_id represents the channel number, frame_seq represents the original audio frame number, frame_num represents the number of audio frames, and payload represents the storage area of ​​the audio frames. The file pointer _audio_file is offset by 40 bytes to point to the starting address of the number of audio data bytes in the data sub-block, and then 4 bytes of data are read down to obtain the size value of the data. Here, AUDIO_FRAME_SIZE_CODEC represents the length of one frame of audio data, which is defined as 1024. Then, based on the sampling depth of the audio file used, the number of audio frames in the entire WAV file is further calculated.

[0024] After calculating the number of audio frames, the audio block is divided into frames. This module periodically retrieves data from the audio block. When the audio data sampling rate is 48000Hz, the playback time of each audio frame is calculated to be 21.333 milliseconds (1 / 48000*1024). The timestamp parameter _next_read_timestamp is set according to this time interval. Each time 21.333 milliseconds have elapsed, a frame of data is read from the file and stored in the payload for the next encoding process. At the same time, the channel number and frame number information of the frame are recorded for subsequent data packet encoding and transmission.

[0025] The method of extracting audio data blocks from WAV files frame by frame is used to simulate the real-time audio stream acquisition process, thereby realizing the data stream output of the data acquisition module.

[0026] Furthermore, the multi-channel playback module: based on the ASIO driver playback buffer double buffer mechanism, it can control and play up to 24 channels of audio. When a large amount of audio data needs to be played in real time, the double buffer includes BufferA and BufferB. BufferA is responsible for receiving the audio data decoded from the upper layer and filling it, while BufferB is responsible for sending the audio data in its buffer to the lower-level kernel driver to make the speaker emit sound. The two buffers work simultaneously. When BufferA is full and BufferB has been sent to the kernel, the functions of the two buffers are interchanged, thereby realizing the streaming playback of data filling and data output. This process is represented by the processing flow of three consecutive frames of audio data of a single channel in the double buffer.

[0027] When audio playback begins, the underlying ASIO driver instructs the upper-layer software to fill the buffer with audio data. At this time, the system automatically calls the callback function bufferSwitchTimeInfo() to fill the playback buffer with audio data. Whenever one side of the data in the dual buffer corresponding to a channel has finished playing, the system plays the other side of the data, and the callback function fills the buffer that has finished playing with data. The parameters of bufferSwitchTimeInfo() also include speaker configuration information, buffer length, and channel sampling type. This application changes these parameters to enable audio playback for different speaker configurations and different sampling formats.

[0028] In 3D audio, the audio data played by each speaker corresponds to the playback buffer of the corresponding channel. 24 double buffers are created, and 24 pointers pszCh[0], pszCh[1], pszCh[2], ..., pszCh

[23] are defined respectively, pointing to buffer 0, buffer 1, buffer 2, ..., buffer 23 respectively. Finally, as long as the channel configuration of the live audio is filled into the corresponding buffer pointer according to the channel configuration, the playback of different combinations of speakers can be realized. The buffer corresponding to each speaker is set to make the specified speaker emit sound and make the playback software compatible with various channel configurations.

[0029] Furthermore, the 3D speaker configuration module specifies the speaker allocation for the multi-channel digital audio signal and the playback buffer for each speaker before audio playback. The position of each speaker in the 3D audio is arranged based on the NHK22.2 multi-channel system, which is divided into three layers: upper, middle and lower. There are 9 speakers in the upper layer, 10 speakers in the middle layer and 3 speakers in the lower layer, plus two subwoofers, for a total of 24 speakers.

[0030] Each of the 24 speakers is assigned a corresponding dual buffer at the system software level. Each channel buffer corresponds one-to-one with each signal channel in the sound card device. The correspondence between each channel buffer and the speaker is consistent with the relationship between each channel of the sound card and the speaker. As long as the channel setting box in the speaker configuration interface is associated with each channel buffer according to this correspondence, the playback and sound output of the specified speaker can be achieved.

[0031] First, based on the channel configuration of the live audio, determine which speakers will be used in the multi-channel playback environment. Then, in the speaker settings interface, select the checkboxes for these speakers in sequence. Assuming that six-channel audio data is being played, select the checkboxes p1, p2, p3, p4, p5, and p6 in sequence according to the correspondence between the channel settings and the speakers. The system then iterates through the speaker settings boxes in the configuration interface. If speaker p is selected, the flag choose[p] is set to 1. At the same time, the system reads the channel information of the original audio and proceeds to the next step of speaker configuration determination. If the number of speakers set is inconsistent with the number of channels in the original audio, a prompt box will pop up and... Return to the speaker settings interface to reset; otherwise, proceed to the next step of allocating the playback buffer. After successfully completing the speaker settings, open the corresponding buffers for these speakers. Check the flag bits choose[0] to choose

[23] of speakers 1 to 24 in sequence to see if they are true. If the values ​​of choose[p1], choose[p2], choose[p3], choose[p4], choose[p5] and choose[p6] are all 1, then open the buffer corresponding to the speaker and close the other buffers with choose[n] = 0, thus completing the association from the speaker settings interface to the specified buffer.

[0032] After the speaker configuration process described above, when the live streaming service is started and data from each channel is sent to the playback buffer, the data stream automatically skips the unselected speakers and is stored sequentially in the buffer of the selected speakers, thus enabling playback control of the speakers at the specified locations.

[0033] Furthermore, the audio encoding and decoding module:

[0034] (1) 3D audio encoding module

[0035] The audio encoding task is divided into five states: Begin, Read, Encode, Send, and Done. These states represent the encoder initialization at the start of the encoding task, reading the raw audio frame data, sending the encoded data packet to the encoder for encoding, and the end of the encoding task.

[0036] Before all the original audio frames are retrieved, the encoder periodically reads the audio frames, encodes the frames, and transmits them to the sending module. After all the original audio frames are retrieved, AudioEncodeTask_Read returns 0, and the status jumps to AudioEncodeTask_Done, ending the encoding task.

[0037] In the AudioEncodeTask_Encode state, the audio encoder interface function is called to encode one frame of audio data. The encoding interface function is defined as follows:

[0038] int Codec::encode(short*frame,unsignedchar*bits)

[0039] The encoding function takes an audio frame and the address of the encoded bitstream as input, and is called as follows:

[0040] _packet_ptr.packet->payload_size=

[0041] _codec.encode(_frame_ptr.frame->payload,

[0042] _packetptr.packet->payload+sizeof(RTPHeader));

[0043] The encoding function _codec.encode() takes the audio frame _frame_ptr.frame as input to the encoder, outputs the encoded bitstream and stores it in _packet_ptr.packet->payload+sizeof(RTPHeader). After the audio frame is encoded, the value of the audio frame sequence number is assigned to the corresponding encoded bitstream data packet.

[0044] (2) Multi-channel decoding module

[0045] The decoding task begins with AudioDecodeTask_Begin, and the decoder is initialized. In the AudioDecodeTask_Decode state, the decoder interface function is called to retrieve a frame of data from the RTP receive buffer and then send it to the decoder for decoding. The corresponding audio frame after decoding is stored in the decoding buffer _decode_buf, ready to be sent to the playback module for processing and playback. The decoding interface function _codec.decode() is defined as follows:

[0046] unsignedlongCodec::decode(unsignedchar*bits,

[0047] unsignedlongbits_len,void*frame)

[0048] The decoding function is called as follows:

[0049] _decode_buflen=

[0050] _codec.decode(audio_packet.packet->payload,

[0051] audiopacket.packet->payload_size,_decode_buf);

[0052] audio_packet.packet->payload stores the data packets before decoding, _decode_buf is the storage address of the decoded audio data frames, and audio_packet.packet->payload_size represents the length of the data packets extracted from decoding one frame.

[0053] The decoding module's states are: Fetch (the decoder retrieves the corresponding data packets from the RTP receive buffer) and Add (the decoded audio frames are added to the decoder's _decode_buf to be processed). The completion of the decoding task is not related to whether there are data packets available in the receive buffer. The end sign is when the _to_be_ended flag is true, at which point the system determines that the decoding task is complete and ends the audio data decoding.

[0054] Furthermore, the IP network audio transmission control module: During the transmission of audio data packets, it first performs initialization calculations for the RTP session, sets session parameters such as payload type and timestamp parameter values, defines transmission port parameters, then adds RTP packet headers to the audio data packets, modifies the SSRC field, finally specifies the client's IP address and base transmission port, calls the send_rtp_packet method to send the audio data packets, and performs session termination calculations after the sending task is completed;

[0055] (1) RTP session initialization

[0056] First, an instance describing the current session is created using the CRTPSendSession class. Then, the Creat() method of this class is called to complete the initialization calculation. The Creat() method sets two parameters: session parameters sessparams and transport parameters transparams. sessparams describes the parameters used by the CRTPSendSession instance, especially setting the appropriate timestamp unit by calling the SetTimestampUnit() method of the RTPSession class. The transparams parameter sets the transport layer UDP parameters based on IPv4.

[0057] The audio data timestamp will increment by 1 in each sampling period. The sampling rate of the live audio data is 48kHz. SetOwnTimestampUnit() is used to set the timestamp unit to the reciprocal of the audio sampling rate, 1 / 48000. This parameter needs to be modified accordingly when the sampling rate of the transmitted audio source changes. status indicates the status feedback of this initialization work. When status is greater than 0, the flag _available value becomes true, completing the initialization of this RTP session.

[0058] (2) Real-time transmission of RTP audio packets

[0059] The sending module packages the acquired audio stream before sending data in this RTP session. The process is as follows: first, the RTP packet header is filled, then the audio data is added to the RTP packet payload, and finally the total RTP length is verified before sending.

[0060] During the process of filling in the standard fields of the RTP header, the Version, Padding, Extension, PayloadType, Length, Marker, SequenceNumber, Timestamp, and SSRC identifier need to be assigned and calculated. The payload is a data stream in a multi-channel digital audio encoding format, and its payload type is 96 as defined in RFC3551.

[0061] Then, the `add_rtp_header()` function is used to add RTP packet header data. Here, `length` represents the length of the payload audio data. If the payload length is less than the fixed length of the RTP header, no RTP header is added to the payload data. The `extension` bit is set to zero, indicating that no extension field is needed, thus improving RTP transmission efficiency. When calling this function, the actual arguments for the last three parameters—`sequence`, `timestamp`, and `ssrc`—are `(unsigned short)(Oxfff&_task_info.channel_id)`, `_packet_ptr.packet->frame_num`, and `_packet_ptr.packet->frame_seq`, respectively. These parameters are used to identify the frame number and frame sequence of the audio data in the RTP packet, which is then used for subsequent identification, classification, and synchronization at the receiving end.

[0062] Before sending the packetized RTP packet by calling send_rtp_packet, the length of the packet is checked.

[0063] Determine the possible value range for the grouped data;

[0064] Before sending an RTP packet, it must be verified that the payload_size does not exceed 1460 bytes and the rtp_packet_length does not exceed 1500 bytes.

[0065] send_rtp_packet(_task_info.decode_server_addr,

[0066] _packet_ptr.packet->payload,

[0067] _packetptr.packet->payload_size);

[0068] Finally, the send_rtp_packet() function is called to specify the address of the receiving end and send the RTP packet data to the destination address, completing the task of packaging and sending the encoded audio stream.

[0069] Furthermore, the IP live audio receiving and processing module instantiates the corresponding RTP receiving process through JRTBLIB to realize data reception and caching. The data receiving module calls the Poll() method in the RTPSession class to receive audio data packets listened to at the transmission port. The receiving mode can be set to RECEIVEMODE_ALL, the default mode. By detecting all received data packets in the transmission port, RTP payload data and RTP packet identifier association information are extracted from them.

[0070] When the live streaming service starts, the receiving end creates an RTP receiving session and a receiving thread. Before implementing the receiving function, it performs initialization calculations on the RTP receiving session and creates a CRTPRecvSession class to implement the receiving session. Its initialization calculations are consistent with those of the sending end.

[0071] The process of receiving RTP data packets and the handling of RTP receive sessions:

[0072] First, the module calls the start() method in CRTPRecvSession() to start the RTP packet reception service on the receiving end. The start() method first reads the local configuration information and initializes it, then creates an RTP reception task thread and creates an RTP reception session _rtp_recv_session() under the thread. When the RTP reception session is initialized, the data reception monitoring port is set to correspond to the sending port, and the base port _rtp_recv_base_port is set to 20000.

[0073] Once the RTP receive session is created and initialized, the callback function on_recv_rtp_packet() will automatically retrieve the RTP packets arriving at the corresponding receive port, perform unpacking calculations, and identify the sequence, timestamp, ssrc, and payload_type identifiers in the RTP packets. Next, it uses the buffer pointer ac to point to the starting address of the audio data within the RTP packet, and finally calls addaudio_packet() to perform data buffering. It reads data of length - sizeof(RTPHeader) from the buffer pointer ac and sorts the data according to the channel number sequence and frame number ssrc to achieve ordered buffering of the audio payload data.

[0074] Audio payload data is cached using a doubly linked list data structure to store audio data. The sequential audio stream is cached by inserting elements into the list, especially inserting elements at the beginning and end, and changing the pointers of the elements.

[0075] When using add_audio_packet to perform data caching calculations, a maximum buffer size is set for each channel buffer, which is the maximum length of the audio data packet buffer queue, _max_audio_packet_num. When the cached data packets exceed this length, the data at the head of the queue will be cleared and discarded to ensure that subsequent data can be received smoothly. The audio data packet queue buffer eliminates some of the jitter problems caused by network transmission.

[0076] The audio data transmission in this application is carried on the basis of IP network. The maximum buffer length _max_audio_packet_num is set to 500. This buffer can cache a maximum of 500 data packets at the same time, eliminating jitter caused by large network fluctuations.

[0077] Compared with existing technologies, the innovations and advantages of this application are as follows:

[0078] First, this application's multi-channel digital audio possesses rich location information and a realistic sense of presence, and solves the problem of real-time transmission and playback of multi-channel digital audio. Addressing the limitations and low effectiveness of traditional broadcast television multi-channel digital audio live streaming, it implements a real-time multi-channel digital audio transmission and playback system based on an IP network through a real-time audio data transmission network protocol. This system achieves several improvements: 1) It constructs a 3D audio data transmission channel based on the high-efficiency RTP / UDP transmission protocol, solving the real-time audio transmission problem; 2) It adopts a multi-channel modular design, independently analyzing, designing, and implementing different functional modules for audio encoding / decoding, data transmission, and buffered playback using IP networks; and 3) It uses multi-threaded live streaming to allocate and control the start / stop of each task thread, improving the flexibility of audio live streaming; 4) Each functional module uses standardized data interfaces, facilitating system function modification and portability; and 5) It designs and implements a speaker configuration module, allowing users to set speakers in specified positions according to any channel configuration, and can simultaneously transmit and play up to 24 channels, further improving the flexibility of the live streaming system. It achieves real-time transmission and playback of multi-channel digital audio supporting different channel configurations, with each functional module designed independently, and the system achieves the goals of real-time performance, flexibility, and stability.

[0079] Secondly, this application designs and implements a multi-channel digital audio live streaming system, achieving modular design for each function. Through the explanation of the characteristics of multi-channel digital audio live streaming and the analysis of network transmission methods, the RTP real-time transmission protocol was ultimately set to realize the function of multi-channel digital audio data transmission over an IP network, establishing a transmission channel for multi-channel digital audio data between different IP addresses. The modular programming approach was adopted to implement the acquisition analog module, encoding / decoding module, speaker configuration module, and ASIO-based audio driver playback control module. Each module has standardized data interfaces, facilitating subsequent system modification and integration. MFC was used to implement the visual interactive interface of this application's live streaming system, allowing channel number settings, client IP address settings, and speaker configuration calculations to be completed on the interactive interface, improving the system's integrity and ease of use.

[0080] Third, this application implements multi-threaded allocation and control of various functional modules in the system. In the design and implementation of the multi-channel digital audio live streaming system, a server / client-based system architecture is adopted, combined with multi-threaded processing technology to achieve allocation and start / stop control of each task thread. Furthermore, for different task functions of the server and client, corresponding triggering conditions and termination flags for task threads are designed to avoid conflicts between different task threads and ensure the stability of the live streaming system. The multi-channel digital audio data remains correct and consistent before and after network transmission, and the transmission module transmits data correctly. The system can maintain stable live streaming service under appropriate pressure and long-term continuous operation, ensuring smooth audio playback and good playback quality. The live streaming system should maintain a short live playback latency; in this application, an overall live streaming latency of less than 2 seconds is used as a real-time performance indicator.

[0081] Fourth, this application expands existing players that support multiple audio signals by connecting them to a multi-channel digital audio transmission module. This changes the original player's single function of only being able to read and play WAV files from the local source, enabling it to support the playback of real-time audio data streams. Ultimately, it combines the encoding / decoding module, transmission module, and speaker configuration module to form a multi-channel digital audio live streaming system. First, even without access to the audio signal acquisition subsystem, the sending end can obtain the audio data to be sent by reading audio files from a specified path. Second, the sending end can configure the number of live audio channels, and the system can read the corresponding number of mono WAV audio files. Third, the sending end can encode and compress the audio data stream and package and send the encoded data. Fourth, the sending end can set the destination IP address and base transmission port to establish an end-to-end RTP transmission session. Fifth, the receiving end can buffer the received audio data packets to solve the jitter problem caused by data transmission over the network. Sixth, the playback end can configure speakers according to the number of channels to play audio data sent to a specified speaker channel. Seventh, the live streaming system of this application supports multi-channel digital audio live streaming playback with any number of channels configured up to 24. All functional requirements are met, and the expected performance indicators are achieved.

[0082] Fifth, based on the large number of audio channels and the large amount of audio data, this application designs and implements a multi-channel digital audio live streaming system supporting 3D audio. It enables real-time transmission and playback of audio with dual channels, 5.1 channels, 7.1 channels, 22.2 channels, and other user-defined channel numbers, while ensuring the system's real-time performance and stability, guaranteeing a good real-time live streaming effect for multi-channel digital audio. This application not only meets the needs of real-time IP network transmission and live streaming of 3D multi-channel digital audio, but also, by combining more efficient 3D audio encoding and decoding technologies and speaker simplification technologies, expands 3D audio live streaming to a wider range of application scenarios, promoting the research and application of core technologies in key 3D audio projects and the high-quality popularization and promotion of 3D audio technology applications. Attached Figure Description

[0083] Figure 1 This is a framework diagram of a multi-channel digital audio real-time IP network transmission live streaming system.

[0084] Figure 2 This is a flowchart of the server thread real-time processing design.

[0085] Figure 3 This is a flowchart of the client thread processing design.

[0086] Figure 4 This is a schematic diagram of the double buffer processing flow of the playback module.

[0087] Figure 5 This is a diagram showing the locations of the 24 speakers in a multi-channel audio live streaming system.

[0088] Figure 6 This is the design flowchart for the data transmission module.

[0089] Figure 7 This is a flowchart of the RTP audio data packet sending process.

[0090] Figure 8 This is a flowchart of the real-time packet transmission process for RTP audio packets. Specific implementation methods

[0091] The technical solution of the 3D multi-channel digital audio real-time IP network transmission live streaming system provided in this application will be further described below with reference to the accompanying drawings, so that those skilled in the art can better understand this application and implement it.

[0092] In today's rapidly developing multimedia information technology landscape, people have high demands for diverse information content and real-time message delivery. In the field of digital audio, traditional mono and stereo audio can no longer satisfy people's pursuit of a sense of presence and immersion. Multichannel digital audio, with its rich spatial information and realistic sense of presence, has become a hot topic. However, how to transmit and play multichannel digital audio in real time has become a pressing problem to be solved.

[0093] This application addresses the limitations and low timeliness of traditional broadcast television multi-channel digital audio live streaming. It implements a real-time multi-channel digital audio transmission and playback system based on an IP network through a real-time audio data transmission network protocol. First, by analyzing the IP network structure live streaming model, a 3D audio data transmission channel is constructed based on the high-efficiency RTP / UDP transmission protocol, solving the real-time audio transmission problem. Second, in the system architecture design, a multi-channel modular design is adopted, with independent IP network analysis, design, and implementation for different functional modules such as audio encoding / decoding, data transmission, and buffered playback. Multi-threaded live streaming is used to allocate and control the start / stop of each task thread, improving the flexibility of audio live streaming. Third, each functional module uses standardized data interfaces, facilitating system function modification and portability. Fourth, considering the different channel configuration characteristics of multi-channel digital audio, a speaker configuration module is designed and implemented. Users can set speakers in specified positions according to any channel configuration, and up to 24 channels can be transmitted and played simultaneously, further improving the flexibility of the live streaming system.

[0094] This application enables real-time transmission and playback of multi-channel digital audio with different channel configurations. The system adopts a server / client architecture, with each functional module designed independently, achieving the goals of real-time performance, flexibility, and stability.

[0095] I. System Development Environment

[0096] The multi-channel digital audio live streaming system of this application is developed in Windows environment using C++ language on the Visual Studio 2019 platform. It also adopts JRTPLIB to support RTP data packet transmission. JRTPLIB is an open source wrapper library for RTP developed using object-oriented programming. It is designed based on RFC1889 and RFC3550 to meet the needs of RTP protocol for real-time audio transmission.

[0097] II. System IP Network Analysis

[0098] This application extends existing players that support multiple audio signals by connecting them to a multi-channel digital audio transmission module. This changes the original player's single function of only being able to read and play WAV files from the local source, enabling it to support the playback of real-time audio data streams. Ultimately, it combines the encoding / decoding module, transmission module, and speaker configuration module to form a multi-channel digital audio live streaming system.

[0099] For an audio live streaming system, the process of an audio signal being transmitted from one end to another and perceived by people can be roughly divided into the following steps: audio signal acquisition, compression encoding, packet transmission, classification and reception, buffering and decoding, and speaker playback. This application focuses on the packet transmission, classification and reception, buffering and parsing, and sending the audio data to the speaker for playback of multi-channel digital audio. Regarding signal acquisition, considering the large number of multi-channel digital audio microphones required and the complexity of actual operation, this application uses a file stream reading method to simulate the sound pickup process.

[0100] The system as a whole is required to transmit multiple audio signals and achieve continuous and stable real-time playback. Therefore, various indicators of the live streaming system are measured to clarify the system's functional and performance requirements, providing reference indicators for subsequent system testing and verifying the feasibility and stability of the live streaming system.

[0101] III. Overall Architecture of Multi-channel Audio System

[0102] Multi-channel digital audio live streaming is a unidirectional data stream transmission. It adopts a C / S architecture to complete the data acquisition and compression encoding tasks with large computational loads on the server side, while the decoding and playback tasks are completed on the client side. This makes full use of the hardware on both ends and distributes the tasks to the client and server sides, reducing the system's resource consumption.

[0103] The overall framework of the multi-channel digital audio live streaming system in this application, as well as the deployment of the server and client functional modules, are as follows: Figure 1 As shown, the upper and lower layers are divided into server and client respectively. On the server side, there are audio source acquisition module, encoding module, and sending control module; on the client side, there are receiving buffer parsing module, decoding module, speaker configuration module, and playback module. During the operation of the live streaming system, once the client starts the live streaming service, it begins to monitor the transmission port and receive data. There is no signaling control, and the connectionless UDP protocol makes the live streaming system not required to establish a handshake connection, making it more flexible and convenient to use.

[0104] After the live streaming service is started on the server side, the associated audio channel configuration initialization calculation is performed first. Then, the acquired audio data is frame-encoded and then enters the sending control module, which includes establishing an RTP session, setting the port number, and defining and calculating the header of each audio frame. The data packets are then sent to the lower-layer IP network for transmission. On the client side, the receiving module traverses all data sources, unpacks the received data packets, stores them in the dejitter buffer for calculation, and finally decodes them and outputs them to the playback device through the speaker configuration module to achieve live streaming of multi-channel digital audio.

[0105] IV. Thread Processing in Digital Audio Systems

[0106] The live streaming system framework is divided into two subsystems: server and client. Each subsystem needs to calculate and manage multiple task threads of different functional modules simultaneously in order to achieve real-time audio live streaming. Multithreading is used to improve system processing efficiency and system resource utilization. The analysis is divided into two independent parts: server and client.

[0107] (I) Real-time processing of server threads

[0108] In the live streaming system, the server reads in multi-channel digital audio data streams, simultaneously encodes, compresses, and sends them. The server then... Figure 2 The thread design flowchart is used to realize the coordinated operation of various functional modules of the server.

[0109] The server's main thread handles system configuration and initialization based on MFC. This thread completes system configuration tasks such as setting the destination IP address and the RTP session base port. The server thread processing flow is as follows: Figure 2 As shown.

[0110] Once the initialization is complete, the live stream begins. The server system creates an encoding and sending thread. Under this thread, the server system completes the tasks of reading data, encoding and compressing, establishing RTP sessions, and sending data to the destination address. After the server system's main thread performs the calculation to shut down the service, the system stops encoding and sending audio data and releases resources for the data encoding and sending thread, thus enabling control over the start and stop of the encoding and sending thread.

[0111] (II) Real-time processing of client threads

[0112] Client thread handling design flow as follows Figure 3As shown. The client system completes the audio driver settings and speaker configuration initialization in the main thread. After the live streaming service is started, the client system starts a data receiving thread, creates an RTP receiving session, and monitors the session transmission port to receive and buffer data packets in sequence. When the number of data packets in the receiving buffer exceeds the set value, the client starts a decoding buffer thread to retrieve data from the receiving buffer and perform decoding calculations.

[0113] After the main thread begins audio playback calculations, the client system opens a new playback thread. This playback thread writes the decoded data into the playback buffer and sends it to the designated speaker for playback. After the main thread performs the stop playback calculation, the client system closes the playback thread. Finally, when the live streaming service is closed in the main thread, the system releases the data receiving and decoding thread, ending the live streaming service.

[0114] V. 3D Audio Processing Module

[0115] The audio processing module includes a series of data processing modules, such as front-end multi-channel digital audio data acquisition, client audio playback, encoding before transmission, and decoding before playback. The acquisition module uses file stream reading for simulation, the playback module uses an ASIO-based multi-channel digital audio interface, and the encoding / decoding module uses a multi-channel digital audio codec. Each sub-module is relatively independent and has high portability.

[0116] (I) Audio Acquisition and Playback Configuration Module

[0117] (1) 3D audio data acquisition module

[0118] The acquisition module uses WAV file streams for simulation. During system configuration, a number of WAV files matching the current channel configuration are prepared and placed in the project's default path. When the live streaming service starts, these WAV audio files are read in, and then a multi-channel digital audio data encoding and sending thread is created, thereby enabling the simultaneous parallel processing of audio data from each channel.

[0119] Streaming audio data: The file pointer _audio_file is set to point to the beginning of the audio data block, and then the data is retrieved frame by frame to achieve streaming input of audio data. `channel_id` represents the channel number, `frame_seq` represents the original audio frame number, `frame_num` represents the number of audio frames, and `payload` represents the storage area for the audio frames.

[0120] Offset the file pointer _audio_file by 40 bytes to point to the starting address of the audio data bytes in the data sub-block, and then read down 4 bytes of data to obtain the size of the data. Here, AUDIO_FRAME_SIZE_CODEC represents the length of one frame of audio data, defined as 1024. Then, based on the sampling depth of the audio file used, the number of audio frames in the entire WAV file is further calculated.

[0121] After calculating the number of audio frames, the audio block is divided into frames. This module periodically retrieves data from the audio block. When the audio data sampling rate is 48000Hz, the playback time of each audio frame is calculated to be 21.333 milliseconds (1 / 48000*1024). The timestamp parameter _next_read_timestamp is set according to this time interval. Each time 21.333 milliseconds have elapsed, a frame of data is read from the file and stored in the payload for the next encoding process. At the same time, the channel number and frame number information of the frame are recorded for subsequent data packet encoding and transmission.

[0122] The method of extracting audio data blocks from WAV files frame by frame is used to simulate the real-time audio stream acquisition process, thereby realizing the data stream output of the data acquisition module.

[0123] (2) Multi-channel playback module

[0124] The playback module is based on an ASIO driver-based double-buffering mechanism, supporting channel control and playback for up to 24 audio channels. When large amounts of audio data need to be played in real-time, it avoids interruptions due to buffers filling up. The double buffering consists of BufferA and BufferB. BufferA receives and fills the decoded audio data from the upper layer, while BufferB sends its own audio data to the lower-level kernel driver to enable sound output from the speakers. These two buffers operate simultaneously. When BufferA is full and BufferB has been sent to the kernel, their functions switch, thus achieving streaming playback with data filling and output. This process can be represented by the processing flow of three consecutive frames of audio data from a single channel within the double buffer, such as... Figure 4 As shown.

[0125] When audio playback begins, the underlying ASIO driver instructs the upper-layer software to fill the buffer with audio data. At this time, the system automatically calls the callback function bufferSwitchTimeInfo() to fill the playback buffer with audio data. Whenever one side of the data in the dual buffer corresponding to a channel has finished playing, the system plays the other side of the data, and the callback function fills the buffer that has finished playing with data. The parameters of bufferSwitchTimeInfo() also include speaker configuration information, buffer length, and channel sampling type. This application modifies these parameters to enable audio playback for different speaker configurations and different sampling formats.

[0126] In 3D audio, the audio data played by each speaker corresponds to the playback buffer of the corresponding channel. 24 double buffers are created, and 24 pointers pszCh[0], pszCh[1], pszCh[2], ..., pszCh

[23] are defined respectively, pointing to buffer 0, buffer 1, buffer 2, ..., buffer 23 respectively. Finally, as long as the channel configuration of the live audio is filled into the corresponding buffer pointer according to the channel configuration, the playback of different combinations of speakers can be realized. The buffer corresponding to each speaker is set to make the specified speaker emit sound and make the playback software compatible with various channel configurations.

[0127] (3) 3D speaker configuration module

[0128] Before audio playback, the speaker allocation for the multi-channel digital audio signal and the playback buffer for each speaker are specified. The positions of the speakers in the 3D audio are based on the NHK22.2 multi-channel system, divided into three layers: top, middle, and bottom. There are 9 speakers in the top layer, 10 speakers in the middle layer, and 3 speakers in the bottom layer, plus two subwoofers, for a total of 24 speakers. Figure 5 As shown.

[0129] Each of the 24 speakers is assigned a corresponding dual buffer at the system software level. Each channel buffer corresponds one-to-one with each signal channel in the sound card device. The correspondence between each channel buffer and the speaker is consistent with the relationship between each channel of the sound card and the speaker. As long as the channel setting box in the speaker configuration interface is associated with each channel buffer according to this correspondence, the playback of sound from the specified speaker can be achieved.

[0130] First, based on the channel configuration of the live audio, determine which speakers will be used in the multi-channel playback environment. Then, in the speaker settings interface, select the checkboxes for these speakers in sequence. Assuming that six-channel audio data is being played, select the checkboxes p1, p2, p3, p4, p5, and p6 in sequence according to the correspondence between the channel settings and the speakers. The system then iterates through the speaker settings boxes in the configuration interface. If speaker p is selected, the flag choose[p] is set to 1. At the same time, the system reads the channel information of the original audio and proceeds to the next step of speaker configuration determination. If the number of speakers set is inconsistent with the number of channels in the original audio, a prompt box will pop up and... Return to the speaker settings interface to reset; otherwise, proceed to the next step of allocating the playback buffer. After successfully completing the speaker settings, open the corresponding buffers for these speakers. Check the flag bits choose[0] to choose

[23] of speakers 1 to 24 in sequence to see if they are true. If the values ​​of choose[p1], choose[p2], choose[p3], choose[p4], choose[p5] and choose[p6] are all 1, then open the buffer corresponding to the speaker and close the other buffers with choose[n] = 0, thus completing the association from the speaker settings interface to the specified buffer.

[0131] After the speaker configuration process described above, when the live streaming service is started and data from each channel is sent to the playback buffer, the data stream automatically skips the unselected speakers and is stored sequentially in the buffer of the selected speakers, thus enabling playback control of the speakers at the specified locations.

[0132] (II) Audio Encoding and Decoding Module

[0133] (1) 3D audio encoding module

[0134] The audio encoding task is divided into five states: Begin, Read, Encode, Send, and Done. These states represent the encoder initialization at the start of the encoding task, reading the raw audio frame data, sending the encoded data packet to the encoder for encoding, sending the encoded data packet to the sending stage, and the end of the encoding task.

[0135] Before all the original audio frames are retrieved, the encoder periodically repeats the process of reading the audio frame, encoding the frame, and transmitting it to the sending module. When all the original audio frames are retrieved, AudioEncodeTask_Read returns 0, and the status jumps to AudioEncodeTask_Done, ending the encoding task.

[0136] In the AudioEncodeTask_Encode state, the audio encoder interface function is called to encode one frame of audio data. The encoding interface function is defined as follows:

[0137] int Codec::encode(short*frame,unsignedchar*bits)

[0138] The encoding function takes an audio frame and the address of the encoded bitstream as input, and is called as follows:

[0139] _packet_ptr.packet->payload_size=

[0140] _codec.encode(_frame_ptr.frame->payload,

[0141] _packetptr.packet->payload+sizeof(RTPHeader));

[0142] The encoding function _codec.encode() takes the audio frame _frame_ptr.frame as input to the encoder, outputs the encoded bitstream and stores it in _packet_ptr.packet->payload+sizeof(RTPHeader). After the audio frame is encoded, the value of the audio frame sequence number is assigned to the corresponding encoded bitstream data packet.

[0143] (2) Multi-channel decoding module

[0144] The decoding task begins with AudioDecodeTask_Begin, and the decoder is initialized. In the AudioDecodeTask_Decode state, the decoder interface function is called to retrieve a frame of data from the RTP receive buffer and then send it to the decoder for decoding. The corresponding audio frame after decoding is stored in the decoding buffer _decode_buf, ready to be sent to the playback module for processing and playback. The decoding interface function _codec.decode() is defined as follows:

[0145] unsignedlongCodec::decode(unsignedchar*bits,

[0146] unsignedlongbits_len,void*frame)

[0147] The decoding function is called as follows:

[0148] _decode_buflen=

[0149] _codec.decode(audio_packet.packet->payload,

[0150] audiopacket.packet->payload_size,_decode_buf);

[0151] audio_packet.packet->payload stores the data packets before decoding, _decode_buf is the storage address of the decoded audio data frames, and audio_packet.packet->payload_size represents the length of the data packets extracted from decoding one frame.

[0152] The decoding module's states are: Fetch (the decoder retrieves the corresponding data packets from the RTP receive buffer) and Add (the decoded audio frames are added to the decoder's _decode_buf to be processed). The completion of the decoding task is not related to whether there are data packets available in the receive buffer. The end sign is when the _to_be_ended flag is true, at which point the system determines that the decoding task is complete and ends the audio data decoding.

[0153] VI. Data Transmission Module

[0154] The data transmission module in the audio live streaming system uses a communication protocol combining RTP and UDP. It handles the RTP packetization and transmission of compressed audio data, as well as the reception, unpacking, buffering, and synchronization calculations before client-side decoding and playback. The transmission module is developed based on the JRTPLIB open-source library, and its compilation, installation, and configuration are completed on the VS2019 platform. The steps are as follows:

[0155] Step 1: Download and decompress the jrtplib and jthread files;

[0156] Step 2: Compile jthread to generate jthread.lib and jthread_d.lib; open cmake, add input and output paths, and complete the configure configuration; click generate to generate the VS2019 project file; open the project file and compile, set Solutionjthread in SolutionExplorer, run RebuildSolution, if there are no compilation errors, then set INSTALL project, run Build, and generate lib and cmake files under lib;

[0157] Step 3: Compile jrtplib to generate jrtplib.lib and jrtplib_d.lib; add the path and configure; click generate to generate the VS2019 project file; open the project file and compile, generating jrtplib_d.lib and jrtplib.lib under debug and release respectively; if the compilation is successful, jrtplib_d.lib, jrtplib.lib and cmake files will be generated under lib.

[0158] Step 4: Copy all the header files in the compiled jrtplib3 and jthread directories to a newly created folder named jrtplib in the project directory. This will make it easier to call the header files in the same directory and complete the installation and configuration of the library.

[0159] After compiling and installing the JRTPLIB library, design an audio data transmission and reception system. Figure 6 This is a flowchart illustrating the design process of the data transmission module.

[0160] Before starting the service, the data transmission module first initializes and configures the audio source and the number of channels, and then enters the data sending or receiving process.

[0161] In the server sending module, the audio data source is first obtained sequentially, the RTP session is initialized, and the target IP address and port number are specified. Then, the audio is processed, one frame of audio data is read from each audio source and encoded, then packetized in RTP format and then a UDP packet is constructed for sending. This is done using the CRTPSendSession() function. Finally, it is determined whether the audio has been sent. If the sending is complete, the sending thread is closed and the system exits. If there is still data to be sent, the previous steps are skipped to continue sending multi-channel digital audio until the live broadcast service ends and all data has been sent.

[0162] In the client's receiving module, the RTP receiving session establishment process first initializes the system, then adds and creates a receiving process. Within the receiving process, audio is received and stored in the buffer. An RTP receiving session is created to receive multi-channel digital audio using the CRTPRecvSession() function. After receiving the audio data, each audio data packet is unpacked and calculated. The audio data frame number is identified and stored in the receiving buffer. Finally, the system service identifier determines whether the live streaming service has ended. If the system end identifier is true, the sending thread is closed and the system exits; otherwise, audio reception continues until the live streaming service stops.

[0163] (I) IP Network Audio Transmission Control Module

[0164] During the transmission of audio data packets, the RTP session is first initialized by setting session parameters such as payload type and timestamp values, and defining transmission port parameters. Then, an RTP header is added to the audio data packet, the SSRC field is modified, and finally, the client's IP address and base transmission port are specified. The `send_rtp_packet` method is called to send the audio data packet. After the sending task is completed, the session termination calculation is performed. This process can be performed by... Figure 6 express.

[0165] Figure 7 This diagram illustrates the process of sending RTP audio data packets. Initializing the RTP session is the primary prerequisite for transmitting audio data over the network. The packetization and sending stage after acquiring the audio stream is the core of the entire sending process. JRTPLIB is used to implement the initialization of the RTP session and the packetization and sending of audio data.

[0166] (1) RTP session initialization

[0167] First, an instance describing the current session is created using the CRTPSendSession class. Then, the Creat() method of this class is called to complete the initialization calculation. The Creat() method sets two parameters: session parameters sessparams and transport parameters transparams. sessparams describes the parameters used by the CRTPSendSession instance, especially setting the appropriate timestamp unit, which is done by calling the SetTimestampUnit() method of the RTPSession class. The transparams parameter sets the transport layer UDP parameters based on IPv4.

[0168] The audio data timestamp will increment by 1 in each sampling period. The sampling rate of the live audio data is 48kHz. SetOwnTimestampUnit() is used to set the timestamp unit to the reciprocal of the audio sampling rate, 1 / 48000. This parameter needs to be modified accordingly when the sampling rate of the transmitted audio source changes. status indicates the status feedback of this initialization work. When status is greater than 0, the flag _available becomes true, completing the initialization of this RTP session.

[0169] (ii) Real-time transmission of RTP audio packets

[0170] The sending module packages the acquired audio stream before sending data in this RTP session. The process is as follows: first, the RTP packet header is padded; then, the audio data is added to the RTP packet payload; finally, the total RTP length is verified before sending. This process can be described by... Figure 8 express.

[0171] During the process of filling in the standard fields of the RTP header, the following parameters need to be calculated and assigned: Version, Padding, Extension, PayloadType, Length, Marker, SequenceNumber, Timestamp, and SSRC. The payload is a data stream in a multi-channel digital audio encoding format, and its payload type is 96 as defined in RFC3551.

[0172] Then, the `add_rtp_header()` function is used to add RTP packet header data. Here, `length` represents the length of the payload audio data. If the payload length is less than the fixed length of the RTP header, no RTP header is added to the payload data. The `extension` bit is set to zero, indicating that no extension field is needed, thus improving RTP transmission efficiency. When calling this function, the actual arguments for the last three parameters—`sequence`, `timestamp`, and `ssrc`—are `(unsigned short)(Oxfff&_task_info.channel_id)`, `_packet_ptr.packet->frame_num`, and `_packet_ptr.packet->frame_seq`, respectively. These parameters are used to identify the frame number and frame sequence of the audio data in the RTP packet, which is then used for subsequent identification, classification, and synchronization at the receiving end.

[0173] Before sending the packetized RTP data using send_rtp_packet, the packet length is validated to determine the acceptable range of values ​​for the packet data.

[0174] The maximum length of audio data in an RTP packet is 1500 bytes minus the size of the headers added in all network structure layers. This can be calculated as follows: the RTP header length is 12 bytes, the UDP header length is 8 bytes, and the IP packet header length without appendices is 20 bytes. Therefore, the maximum length of audio data in an RTP packet is 1500 bytes minus (12 bytes + 8 bytes + 20 bytes) = 1460 bytes.

[0175] Before sending an RTP packet, it must be verified that the payload_size does not exceed 1460 bytes and the rtp_packet_length does not exceed 1500 bytes.

[0176] send_rtp_packet(_task_info.decode_server_addr,

[0177] _packet_ptr.packet->payload,

[0178] _packetptr.packet->payload_size);

[0179] Finally, the send_rtp_packet() function is called to specify the address of the receiving end and send the RTP packet data to the destination address, completing the task of packaging and sending the encoded audio stream.

[0180] (II) IP Live Audio Reception and Processing Module

[0181] The corresponding RTP receiving process is instantiated using JRTBLIB to achieve data reception and buffering. The data receiving module calls the Poll() method in the RTPSession class to receive audio data packets listened to on the transmission port. The receiving mode can be set to RECEIVEMODE_ALL, the default mode. By detecting all received data packets in the transmission port, RTP payload data and RTP packet identifier association information are extracted from them.

[0182] When the live streaming service starts, the receiving end creates an RTP receiving session and a receiving thread. Before implementing the receiving function, it performs initialization calculations on the RTP receiving session and creates a CRTPRecvSession class to implement the receiving session. Its initialization calculations are consistent with those of the sending end.

[0183] The process of receiving RTP data packets and the handling of RTP receive sessions:

[0184] First, the module calls the start() method in CRTPRecvSession() to start the RTP packet reception service on the receiving end. The start() method first reads the local configuration information and initializes it, then creates an RTP reception task thread and creates an RTP reception session _rtp_recv_session() under the thread. When the RTP reception session is initialized, the data reception monitoring port is set to correspond to the sending port, and the base port _rtp_recv_base_port is set to 20000.

[0185] Once the RTP receive session is created and initialized, the callback function on_recv_rtp_packet() will automatically retrieve the RTP packets arriving at the corresponding receive port, unpack and calculate them, and identify the sequence, timestamp, ssrc, and payload_type identifiers in the RTP packets. Next, the cache pointer ac points to the starting address of the audio data within the RTP packet, and finally, addaudio_packet() is called to perform data caching. Data of length - sizeof(RTPHeader) is read from the cache pointer ac, and the data is sorted according to the channel number sequence and frame number ssrc to achieve ordered caching of the audio payload data.

[0186] Audio payload data is cached using a doubly linked list data structure to store audio data. The sequential audio stream is cached by inserting elements into the list, especially inserting elements at the beginning and end, and changing the pointers of the elements.

[0187] When using add_audio_packet for data caching calculation, a maximum buffer size is set for each channel buffer, which is the maximum length of the audio data packet buffer queue, _max_audio_packet_num. When the cached data packets exceed this length, the data at the head of the queue will be cleared and discarded to ensure that subsequent data can be received smoothly. The audio data packet queue buffer eliminates some of the jitter problems caused by network transmission.

[0188] The audio data transmission in this application is carried on the basis of IP network. The maximum buffer length _max_audio_packet_num is set to 500. This buffer can cache a maximum of 500 data packets at the same time to eliminate jitter caused by large network fluctuations.

Claims

1. A 3D multi-channel digital audio real-time IP network transmission live streaming system, characterized in that, A real-time multi-channel digital audio transmission and playback system based on an IP network is realized through a real-time audio data transmission network protocol. One approach is to solve the real-time audio transmission problem by analyzing the live streaming model of the IP network structure and constructing a 3D audio data transmission channel based on the RTP / UDP high-efficiency transmission protocol. Secondly, the system architecture design adopts a multi-channel modular design, with independent IP network analysis, design, and implementation for different functional modules such as audio encoding / decoding, data transmission, and buffered playback. Multi-threaded live streaming is used to allocate and control the start / stop of each task thread, improving the flexibility of audio live streaming. Thirdly, each functional module uses standardized data interfaces, facilitating system function modification and porting. Fourthly, based on the different channel configuration characteristics of multi-channel digital audio, a speaker configuration module is designed and implemented. Users can set speakers in specified positions according to any channel configuration, and up to 24 channels can be simultaneously transmitted and played, further improving the flexibility of the live streaming system. First, the RTP real-time transmission protocol is configured to enable multi-channel digital audio data transmission over an IP network, establishing a transmission channel for multi-channel digital audio data between different IP addresses. Then, multi-channel modular programming is used to implement the acquisition analog module, encoding / decoding module, speaker configuration module, and ASIO-based audio driver-based playback control module, with standardized data interfaces set for each module. Next, MFC is used to implement the visual interaction of the live streaming system, allowing channel number settings, client IP address settings, and speaker configuration calculations to be completed within the interactive interface. Finally, multi-threading allocation and control for each module are configured, adopting a server / client-based system architecture. Multi-threading is used to allocate and control the start and stop of each task thread, designing corresponding trigger conditions and termination flags for different tasks on the server and client sides to avoid conflicts between different task threads.

2. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, Multi-channel audio system overall architecture: Multi-channel digital audio live streaming is a unidirectional data stream transmission. It adopts a C / S architecture to complete the data acquisition and compression encoding tasks with large computational loads on the server side, while the decoding and playback tasks are completed on the client side. This makes full use of the hardware at both ends and distributes the tasks to the client and server sides, reducing the system's resource consumption. The overall framework of the multi-channel digital audio live streaming system and the deployment of server and client functional modules are as follows: The upper and lower layers are divided into server and client respectively. On the server side, there are audio source acquisition module, encoding module, and sending control module; on the client side, there are receiving buffer parsing module, decoding module, speaker configuration module, and playback module. During the operation of the live streaming system, once the client starts the live streaming service, it begins to monitor the transmission port and receive data. There is no signaling control, and the connectionless UDP protocol eliminates the need for a handshake connection establishment process, making it more flexible and convenient to use. After the live streaming service is started on the server, the associated audio channel configuration initialization calculation is performed first. Then, the acquired audio data is frame-encoded and then enters the sending control module, including establishing an RTP session, setting the port number, and defining and calculating the packet header for each audio frame. The data packets are then sent to the lower-layer IP network for transmission. In the client receiving module, the receiving and processing module traverses all data sources, unpacks the received data packets, stores them in the dejitter buffer for calculation, and finally decodes them and outputs them to the playback device through the speaker configuration module to realize live playback of multi-channel digital audio.

3. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, The server thread handles real-time processing: reading in multi-channel digital audio data streams and simultaneously encoding, compressing, and sending them; the server's main thread performs system configuration and initialization work based on MFC calculations, and completes system configuration such as setting the destination IP address and setting the RTP session base transmission port. Once the initialization is complete, the live stream begins. The server system creates an encoding and sending thread. Under this thread, the server system completes the tasks of reading data, encoding and compressing, establishing RTP sessions, and sending data to the destination address. After the server system's main thread performs the calculation to shut down the service, the system stops encoding and sending audio data and releases resources for the data encoding and sending thread, thus enabling control over the start and stop of the encoding and sending thread.

4. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, Real-time processing on the client thread: The client system completes the audio driver settings and speaker configuration initialization in the main thread. When the live streaming service is started, the client system starts the data receiving thread, creates an RTP receiving session, and monitors the session transmission port to complete the receiving and buffering of data packets in sequence. When the number of data packets in the receiving buffer exceeds the set value, the client starts the decoding buffer thread, retrieves the data from the receiving buffer, and performs decoding calculations. After the main thread begins audio playback calculations, the client system opens a new playback thread. This playback thread writes the decoded data into the playback buffer and sends it to the designated speaker for playback. After the main thread performs the stop playback calculation, the client system closes the playback thread. Finally, when the live streaming service is closed in the main thread, the system releases the data receiving and decoding thread, ending the live streaming service.

5. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, 3D audio data acquisition module: The acquisition module uses WAV file streams for simulation. During system configuration, a number of WAV files matching the current channel configuration are prepared and placed in the project's default path. When the live streaming service starts, these WAV audio files are read in, and then a multi-channel digital audio data encoding and sending thread is created, thereby enabling the simultaneous parallel processing of audio data from each channel. Audio data streaming processing: The file pointer _audio_file is set to point to the beginning of the audio data block, and then the data is retrieved frame by frame to achieve streaming input of audio data. channel_id represents the channel number, frame_seq represents the original audio frame number, frame_num represents the number of audio frames, and payload represents the storage area of ​​the audio frames. The file pointer _audio_file is offset by 40 bytes to point to the starting address of the number of audio data bytes in the data sub-block, and then 4 bytes of data are read down to obtain the size value of the data. Here, AUDIO_FRAME_SIZE_CODEC represents the length of one frame of audio data, which is defined as 1024. Then, based on the sampling depth of the audio file used, the number of audio frames in the entire WAV file is further calculated. After calculating the number of audio frames, the audio block is divided into frames. This module periodically retrieves data from the audio block. When the audio data sampling rate is 48000Hz, the playback time of each audio frame is calculated to be 21.333 milliseconds (1 / 48000*1024). The timestamp parameter _next_read_timestamp is set according to this time interval. Each time 21.333 milliseconds have elapsed, a frame of data is read from the file and stored in the payload for the next encoding process. At the same time, the channel number and frame number information of the frame are recorded for subsequent data packet encoding and transmission. The method of extracting audio data blocks from WAV files frame by frame is used to simulate the real-time audio stream acquisition process, thereby realizing the data stream output of the data acquisition module.

6. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, Multi-channel playback module: Based on the ASIO driver playback buffer double buffer mechanism, it can control and play up to 24 channels of audio. When a large amount of audio data needs to be played in real time, the double buffer includes BufferA and BufferB. BufferA is responsible for receiving the audio data decoded by the upper layer and filling it, while BufferB is responsible for sending the audio data in its buffer to the lower layer kernel driver to make the speaker emit sound. The two buffers work simultaneously. When BufferA is full and BufferB has been sent to the kernel, the functions of the two buffers are interchanged, thereby realizing the streaming playback of data filling and data output. This process is represented by the processing flow of three consecutive frames of audio data of a single channel in the double buffer. When audio playback begins, the underlying ASIO driver instructs the upper-layer software to fill the buffer with audio data. At this time, the system automatically calls the callback function bufferSwitchTimeInfo() to fill the playback buffer with audio data. Whenever one side of the data in the dual buffer corresponding to a channel has finished playing, the system plays the other side of the data, and the callback function fills the buffer that has finished playing with data. The parameters of bufferSwitchTimeInfo() also include speaker configuration information, buffer length, and channel sampling type. This application changes these parameters to enable audio playback for different speaker configurations and different sampling formats. In 3D audio, the audio data played by each speaker corresponds to the playback buffer of the corresponding channel. 24 double buffers are created, and 24 pointers pszCh[0], pszCh[1], pszCh[2], ..., pszCh[23] are defined respectively, pointing to buffer 0, buffer 1, buffer 2, ..., buffer 23 respectively. Finally, as long as the channel configuration of the live audio is filled into the corresponding buffer pointer according to the channel configuration, the playback of different combinations of speakers can be realized. The buffer corresponding to each speaker is set to make the specified speaker emit sound and make the playback software compatible with various channel configurations.

7. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, 3D Speaker Configuration Module: Before audio playback, the speaker allocation for the multi-channel digital audio signal and the playback buffer for each speaker are specified. The position of each speaker in the 3D audio is based on the NHK22.2 multi-channel system and is divided into three layers: upper, middle and lower. There are 9 speakers in the upper layer, 10 speakers in the middle layer and 3 speakers in the lower layer, plus two subwoofers, for a total of 24 speakers. Each of the 24 speakers is assigned a corresponding dual buffer at the system software level. Each channel buffer corresponds one-to-one with each signal channel in the sound card device. The correspondence between each channel buffer and the speaker is consistent with the relationship between each channel of the sound card and the speaker. As long as the channel setting box in the speaker configuration interface is associated with each channel buffer according to this correspondence, the playback and sound output of the specified speaker can be achieved. First, based on the channel configuration of the live audio, determine which speakers will be used in the multi-channel playback environment. Then, in the speaker settings interface, select the checkboxes for these speakers in sequence. Assuming that six-channel audio data is being played, select the checkboxes p1, p2, p3, p4, p5, and p6 in sequence according to the correspondence between the channel settings box and the speakers. The system then iterates through the speaker settings box in the configuration interface. If speaker p is selected, set the flag choose[p] to 1. At the same time, the system reads the channel information of the original audio and proceeds to the next step of speaker configuration determination. If the number of speakers set is inconsistent with the number of channels of the original audio, a prompt box will pop up and return to the speaker settings interface to reset. Otherwise, proceed to the next step of playback buffer allocation. After successfully completing the speaker settings, open the corresponding buffers for these speakers, and check whether the flag bits choose[0] to choose[23] of speakers 1 to 24 are true. If the values ​​of choose[p1], choose[p2], choose[p3], choose[p4], choose[p5] and choose[p6] are all 1, then open the buffer corresponding to the speaker and close the other buffers with choose[n] = 0, thus completing the association from the speaker settings interface to the specified buffer. After the speaker configuration process described above, when the live streaming service is started and data from each channel is sent to the playback buffer, the data stream automatically skips the unselected speakers and is stored sequentially in the buffer of the selected speakers, thus enabling playback control of the speakers at the specified locations.

8. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, Audio encoding and decoding module: (1) 3D audio encoding module The audio encoding task is divided into five states: Begin, Read, Encode, Send, and Done. These states represent the encoder initialization at the start of the encoding task, reading the raw audio frame data, sending the encoded data packet to the encoder for encoding, and the end of the encoding task. Before all the original audio frames are retrieved, the encoder periodically reads the audio frames, encodes the frames, and transmits them to the sending module. After all the original audio frames are retrieved, AudioEncodeTask_Read returns 0, and the status jumps to AudioEncodeTask_Done, ending the encoding task. In the AudioEncodeTask_Encode state, the audio encoder interface function is called to encode one frame of audio data. The encoding interface function is defined as follows: int Codec::encode(short*frame,unsignedchar*bits) The encoding function takes an audio frame and the address of the encoded bitstream as input, and is called as follows: _packet_ptr.packet->payload_size= _codec.encode(_frame_ptr.frame->payload, _packetptr.packet->payload+sizeof(RTPHeader)); The encoding function _codec.encode() takes the audio frame _frame_ptr.frame as input to the encoder, outputs the encoded bitstream and stores it in _packet_ptr.packet->payload+sizeof(RTPHeader). After the audio frame is encoded, the value of the audio frame sequence number is assigned to the corresponding encoded bitstream data packet. (2) Multi-channel decoding module The decoding task begins with AudioDecodeTask_Begin, and the decoder is initialized. In the AudioDecodeTask_Decode state, the decoder interface function is called to retrieve a frame of data from the RTP receive buffer and then send it to the decoder for decoding. The corresponding audio frame after decoding is stored in the decoding buffer _decode_buf, ready to be sent to the playback module for processing and playback. The decoding interface function _codec.decode() is defined as follows: unsignedlongCodec::decode(unsignedchar*bits, unsignedlongbits_len,void*frame) The decoding function is called as follows: _decode_buflen= _codec.decode(audio_packet.packet->payload, audiopacket.packet->payload_size,_decode_buf); audio_packet.packet->payload stores the data packets before decoding, _decode_buf is the storage address of the decoded audio data frames, and audio_packet.packet->payload_size represents the length of the data packets extracted from decoding one frame; The decoding module's states are: Fetch (the decoder retrieves the corresponding data packets from the RTP receive buffer) and Add (the decoded audio frames are added to the decoder's _decode_buf to be processed). The completion of the decoding task is not related to whether there are data packets available in the receive buffer. The end sign is when the _to_be_ended flag is true, at which point the system determines that the decoding task is complete and ends the audio data decoding.

9. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, IP network audio transmission control module: During the transmission of audio data packets, the RTP session is first initialized and calculated, and session parameters such as payload type and timestamp parameter values ​​are set. The transmission port transmission parameters are defined. Then, the audio data packets are given an RTP header and the SSRC field is modified. Finally, the client's IP address and the base transmission port are specified, and the send_rtp_packet method is called to send the audio data packets. After the sending task is completed, the session termination calculation is performed. (1) RTP session initialization First, an instance describing the current session is created using the CRTPSendSession class. Then, the Creat() method of this class is called to complete the initialization calculation. The Creat() method sets two parameters: session parameters sessparams and transport parameters transparams. sessparams describes the parameters used by the CRTPSendSession instance, especially setting the appropriate timestamp unit by calling the SetTimestampUnit() method of the RTPSession class. The transparams parameter sets the transport layer UDP parameters based on IPv4. The audio data timestamp will increment by 1 in each sampling period. The sampling rate of the live audio data is 48kHz. SetOwnTimestampUnit() is used to set the timestamp unit to the reciprocal of the audio sampling rate, 1 / 48000. This parameter needs to be modified accordingly when the sampling rate of the transmitted audio source changes. status indicates the status feedback of this initialization work. When status is greater than 0, the flag _available value becomes true, completing the initialization of this RTP session. (2) Real-time transmission of RTP audio packets The sending module packages the acquired audio stream before sending data in this RTP session. The process is as follows: first, the RTP packet header is filled, then the audio data is added to the RTP packet payload, and finally the total RTP length is verified before sending. During the process of filling in the standard fields of the RTP header, the Version, Padding, Extension, PayloadType, Length, Marker, SequenceNumber, Timestamp, and SSRC identifier need to be assigned and calculated. The payload is a data stream in a multi-channel digital audio encoding format, and its payload type is 96 as defined in RFC3551. Then, the `add_rtp_header()` function is used to add RTP packet header data. Here, `length` represents the length of the payload audio data. If the payload length is less than the fixed length of the RTP header, no RTP header is added to the payload data. The `extension` bit is set to zero, indicating that no extension field is needed, thus improving RTP transmission efficiency. When calling this function, the actual arguments for the last three parameters—`sequence`, `timestamp`, and `ssrc`—are `(unsigned short)(Oxfff&_task_info.channel_id)`, `_packet_ptr.packet->frame_num`, and `_packet_ptr.packet->frame_seq`, respectively. These parameters are used to identify the frame number and frame sequence of the audio data in the RTP packet, which is then used for subsequent identification, classification, and synchronization at the receiving end. Before sending the packetized RTP packet by calling send_rtp_packet, the length of the packet is checked to determine the range of possible values ​​for the packet data. Before sending an RTP packet, it must be verified that the payload_size does not exceed 1460 bytes and the rtp_packet_length does not exceed 1500 bytes. send_rtp_packet(_task_info.decode_server_addr, _packet_ptr.packet->payload, _packetptr.packet->payload_size); Finally, the send_rtp_packet() function is called to specify the address of the receiving end and send the RTP packet data to the destination address, completing the task of packaging and sending the encoded audio stream.

10. The 3D multi-channel digital audio real-time IP network transmission live streaming system according to claim 1, characterized in that, IP live audio receiving and processing module: The corresponding RTP receiving process is instantiated through JRTBLIB to realize data reception and buffering. The data receiving module calls the Poll() method in the RTPSession class to receive audio data packets listened to on the transmission port. The receiving mode can be set to RECEIVEMODE_ALL, the default mode. By detecting all received data packets in the transmission port, RTP payload data and RTP packet identifier association information are extracted from them. When the live streaming service starts, the receiving end creates an RTP receiving session and a receiving thread. Before implementing the receiving function, it performs initialization calculations on the RTP receiving session and creates a CRTPRecvSession class to implement the receiving session. Its initialization calculations are consistent with those of the sending end. The process of receiving RTP data packets and the handling of RTP receive sessions: First, the module calls the start() method in CRTPRecvSession() to start the RTP packet reception service on the receiving end. The start() method first reads the local configuration information and initializes it, then creates an RTP reception task thread and creates an RTP reception session _rtp_recv_session() under the thread. When the RTP reception session is initialized, the data reception monitoring port is set to correspond to the sending port, and the base port _rtp_recv_base_port is set to 20000. Once the RTP receive session is created and initialized, the callback function on_recv_rtp_packet() will automatically retrieve the RTP packets arriving at the corresponding receive port, perform unpacking calculations, and identify the sequence, timestamp, ssrc, and payload_type identifiers in the RTP packets. Next, it uses the buffer pointer ac to point to the starting address of the audio data within the RTP packet, and finally calls addaudio_packet() to perform data buffering. It reads data of length - sizeof(RTPHeader) from the buffer pointer ac and sorts the data according to the channel number sequence and frame number ssrc to achieve ordered buffering of the audio payload data. Audio payload data is cached using a doubly linked list data structure to store audio data. The sequential audio stream is cached by inserting elements into the list, especially inserting elements at the beginning and end, and changing the pointers of the elements. When using add_audio_packet to perform data caching calculations, a maximum buffer size is set for each channel buffer, which is the maximum length of the audio data packet buffer queue, _max_audio_packet_num. When the cached data packets exceed this length, the data at the head of the queue will be cleared and discarded to ensure that subsequent data can be received smoothly. The audio data packet queue buffer eliminates some of the jitter problems caused by network transmission. The audio data transmission in this application is carried on the basis of IP network. The maximum buffer length _max_audio_packet_num is set to 500. This buffer can cache a maximum of 500 data packets at the same time, eliminating jitter caused by large network fluctuations.