Converged communication method and device based on SIP (Session Initiation Protocol) and half-duplex TCP (Transmission Control Protocol)

By adopting a converged communication method based on SIP protocol and half-duplex TCP protocol in the command and dispatch platform, the problem of difficulty in realizing seamless communication between multiple types of clients is solved in the prior art, and efficient, secure and reliable communication and audio data transmission are achieved.

CN119996383APending Publication Date: 2025-05-13SHANGHAI SHUGUO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411992251.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the face of complex application scenarios, it is difficult to achieve seamless communication between multiple types of clients, especially when efficient, secure and reliable information transmission is required.

Method used

The fusion communication method based on SIP protocol and half-duplex TCP protocol is adopted. By establishing a session connection between the SIP client and the TCP client and the server, access rights of the sound sensor are obtained, audio data is collected and preprocessed, and audio data is encoded into a voice packet of a specified format for transmission.

Benefits of technology

It realizes seamless communication between clients under different communication protocols, improves the communication efficiency and flexibility of the command and dispatch platform, enhances the security and privacy protection of communication, and ensures the efficient transmission and playback quality of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996383A_ABST
    Figure CN119996383A_ABST
Patent Text Reader

Abstract

The invention discloses a converged communication method and device based on an SIP protocol and a half-duplex TCP protocol, and relates to the technical field of communication. The method comprises the following steps: respectively establishing session connections between an SIP client and a server and between a TCP client and the server; acquiring original audio data through a sound sensor of the first client, and preprocessing the original audio data to obtain audio data; encoding the audio data into an audio packet in a specified format and sending the audio packet to a server, processing the audio packet through the server and then sending the processed audio packet to a second client, the first client and the second client being one of an SIP client and a TCP client, and the second client being different from the first client; and carrying out legality detection on the processed audio packet through the second client, and playing the audio packet through an audio player of the second client after the detection is passed. By implementing the technical scheme provided by the invention, efficient, safe and reliable communication between different clients in the command and dispatch platform is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of communication technology, and in particular to a fusion communication method and device based on SIP protocol and half-duplex TCP protocol. Background Art

[0002] With the rapid development of information technology, various communication protocols and technologies continue to emerge, enabling the command and dispatch platform to achieve more efficient and reliable information transmission between multiple points.

[0003] Existing technical solutions are mainly based on a single protocol communication method, which realizes data transmission between the client and the server, that is, directly establishes a connection between the client and the server, and performs real-time audio and video communication through a specific protocol. Although the single protocol method is simple and easy to implement, it is inadequate when facing complex application scenarios, especially when it is necessary to support multiple types of clients at the same time.

[0004] Therefore, a new converged communication method is urgently needed to overcome these problems in the prior art. Summary of the invention

[0005] The present application provides a fusion communication method and device based on the SIP protocol and the half-duplex TCP protocol, which realizes efficient, secure and reliable communication between different clients in a command and dispatch platform.

[0006] In a first aspect of the present application, a fusion communication method based on the SIP protocol and the half-duplex TCP protocol is provided, which is applied to a command and dispatch platform, wherein the command and dispatch platform includes a server, a SIP client and a TCP client, and the method includes: Establishing session connections between the SIP client, the TCP client and the server respectively, and obtaining access rights to the sound sensors of the SIP client and the TCP client; Collecting raw audio data through the sound sensor of the first client, and preprocessing the raw audio data to obtain audio data, the first client is one of the SIP client and the TCP client, and the preprocessing includes format conversion and merging compression; Encoding the audio data into a voice packet in a specified format and sending the voice packet to the server, and processing the voice packet through the server and sending the voice packet to a second client, where the second client is one of the SIP client and the TCP client, and the second client is different from the first client; The processed audio package is subjected to a legality check through the second client, and the audio package is played through the audio player of the second client after the check passes.

[0007] 2 Optionally, the preprocessing of the original audio data to obtain audio data includes: Convert the original audio data into first audio data in ALAW format, and convert the first audio data into second audio data in PCM format; The second audio data is combined and compressed to be converted into audio data with a preset sampling rate and sampling bit number.

[0008] 3 Optionally, encoding the audio data into a voice packet in a specified format and sending the voice packet to the server includes: The audio data is converted into an audio packet in Opus format by Opus encoding, the audio packet in Opus format is encapsulated into an RTP voice packet by using the RTP protocol, and the RTP voice packet is sent to the server at a fixed time interval.

[0009] 4 Optionally, the step of encapsulating the audio packet in Opus format into an RTP voice packet using the RTP protocol includes: Add a data length field having a first number of bytes in length and a data type field having a second number of bytes in length before an audio packet in Opus format to form an RTP data packet header having a third number of bytes, where the third number is the sum of the first number and the second number; Set the maximum allowable size of each RTP voice packet based on network bandwidth, latency requirements, and terminal device processing capabilities; The audio packet in Opus format with the RTP data packet header is packetized according to the maximum allowed size to obtain multiple RTP voice packets, and each RTP voice packet is marked with a sequence number.

[0010] 5 Optionally, the processing of the audio packet by the server and then sending the processed audio packet to the second client comprises: Remove the packet header and RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and convert the voice packet in Opus format into a PCM voice packet with the preset sampling rate and sampling number; The PCM voice packets are converted into voice packets in ALAW format according to a preset offset, the voice packets in ALAW format are sorted through data splicing, and sent to the second client.

[0011] 6 Optionally, removing the packet header and the RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and converting the voice packet in Opus format to a PCM voice packet with the preset sampling rate and sampling number includes: Deleting the RTP data packet header of the third number of bytes and the RTP header of the fourth number of bytes from the received RTP voice packet, and restoring the audio data segment in Opus format; Using an Opus decoding algorithm, the Opus format audio data segment is decoded to convert it into uncompressed PCM format audio data; According to the preset sampling rate and sampling bit number, the PCM audio data is resampled and the bit number is converted to obtain a PCM voice packet.

[0012] 7 Optionally, converting the PCM voice packet into a voice packet in ALAW format according to a preset offset includes: Mapping each audio sample value in the PCM voice packet into a coded value in ALAW format according to a preset conversion rule; The converted ALAW code values ​​are packaged to form continuous ALAW format voice packets.

[0013] In a second aspect of the present application, a fusion communication system based on the SIP protocol and the half-duplex TCP protocol is provided, comprising a connection module, a collection module, an execution module and a playback module, wherein: A connection module, configured to establish a session connection between a SIP client, a TCP client and a server, and obtain access rights to the sound sensors of the SIP client and the TCP client; A collection module, configured to collect raw audio data through the sound sensor of a first client, and pre-process the raw audio data to obtain audio data, wherein the first client is one of the SIP client and the TCP client, and the pre-processing includes format conversion and merging and compression; an execution module configured to encode the audio data into a voice packet in a specified format and send the voice packet to the server, and to process the voice packet through the server and send the packet to a second client, wherein the second client is one of the SIP client and the TCP client, and the second client is different from the first client; The playing module is configured to perform a legality check on the processed audio package through the second client, and play the audio package through the audio player of the second client after the check passes.

[0014] In a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes any one of the methods described above.

[0015] 10 In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any of the methods described above is executed.

[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By combining the SIP protocol and the half-duplex TCP protocol, seamless communication between clients under different communication protocols (SIP clients and TCP clients) is achieved. This improves the communication efficiency and flexibility of the command and dispatch platform, allowing information to be smoothly transmitted in different devices and network environments; 2. By obtaining access rights to the sound sensors of the SIP client and TCP client, it ensures that only authorized devices can collect and transmit audio data, thereby enhancing the security and privacy protection of communications; 3. Performing pre-processing operations such as format conversion and merging and compressing the original audio data not only reduces the bandwidth requirements for data transmission, but also improves the transmission efficiency and playback quality of audio data. This is especially important for command and dispatch platforms that need to transmit and process large amounts of audio data in real time; 4. Encode the audio data into an audio package of a specified format, and transmit and process it through the server, so that the audio data can be shared and played between different clients. This encoding and decoding process ensures the integrity and consistency of the audio data and avoids communication failures caused by incompatible formats; 5. Performing a legitimacy check on the received audio package on the second client can further ensure the security and reliability of communication. After the check passes, the audio package is played through the audio player of the second client, so that the user can hear the audio information from the first client in real time, thereby improving the real-time and accuracy of command and dispatch. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of a fusion communication method based on the SIP protocol and the half-duplex TCP protocol disclosed in an embodiment of the present application; Figure 2 A schematic diagram of the principle of implementing a fusion communication method based on the SIP protocol and the half-duplex TCP protocol on the command and dispatch platform disclosed in the embodiment of the present application; Figure 3 A flow chart of the SIP client disclosed in the embodiment of the present application as the initiator; Figure 4 A schematic diagram of a process in which a TCP client as an initiator is disclosed in an embodiment of the present application; Figure 5A schematic diagram of the process of transmitting and playing voice as a receiver using a TCP client disclosed in an embodiment of the present application; Figure 6 A schematic diagram of the process of transmitting and playing voice as a receiver in the SIP client disclosed in the embodiment of the present application; Figure 7 It is a module schematic diagram of a fusion communication system based on the SIP protocol and the half-duplex TCP protocol disclosed in an embodiment of the present application; Figure 8 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present application.

[0018] Explanation of the reference numerals: 701, connection module; 702, acquisition module; 703, execution module; 704, playback module; 801, processor; 802, communication bus; 803, user interface; 804, network interface; 805, memory. DETAILED DESCRIPTION

[0019] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0020] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.

[0021] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0022] This embodiment discloses a fusion communication method based on the SIP protocol and the half-duplex TCP protocol, which is applied to a command and dispatch platform. The command and dispatch platform includes a server, a SIP client, and a TCP client. Figure 1is a flow chart of a fusion communication method based on the SIP protocol and the half-duplex TCP protocol disclosed in an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: S101, establishing session connections between a SIP client, a TCP client and a server respectively, and obtaining access rights to sound sensors of the SIP client and the TCP client; S102, collecting original audio data through the sound sensor of the first client, and preprocessing the original audio data to obtain audio data, the first client is one of the SIP client and the TCP client, and the preprocessing includes format conversion and merging compression; S103, encoding the audio data into a voice packet in a specified format and sending the voice packet to the server, and processing the voice packet through the server and sending the voice packet to a second client, where the second client is one of the SIP client and the TCP client, and the second client is different from the first client; S104: Perform a legitimacy check on the processed audio package through the second client, and play the audio package through the audio player of the second client after passing the check.

[0023] The SIP client and TCP client need to establish a session connection with the server respectively. For the SIP client, this usually involves the signaling exchange of the SIP protocol, including steps such as registration, invitation, and confirmation, to establish a stable communication session. For the TCP client, a connection is established with the server through the three-way handshake process of the TCP protocol. While or after the session connection is established, the server needs to request access rights to the sound sensor from the SIP client and the TCP client. This usually involves authentication and permission checks to ensure that only authorized devices can access and use the sound sensor. The SIP protocol uses SIP URI (Uniform Resource Identifier) ​​to identify users and services, and establishes and maintains sessions through SIP messages (such as INVITE, ACK, etc.). The TCP protocol uses port numbers and IP addresses to establish connections, and confirms the establishment of connections through a three-way handshake process (SYN-SYN / ACK-ACK). Access rights to the sound sensor may be managed through access control lists (ACLs), OAuth, or other authentication mechanisms. When the first client (SIP client or TCP client) obtains access rights to the sound sensor, it begins to collect raw audio data through the sound sensor. This usually involves the digitization process of analog signals, that is, converting sound signals into digital signals. The collected raw audio data may need to be pre-processed, such as format conversion and merging and compression. Format conversion may involve converting audio data from one encoding format to another format that is more suitable for transmission and processing. Merging and compression may involve merging multiple audio frames into a data packet and compressing it to reduce the data size. Common audio encoding formats include PCM (Pulse Code Modulation), ALAW, etc. Merging and compression may involve splicing audio frames and applying compression algorithms, such as audio compression standards such as G.711 and G.729. The pre-processed audio data needs to be encoded into a voice packet in a specified format. This usually involves the application of audio encoding algorithms, such as Opus, G.723.1, etc. The encoded voice packet has a smaller data size, a higher compression rate, and better sound quality. The encoded voice packet is sent to the server through the connection between the first client and the server. For a SIP client, this may involve the use of RTP (Real-time Transport Protocol) to transmit audio data. For a TCP client, data is sent directly over a TCP connection. The RTP protocol is usually used together with RTCP (Real-time Transport Control Protocol) to provide real-time transmission of audio data, error detection, flow control and other functions. The TCP protocol provides a reliable transmission service, ensuring the integrity and sequence of data through confirmation mechanism, retransmission mechanism and flow control mechanism. After the server receives the audio packet sent by the first client, it may need to perform processing operations such as decoding, format conversion, and resampling to meet the playback requirements of the second client (SIP client or TCP client).If the first client is a SIP client, the second client is a TCP client; if the first client is a TCP client, the second client is a SIP client. The processed audio packet is sent to the second client through the connection between the server and the second client. This may also involve the use of the RTP protocol (for SIP clients) or the TCP protocol (for TCP clients). The server may need to support multiple audio encoding formats and decoding algorithms to meet the needs of different clients. For audio data that needs to be played in real time, the server may need to implement a low-latency transmission and processing mechanism. After the second client receives the audio packet sent by the server, it first performs a legitimacy check. This usually involves checking the integrity, signature or encryption status of the audio packet to ensure the security and authenticity of the audio data. If the legitimacy check passes, the second client will use its audio player to play the received audio packet. This involves operations such as audio decoding, format conversion and playback control. The legitimacy check may involve security mechanisms such as digital signature verification and the application of encryption and decryption algorithms. The audio player may need to be configured and adjusted according to parameters such as the encoding format and sampling rate of the audio data to ensure the correct playback effect.

[0024] By establishing session connections between the SIP client and the TCP client and the server, the interconnection of devices under different communication protocols is realized, and the compatibility and flexibility of the communication system are enhanced. Both the instant communication needs based on the SIP protocol and the data transmission needs based on the TCP protocol can be met. The original audio data is collected by the sound sensor of the first client, and pre-processing operations such as format conversion and merging and compression are performed, which effectively improves the quality and transmission efficiency of the audio data. Format conversion ensures the compatibility of audio data in different devices and network environments, while merging and compression reduces the bandwidth required for data transmission and reduces communication costs. The pre-processed audio data is encoded into a voice packet in a specified format, and transmitted and processed by the server, realizing efficient and reliable voice communication. This encoding method not only ensures the integrity and consistency of voice data, but also reduces the complexity of data transmission and improves the overall performance of the communication system. The second client performs a legitimacy check on the received audio packet, effectively preventing the intrusion and malicious attacks of illegal data, and ensuring the security and reliability of the communication system. At the same time, the processed audio packet is played by the audio player of the second client, so that the user can hear the voice information from the first client in real time, further improving the real-time and accuracy of communication. It is particularly suitable for scenarios that require efficient and real-time communication, such as command and dispatch platforms. By achieving seamless communication between SIP clients and TCP clients, as well as rapid collection, processing and transmission of audio data, this method can significantly improve the efficiency and accuracy of command and dispatch, and provide timely and accurate information support for decision makers.

[0025] Figure 2 The schematic diagram is a schematic diagram of the principle of implementing a fusion communication method based on the SIP protocol and the half-duplex TCP protocol on the command and dispatch platform disclosed in the embodiment of the present application. Figure 2 As shown, after the SIP&TCP service (i.e., the server) receives the SIP request sent by the SIP client, it processes the SIP service and the TCP service and sends the audio data or message to the corresponding TCP terminal. Similarly, after the SIP&TCP service receives the TCP request sent by the TCP client, it processes the TCP service and the SIP service and sends the audio data or message to the SIP terminal.

[0026] Optionally, the preprocessing the original audio data to obtain audio data includes: Convert the original audio data into first audio data in ALAW format, and convert the first audio data into second audio data in PCM format; The second audio data is combined and compressed to be converted into audio data with a preset sampling rate and sampling bit number.

[0027] Read the raw audio data, which is usually stored in a linear PCM format, i.e., each sample value represents the instantaneous amplitude of the audio signal, ranging from the lowest negative value to the highest positive value (e.g., -32768 to 32767 for 16-bit PCM). Apply the ALAW compression algorithm to convert each PCM sample value to its corresponding ALAW encoded value. This conversion process is non-linear and is intended to reduce the amount of variation in low-amplitude signals, thereby saving storage space and reducing transmission bandwidth requirements. The result is first audio data encoded in the ALAW format. Apply the ALAW decoding algorithm to convert each ALAW encoded value back to its corresponding linear PCM sample value. This decoding process is the reverse of the previous compression process and is intended to restore the dynamic range and details of the original audio signal. The result is second audio data encoded in the PCM format, which is now consistent in amplitude with the original audio data, but may have been requantized or sample rate converted (if required in subsequent steps). If the audio data's sample rate or sample bit count does not match the desired format, resampling or bit conversion is performed. Resampling involves changing the sample rate of the audio data to match the processing capabilities of the target device or network bandwidth limitations. Bit conversion involves changing the number of bits per sample (such as from 16 bits to 8 bits) to further reduce the amount of data. In some cases, it may also be necessary to merge the audio data, such as merging multiple audio channels into a mono signal, or splicing multiple audio clips into a continuous file. Finally, apply an appropriate compression algorithm (such as lossless compression or lossy compression, depending on the trade-off between sound quality and file size) to further reduce the amount of data and optimize storage and transmission efficiency. The result is audio data that meets the preset sampling rate and sampling bit number, which is now ready for subsequent encoding, transmission or processing steps.

[0028] The original audio data may be encoded in different formats, which may have compatibility issues in different devices and environments. By converting the original audio data into the ALAW format, the compatibility of the audio data during transmission and storage can be ensured, because ALAW is a widely used audio compression format suitable for a variety of communication systems and devices. Further converting the first audio data in the ALAW format into the second audio data in the PCM format can make full use of the advantages of the PCM format in audio quality. PCM (Pulse Code Modulation) is an uncompressed digital audio format that can provide higher quality audio signals and is suitable for application scenarios with high requirements for sound quality. After converting the audio data from the ALAW format to the PCM format, the audio data may need to be merged and compressed. The purpose of this step is to reduce the redundancy of audio data and improve storage efficiency. By merging multiple audio data segments, data fragmentation can be reduced and data access speed can be increased. At the same time, through compression processing, the storage space occupied by the audio data can be further reduced, and the storage cost can be reduced. During the merging and compression process, the audio data can be converted to a preset sampling rate and sampling bit number. The sampling rate and sampling bit number are key factors in determining audio quality. By controlling these two parameters, it can be ensured that the audio data has consistent sound quality performance in different devices and environments. For example, in some application scenarios, the sampling rate of audio data may need to be set to a specific value (such as 44.1 kHz or 48 kHz) to ensure the accuracy and stability of the audio signal.

[0029] Optionally, encoding the audio data into a voice packet in a specified format and sending the voice packet to the server includes: The audio data is converted into an audio packet in Opus format by Opus encoding, the audio packet in Opus format is encapsulated into an RTP voice packet by using the RTP protocol, and the RTP voice packet is sent to the server at a fixed time interval.

[0030] Opus (Opus Interactive Audio Codec) is an open, copyright-free audio codec format that aims to provide a low-latency, high-quality audio transmission and storage solution. The core principle of the Opus codec is a mixture of multiple audio codec technologies, including linear predictive coding (LPC), MDCT transform, vector quantization, entropy coding, etc. The original audio signal undergoes preprocessing steps, including filtering, resampling, audio gain adjustment, etc., to meet the requirements of the encoder. Framing: The audio signal is divided into shorter time segments, usually 20 milliseconds to 60 milliseconds. Each time segment is called a frame, which is used for subsequent processing and encoding. Feature extraction: Feature extraction is performed on each frame. Common features include short-time spectrum, cepstral coefficients, linear prediction coefficients, etc., which are used for acoustic modeling and encoding. Multiple coding techniques are used to encode the extracted features. Opus uses a hybrid coding method, including MDCT transform, vector quantization, residual coding, etc. During the encoding process, the appropriate coding algorithm and parameters are selected according to the characteristics of the audio signal. Opus supports multiple encoding modes, including VOIP mode (for real-time communication), music mode (for music and high-fidelity audio), speech mode (for speech and speech recognition), etc. Entropy encoding is performed on the encoded data to reduce the number of bits required for data representation. Opus uses a variety of entropy coding techniques, such as arithmetic coding and Huffman coding. The encoded and entropy-encoded data are packaged into audio frames, including audio data, control information, frame synchronization flags, etc. The packaged data forms an audio packet in Opus format, which can be transmitted or stored. Real-time Transport Protocol (RTP) is a transmission protocol for multimedia data streams on the Internet. RTP is designed to provide time information and achieve stream synchronization. RTP usually runs on top of the User Datagram Protocol (UDP), and the two together complete the function of the transport layer. The RTP packet consists of an RTP header and payload data. Among them, the RTP header contains information such as data type and encoding method, data packet sequence number, data transmission timestamp, and data source identifier. The receiving end can correctly reconstruct the original signal based on this information. Use the audio packet in Opus format as the payload data of the RTP group and add the corresponding RTP header information to encapsulate it into an RTP voice packet. Create an RTP socket: Create an RTP socket on the sender to send the RTP voice packet. Set a fixed time interval according to actual needs, such as sending an RTP voice packet every 20 milliseconds. At a fixed time interval, send the encapsulated RTP voice packet to the server through the RTP socket. After receiving the RTP voice packet, the server can perform corresponding decoding and processing.

[0031] Opus is an open, royalty-free audio codec format designed to provide a low-latency, high-quality audio transmission and storage solution. It is able to provide high-quality audio transmission while maintaining low latency, making it ideal for real-time communication applications. Opus supports a wide range of bit rates, from very low bit rates (e.g., 6kbps) to very high bit rates (e.g., 512kbps), which makes it suitable for different application requirements, from low-bandwidth network environments to high-quality audio storage. Opus has excellent compression performance and can provide high-quality audio at lower bit rates, thereby saving transmission bandwidth and storage space. With Opus encoding, audio data can be efficiently compressed and encoded, reducing redundant information and improving transmission efficiency. The Opus encoder also supports mixing and splitting of multiple audio streams, which is very useful for multi-channel audio transmission and processing, such as multiple participants in an audio conference. Encapsulating audio packets in Opus format into RTP voice packets using the RTP protocol can ensure real-time transmission of audio data. RTP voice packets contain information such as the sequence number and timestamp of the audio data, which helps the receiving end to reassemble, synchronize, and detect errors in data packets. By sending RTP voice packets at fixed time intervals, the continuity and real-time nature of audio data can be maintained, and transmission delay and jitter can be reduced. Opus has a certain degree of fault tolerance and can still provide good audio quality in the event of network packet loss or partial data loss. It uses error correction coding and forward error correction technology to recover lost audio data through methods such as resampling, interpolation, and hiding lost data.

[0032] Optionally, the step of encapsulating the audio packet in the Opus format into an RTP voice packet using the RTP protocol includes: Add a data length field having a first number of bytes in length and a data type field having a second number of bytes in length before an audio packet in Opus format to form an RTP data packet header having a third number of bytes, where the third number is the sum of the first number and the second number; Set the maximum allowable size of each RTP voice packet based on network bandwidth, latency requirements, and terminal device processing capabilities; The audio packet in Opus format with the RTP data packet header is packetized according to the maximum allowed size to obtain multiple RTP voice packets, and each RTP voice packet is marked with a sequence number.

[0033] The data length field is used to indicate the length of the data carried in the RTP packet, and the data type field is used to indicate the data type carried in the RTP packet. For Opus audio, it will have a specific load type code, which is negotiated and determined by the RTCP protocol or other non-RTP mechanisms in the initial stage of the RTP session. The receiving end uses this field to identify and correctly process the received data. The total length of the RTP header is composed of the data length field and the data type field. In the embodiment of the present application, the data length field with a length of 4 bytes and the data type field with a length of 4 bytes, a total of 8 bytes, constitute the data packet header. The maximum allowable size of the RTP voice packet ensures that the RTP packet can be effectively transmitted and processed under different network conditions. If the packet is too large, it may cause increased transmission delay, increased packet loss rate, or the terminal device cannot process it. Therefore, when encapsulating Opus audio data, it is necessary to set a reasonable maximum packet size based on these factors. According to the set maximum allowable size, the audio packet in Opus format with the RTP data packet header is packetized. This usually means that the longer Opus audio frame is divided into multiple smaller fragments, each of which can be transmitted as a separate RTP packet. During the packetization process, it is necessary to ensure that each fragment is part of a complete Opus audio frame (or at least a fragment that can be correctly reassembled by the receiver). In addition, it is necessary to add appropriate sequence numbers and timestamps and other information to the header of each RTP packet to help the receiver correctly reassemble and play the audio data. The sequence number is used to identify the order of RTP packets in the session. The receiver uses this sequence number to detect problems such as packet loss, reordering, and duplicate packets, and performs error recovery or discards redundant data accordingly. When encapsulating each RTP packet, it needs to be assigned a unique sequence number (usually an integer that increases from an initial value). This sequence number is continuous during the RTP session and is unique for each sender.

[0034] By encapsulating Opus audio data through the RTP protocol, the low latency and high bandwidth characteristics of RTP can be used to achieve efficient transmission of audio data. This helps to ensure the real-time and integrity of audio data and improve user experience. By setting the maximum allowable size of each RTP voice packet, the transmission strategy of audio data can be flexibly adjusted according to actual conditions. This helps to maintain smooth transmission of audio data when network bandwidth is limited or the processing power of terminal devices is insufficient. Marking each RTP voice packet with a sequence number helps the receiving end to reorganize and sort the received audio data. This simplifies the processing flow of audio data and reduces the difficulty of management.

[0035] Optionally, the processing the audio packet by the server and then sending it to the second client includes: Remove the packet header and RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and convert the voice packet in Opus format into a PCM voice packet with the preset sampling rate and sampling number; The PCM voice packets are converted into voice packets in ALAW format according to a preset offset, the voice packets in ALAW format are sorted through data splicing, and sent to the second client.

[0036] When the server receives an RTP voice packet, it first removes its packet header and RTP header. The packet header usually contains information such as data length and data type, while the RTP header contains RTP protocol-related information such as sequence number, timestamp, and synchronization source identifier (SSRC). After removing these header information, the original Opus format voice packet can be obtained. After removing the header information, the server will restore the remaining byte stream to the Opus format voice packet according to the previously set encoding rules or protocols. This process usually involves parsing and reassembling the byte stream to ensure that the restored Opus voice packet is consistent with the original audio data. Before format conversion, the server will set the preset sampling rate and number of samples according to requirements. The sampling rate determines the sampling frequency of the audio data, while the number of samples determines the number of bits per sampling point. The choice of these parameters will affect the quality and size of the converted PCM voice packet. Use appropriate audio processing tools or libraries (such as FFmpeg, SoX, etc.) to convert the Opus format voice packet to the PCM format voice packet. This process involves steps such as decoding, resampling, and encoding of audio data to ensure that the converted PCM voice packets meet the preset sampling rate and number of samples. Before converting the PCM to ALAW format, the server may set a preset offset. This offset is usually used to adjust the numerical range of the PCM voice packet to make it more suitable for the requirements of ALAW encoding. Use the relevant functions or methods in the audio processing tool or library to convert the PCM voice packet into an ALAW format voice packet. This process involves nonlinear quantization of each sampling point in the PCM voice packet to generate an ALAW encoded voice packet. Before sending the ALAW format voice packet to the second client, the server may perform data splicing on it. Data splicing usually involves merging multiple ALAW voice packets in time order or other rules to generate a complete audio data stream. Finally, the server sends the spliced ​​ALAW format voice packet to the second client through an appropriate network protocol (such as TCP, UDP, etc.). During the sending process, the server may consider factors such as network bandwidth, delay, and packet loss rate to ensure the real-time and integrity of the audio data.

[0037] By removing the packet header and RTP header of the RTP voice packet, the server can accurately restore the original Opus format voice packet. Subsequently, the Opus format voice packet is converted into a PCM voice packet with a preset sampling rate and number of samples. This step ensures that the format of the audio data matches the decoding capability of the second client, thereby ensuring the smooth playback of the audio. The PCM voice packet is converted into an ALAW format voice packet to enhance the compatibility of the audio data. By converting to the ALAW format, the server can ensure that the audio data is widely supported on the second client, improving the reliability of audio communication. The ALAW format voice packets are sorted by data splicing to ensure the order of the audio data. In real-time audio communication, the order of audio data is crucial, which is related to the continuity and fluency of the audio. The sorted voice packets can ensure that the second client can receive and play the audio in the correct order, thereby avoiding the discontinuity and confusion of the audio. Although the ALAW format voice packet may have a certain compression loss relative to the PCM format, this loss is within an acceptable range and is exchanged for a smaller data packet size and higher transmission efficiency. Especially when network bandwidth is limited, using the ALAW format can reduce the amount of data transmitted, thereby reducing network latency and packet loss rate and improving the quality of audio communication.

[0038] Optionally, removing the packet header and the RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and converting the voice packet in Opus format to a PCM voice packet with the preset sampling rate and sampling number includes: Deleting the RTP data packet header of the third number of bytes and the RTP header of the fourth number of bytes from the received RTP voice packet, and restoring the audio data segment in Opus format; Using an Opus decoding algorithm, the Opus format audio data segment is decoded to convert it into uncompressed PCM format audio data; According to the preset sampling rate and sampling bit number, the PCM audio data is resampled and the bit number is converted to obtain a PCM voice packet.

[0039] The RTP packet header contains key information such as sequence number, timestamp, and synchronization source identifier (SSRC) for data transmission, synchronization, and reassembly. According to the RTP protocol specification, the length of the packet header is fixed, but the specific number of bytes may vary depending on the implementation. Generally, the length of the RTP packet header does not exceed a certain range (such as 12 to 20 bytes, including the IP, UDP, and RTP headers). Here, the "third number of bytes" refers to the specific length of the RTP packet header, which needs to be determined according to the actual situation. In addition to the RTP packet header, Opus audio data may also contain additional encapsulation format headers (such as OGG encapsulation format) when encapsulated into RTP packets. These header information is not required for the decoder, so it needs to be removed before decoding. Here, the "fourth number of bytes" refers to the length of the RTP payload header or encapsulation format header, which also needs to be determined according to the actual situation. After deleting the RTP packet header and RTP header, the original Opus format audio data segments can be obtained. These data segments are the output of the Opus encoder and contain compressed representations of the audio. The Opus decoder can be a software library (such as opusdec in opus-tools) or a hardware accelerator. The embodiment of the present application can use opusdec in opus-tools as a decoder. The recovered Opus format audio data segment is input into the Opus decoder. The decoder will output uncompressed PCM format audio data. These data are original audio signals and can be used directly for playback, processing or storage. PCM (Pulse Code Modulation) is a digital representation method for representing analog signals. In digital audio processing, PCM audio data usually has a fixed sampling rate and sampling bit number. The preset sampling rate and sampling bit number are determined according to the application scenario and requirements. For example, some audio processing algorithms may require a specific sampling rate and sampling bit number as input. If the sampling rate of the PCM audio data does not match the preset sampling rate, a resampling operation is required. Resampling refers to changing the sampling rate of audio data by methods such as interpolation or decimation. Specialized audio processing libraries (such as FFmpeg, Libresample, etc.) can be used to perform resampling operations. These libraries provide efficient resampling algorithms and interfaces. If the sampling bit number of the PCM audio data does not match the preset sampling bit number, a bit conversion operation is required. Bit conversion refers to changing the quantization bit number of the audio data (such as from 16 bits to 8 bits or from 8 bits to 16 bits). Bit conversion can be achieved through simple bit operations or more complex quantization algorithms. Similarly, a dedicated audio processing library can be used to perform bit conversion operations. After completing the resampling and bit conversion, PCM voice packets that meet the preset requirements can be obtained. These voice packets can be directly used for subsequent audio processing or storage operations.

[0040] By removing the RTP packet header and the RTP header (which occupy the third and fourth number of bytes respectively), the server can accurately recover the audio data segments in Opus format from the RTP voice packets. This process ensures the integrity of the audio data and avoids the degradation of audio quality due to data loss or damage. The Opus format audio data segments are decoded and converted into uncompressed PCM format audio data using the Opus decoding algorithm. The Opus decoding algorithm is known for its high efficiency and low latency, which can ensure the decoding speed and quality of audio data to meet the needs of real-time audio communication. The PCM audio data is resampled and the bit number is converted according to the preset sampling rate and sampling bit number. This process enables the server to flexibly adjust the sampling rate and sampling bit number of the PCM voice packet according to the decoding capability of the second client or the needs of a specific application scenario. This flexibility helps ensure the compatibility of audio data in different devices and network environments. Through accurate decoding and flexible sampling rate and sampling bit number conversion, the server can generate high-quality PCM voice packets. After being transmitted to the second client, these voice packets can present clear and coherent audio effects, improving the user's listening experience. When decoding and converting on the server side, processing delays can be reduced through optimization of algorithms and hardware acceleration. This helps ensure real-time transmission and playback of audio data and reduces audio interruptions or delays caused by processing delays.

[0041] Optionally, converting the PCM voice packet into a voice packet in ALAW format according to a preset offset includes: Mapping each audio sample value in the PCM voice packet into a coded value in ALAW format according to a preset conversion rule; The converted ALAW code values ​​are packaged to form continuous ALAW format voice packets.

[0042] Each audio sample value in a PCM voice packet is usually stored in linear PCM format, that is, the sample value directly represents the amplitude of the audio signal. However, ALAW (A-law algorithm) is an audio compression algorithm that achieves compression by mapping linear PCM sample values ​​to a set of finite number of code values. In this process, the server maps each audio sample value in the PCM voice packet to a code value in ALAW format according to a preset conversion rule. This conversion rule is usually a lookup table or a mathematical function that maps the PCM sample value to a specific value in the ALAW code space based on its size. When all PCM sample values ​​are converted to ALAW code values, the server packages these code values ​​into continuous ALAW format voice packets. In the ALAW format, each code value usually occupies 8 bits (i.e. 1 byte), so the packaging process is to arrange these code values ​​in sequence into a continuous byte stream. It should be noted that ALAW format voice packets may also include some additional metadata, such as packet headers, timestamps, etc., which are used to identify the structure and timing information of the voice packet. However, these metadata are not part of the PCM to ALAW conversion process, but are added as needed during the packaging process.

[0043] Although the PCM (Pulse Code Modulation) format has excellent sound quality, it has a large amount of data, which is not conducive to storage and transmission. The ALAW format, on the other hand, uses non-uniform quantization to retain more data for the low-volume part and less data for the high-volume part, thereby achieving efficient audio compression. This compression method not only reduces storage and transmission costs, but also retains sufficient audio details and ensures sound quality. Converting PCM voice packets to ALAW format helps to achieve audio data interoperability between different devices and different network environments. In addition, the standardization of the ALAW format also simplifies the processing flow of audio data and reduces the complexity of system integration. Compared with the PCM format, the ALAW format has a smaller amount of data, which means that more audio data can be stored or transmitted under the same storage space or transmission bandwidth. This is undoubtedly a huge advantage for application scenarios that need to process a large amount of audio data (such as speech recognition, speech synthesis, etc.). In the process of audio data processing, compression and decompression are two important links. The compression algorithm of the ALAW format is relatively simple, and the decompression speed is also fast, which helps to reduce processing delays and improve the real-time nature of audio data. This is especially important for real-time audio communication applications. Although the ALAW format is a lossy compression method, it uses a non-uniform quantization method and performs fine quantization processing on the low-volume part, so it can retain the audio details and sound quality as much as possible while ensuring compression efficiency. This makes the ALAW format superior to some other lossy compression formats in terms of audio quality. After converting the PCM voice packet to the ALAW format, the subsequent audio processing process can be simplified. For example, in the storage, transmission and playback of audio data, ALAW-encoded audio data can be used directly without additional decoding processing. This reduces the complexity of the system and improves processing efficiency.

[0044] Figure 3 This is a flow chart of the SIP client disclosed in the embodiment of the present application as the initiator, as shown in FIG. Figure 3 As shown, the process includes: S301, SIP client performs authentication, S302, prompts an error message when authentication fails, S303, establishes a session when authentication succeeds, S304, collects original audio data (ALAW format), S305, enters SIP service, S306, converts original audio data into PCM audio data, S307, OPUS compression, S308, splices and sorts voice packets, S309, enters TCP service, and S310, transmits to the target TCP client.

[0045] Figure 4 This is a flow chart of the TCP client as the initiator disclosed in the embodiment of the present application, such as Figure 4As shown, the process includes: S401, TCP client performs authentication, S402, prompts an error message when authentication fails, S403, establishes a session when authentication succeeds, S404, collects original audio data (PCM format), S405, opus compression, S406, splices and sorts voice packets, S407, enters TCP service, S408, opus decoding is PCM audio packet, S409, PCM audio packet is converted into ALAW audio packet, S410, enters SIP service, S411, transmits to the target SIP client.

[0046] Figure 5 The following is a flow chart of the TCP client as the receiving party for voice transmission and playback disclosed in the embodiment of the present application, as shown in FIG. Figure 5 As shown, the process includes: S501, the SIP client performs authentication, S502, prompts an error message when the authentication fails, S503, establishes a session when the authentication succeeds, S504, obtains microphone permission, S505, prompts an error message when the permission fails to be obtained, S506, collects the alaw audio stream when the permission is successfully obtained, S507, enters the SIP service, S508, converts the alaw audio data into pcm audio data, S509, merges and compresses the audio data, S510, subpackets the audio data, S511, opus compression, S512, encapsulates the voice packet in RTP format, S513, sends it to the TCP client, and the TCP client plays it.

[0047] Figure 6 The following is a flow chart of the SIP client disclosed in the embodiment of the present application as a receiver for voice transmission and playback, as shown in FIG. Figure 6 As shown, the process includes: S601, TCP client performs authentication, S602, prompts an error message when authentication fails, S603, establishes a session when authentication succeeds, S604, obtains microphone permission, S605, prompts an error message when permission acquisition fails, S606, collects PCM audio stream when permission acquisition succeeds, S607, merges and compresses audio data, S608, sub-packets audio data, S609, opus compression, S610, RTP format encapsulates voice packets, S611, enters TCP service, S612, RTP voice packet restoration, S613, opus decompression, S614, PCM audio data is converted into alaw audio data, S615, audio data is sub-packetized, S616, sent to SIP client, and SIP client plays.

[0048] This embodiment also discloses a fusion communication system based on the SIP protocol and the half-duplex TCP protocol. Figure 7 Schematic diagram of a module of a fusion communication system based on SIP protocol and half-duplex TCP protocol disclosed in an embodiment of the present application. Figure 7As shown, the system includes a connection module 701, a collection module 702, an execution module 703 and a playback module 704, wherein: A connection module 701 is configured to respectively establish a session connection between a SIP client, a TCP client and a server, and obtain access rights to the sound sensors of the SIP client and the TCP client; A collection module 702 is configured to collect raw audio data through the sound sensor of a first client, and pre-process the raw audio data to obtain audio data, wherein the first client is one of the SIP client and the TCP client, and the pre-processing includes format conversion and merging and compression; The execution module 703 is configured to encode the audio data into a voice packet in a specified format and send the voice packet to the server, and the server processes the voice packet and sends the voice packet to a second client, where the second client is one of the SIP client and the TCP client, and the second client is different from the first client; The playing module 704 is configured to perform a legitimacy check on the processed audio package through the second client, and play the audio package through the audio player of the second client after the check passes.

[0049] Optionally, the acquisition module 702 is configured to: Convert the original audio data into first audio data in ALAW format, and convert the first audio data into second audio data in PCM format; The second audio data is combined and compressed to be converted into audio data with a preset sampling rate and sampling bit number.

[0050] Optionally, the execution module 703 is configured to: The audio data is converted into an audio packet in Opus format by Opus encoding, the audio packet in Opus format is encapsulated into an RTP voice packet by using the RTP protocol, and the RTP voice packet is sent to the server at a fixed time interval.

[0051] Optionally, the execution module 703 is configured to: Add a data length field having a first number of bytes in length and a data type field having a second number of bytes in length before an audio packet in Opus format to form an RTP data packet header having a third number of bytes, where the third number is the sum of the first number and the second number; Set the maximum allowable size of each RTP voice packet based on network bandwidth, latency requirements, and terminal device processing capabilities; The audio packet in Opus format with the RTP data packet header is packetized according to the maximum allowed size to obtain multiple RTP voice packets, and each RTP voice packet is marked with a sequence number.

[0052] Optionally, the execution module 703 is configured to: Remove the packet header and RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and convert the voice packet in Opus format into a PCM voice packet with the preset sampling rate and sampling number; The PCM voice packets are converted into voice packets in ALAW format according to a preset offset, the voice packets in ALAW format are sorted through data splicing, and sent to the second client.

[0053] Optionally, the execution module 703 is configured to: Deleting the RTP data packet header of the third number of bytes and the RTP header of the fourth number of bytes from the received RTP voice packet, and restoring the audio data segment in Opus format; Using an Opus decoding algorithm, the Opus format audio data segment is decoded to convert it into uncompressed PCM format audio data; According to the preset sampling rate and sampling bit number, the PCM audio data is resampled and the bit number is converted to obtain a PCM voice packet.

[0054] Optionally, the execution module 703 is configured to: Mapping each audio sample value in the PCM voice packet into a coded value in ALAW format according to a preset conversion rule; The converted ALAW code values ​​are packaged to form continuous ALAW format voice packets.

[0055] It should be noted that: when the device provided in the above embodiment realizes its function, only the division of the above functional modules is used as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0056] This embodiment also discloses an electronic device, referring to Figure 8 The electronic device may include: at least one processor 801 , at least one communication bus 802 , a user interface 803 , a network interface 804 , and at least one memory 805 .

[0057] The communication bus 802 is used to realize the connection and communication between these components.

[0058] The user interface 803 may include a display screen (Display) and a camera (Camera). The optional user interface 803 may also include a standard wired interface and a wireless interface.

[0059] The network interface 804 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0060] Among them, the processor 801 may include one or more processing cores. The processor 801 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 805, and calling data stored in the memory 805. Optionally, the processor 801 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 801 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 801, and it can be implemented separately through a chip.

[0061] Among them, the memory 805 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 805 includes a non-transitory computer-readable storage medium. The memory 805 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 805 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 805 may also be optionally at least one storage device located away from the aforementioned processor 801. As Figure 8 As shown, the memory 805 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program of a converged communication method based on the SIP protocol and the half-duplex TCP protocol.

[0062] exist Figure 8 In the electronic device shown, the user interface 803 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 801 can be used to call the application program of the fusion communication method based on the SIP protocol and the half-duplex TCP protocol stored in the memory 805. When executed by one or more processors 801, the electronic device executes one or more methods in the above-mentioned embodiments.

[0063] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.

[0064] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0065] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0066] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0067] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0068] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory 805, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory 805 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.

[0069] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A fusion communication method based on SIP protocol and half-duplex TCP protocol, characterized in that: Applied to a command and dispatch platform, the command and dispatch platform includes a server, a SIP client and a TCP client, and the method includes: Establishing session connections between the SIP client, the TCP client and the server respectively, and obtaining access rights to the sound sensors of the SIP client and the TCP client; Collecting raw audio data through the sound sensor of the first client, and preprocessing the raw audio data to obtain audio data, the first client is one of the SIP client and the TCP client, and the preprocessing includes format conversion and merging compression; Encoding the audio data into a voice packet in a specified format and sending the voice packet to the server, and processing the voice packet through the server and sending the voice packet to a second client, where the second client is one of the SIP client and the TCP client, and the second client is different from the first client; The processed audio package is subjected to a legality check through the second client, and the audio package is played through the audio player of the second client after the check passes.

2. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 1 is characterized in that: The preprocessing of the original audio data to obtain audio data comprises: Convert the original audio data into first audio data in ALAW format, and convert the first audio data into second audio data in PCM format; The second audio data is combined and compressed to be converted into audio data with a preset sampling rate and sampling bit number.

3. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 2 is characterized in that: The step of encoding the audio data into a voice packet in a specified format and sending the voice packet to the server includes: The audio data is converted into an audio packet in Opus format by Opus encoding, the audio packet in Opus format is encapsulated into an RTP voice packet by using the RTP protocol, and the RTP voice packet is sent to the server at a fixed time interval.

4. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 3 is characterized in that: The method of using the RTP protocol to encapsulate the audio packet in the Opus format into the RTP voice packet includes: Add a data length field having a first number of bytes in length and a data type field having a second number of bytes in length before an audio packet in Opus format to form an RTP data packet header having a third number of bytes, where the third number is the sum of the first number and the second number; Set the maximum allowable size of each RTP voice packet based on network bandwidth, latency requirements, and terminal device processing capabilities; The audio packet in Opus format with the RTP data packet header is packetized according to the maximum allowed size to obtain multiple RTP voice packets, and each RTP voice packet is marked with a sequence number.

5. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 3 is characterized in that: The processing of the audio packet by the server and then sending the audio packet to the second client comprises: Remove the packet header and RTP header from the RTP voice packet to restore it to a voice packet in Opus format, and convert the voice packet in Opus format into a PCM voice packet with the preset sampling rate and sampling number; The PCM voice packets are converted into voice packets in ALAW format according to a preset offset, the voice packets in ALAW format are sorted through data splicing, and sent to the second client.

6. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 5, characterized in that: The step of removing the packet header and the RTP header from the RTP voice packet to restore the voice packet in Opus format, and converting the voice packet in Opus format into a PCM voice packet with the preset sampling rate and sampling number includes: Deleting the RTP data packet header of the third number of bytes and the RTP header of the fourth number of bytes from the received RTP voice packet, and restoring the audio data segment in Opus format; Using an Opus decoding algorithm, the Opus format audio data segment is decoded to convert it into uncompressed PCM format audio data; According to the preset sampling rate and sampling bit number, the PCM audio data is resampled and the bit number is converted to obtain a PCM voice packet.

7. The fusion communication method based on SIP protocol and half-duplex TCP protocol according to claim 5, characterized in that: The converting of the PCM voice packet into a voice packet in ALAW format according to a preset offset comprises: Mapping each audio sample value in the PCM voice packet into a coded value in ALAW format according to a preset conversion rule; The converted ALAW code values ​​are packaged to form continuous ALAW format voice packets.

8. A fusion communication system based on SIP protocol and half-duplex TCP protocol, characterized in that: It includes a connection module, a collection module, an execution module and a playback module, among which: A connection module, configured to establish a session connection between a SIP client, a TCP client and a server, and obtain access rights to the sound sensors of the SIP client and the TCP client; A collection module, configured to collect raw audio data through the sound sensor of a first client, and pre-process the raw audio data to obtain audio data, wherein the first client is one of the SIP client and the TCP client, and the pre-processing includes format conversion and merging and compression; an execution module configured to encode the audio data into a voice packet in a specified format and send the voice packet to the server, and to process the voice packet through the server and send the packet to a second client, wherein the second client is one of the SIP client and the TCP client, and the second client is different from the first client; The playing module is configured to perform a legality check on the processed audio package through the second client, and play the audio package through the audio player of the second client after the check passes.

9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Converged communication method and apparatus based on sip protocol and half-duplex TCP protocol

    WO2026144086A1