Voice data transmission method and device, equipment, storage medium and program product

By carrying only the encoding parameter set of the first encoded frame in the frame header of the encoded frame, and reusing the frame header of the first encoded frame in other encoded frames, the problem of high bandwidth overhead in data packet transmission in bandwidth-constrained scenarios is solved, and more efficient data transmission is achieved.

CN121842162APending Publication Date: 2026-04-10UNISOC CHONGQING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNISOC CHONGQING TECH CO LTD
Filing Date
2026-01-04
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In bandwidth-constrained communication environments, existing voice coding technologies incur significant bandwidth overhead for data packet transmission, failing to meet the needs of low-bandwidth scenarios.

Method used

By carrying only the encoding parameter set of the first encoded frame in the frame header of the encoded frame, and reusing the frame header of the first encoded frame in other encoded frames, the frame header structure is set up in a differential frame header manner, thereby reducing the frame header overhead.

Benefits of technology

It effectively reduces the bandwidth overhead of real-time transmission protocol voice packets and improves data transmission efficiency in low-bandwidth scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842162A_ABST
    Figure CN121842162A_ABST
Patent Text Reader

Abstract

The invention discloses a voice data transmission method and device, equipment, a storage medium and a program product, and relates to the technical field of voice coding, and the method comprises the steps: determining a frame structure of a second coding frame according to a coding parameter set corresponding to a first coding frame; wherein under the condition that the coding parameter set corresponding to the second coding frame is the same as the coding parameter set corresponding to the first coding frame, the frame header of the second coding frame does not comprise a field bearing the coding parameter set; the second coding frame and the first coding frame are coding frames of the same type; generating the second coded frame according to the frame structure of the second coded frame; and transmitting a real-time transport protocol (RTP) voice packet, wherein the load of the RTP voice packet comprises at least one second coding frame. In the invention, for the coding frames using the same coding parameter set, the coding parameter set is only carried in the frame header of the first coding frame, the setting of the differential frame header is realized, and the frame header overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech coding technology, specifically to a speech data transmission method, apparatus, device, storage medium, and program product. Background Technology

[0002] Ultra-low bitrate codec (ULBC) is a speech coding technology specifically designed for extremely low-bandwidth communication environments. It is primarily suitable for scenarios with high latency, moderate error rates, and limited bandwidth. ULBC applications mainly include geostationary Earth Orbit (GEO) satellite communication and base station-free terrestrial communication such as Bluetooth and Wi-Fi. In practical applications, reducing the bandwidth overhead of data packets to suit bandwidth-constrained scenarios is a key challenge that needs to be addressed. Summary of the Invention

[0003] This application provides a voice data transmission method, apparatus, device, storage medium, and program product, which solves the problem that current data packet transmission has high bandwidth overhead and cannot meet the needs of bandwidth-limited scenarios.

[0004] Firstly, a method for transmitting voice data is provided, the method comprising:

[0005] The frame structure of the second encoded frame is determined based on the encoding parameter set corresponding to the first encoded frame; wherein, if the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the field carrying the encoding parameter set; the second encoded frame and the first encoded frame are encoded frames of the same type;

[0006] The second encoded frame is generated according to the frame structure of the second encoded frame;

[0007] Transmit Real-Time Transport Protocol (RTP) voice packets, wherein the payload of the RTP voice packets includes at least one of the second encoded frames.

[0008] By implementing the first aspect of the method, when different encoded frames correspond to the same set of encoded parameters, the encoded parameter set is carried only in the header of the first encoded frame. Other encoded frames reuse the header of the first encoded frame without carrying the encoded parameter set. This enables the setting of differential headers for encoded frames, reduces header overhead, and thus reduces bandwidth overhead for RTP voice packet transmission.

[0009] In one possible implementation, the frame header of the second encoded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0010] In one possible implementation, if the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0011] In one possible implementation, when the second encoded frame is a silence insertion descriptor (SID), the first encoded frame and the second encoded frame are in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

[0012] In one possible implementation, generating the second encoded frame according to the frame structure of the second encoded frame includes:

[0013] When the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a first value. The first value is used to indicate that the encoding parameter set corresponding to the second encoded frame does not need to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0014] In one possible implementation, generating the second encoded frame according to the frame structure of the second encoded frame includes:

[0015] When the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a second value, which is used to indicate that the encoding parameter set corresponding to the second encoded frame needs to be updated on the encoding parameter set corresponding to the first encoded frame.

[0016] The encoding parameter set corresponding to the second encoded frame is encapsulated in the encoding parameter set field in the frame header of the second encoded frame;

[0017] The encoded data corresponding to the second encoded frame is encapsulated within the frame payload of the second encoded frame.

[0018] Secondly, a voice data transmission method is provided, the method comprising:

[0019] Receive Real-Time Transport Protocol (RTP) voice packets; wherein the payload of the RTP voice packets carries at least one second coded frame, and when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type;

[0020] The frame header of the second encoded frame carried in the payload of the RTP voice packet is parsed to obtain the frame header parsing result;

[0021] Based on the frame header parsing result, the audio signal corresponding to the second encoded frame is obtained.

[0022] By implementing the second aspect of the method, if the receiving end determines that the header of the received second encoded frame does not carry an encoding parameter set, it can reuse the encoding parameter set carried in the header of an encoded frame of the same type as the second encoded frame for decoding. This reduces the header overhead while ensuring that the received encoded frame can be correctly decoded.

[0023] In one possible implementation, the frame header of the second encoded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0024] In one possible implementation, if the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0025] In one possible implementation, when the second encoded frame is a silence insertion descriptor (SID), the first encoded frame and the second encoded frame are in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

[0026] In one possible implementation, parsing the header of the second coded frame carried within the payload of the RTP voice packet to obtain the header parsing result includes:

[0027] The first flag bit of the frame header is parsed to obtain the parsing result of the first flag bit, wherein the frame header parsing result includes the parsing result of the first flag bit.

[0028] In one possible implementation, based on the frame header parsing result, the audio signal corresponding to the second encoded frame is obtained, including:

[0029] If the parsing result of the first flag bit is the first value and the second encoded frame has no frame payload, the audio signal corresponding to the first encoded frame is used as the audio signal corresponding to the second encoded frame, wherein the audio signal is a silence signal.

[0030] In one possible implementation, obtaining the audio signal corresponding to the second encoded frame based on the frame header parsing result includes:

[0031] If the parsing result of the first flag bit is a first value, and the first encoded frame includes a frame payload, the frame payload is parsed to obtain the first encoded data;

[0032] Using the encoding parameter set corresponding to the first encoded frame, the first encoded data is decoded to obtain the audio signal corresponding to the second encoded frame, wherein the audio signal is a speech signal.

[0033] In one possible implementation, parsing the header of the second coded frame carried within the payload of the RTP voice packet to obtain the header parsing result further includes:

[0034] If the parsing result of the first flag bit is the second value, the encoding parameter set field of the frame header is parsed to obtain the encoding parameter set of the second encoded frame, wherein the frame header parsing result also includes the encoding parameter set of the second encoded frame.

[0035] Thirdly, a voice data transmission device is provided, the device comprising:

[0036] The determining module is used to determine the frame structure of the second encoded frame based on the encoding parameter set corresponding to the first encoded frame; wherein, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the field carrying the encoding parameter set; the second encoded frame and the first encoded frame are encoded frames of the same type;

[0037] The generation module is used to generate the second encoded frame according to the frame structure of the second encoded frame;

[0038] A transmission module is used to transmit Real-Time Transport Protocol (RTP) voice packets, wherein the payload of the RTP voice packets includes at least one second encoded frame.

[0039] Fourthly, a voice data transmission device is provided, the device comprising:

[0040] A receiving module is used to receive Real-Time Transport Protocol (RTP) voice packets; wherein the payload of the RTP voice packet carries at least one second coded frame, and when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type;

[0041] The parsing module is used to parse the frame header of the second encoded frame carried in the payload of the RTP voice packet and obtain the frame header parsing result;

[0042] The acquisition module is used to obtain the audio signal corresponding to the second encoded frame based on the frame header parsing result.

[0043] Fifthly, a communication device is provided, the communication device including a processor and a memory interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor invokes the program instructions to execute the method as described in the first aspect, or to execute the method as described in the second aspect.

[0044] In a sixth aspect, a chip is provided, the chip including a processor and an interface, the processor and the interface being coupled; the interface is used to receive or output signals, and the processor is used to execute code instructions to perform the method as described in the first aspect, or to perform the method as described in the second aspect.

[0045] In a seventh aspect, a module device is provided, the module device comprising a communication module, a power module, a storage module, and a chip module, wherein: the power module is used to provide electrical energy to the module device; the storage module is used to store data and / or instructions; the communication module is used to communicate with external devices; and the chip module is used to invoke the data and / or instructions stored in the storage module, and in conjunction with the communication module, to execute the method as described in the first aspect, or to execute the method as described in the second aspect.

[0046] Eighthly, a computer-readable storage medium is provided, on which a program or instructions are stored, which, when executed by a computer, implement the steps of the voice data transmission method as described in the first aspect, or, when executed by a processor, implement the steps of the voice data transmission method as described in the second aspect.

[0047] A ninth aspect provides a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the voice data transmission method described in the first aspect, or, when executed by the processor, implement the steps of the voice data transmission method described in the second aspect.

[0048] In the embodiments of this application, when encapsulating encoded data using the same encoding parameter set into encoded frames, only the header of the first encoded frame may carry the encoding parameter set, while other encoded frames may not carry the encoding parameter set. This reduces the overhead of the header and thus reduces the bandwidth overhead of RTP voice packet transmission. Attached Figure Description

[0049] Figure 1 This is one of the schematic diagrams of AMR encoding in related technologies;

[0050] Figure 2 This is the second schematic diagram of AMR encoding in related technologies;

[0051] Figure 3 This is one of the flowcharts illustrating the voice data transmission method provided in the embodiments of this application;

[0052] Figure 4 This is a schematic diagram of the frame structure of a speech frame carrying a set of encoding parameters in this application;

[0053] Figure 5 This is a schematic diagram of the frame structure of a speech frame that does not carry a set of encoding parameters in this application;

[0054] Figure 6 This is a schematic diagram of the frame structure of the SID carrying the encoding parameter set in this application;

[0055] Figure 7 This is a schematic diagram of the frame structure of the SID without the encoding parameter set in this application;

[0056] Figure 8 This is one of the schematic diagrams of the first and second coded frames in this application;

[0057] Figure 9 This is a second schematic diagram of the first and second coded frames in this application;

[0058] Figure 10 This is a comparative diagram of the complete speech frame structure and the differential speech frame structure in this application;

[0059] Figure 11 This is a comparative schematic diagram of the complete silence frame structure and the differential silence frame structure in this application;

[0060] Figure 12This is a second schematic flowchart of the voice data transmission method provided in the embodiments of this application;

[0061] Figure 13 This is one of the structural schematic diagrams of the voice data transmission device provided in the embodiments of this application;

[0062] Figure 14 This is a second schematic diagram of the structure of the voice data transmission device provided in the embodiments of this application;

[0063] Figure 15 This is a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0064] Figure 16 This is a schematic diagram of the structure of a chip module provided in an embodiment of this application. Detailed Implementation

[0065] In the embodiments of this application, the terms "first," "second," and "third" are used to distinguish identical or similar items with essentially the same function and purpose. Those skilled in the art will understand that the terms "first," "second," and "third" do not limit the quantity or execution order, nor are they necessarily required to be different. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0066] It should be understood that in the embodiments of this application, "at least one" refers to one or more; "multiple" refers to two or more. Furthermore, the word "equal to" in this application can be used in conjunction with "greater than" or "less than". When "equal to" and "greater than" are used together, the technical solution of "greater than" is adopted; when "equal to" and "less than" are used together, the technical solution of "less than" is adopted.

[0067] In the embodiments of this application, the terms "of," "corresponding (relevant)," "corresponding," "associated (related)," and "mapped" may sometimes be used interchangeably. It should be noted that when no distinction is emphasized, the concepts or meanings expressed are consistent.

[0068] Among related technologies, the encoder with the lowest bitrate is the Adaptive Multi-Rate Codec (AMR). The AMR encoding bitrate ranges from 4.75 to 12.2 kbps, with a minimum bitrate of 4.75 kbps. The frame header information of the AMR encoder is 5 bits, including two parts: Frame Type (FT) and Quality (Q). Specifically, the FT is 4 bits, representing 16 modes, such as speech mode, Silence Insertion Descriptor (SID), and no data; the Q is 1 bit, indicating whether the frame is a "good" or "bad" frame.

[0069] If Real-time Transport Protocol (RTP) is used to transmit AMR, two additional pieces of information are required, totaling 5 bits. These two additional pieces of information are: Codec Mode Request (CMR) and the F flag (follow). The CMR is used to request the receiver to set the next transmission rate, with CMR=15 indicating no change. The F flag indicates whether this frame is the only frame or the last frame, with F=1 indicating that there are other frames following this frame.

[0070] An example of using AMR encoding is as follows: Figure 1 As shown, where, Figure 1 This represents the payload structure of an RTP packet with an encoding bitrate of 7.4 kbps (FT=4). CMR=15 indicates that the encoder parameters remain unchanged, FT=4 indicates a bitrate of 7.4 kbps, Q=1, followed by the 148 encoded bits, and the last two placeholders are padded.

[0071] An example of using AMR encoding is as follows: Figure 2 As shown, where, Figure 2 For an RTP packet containing four frames, the headers of the four AMR frames are placed in the header of the payload. CMR=1 indicates that the encoding bitrate should be changed to 8.85kbps. FT=0 indicates that the bitrate of the first frame is 6.6kbps (FT=0), the second frame is a SID frame (FT=9), the third frame has no data (FT=15), and the bitrate of the fourth frame is 8.85kbps (FT=1).

[0072] As can be seen, in related technologies, the RTP packet payload, excluding CMR information, contains a total of 6 bits, of which the AMR encoder frame header has 5 bits. For a 20 ms frame length, this results in an overhead of 250 bps. At a bitrate of 4.75 kbps, this 250 bps accounts for approximately 5% (250 / 5000) of the overhead; while at a bitrate of 800 bps, the 250 bps accounts for approximately 23.8% (250 / 1050) of the overhead.

[0073] While ULBC may support fewer bitrates than AMR, it supports more frame lengths and sampling rates, thus providing more information for the frame header. However, ULBC's bitrate is lower than AMR's. If AMR's frame header processing methods were used, this overhead would be significant for ULBC.

[0074] In view of this, embodiments of this application provide a voice data transmission method to solve the problem of large frame header overhead.

[0075] The voice data transmission method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings. The execution entity in the embodiments of this application can be a terminal device. Alternatively, the execution entity in the embodiments of this application can be a device matched with the terminal device, such as a processor, chip, or chip module. The following description uses a terminal device as an example. For example, terminal equipment can also be called user equipment (UE). Terminal equipment can be mobile phones, tablet personal computers, laptop computers, laptops, personal digital assistants (PDAs), handheld computers, netbooks, ultra-mobile personal computers (UMPCs), mobile internet devices (MIDs), augmented reality (AR) devices, virtual reality (VR) devices, robots, wearable devices, flight vehicles, vehicle user equipment (VUEs), shipboard equipment, pedestrian user equipment (PUEs), smart home devices (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game consoles, personal computers (PCs), ATMs, or self-service machines, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, and smart clothing. In-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units.

[0076] The voice transmission method provided in the embodiments of this application is executed by a first communication device, such as a terminal device. Figure 3 As shown, the voice transmission method includes:

[0077] Step 301: The first communication device determines the frame structure of the second coded frame based on the encoding parameter set corresponding to the first coded frame; wherein, if the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the first coded frame and the second coded frame are coded frames of the same type.

[0078] For example, taking ULBC as an example, the above encoding parameter set may include commonly used ULBC encoding items, as shown in Table 1 below:

[0079] Table 1 Common ULBC Encoding Options

[0080] Encoding options For example Number of bits required Bitrate 1~3kbps, for example 800bps / 1200bps / 2400bps / SID 2 Sampling rate 8 / 16kHz 1 Frame length 20 / 60ms 1

[0081] It should be noted that Table 1 above is only used to illustrate the parameters included in the encoding parameter set, and is not intended to limit the encoding parameter set of this application.

[0082] For example, the encoding parameters in the above encoding parameter set are the parameters used in the process of generating the encoded data in the frame payload. In other words, the first communication device uses the above encoding parameter set to encode the audio signal to generate the corresponding encoding parameter set, and the encoded data is carried in the frame payload of the encoded frame.

[0083] For example, the first encoded frame described above can be a speech frame or a silence frame, and correspondingly, the second encoded frame can also be a speech frame or a silence frame. In other words, the first encoded frame and the second encoded frame are of the speech type, or the first encoded frame and the second encoded frame are of the silence type.

[0084] For example, the first encoded frame is located before the second encoded frame, that is: the first communication device first generates the first encoded frame and then generates the second encoded frame.

[0085] For example, if the header of the second encoded frame does not carry an encoding parameter set, the receiving device decodes the encoded data in the payload of the second encoded frame according to the encoding parameter set carried in the header of the first encoded frame, in accordance with the protocol agreement, predefined or preconfigured method.

[0086] For example, when the second encoded frame is a speech type encoded frame, the frame structure of the second encoded frame includes a frame header and a frame payload; wherein, when the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame includes the encoding parameter set, and when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the encoding parameter set.

[0087] For example, when the second encoded frame is a silent type encoded frame, if the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, then the frame structure of the second encoded frame includes a frame header and a frame payload, and the frame header includes the encoding parameter set; if the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, then the frame structure of the second encoded frame only includes a frame header, and the frame header does not include the encoding parameter set.

[0088] For example, the encoding parameter set corresponding to the second encoded frame of the first communication device can be obtained from the receiving device (such as the second communication device). For example, the receiving device sends a first indication information indicating that the encoding parameter set remains unchanged. In this way, the first communication device determines that the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame. Or, for example, the receiving device sends a second indication information indicating that some parameters in the encoding parameter set need to be updated. In this case, the second indication information can carry at least the updated encoding parameters. In this way, the first communication device determines that the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame.

[0089] For example, the first and second encoded frames mentioned above can be encoded frames within a single packet. The packet length for grouping the encoded frames can be the period length of the speech binding period. This allows the encoded data in an RTP speech packet to share the same group frame header information as much as possible. The speech binding period can be 80 / 160 / 240ms. Two bits can be used in an RTP speech packet to represent the four periods: single frame / 80 / 160 / 240ms.

[0090] Step 302: The first communication device generates the second encoded frame according to the frame structure of the second encoded frame.

[0091] Step 303: The first communication device transmits an RTP voice packet, the payload of which includes at least one second coded frame.

[0092] For example, an RTP voice packet may also include the aforementioned first encoded frame.

[0093] It should be noted that the number of coded frames included in the payload of an RTP voice packet can be determined based on the packet length; that is, a coded frame within a packet can be encapsulated within an RTP voice packet. Of course, an RTP voice packet can also include coded frames from multiple packets, and this application does not impose any specific limitations.

[0094] In the embodiments of this application, firstly, the first communication device determines the frame structure of the second coded frame based on the encoding parameter set corresponding to the first coded frame; wherein, when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the first coded frame and the second coded frame are coded frames of the same type. Secondly, the first communication device generates the second coded frame based on the frame structure of the second coded frame. This achieves differential frame header setting, thereby reducing frame header overhead. Thirdly, the first communication device transmits RTP voice packets, and the payload of the RTP voice packets includes at least one second coded frame. Here, because the second coded frame sets the frame header differentially, the bandwidth overhead of the RTP voice packets can be reduced.

[0095] In some embodiments, the frame header of the second encoded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0096] For example, the first flag bit can be a 1-bit field, such as a 1-bit FLAG field, used to indicate whether the current frame header has changed. For instance, when the first flag bit is set to 1, it indicates that the current frame header has changed, meaning that the encoding parameter set corresponding to the encoded frame containing the first flag bit has changed compared to the encoding parameter set corresponding to a previous encoded frame of the same type. In this case, the receiving device, such as the second communication device, needs to further decode the field carrying the encoding parameter set in the frame header to obtain the encoding parameter set corresponding to the encoded frame, and then decode the encoded data in the frame payload based on the decoded encoding parameter set. When the first flag bit is set to 0, it indicates that the current frame header has not changed, meaning that the encoding parameter set corresponding to the encoded frame containing the first flag bit has not changed compared to the encoding parameter set corresponding to a previous encoded frame of the same type. In this case, the receiving device can reuse the encoded parameter set decoded from the previous encoded frame to decode the encoded frame.

[0097] In some embodiments, when the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0098] In other words, when encapsulating encoded frames, for voice frames, the frame structure of the currently encapsulated encoded frame can be determined by whether the encoding parameter set of the currently encapsulated encoded data is the same as the encoding parameter set of the adjacent previously encapsulated encoded data. For example, if the corresponding encoding parameter sets are the same, the currently encapsulated encoded frame does not include the field carrying the encoding parameter set; if the corresponding encoding parameter sets are different, the currently encapsulated encoded frame includes the field carrying the encoding parameter set. This achieves the setting of a differential frame header, reduces the overhead of the frame header, and further reduces the transmission bandwidth overhead of RTP voice packets.

[0099] For example, based on the above embodiments, the frame structure of the first encoded frame of the speech frame type can be as follows: Figure 4 As shown, it includes a frame header and a frame payload. The frame header includes a first flag bit / FLAG field (1 bit), a bit rate field (2 bits), a sampling rate field (1 bit), and a frame length field (1 bit). The frame payload carries encoded data.

[0100] For example, based on the above embodiments, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame structure of the second encoded frame is as follows: Figure 5 As shown, it includes a frame header and a frame payload. The frame header includes a first flag bit / FLAG flag field (1 bit), and the frame payload carries encoded data.

[0101] Combination Figure 4 and Figure 5 As can be seen, in the embodiments of this application, when the encoding parameter sets corresponding to adjacent voice frames are the same, using differential frame headers to encapsulate the encoded frames can significantly reduce the frame header overhead, thereby further reducing the bandwidth overhead of RTP voice packets.

[0102] In some embodiments, when the second coded frame is a SID, the first and second coded frames are within the same SID update cycle, and the header of the first coded frame includes a field carrying a set of coded parameters. The length of the SID update cycle is the same as the length of the voice bundling time. This allows the receiving device to obtain the SID update cycle according to pre-configuration and protocol agreements, without requiring the first communication device to use additional resources to inform the receiving device of the SID update cycle. For example, the SID update cycle can be 240ms, meaning the noise coding is updated every 240ms.

[0103] In other words, for a SID type encoded frame, the frame header of the second encoded frame does not include the field carrying the encoded parameter set if the first encoded parameter set and the second encoded parameter set satisfy the following conditions: the first encoded frame and the second encoded frame are in the same SID update cycle; the frame header of the first encoded frame includes the field carrying the encoded parameter set; and the encoded parameter set of the first encoded frame and the encoded parameter set of the second encoded frame are the same.

[0104] For example, the first encoded frame can be the first encoded frame of a SID update cycle.

[0105] It's important to note that silence frames do not encode actual speech; they only need to include the speech features that generate comfort noise. The number of bytes of encoded data corresponding to these comfort noise features is only half the number of bytes of encoded data in the smallest speech frame. Furthermore, since the comfort noise in silence frames needs to be updated periodically, the speech features carried by silence frames within an update cycle should be identical. Therefore, the encoding parameter set and speech features corresponding to a SID within a SID update cycle are the same. Consequently, within a SID update cycle, only the frame header of the first SID needs to encapsulate the encoding parameter set and the encoded data corresponding to the speech features; other SIDs can reuse the encoding parameter set and encoded data from that first SID.

[0106] Therefore, in the above embodiments, the frame header of the first coded frame carries a field carrying the encoding parameter set, and the frame payload of the first coded frame carries the encoded data obtained by encoding the silence signal according to the encoding parameter set. The frame header of the second coded frame does not include the field carrying the encoding parameter set, nor does it include the frame payload. In other words, the second coded frame only includes a frame header, which only includes a first flag bit. The first flag bit indicates that the encoding parameter set corresponding to the second coded frame does not need to be updated based on the encoding parameter set corresponding to the first coded frame. This further reduces overhead.

[0107] For example, based on the above embodiments, the frame structure of the first encoded frame of the silence frame type is as follows: Figure 6 As shown, the frame includes a frame header and a frame payload. The frame header includes a first flag field (FLAG, 1 bit), a bit rate field (2 bits), a sampling rate field (1 bit), and a frame length field (1 bit). The frame payload carries the encoded data. Since the bit rate of the SID is less than the minimum speech coding bit rate (e.g., if the minimum speech bit rate is 800 bps, then the SID bit rate must be less than 800 bps, such as 400 bps), the bit rate field in the SID indicates that the encoded frame is a SID. The encoded data is the result of encoding some speech features of comfort noise.

[0108] For example, based on the above embodiments, the frame structure of the second encoded frame of the SID type is as follows: Figure 7 The image shown only includes the frame header.

[0109] In some embodiments, step 302, generating a second encoded frame according to the frame structure of the second encoded frame, includes:

[0110] If the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to the first value. The first value is used to indicate that the encoding parameter set corresponding to the second encoded frame does not need to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0111] For example, the first value is 0.

[0112] For example, if the second encoded frame and the first encoded frame are speech frames, after configuring the first flag bit to the first value, step 302 further includes: encapsulating the encoded data corresponding to the second encoded frame within the frame payload of the second encoded frame.

[0113] Here, also Figure 4 and Figure 5 Taking the frame structure as an example, assuming the second coded frame has a bit rate of 800bps, a sampling rate of 8kHz, and a frame length of 20ms, then compared with... Figure 4 The first encoded frame corresponding to the frame structure is as follows: Figure 8 As shown in the first encoded frame, correspondingly, with Figure 5 The corresponding second encoded frame is as follows Figure 8 The second encoded frame is shown in the image.

[0114] For example, if the first coded frame and the second coded frame are SIDs, assuming the sampling rate of the first coded frame is 8kHz and the frame length is 20ms, then... Figure 6 The first encoded frame corresponding to the frame structure is as follows: Figure 9 As shown in the first encoded frame, with Figure 7 The corresponding frame structure corresponds to the second encoded frame, such as Figure 9 The second encoded frame is shown in the image.

[0115] In some embodiments, step 302, generating a second encoded frame according to the frame structure of the second encoded frame, includes:

[0116] Step A1: If the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to the second value. The second value is used to indicate that the encoding parameter set corresponding to the second encoded frame needs to be updated on the encoding parameter set corresponding to the first encoded frame.

[0117] For example, the second value is 1.

[0118] Step A2: Encapsulate the encoding parameter set corresponding to the second encoded frame into the encoding parameter set field in the frame header of the second encoded frame.

[0119] Step A3: Encapsulate the encoded data corresponding to the second encoded frame within the frame payload of the second encoded frame.

[0120] For example, when the second encoded frame is a speech frame, the encoded data is the encoded data obtained by encoding the speech signal using the encoding parameter set corresponding to the second encoded frame; when the second encoded frame is a SID, the encoded data is the data obtained by encoding some speech features of comfort noise using the encoding parameter set corresponding to the second encoded frame.

[0121] In other words, when the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame should include a field carrying the encoding parameter set. In this way, the receiving device can decapsulate the frame header of the second encoded frame to obtain the encoding parameter set, and then decode the encoded data in the frame payload of the second encoded frame based on the encoding parameter set to obtain the original audio signal.

[0122] It should be noted that in the above embodiments, when the second encoded frame is a speech frame and the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame structure of the second encoded frame is as follows: Figure 4 As shown, when the second coded frame is a silent frame and the coding parameter set corresponding to the second coded frame is different from the coding parameter set corresponding to the first coded frame, the frame structure of the second coded frame is as follows. Figure 6 As shown.

[0123] Based on the above embodiments, the frame header of the second encoded frame further includes a second flag bit, which is used to indicate whether the second encoded frame is the last frame of an RTP voice packet. Based on this, step 302, generating the second encoded frame according to its frame structure, includes:

[0124] In the case that the second encoded frame is the last pin of an RTP voice packet, the second flag bit is configured to the first value;

[0125] Alternatively, if the second encoded frame is not the last frame of the RTP voice packet, the second flag bit can be configured to the second value.

[0126] For example, the first value is 0, and the corresponding second value is 1.

[0127] For example, the second flag can also be called the F (Follow) flag.

[0128] As an example, based on the above, the frame structure of a complete speech frame is as follows: Figure 10The complete speech frame structure in the text, using a differential frame header, is as follows: Figure 10 The differential speech frame structure is a speech frame structure that has the same set of coding parameters as the preceding speech frame.

[0129] by Figure 10 For example, if the first coded frame is a complete speech frame structure (i.e., the coding parameter set of the first coded frame is different from that of its adjacent preceding frame, the first and second coded frames are encapsulated in the same RTP speech packet, and the coding parameter set of the second coded frame is the same as that of the first coded frame), then the second flag bit / F flag of the first coded frame is set to 1, the first flag bit / FLAG flag is set to 1, and fields such as bit rate, sampling rate, and frame length are encapsulated according to the coding parameter set of the first coded frame, and the corresponding coded data is encapsulated in the frame payload. Based on this, the second coded frame is a differential speech frame structure. The first flag bit / FLAG flag of the second coded frame is set to 0, indicating that the second coded frame reuses the coding parameter set of the first coded frame. If the second coded frame is the last coded frame in the RTP speech packet, then the second flag bit / F flag of the second coded frame is set to 0; if the second coded frame is not the last coded frame in the RTP speech packet, then the second flag bit / F flag of the second coded frame is set to 1, and the corresponding coded data is encapsulated in the frame payload.

[0130] As another example, based on the above, the frame structure of a complete silent frame / the frame structure of the first SID in a SID update cycle is as follows: Figure 11 The complete silence frame structure in the document includes the structure of a silence frame with a differential frame header, and the frame structure of other SIDs besides the first SID in a SID update cycle, as shown below. Figure 11 The differential silence frame structure in the text.

[0131] by Figure 11For example, if the first encoded frame is a complete voice frame structure, such as the first SID in a SID update cycle, or if the receiving device, such as the second communication device, indicates that the encoding parameter set corresponding to the first encoded frame has changed, the first flag bit / FLAG bit of the first encoded frame is set to 1. If the first encoded frame is the last encoded frame in its RTP voice packet, the second flag bit / F bit of the first encoded frame is set to 0; otherwise, it is set to 1. Fields such as bit rate, sampling rate, and frame length are encapsulated according to the encoding parameter set corresponding to the first encoded frame, and the corresponding encoded data is encapsulated in the frame payload. Based on this, if the second coded frame and the first coded frame are in the same SID update cycle, and the receiving device, such as the second communication device, indicates that the encoding parameter set corresponding to the second coded frame has not been updated, then the second coded frame is a differential silence frame structure. The first flag bit / FLAG flag of the second coded frame is set to 0, indicating that the second coded frame reuses the encoding parameter set in the first coded frame. If the second coded frame is the last coded frame in the RTP voice packet, then the second flag bit / F flag of the second coded frame is set to 0; otherwise, it is set to 1. The second coded frame has no frame payload, but reuses the encoded data in the frame payload of the first coded frame.

[0132] The encoding process of the second encoded frame will be explained below, taking the second encoded frame as an example, which is a speech frame:

[0133] First, the encoding parameter set and the voice signal are acquired. The encoding parameter set can be acquired based on the instructions of the receiving device, such as the second communication device. For example, the receiving device can instruct to reuse the first encoding parameter set corresponding to the previous voice frame, or it can instruct to use a new second encoding parameter set.

[0134] Second, the speech signal is encoded using the determined set of encoding parameters to obtain encoded data.

[0135] Third, determine the frame structure. When the determined encoding parameter set is the first encoding parameter set, the frame structure is a differential speech frame structure, that is, the frame header does not include the fields carrying the encoding parameter set; when the determined encoding parameter set is the second encoding parameter set, the frame structure is a complete speech frame structure, that is, the frame header includes the fields carrying the encoding parameter set.

[0136] Fourth, determine whether the encoded frame is the last frame in the RTP voice packet.

[0137] Fifth, encapsulate the encoded frame: First, encapsulate the encoded data in the frame payload. Second, when the frame structure is a complete speech frame structure, set the first flag to 1 and encapsulate the encoded parameter set in the frame header. When the frame structure is a differential frame structure, set the first flag to 0. Third, when the encoded frame is the last frame in an RTP speech packet, set the second flag in the frame header to 0. When the encoded frame is not the last frame in an RTP speech packet, set the second flag in the frame header to 1.

[0138] Sixth, encapsulate RTP voice packets based on the encapsulated encoded frames.

[0139] For example, when an RTP voice packet includes four coded frames, it can be represented as: |+F=1(1bit)+||+FLAG=1(1bit)+||+BR(2bit)+|+SR(1bit)+||+FL(1bit)+||+F=1+||+FLAG=0+||+F=1+||+FLAG=0||+F=0+||+FLAG=0+||; when an RTP voice packet includes one coded frame, it can be represented as: |+F=0+||+FLAG=1+||+BR+|+SR+||+FL+||; where F represents the second flag bit, FLAG represents the first flag bit, BR represents the bit rate, SR represents the sampling rate, and FL represents the frame length.

[0140] The encoding process of the second encoded frame will be explained below, taking the second encoded frame as a silent frame as an example:

[0141] First, obtain the encoding parameter set, which can be obtained based on the instruction of the receiving device, such as the second communication device. For example, the receiving device can instruct to reuse the third encoding parameter set corresponding to the previous silence frame, or it can instruct to obtain a new fourth encoding parameter set.

[0142] Second, determine whether the second encoded frame is the first SID of a SID update cycle. If yes, determine that the frame structure of the second encoded frame is a complete frame structure and perform the following "third" step. If no, if the encoding parameter set corresponding to the second encoded frame indicated by the second communication device is the fourth encoding parameter set, determine that the frame structure of the second encoded frame is a complete frame structure and perform the following "third" step. Alternatively, if the encoding parameter set corresponding to the second encoded frame indicated by the second communication device is the third encoding parameter set, determine that the frame structure of the second encoded frame is a differential frame structure and perform the following "fifth" step.

[0143] Third, when the second coding frame is a complete silence frame structure, the partial speech features of the comfort noise that needs to be encoded are obtained, and the partial speech features of the comfort noise are encoded using the coding parameter set corresponding to the second coding frame to obtain the encoded data.

[0144] Fourth, encapsulate the encoded frame: First, encapsulate the encoded data in the frame payload; second, set the first flag to 1 and encapsulate the encoded parameter set in the frame header; third, when the second encoded frame is the last frame in the RTP voice packet, set the second flag in the frame header to 0; when the second encoded frame is not the last frame in the RTP voice packet, set the second flag in the frame header to 1.

[0145] Fifth, if the second encoded frame is a differential silence frame, set the first flag in the frame header to 0, and determine whether the second encoded frame is the last voice frame in the RTP voice packet. If it is, set the second flag in the frame header to 0; otherwise, set the second flag in the frame header to 1. At this point, the frame encapsulation process of the second encoded frame is completed.

[0146] Sixth, based on the encapsulated encoded frame, i.e. the encoded frame encapsulated in step "four" or step "fifth", encapsulate RTP voice packets.

[0147] In the above-described voice data transmission method of this application, different frame structures are adopted based on whether the encoding parameter sets corresponding to different encoded frames are the same. Specifically, a field is set in the frame header to indicate whether the encoding parameter sets are the same. When the encoding parameter sets are the same, on the one hand, this field is set to 0, so that the receiving device can determine the encoding parameter set of the previous encoded frame to be reused based on this field. On the other hand, no field for carrying the encoding parameter set is set in the frame header, thus reducing the overhead of the frame header.

[0148] Embodiments of this application also provide a voice data transmission method, which is applied to, or executed by, a second communication device, such as a terminal / UE, a terminal device, etc. Figure 12 As shown, the method includes:

[0149] Step 1201: The second communication device receives an RTP voice packet; wherein the payload of the RTP voice packet carries at least one second coded frame, and if the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type.

[0150] For example, an RTP voice packet may also include the aforementioned first encoded frame.

[0151] Step 1202: The second communication device parses the frame header of the second coded frame carried in the payload of the RTP voice packet to obtain the frame header parsing result.

[0152] For example, taking ULBC as an example, the above encoding parameter set may include commonly used ULBC encoding items, such as bit rate, sampling rate, frame length, etc. in Table 1 above, but is not limited to this.

[0153] For example, the encoding parameters in the aforementioned encoding parameter set are the parameters used in the process of generating the encoded data in the frame payload. In other words, the first communication device uses the aforementioned encoding parameter set to encode the audio signal to generate the corresponding encoding parameter set, and the encoded data is carried in the frame payload of the encoded frame. Correspondingly, the second communication device uses the aforementioned encoding parameter set to decode the encoded data.

[0154] For example, the first encoded frame described above can be a speech frame or a silence frame, and correspondingly, the second encoded frame can also be a speech frame or a silence frame. In other words, the first encoded frame and the second encoded frame can be of the speech type, or the first encoded frame and the second encoded frame can be of the silence type.

[0155] For example, the first encoded frame is located before the second encoded frame, that is, the second communication device decodes the first encoded frame first, and then decodes the second encoded frame.

[0156] For example, if the header of the second encoded frame does not carry an encoding parameter set, the second communication device decodes the encoded data in the payload of the second encoded frame according to the encoding parameter set carried in the header of the first encoded frame, in accordance with the protocol agreement, predefined or preconfigured method.

[0157] For example, when the second encoded frame is a speech type encoded frame, the frame structure of the second encoded frame includes a frame header and a frame payload; wherein, when the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame includes the encoding parameter set, and when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the encoding parameter set.

[0158] For example, when the second encoded frame is a silent type encoded frame, if the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, then the frame structure of the second encoded frame includes a frame header and a frame payload, and the frame header includes the encoding parameter set; if the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, then the frame structure of the second encoded frame only includes a frame header, and the frame header does not include the encoding parameter set.

[0159] For example, the first and second encoded frames mentioned above can be encoded frames within a single packet. The packet length for grouping the encoded frames can be the period length of the speech binding period. This allows the encoded data in an RTP speech packet to share the same group frame header information as much as possible. The speech binding period can be 80 / 160 / 240ms. Two bits can be used in an RTP speech packet to represent the four periods: single frame / 80 / 160 / 240ms.

[0160] Step 1203: The second communication device obtains the audio signal corresponding to the second encoded frame based on the frame header parsing result.

[0161] For example, if the frame header parsing result includes a set of encoding parameters, the second communication device decodes the encoded data in the frame payload of the second encoded frame according to the set of encoding parameters to obtain an audio signal.

[0162] For example, if the frame header parsing result does not include the encoding parameter set, and the second encoded frame is a speech frame, the second communication device decodes the encoded data in the frame payload of the second encoded frame according to the encoding parameters corresponding to the first encoded frame to obtain the audio signal / speech signal. If the second encoded frame is a silence frame, the second communication device reuses the audio signal / comfortable noise speech features corresponding to the first encoded frame.

[0163] In the embodiments of this application, firstly, the second communication device receives RTP voice packets; secondly, the second communication device parses the frame header of the second coded frame carried in the payload of the RTP voice packet to obtain the frame header parsing result; wherein, the payload of the RTP voice packet carries at least one second coded frame, and if the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type; thus, the frame header overhead can be reduced, thereby reducing the bandwidth overhead of RTP voice packet transmission. Thirdly, the second communication device obtains the audio signal corresponding to the second coded frame based on the frame header parsing result. In this way, while reducing overhead, the second communication device can also successfully decode the coded frame and obtain the corresponding audio signal.

[0164] In some embodiments, the frame header of the second encoded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0165] For example, the first flag bit can be a 1-bit field, such as a 1-bit FLAG field, used to indicate whether the current frame header has changed. For instance, when the first flag bit is set to 1, it indicates that the current frame header has changed, that is, the encoding parameter set corresponding to the encoded frame where the first flag bit is located has changed compared to the encoding parameter set corresponding to a previous encoded frame of the same type. In this case, the second communication device can further decode the field carrying the encoding parameter set in the frame header to obtain the encoding parameter set corresponding to the encoded frame, and decode the encoded data in the frame payload according to the decoded encoding parameter set. When the first flag bit is set to 0, it indicates that the current frame header has not changed, that is, the encoding parameter set corresponding to the encoded frame where the first flag bit is located has not changed compared to the encoding parameter set corresponding to a previous encoded frame of the same type. In this case, if there is a frame payload in the encoded frame, the second communication device can reuse the encoding parameter set decoded from the previous encoded frame to decode the encoded frame.

[0166] In some embodiments, when the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0167] In other words, for a voice frame, the frame structure of the current encoded frame is related to whether its encoding parameter set is the same as that of the preceding encoded frame. For example, if the corresponding encoding parameter sets are the same, the currently encapsulated encoded frame does not include the field carrying the encoding parameter set; if the corresponding encoding parameter sets are different, the currently encapsulated encoded frame includes the field carrying the encoding parameter set. This enables the setting of a differential frame header, reduces the overhead of the frame header, and further reduces the transmission bandwidth overhead of RTP voice packets.

[0168] For example, based on the above embodiments, the frame structure of the first encoded frame of the speech frame type can be as follows: Figure 4 As shown, it includes a frame header and a frame payload. The frame header includes a first flag bit / FLAG field (1 bit), a bit rate field (2 bits), a sampling rate field (1 bit), and a frame length field (1 bit). The frame payload carries encoded data.

[0169] For example, based on the above embodiments, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame structure of the second encoded frame is as follows: Figure 5 As shown, it includes a frame header and a frame payload. The frame header includes a first flag bit / FLAG flag field (1 bit), and the frame payload carries encoded data.

[0170] Combination Figure 4 and Figure 5As can be seen, in the embodiments of this application, when the encoding parameter sets corresponding to adjacent voice frames are the same, using differential frame headers to encapsulate the encoded frames can significantly reduce the frame header overhead, thereby further reducing the bandwidth overhead of RTP voice packets.

[0171] In some embodiments, when the second coded frame is a SID, the first and second coded frames are within the same SID update cycle, and the header of the first coded frame includes a field carrying a set of coded parameters. The length of the SID update cycle is the same as the length of the voice binding cycle. Thus, the second communication device can obtain the SID update cycle according to pre-configuration and protocol agreements without the first communication device needing to use additional resources to inform the receiving device of the SID update cycle. For example, the SID update cycle can be 240ms, meaning the noise code is updated every 240ms.

[0172] In other words, for a SID type encoded frame, the frame header of the second encoded frame does not include the field carrying the encoded parameter set if the first encoded parameter set and the second encoded parameter set satisfy the following conditions: the first encoded frame and the second encoded frame are in the same SID update cycle; the frame header of the first encoded frame includes the field carrying the encoded parameter set; and the encoded parameter set of the first encoded frame and the encoded parameter set of the second encoded frame are the same.

[0173] For example, the first encoded frame can be the first encoded frame of a SID update cycle.

[0174] It's important to note that silence frames do not encode actual speech; they only need to include the speech features that generate comfort noise. The number of bytes of encoded data corresponding to these comfort noise features is only half the number of bytes of encoded data in the smallest speech frame. Furthermore, since the comfort noise in silence frames needs to be updated periodically, the speech features carried by silence frames within an update cycle should be identical. Therefore, the encoding parameter set and speech features corresponding to a SID within a SID update cycle are the same. Consequently, within a SID update cycle, only the frame header of the first SID needs to encapsulate the encoding parameter set and the encoded data corresponding to the speech features; other SIDs can reuse the encoding parameter set and encoded data from that first SID.

[0175] Therefore, in the above embodiments, the frame header of the first coded frame carries a field carrying the encoding parameter set, and the frame payload of the first coded frame carries the encoded data obtained by encoding the silence signal according to the encoding parameter set. The frame header of the second coded frame does not include the field carrying the encoding parameter set, nor does it include the frame payload. In other words, the second coded frame only includes a frame header, which only includes a first flag bit. The first flag bit indicates that the encoding parameter set corresponding to the second coded frame does not need to be updated based on the encoding parameter set corresponding to the first coded frame. This further reduces overhead.

[0176] For example, based on the above embodiments, the frame structure of the first encoded frame of the silence frame type is as follows: Figure 6 As shown, the frame includes a frame header and a frame payload. The frame header includes a first flag field (FLAG, 1 bit), a bit rate field (2 bits), a sampling rate field (1 bit), and a frame length field (1 bit). The frame payload carries the encoded data. Since the bit rate of the SID is less than the minimum speech coding bit rate (e.g., if the minimum speech bit rate is 800 bps, then the SID bit rate must be less than 800 bps, such as 400 bps), the bit rate field in the SID indicates that the encoded frame is a SID. The encoded data is the result of encoding some speech features of comfort noise.

[0177] For example, based on the above embodiments, the frame structure of the second encoded frame of the SID type is as follows: Figure 7 The image shown only includes the frame header.

[0178] In some embodiments, step 1202, parsing the frame header of the second coded frame carried within the payload of the RTP voice packet to obtain the frame header parsing result, includes:

[0179] The first flag bit of the frame header is parsed to obtain the parsing result of the first flag bit. The frame header parsing result includes the parsing result of the first flag bit.

[0180] For example, the parsing result of the first flag bit is a first value or a second value, for example, the first value is 0, and the corresponding second value is 1.

[0181] In some embodiments, step 1203, obtaining the audio signal corresponding to the second coded frame based on the frame header parsing result, includes:

[0182] If the parsing result of the first flag bit is the first value and the second coded frame has no frame payload, the audio signal corresponding to the first coded frame is used as the audio signal corresponding to the second coded frame, wherein the audio signal is a silence signal.

[0183] It should be noted that in the above embodiment, the type of the second encoded frame is a silence frame, that is, the second encoded frame is a SID. Furthermore, since the first flag bit is a first value, the second encoded frame is not the first encoded frame in a SID update cycle. As can be seen from the previous example, in this case, the second encoded frame only includes a frame header, and the frame header only includes the first flag bit. Thus, the second communication device does not need to parse the second encoded frame again, but can directly reuse the silence signal corresponding to the first encoded frame.

[0184] In some embodiments, step 1203, obtaining the audio signal corresponding to the second coded frame based on the frame header parsing result, includes:

[0185] If the parsing result of the first flag bit is a first value, and the first coded frame includes a frame payload, the frame payload is parsed to obtain the first coded data. Then, using the encoding parameter set corresponding to the first coded frame, the first coded data is decoded to obtain the audio signal corresponding to the second coded frame, wherein the audio signal is a speech signal.

[0186] It should be noted that when the first flag is a first value indicating that the encoding parameter set corresponding to the second encoded frame is the same as that corresponding to the first encoded frame, and the second encoded frame also includes a frame payload, the frame type of the second encoded frame is a speech frame. In this way, the second communication device can reuse the encoding parameter set corresponding to the first encoded frame to decode the encoded data in the frame payload of the second encoded frame to obtain the corresponding speech signal.

[0187] In some embodiments, step 1202, parsing the frame header of the second coded frame carried within the payload of the RTP voice packet to obtain the frame header parsing result, includes:

[0188] If the parsing result of the first flag bit is the second value, the encoding parameter set field of the frame header is parsed to obtain the encoding parameter set of the second encoded frame, wherein the frame header parsing result is also the encoding parameter set of the second encoded frame.

[0189] In other words, if the first flag bit in the frame header is determined to be a second value indicating that the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame will also carry the encoding parameter set. Therefore, the second communication device can further parse the field carrying the encoding parameter set in the frame header of the second encoded frame to obtain the encoding parameter set corresponding to the second encoded frame.

[0190] Based on the above embodiments, step 1203, obtaining the audio signal corresponding to the second encoded frame according to the frame header parsing result, includes:

[0191] The frame payload of the second encoded frame is parsed to obtain the second encoded data;

[0192] Based on the encoding parameter set corresponding to the first encoded frame, the second encoded data is decoded to obtain the audio signal corresponding to the first encoded frame, wherein the audio signal is a silence signal or a speech signal.

[0193] In other words, when the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame will carry the encoding parameter set corresponding to the second encoded frame. After parsing the frame header of the second encoded frame, the second communication device can continue to use the encoding parameter set corresponding to the second encoded frame to decode the second encoded data in the frame payload of the second encoded frame to obtain the corresponding silence signal or voice signal.

[0194] In some embodiments, the frame header of the second encoded frame further includes a second flag bit, which is used to indicate whether the second encoded frame is the last frame of an RTP voice packet;

[0195] Step 1202, which involves parsing the header of the second coded frame carried within the payload of the RTP voice packet to obtain the header parsing result, also includes:

[0196] The second flag bit is parsed to obtain the parsing result. Specifically, if the second encoded frame is the last frame of the RTP voice packet, the second flag bit has a first value; if the second encoded frame is not the last frame of the RTP voice packet, the second flag bit has a second value. Thus, the second communication device can determine whether all encoded frames in the RTP voice packet have been decoded based on the second flag bit.

[0197] For example, the first value is 0, and the corresponding second value is 1.

[0198] For example, the second flag can also be called the F (Follow) flag.

[0199] It should be noted that the decoding process of the second coded frame by the second communication device is the reverse process of the encoding process of the second coded frame by the first communication device. Therefore, the decoding process of the second coded frame by the second communication device can refer to the encoding process of the second coded frame by the first communication device, and will not be repeated here.

[0200] The following is an example illustrating how the above-described method of this application can be encoded to save overhead.

[0201] Assuming ULBC operates in 2400bps, 8kHz, 20ms mode, with a packet length of 4 frames and unchanged frame header information, if the AMR frame header representation method is used, the frame header overhead is 20 bits.

[0202] However, the method described in the embodiments of this application is as follows:

[0203] Frame 1: F=0 (1 bit, indicating that there are more frames to follow) FLAG=1 (1 bit, indicating a new frame header), frame header = "bitrate: 2400bps (10), frame length: 20ms (0), sampling rate: 8kHz (0)" (4 bits in total).

[0204] Frames 2-3: F=0 (1 bit, indicating that there are more frames to follow) FLAG=0 (1 bit, indicating that the frame header is reused).

[0205] In frame 4, F=1 (1 bit, indicating the last frame) FLAG=0 (1 bit, indicating the multiplexed frame header).

[0206] As can be seen, by using the method described in the embodiments of this application, the frame header overhead is 12 bits, saving 40% of the overhead.

[0207] Please refer to Figure 13 This application provides a voice data transmission device 1300, which may be a processor, chip, chip module, or terminal device. The voice data transmission device 1300 includes:

[0208] The determining module 1301 is used to determine the frame structure of the second encoded frame based on the encoding parameter set corresponding to the first encoded frame; wherein, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the field carrying the encoding parameter set; the second encoded frame and the first encoded frame are encoded frames of the same type;

[0209] The generation module 1302 is used to generate the second encoded frame according to the frame structure of the second encoded frame;

[0210] The transmission module 1303 is used to transmit Real-Time Transport Protocol (RTP) voice packets, wherein the payload of the RTP voice packets includes at least one second encoded frame.

[0211] In the embodiments of this application, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame may not carry the encoding parameter set, but may reuse the encoding parameter set corresponding to the first encoded frame, which can reduce the frame header overhead.

[0212] In some embodiments, optionally, the frame header of the second coded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second coded frame needs to be updated based on the encoding parameter set corresponding to the first coded frame.

[0213] In some embodiments, optionally, if the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0214] In some embodiments, optionally, when the second encoded frame is a SID, the first encoded frame and the second encoded frame are located in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

[0215] In some embodiments, optionally, generating the second encoded frame according to the frame structure of the second encoded frame includes:

[0216] When the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a first value. The first value is used to indicate that the encoding parameter set corresponding to the second encoded frame does not need to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0217] In some embodiments, optionally, generating the second encoded frame according to the frame structure of the second encoded frame includes:

[0218] When the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a second value, which is used to indicate that the encoding parameter set corresponding to the second encoded frame needs to be updated on the encoding parameter set corresponding to the first encoded frame.

[0219] The encoding parameter set corresponding to the second encoded frame is encapsulated in the encoding parameter set field in the frame header of the second encoded frame;

[0220] The encoded data corresponding to the second encoded frame is encapsulated within the frame payload of the second encoded frame.

[0221] For details regarding this implementation method, please refer to the above. Figure 3 The relevant content of the method embodiments shown is not described in detail here. The embodiments of this application and the above method embodiments are based on the same concept and have the same technical effects. For specific principles, please refer to the description of the above method embodiments, which will not be repeated here.

[0222] Embodiments of this application also provide a voice data transmission device. The voice data transmission device 1400 may be a processor, a chip, a chip module, or a terminal device, such as... Figure 14 As shown, the voice data transmission device 1400 includes:

[0223] The receiving module 1401 is used to receive Real-Time Transport Protocol (RTP) voice packets; wherein the payload of the RTP voice packet carries at least one second coded frame, and when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type.

[0224] The parsing module 1402 is used to parse the frame header of the second encoded frame carried in the payload of the RTP voice packet and obtain the frame header parsing result;

[0225] The acquisition module 1403 is used to obtain the audio signal corresponding to the second encoded frame based on the frame header parsing result.

[0226] In the embodiments of this application, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame may not carry the encoding parameter set, but may reuse the encoding parameter set corresponding to the first encoded frame, which can reduce the frame header overhead.

[0227] In some embodiments, optionally, the frame header of the second encoded frame includes a first flag bit, which is used to identify whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

[0228] In some embodiments, optionally, when the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

[0229] In some embodiments, optionally, when the second encoded frame is a silence insertion descriptor (SID), the first encoded frame and the second encoded frame are located in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

[0230] In some embodiments, optionally, when the parsing module 1402 obtains the frame header parsing result for the second coded frame carried within the payload of the RTP voice packet, it specifically performs the following:

[0231] The first flag bit of the frame header is parsed to obtain the parsing result of the first flag bit, wherein the frame header parsing result includes the parsing result of the first flag bit.

[0232] In some embodiments, optionally, when the acquisition module 1403 obtains the audio signal corresponding to the second coded frame based on the frame header parsing result, it is specifically used for:

[0233] If the parsing result of the first flag bit is the first value and the second encoded frame has no frame payload, the audio signal corresponding to the first encoded frame is used as the audio signal corresponding to the second encoded frame, wherein the audio signal is a silence signal.

[0234] In some embodiments, optionally, when the acquisition module 1403 obtains the audio signal corresponding to the second coded frame based on the frame header parsing result, it is specifically used for:

[0235] If the parsing result of the first flag bit is the first value, and the first coded frame includes the frame payload, the frame payload is parsed to obtain the first coded data;

[0236] Using the encoding parameter set corresponding to the first encoded frame, the first encoded data is decoded to obtain the audio signal corresponding to the second encoded frame, wherein the audio signal is a speech signal.

[0237] In some embodiments, optionally, when the parsing module 1402 obtains the frame header parsing result of the second coded frame carried in the payload of the RTP voice packet, it is also used to:

[0238] If the parsing result of the first flag bit is the second value, the encoding parameter set field of the frame header is parsed to obtain the encoding parameter set of the second encoded frame, wherein the frame header parsing result also includes the encoding parameter set of the second encoded frame.

[0239] For details regarding this implementation method, please refer to the above. Figure 12 The relevant content of the method embodiments shown is not described in detail here. The embodiments of this application and the above method embodiments are based on the same concept and have the same technical effects. For specific principles, please refer to the description of the above method embodiments, which will not be repeated here.

[0240] Please see Figure 15 , Figure 15 This is a schematic diagram of a communication device provided in an embodiment of this application. The communication device 1500 can be a terminal device or a device compatible with a terminal device, such as a processor, chip, or chip module. The communication device 1500 may include a processor 1501. Optionally, the communication device 1500 may further include a memory 1502 and a computer program or instructions stored on the memory 1502. Figure 15 (Not shown in the diagram). The processor 1501 and memory 1502 are interconnected. Optionally, the communication device 150 may further include a transceiver 1503. The processor 1501, memory 1502, and transceiver 1503 can be connected via a bus 1504 or other means. The bus is in... Figure 15The connections between other components are shown in bold lines only and are not intended to be limiting. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, Figure 15 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0241] The coupling in this application embodiment is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. This application embodiment does not limit the specific connection medium between the processor 1501, memory 1502, and transceiver 1503. Memory 1502 may include read-only memory and random access memory, and provides instructions and data to processor 1501. A portion of memory 1502 may also include non-volatile random access memory.

[0242] Processor 1501 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor; optionally, processor 1501 can also be any conventional processor.

[0243] Transceiver 1503 is used to receive or send data.

[0244] In one implementation, memory 1502 is used to store computer programs or instructions; processor 1501 is used to call the computer programs or instructions stored in memory 1502 for execution. Figure 3 and Figure 12 The steps performed by the first communication device and the second communication device in the corresponding method embodiment.

[0245] In the embodiments of this application, the methods provided in the embodiments of this application can be implemented by running a computer program (including program code or instructions) capable of performing the steps involved in the above-described methods on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a CPU, random access memory (RAM), and read-only memory (ROM). The computer program or instructions can be recorded on, for example, a computer-readable recording medium, loaded into the aforementioned computing device through the computer-readable recording medium, and executed therein.

[0246] Based on the same inventive concept, the communication device 1500 provided in the embodiments of this application solves the problem in the same way and with the same beneficial effects as this application. Figure 3 and Figure 12 The principles and beneficial effects of solving the problem in the illustrated embodiments are similar. Please refer to the implementation principles and beneficial effects of the method. For the sake of brevity, they will not be repeated here.

[0247] The aforementioned communication device may be, for example, a chip or a chip module.

[0248] This application also provides a chip, which includes a processor capable of executing the steps of the first or second communication device in the foregoing method embodiments. The specific implementation of the first or second communication device can be found in the description of the relevant content in the foregoing method embodiments, and will not be repeated here.

[0249] In one alternative implementation, the chip further includes at least one first memory and at least one second memory; the at least one first memory and the processor are interconnected via a circuit, and the first memory stores instructions; the at least one second memory and the processor are interconnected via a circuit, and the second memory stores data that needs to be stored in the above method embodiments.

[0250] Please see Figure 16 , Figure 16 This is a schematic diagram of the structure of a chip module provided in an embodiment of this application. The chip module 1600 can execute the relevant steps of the voice data transmission method executed by the first communication device or the second communication device in the aforementioned method embodiments. The chip module 1600 includes: a communication interface 1601 and a chip 1602.

[0251] The communication interface 1601 is used for internal communication within the chip module, or for communication between the chip module and external devices. The communication interface 1601 can also be described as a communication module. The chip 1602 includes a processor (…). Figure 16(Not shown in the image). Chip 1602 is used to implement the functions of the first communication device or the second communication device in the embodiments of this application, that is, the processor of chip 1602 is used to execute the relevant steps of the first communication device or the second communication device in the foregoing method embodiments. The specific implementation of the first communication device or the second communication device can be referred to the description of the relevant content in the foregoing method embodiments, and will not be repeated here.

[0252] Optionally, chip 1602 may also include memory ( Figure 16 (not shown in the image) and computer programs or instructions stored in memory (not shown in the image) Figure 16 (Not shown in the diagram), the processor executes the computer program or instructions to implement the relevant steps performed by the first communication device or the second communication device as described in the above method embodiments. Specific implementations of the first communication device or the second communication device can be found in the description of the relevant content in the foregoing method embodiments, and will not be repeated here.

[0253] Optionally, the chip 1602 and the communication interface 1601 are interconnected via a line; through the communication interface 1601, the chip module 1600 can exchange data with other chip modules, other terminal devices, servers and other modules or devices.

[0254] Optionally, the chip module 1600 may also include a storage module 1603 and a power module 1604. The storage module 1603 is used to store data and instructions. The power module 1604 is used to provide power to the chip module.

[0255] For various devices and products applied to or integrated into chip modules, each of its modules can be implemented using hardware methods such as circuits. Different modules can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module. Alternatively, at least some modules can be implemented using software programs that run on the processor integrated inside the chip module, while the remaining (if any) modules can be implemented using hardware methods such as circuits.

[0256] This application also provides a computer-readable storage medium storing a computer program or instructions. When the computer program or instructions are executed, for example, when the computer program or instructions are executed by a processor or computer, the method flow of the method embodiment executed by the first or second communication device described above will be implemented. Specific implementations of the first or second communication device can be found in the descriptions of the relevant content in the foregoing embodiments, and will not be repeated here. It is understood that the computer storage medium here may include the built-in storage medium in the first or second communication device, or it may include extended storage media supported by the first or second communication device. The computer storage medium provides storage space, which stores the operating system of the first or second communication device. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer storage medium here may be a high-speed RAM memory, or a non-volatile memory, such as at least one disk storage device, or Flash memory; optionally, it may also be at least one computer storage medium located remotely from the aforementioned processor. The specific implementation of the first or second communication device can be found in the description of the relevant content in the foregoing method embodiments, and will not be repeated here.

[0257] This application also provides a computer program product, including a computer program or instructions, which, when executed, such as when the computer program or instructions are executed by a processor or computer, cause the processor or computer to perform the method flow of the method embodiment described above for the first communication device or the second communication device.

[0258] This application provides a communication system that may include a first communication device that performs the method described in the above method embodiments, and a second communication device that performs the method described in the above method embodiments.

[0259] It should be noted that, for the sake of simplicity, the above embodiments are all described as a series of actions. Those skilled in the art should understand that this application is not limited to the described order of actions, as some steps in the embodiments of this application can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of this application.

[0260] In the above embodiments, the descriptions of each embodiment in this application have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0261] The steps of the methods or algorithms described in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, read-only optical discs (CD-ROMs), or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Additionally, the ASIC can reside in a network device or terminal device. Alternatively, the processor and storage medium can exist as discrete components in the network device or terminal device.

[0262] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0263] Regarding the modules / units included in the various devices and products described in the above embodiments, they can be software modules / units, hardware modules / units, or a combination of both. For example, for devices and products applied to or integrated into a chip, all modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs running on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits. For devices and products applied to or integrated into a chip module, all modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The implementation is achieved through a software program that runs on a processor integrated within the chip module. The remaining modules / units (if any) can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into terminal equipment, each of their modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components within the terminal equipment. Alternatively, at least some modules / units can be implemented using a software program that runs on a processor integrated within the terminal equipment, while the remaining modules / units (if any) can be implemented using hardware methods such as circuits.

[0264] The above detailed embodiments further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this application. It should be understood that the above are merely specific embodiments of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A voice data transmission method, characterized in that, The method includes: The frame structure of the second encoded frame is determined based on the encoding parameter set corresponding to the first encoded frame; wherein, if the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the field carrying the encoding parameter set; the second encoded frame and the first encoded frame are encoded frames of the same type; The second encoded frame is generated according to the frame structure of the second encoded frame; Transmit Real-Time Transport Protocol (RTP) voice packets, wherein the payload of the RTP voice packets includes at least one of the second encoded frames.

2. The method according to claim 1, characterized in that, The header of the second encoded frame includes a first flag bit, which is used to indicate whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

3. The method according to claim 1 or 2, characterized in that, In the case where the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

4. The method according to claim 1 or 2, characterized in that, When the second encoded frame is a silence insertion descriptor (SID), the first encoded frame and the second encoded frame are in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

5. The method according to claim 2, characterized in that, Generating the second encoded frame according to the frame structure of the second encoded frame includes: When the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a first value. The first value is used to indicate that the encoding parameter set corresponding to the second encoded frame does not need to be updated based on the encoding parameter set corresponding to the first encoded frame.

6. The method according to claim 2, characterized in that, Generating the second encoded frame according to the frame structure of the second encoded frame includes: When the encoding parameter set corresponding to the second encoded frame is different from the encoding parameter set corresponding to the first encoded frame, the first flag bit is configured to a second value, which is used to indicate that the encoding parameter set corresponding to the second encoded frame needs to be updated on the encoding parameter set corresponding to the first encoded frame. The encoding parameter set corresponding to the second encoded frame is encapsulated in the encoding parameter set field in the frame header of the second encoded frame; The encoded data corresponding to the second encoded frame is encapsulated within the frame payload of the second encoded frame.

7. A method for transmitting voice data, characterized in that, The method includes: Receive Real-Time Transport Protocol (RTP) voice packets; wherein the payload of the RTP voice packets carries at least one second coded frame, and when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type; The header of the second coded frame carried in the payload of the RTP voice packet is parsed to obtain the header parsing result; Based on the frame header parsing result, the audio signal corresponding to the second encoded frame is obtained.

8. The method according to claim 7, characterized in that, The header of the second encoded frame includes a first flag bit, which is used to indicate whether the encoding parameter set corresponding to the second encoded frame needs to be updated based on the encoding parameter set corresponding to the first encoded frame.

9. The method according to claim 7 or 8, characterized in that, In the case where the second encoded frame is a speech frame, the first encoded frame is the preceding encoded frame adjacent to the second encoded frame.

10. The method according to claim 7 or 8, characterized in that, When the second encoded frame is a silence insertion descriptor (SID), the first encoded frame and the second encoded frame are in the same SID update cycle, and the frame header of the first encoded frame includes a field carrying a set of encoded parameters, wherein the cycle length of the SID update cycle is the same as the cycle length of the voice binding cycle.

11. The method according to claim 8, characterized in that, Parse the header of the second coded frame carried within the payload of the RTP voice packet to obtain the header parsing result, including: The first flag bit of the frame header is parsed to obtain the parsing result of the first flag bit, wherein the frame header parsing result includes the parsing result of the first flag bit.

12. The method according to claim 11, characterized in that, Based on the frame header parsing result, the audio signal corresponding to the second encoded frame is obtained, including: If the parsing result of the first flag bit is the first value and the second encoded frame has no frame payload, the audio signal corresponding to the first encoded frame is used as the audio signal corresponding to the second encoded frame, wherein the audio signal is a silence signal.

13. The method according to claim 11, characterized in that, Based on the frame header parsing result, the audio signal corresponding to the second encoded frame is obtained, including: If the parsing result of the first flag bit is a first value, and the first encoded frame includes a frame payload, the frame payload is parsed to obtain the first encoded data; Using the encoding parameter set corresponding to the first encoded frame, the first encoded data is decoded to obtain the audio signal corresponding to the second encoded frame, wherein the audio signal is a speech signal.

14. The method according to claim 11, characterized in that, Parsing the frame header of the second coded frame carried within the payload of the RTP voice packet to obtain the frame header parsing result also includes: If the parsing result of the first flag bit is the second value, the encoding parameter set field of the frame header is parsed to obtain the encoding parameter set of the second encoded frame, wherein the frame header parsing result also includes the encoding parameter set of the second encoded frame.

15. A voice data transmission device, characterized in that, The device includes: The determining module is used to determine the frame structure of the second encoded frame based on the encoding parameter set corresponding to the first encoded frame; wherein, when the encoding parameter set corresponding to the second encoded frame is the same as the encoding parameter set corresponding to the first encoded frame, the frame header of the second encoded frame does not include the field carrying the encoding parameter set; the second encoded frame and the first encoded frame are encoded frames of the same type; The generation module is used to generate the second encoded frame according to the frame structure of the second encoded frame; A transmission module is used to transmit Real-Time Transport Protocol (RTP) voice packets, wherein the payload of the RTP voice packets includes at least one second encoded frame.

16. A voice data transmission device, characterized in that, The device includes: A receiving module is used to receive Real-Time Transport Protocol (RTP) voice packets; wherein the payload of the RTP voice packet carries at least one second coded frame, and when the encoding parameter set corresponding to the second coded frame is the same as the encoding parameter set corresponding to the first coded frame, the frame header of the second coded frame does not include the field carrying the encoding parameter set; the second coded frame and the first coded frame are coded frames of the same type; The parsing module is used to parse the frame header of the second encoded frame carried in the payload of the RTP voice packet and obtain the frame header parsing result; The acquisition module is used to obtain the audio signal corresponding to the second encoded frame based on the frame header parsing result.

17. A communication device, characterized in that, The communication device includes a processor and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor invokes the program instructions to execute the method as described in any one of claims 1 to 6, or to execute the method as described in any one of claims 7 to 14.

18. A chip, characterized in that, The chip includes a processor and an interface, the processor and the interface being coupled; the interface is used to receive or output signals, and the processor is used to execute code instructions to perform the method as described in any one of claims 1 to 6, or to perform the method as described in any one of claims 7 to 14.

19. A module device, characterized in that, The module device includes a communication module, a power module, a storage module, and a chip module, wherein: the power module is used to provide power to the module device; the storage module is used to store data and / or instructions; the communication module is used to communicate with external devices; and the chip module is used to call the data and / or instructions stored in the storage module, and in conjunction with the communication module, to perform the method as described in any one of claims 1 to 6, or to perform the method as described in any one of claims 7 to 14.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a computer, implement the steps of the voice data transmission method as described in any one of claims 1 to 6, or, when executed by the processor, implement the steps of the voice data transmission method as described in any one of claims 7 to 14.

21. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the voice data transmission method according to any one of claims 1 to 6; or, when the program or instruction is executed by the processor, it implements the steps of the voice data transmission method according to any one of claims 7 to 14.