Audio processing method and device, electronic equipment, terminal, chip and storage medium

In VoLTE audio processing, voice frames are obtained from the storage area and playback is determined based on the effectiveness and delay results, and the acceleration or deceleration algorithm is used to adjust the playback sequence, the problem of long audio playback delay is solved, and the continuity and quality of the audio are improved.

CN120375840APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410544704.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, VoLTE audio playback delay is long, affecting the continuity and quality of audio playback.

Method used

By acquiring the voice frame from the first storage area, the validity and delay of the voice frame are determined, and whether to play the voice frame is determined based on the results, the playback sequence is adjusted using an acceleration or deceleration algorithm to reduce the end-to-end communication delay.

Benefits of technology

Effectively reduce the delay in audio playback and improve the continuity and quality of audio playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375840A_ABST
    Figure CN120375840A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and device, electronic equipment, a terminal, a chip and a storage medium, and the method comprises the steps: obtaining a first voice frame from a first storage region; determining whether each voice frame acquired from the first storage area in the first time period is a valid voice frame to obtain a first result, and / or determining whether the number of invalid voice frames acquired from the first storage area in the second time period is smaller than a preset number to obtain a second result, and / or determining whether the maximum time delay of the effective voice frames in the first time period or the second time period is greater than a time delay threshold value or not, and obtaining a third result; and determining whether to play the first voice frame according to the first result and / or the second result and / or the third result. The technical problems that the audio playing time delay is long, the audio playing continuity is reduced, and the audio quality is affected in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technologies, and in particular, to an audio processing method, apparatus, electronic device, terminal, chip, and storage medium. Background Art

[0002] Voice over Long-Term Evolution (VoLTE) is a data transmission technology based on the Internet Protocol (IP). The audio quality based on VoLTE is mainly affected by factors such as latency, packet loss rate, and jitter. In related technologies, a jitter buffer is set at the receiving end of voice data packets to reduce latency jitter. When playing back a voice data packet popped out from the Jitter Buffer, it is necessary to decode the voice data packet and perform processing such as sorting on the decoded voice frames.

[0003] In this way, the audio playback latency is relatively long, the continuity of audio playback decreases, and the audio quality is affected. Summary of the Invention

[0004] The present disclosure aims to solve one of the technical problems in related technologies to a certain extent.

[0005] To this end, the present disclosure provides an audio processing method, apparatus, electronic device, terminal, chip, and storage medium, which can effectively shorten the audio playback latency, improve the continuity of audio playback, and improve the audio quality.

[0006] To achieve the above object, an embodiment of the first aspect of the present disclosure provides an audio processing method, including: obtaining a first voice frame from a first storage area, where the first storage area is used to cache at least some voice frames, the first voice frame belongs to at least some voice frames, and at least some voice frames belong to a plurality of voice frames obtained by decoding at least one packet payload; determining whether each voice frame obtained from the first storage area within a first time period is a valid voice frame to obtain a first result, and / or determining whether the number of invalid voice frames obtained from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determining whether the maximum latency of valid voice frames within the first time period or the second time period is greater than a latency threshold to obtain a third result; and determining whether to play the first voice frame according to the first result and / or the second result and / or the third result.

[0007] To achieve the above object, an embodiment of the second aspect of the present disclosure provides an audio processing device, including: an acquisition module, configured to acquire a first speech frame from a first storage area, where the first storage area is used to cache at least part of speech frames, the first speech frame belongs to at least part of speech frames, and at least part of speech frames belong to multiple speech frames obtained by decoding at least one packet payload; a first determination module, configured to determine whether each speech frame acquired from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determine whether the number of invalid speech frames acquired from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determine whether the maximum delay of valid speech frames within the first time period or the second time period is greater than a delay threshold to obtain a third result; and a second determination module, configured to determine whether to play the first speech frame according to the first result and / or the second result and / or the third result.

[0008] To achieve the above object, an embodiment of the third aspect of the present disclosure provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the audio processing method as provided in the embodiment of the first aspect of the present disclosure.

[0009] To achieve the above object, an embodiment of the fourth aspect of the present disclosure provides a chip, where the chip includes a processing circuit configured to execute the audio processing method as provided in the embodiment of the first aspect of the present disclosure.

[0010] To achieve the above object, an embodiment of the fifth aspect of the present disclosure provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described above.

[0011] The audio processing method, device, electronic device, terminal, chip, and storage medium provided by the present disclosure can effectively reduce the audio playback delay, improve the continuity of audio playback, and improve the audio quality by acquiring a first speech frame from a first storage area, where the first storage area is used to cache at least part of speech frames, the first speech frame belongs to at least part of speech frames, and at least part of speech frames belong to multiple speech frames obtained by decoding at least one packet payload; determining whether each speech frame acquired from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determining whether the number of invalid speech frames acquired from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determining whether the maximum delay of valid speech frames within the first time period or the second time period is greater than a delay threshold to obtain a third result; and determining whether to play the first speech frame according to the first result and / or the second result and / or the third result.

[0012] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Brief Description of the Drawings

[0013] The above-mentioned and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0014] Figure 1 is a schematic structural diagram of a communication system shown according to an embodiment of the present disclosure;

[0015] Figure 2 is a schematic diagram of an application scenario in an embodiment of the present disclosure;

[0016] Figure 3 is a schematic flowchart of an audio processing method provided by an embodiment of the present disclosure;

[0017] Figure 4 is a schematic diagram of a receiving end processing voice data packets in an embodiment of the present disclosure;

[0018] Figure 5 is a schematic flowchart of another audio processing method provided by an embodiment of the present disclosure;

[0019] Figure 6 is a schematic diagram of the application of an acceleration algorithm in an embodiment of the present disclosure;

[0020] Figure 7 is a schematic flowchart of another audio processing method provided by an embodiment of the present disclosure;

[0021] Figure 8 is a schematic diagram of the application of a deceleration algorithm in an embodiment of the present disclosure;

[0022] Figure 9 is a schematic structural diagram of an audio processing device provided by an embodiment of the present disclosure;

[0023] Figure 10 shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure;

[0024] Figure 11 is a schematic structural diagram of a chip proposed in an embodiment of the present disclosure. Detailed Embodiments

[0025] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure and should not be construed as limiting the present disclosure.

[0026] Figure 1 is a schematic architecture diagram of a communication system shown according to an embodiment of the present disclosure. As Figure 1 shown, the communication system 100 may include a terminal 101 and a network device 102. The network device 102 may include at least one of an access network device and a core network device.

[0027] In some embodiments, the terminal 101 includes, for example, at least one of a mobile phone, a wearable device, an Internet of Things device, a vehicle with communication function, a smart vehicle, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, and a wireless terminal in smart home, but is not limited thereto.

[0028] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include at least one of an evolved NodeB (eNB) in a 5G communication system, a next-generation evolved NodeB (ng-eNB), a next-generation NodeB (gNB), a NodeB (NB), a home NodeB (HNB), a home evolved NodeB (HeNB), a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a 6G communication system, an Open RAN, a Cloud RAN, a base station in other communication systems, and an access node in a WiFi system, but is not limited thereto.

[0029] In some embodiments, the technical solutions of the present disclosure can be applied to the Open RAN architecture. At this time, the interfaces between access network devices or within access network devices involved in the embodiments of the present disclosure can become the internal interfaces of Open RAN, and the processes and information interactions between these internal interfaces can be implemented through software or programs.

[0030] In some embodiments, the access network device can be composed of a central unit (CU) and a distributed unit (DU). Among them, the CU can also be called a control unit. Adopting the CU-DU structure can split the protocol layer of the access network device. The functions of some protocol layers are centrally controlled by the CU, and the functions of the remaining part or all protocol layers are distributed in the DU. The CU centrally controls the DU, but is not limited thereto.

[0031] In some embodiments, the core network device can be a single device including one or more network elements, or multiple devices or device groups, each including all or part of one or more network elements. The network elements can be virtual or physical. The core network, for example, includes at least one of the Evolved Packet Core (EPC), 5G Core Network (5GCN), and Next Generation Core (NGC).

[0032] It can be understood that the communication system described in the embodiments of the present disclosure is to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and does not constitute a limitation on the technical solutions proposed in the embodiments of the present disclosure. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new service scenarios, the technical solutions proposed in the embodiments of the present disclosure are equally applicable to similar technical problems.

[0033] The following embodiments of the present disclosure can be applied to Figure 1 the communication system 100 shown, or a part of the main body, but is not limited thereto. Figure 1 The main bodies shown are examples. The communication system can include Figure 1 all or part of the main bodies, or can include Figure 1 other main bodies outside. The number and form of each main body are arbitrary. The connection relationships between the main bodies are examples. The main bodies can be not connected or connected. Their connection can be in any way, either directly connected or indirectly connected, either wired or wireless.

[0034] Embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, the 4th generation mobile communication system (4G), the 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) network, Device-to-Device (D2D) system, Machine to Machine (M2M) system, Internet of Things (IoT) system, Vehicle-to-Everything (V2X), systems using other communication methods, next-generation systems extended based on them, etc. In addition, multiple systems can be combined (for example, a combination of LTE or LTE-A and 5G, etc.) and applied.

[0035] The audio processing method in the embodiments of the present disclosure can be applied in an electronic device, such as a network device or a terminal. Exemplarily, it can be specifically applied in an electronic device that receives and processes audio, and there is no limitation thereto.

[0036] As Figure 2 shown, Figure 2It is a schematic diagram of an application scenario in an embodiment of the present disclosure. The sending end 21 can send one or more voice data packets to the receiving end 23 through a network device (base station 22). After the receiving end 23 receives one or more voice data packets, it uses the audio processing method provided in the embodiment of the present disclosure to play each voice frame carried by the voice data packet.

[0037] The following describes the audio processing method, apparatus, electronic device, terminal, chip, and storage medium according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0038] Figure 3 It is a schematic flowchart of an audio processing method provided by an embodiment of the present disclosure.

[0039] In this embodiment, the audio processing method is configured in an audio processing apparatus as an example. In this embodiment, the audio processing method can be configured in the audio processing apparatus, and the audio processing apparatus can be set in an electronic device. Specifically, for example, it is set in a terminal. It should be noted that the execution subject of the embodiment of the present disclosure can be, for example, the central processing unit (CPU) of the electronic device in terms of hardware, and can be, for example, the relevant background service of the electronic device in terms of software, and this is not limited.

[0040] As Figure 3 shown, the audio processing method includes the following steps:

[0041] Step S301, obtain a first voice frame from a first storage area, where the first storage area is used to cache at least part of the voice frames, the first voice frame belongs to at least part of the voice frames, and at least part of the voice frames belong to multiple voice frames obtained by decoding the payloads of at least one data packet.

[0042] Among them, the first storage area can be, for example, a jitter buffer (JB). In an embodiment of the present disclosure, the data structure used by the first storage area to store voice frames can be a dynamically allocated linked list, and the length of the linked list can be dynamically allocated according to the size of the data packet. In this allocation method, the depth (size) of the jitter buffer can be dynamically changed. Therefore, in the embodiment of the present disclosure, it is possible to dynamically allocate memory according to the actual size of the received data packet, improving the flexibility of voice frame caching.

[0043] Taking the first storage area as a jitter buffer as an example, as Figure 4 shown, Figure 4It is a schematic diagram of the receiving end processing voice data packets in the embodiments of the present disclosure. Taking the voice data packet as a Real-time Transport Protocol (RTP) packet (RTP packets) for example, the RTP Depacker is a tool or software component for processing RTP packets. The modem can push the RTP packets to the RTP Depacker, and the RTP Depacker can decode the voice data packets to obtain multiple voice frames, and push the multiple voice frames to the AMR / AMR-WB / EVS Decoder (AMR / AMR-WB / EVS decoder, which is a software or hardware component for decoding audio data encoded by Adaptive Multi-Rate (AMR), the wideband version of AMR (Adaptive Multi-Rate Wideband, AMR-WB), and Enhanced Voice Services (EVS)). The Jitter Buffer Management (JBM) caches each voice frame into the Jitter Buffer (JB) based on a linked list. The Acoustic Frontend can pull the Pulse Code Modulation (PCM) buffers containing each voice frame from the JB and play each voice frame.

[0044] Among them, the currently read voice frame from the first storage area can be referred to as the first voice frame.

[0045] In some embodiments, at least some of the multiple voice frames obtained by decoding at least one packet payload can be cached in the first storage area. During the audio processing, the voice frames can be read one by one in sequence from the first storage area and the read voice frames can be played, and there is no limitation on this.

[0046] The voice frames cached in the first storage area in the embodiments of the present disclosure can be the voice frames obtained by decoding the packet payloads of the voice data packets. Thus, since at least one packet payload has been decoded and at least some of the decoded voice frames have been cached in the first storage area, when reading each voice frame from the first storage area, there is no need to perform decoding processing again, and it can support directly playing each read voice frame, thereby effectively reducing the audio playback delay.

[0047] Step S302: Determine whether each voice frame obtained from the first storage area within the first time period is a valid voice frame to obtain a first result, and / or determine whether the number of invalid voice frames obtained from the first storage area within the second time period is less than a preset number to obtain a second result, and / or determine whether the maximum time delay of the valid voice frames within the first time period or the second time period is greater than a time delay threshold to obtain a third result.

[0048] Among them, multiple voice frames are stored in the first storage area, and each voice frame corresponds to a time delay. The maximum time delay among the multiple time delays corresponding to the multiple valid voice frames can be referred to as the maximum time delay of the valid voice frames. In addition, the magnitude of the time delay of the voice frames in the first storage area is related to the caching order. The time delay of the valid voice frames with a higher caching order is smaller, and the time delay of the valid voice frames with a lower caching order is larger. Therefore, the maximum time delay of the valid voice frames is the time delay of the last valid voice frame stored in the first storage area. The last valid voice frame can also be referred to as the penultimate valid voice frame, and there is no restriction on this.

[0049] Among them, a valid voice frame refers to a voice frame containing valid voice information. An invalid voice frame refers to a voice frame that does not contain valid voice information.

[0050] Among them, the first result indicates whether each voice frame obtained from the first storage area within the first time period is a valid voice frame. The first time period is, for example, 1 second. By way of example, it can be determined whether each voice frame obtained from the first storage area within each second is a valid voice frame. Assuming that 5 voice frames can be read from the first storage area within 1 second, then analyze whether these 5 voice frames are all valid voice frames.

[0051] Among them, the second result indicates whether the number of invalid voice frames obtained from the first storage area within the second time period is less than a preset number. The second time period is, for example, 0.2 seconds. By way of example, it can be determined whether the number of invalid voice frames obtained from the first storage area within every 0.2 seconds is less than a preset number (such as 5). Assuming that 10 voice frames can be read from the first storage area within 0.2 seconds, then analyze whether there are less than 5 invalid voice frames among these 10 voice frames.

[0052] Among them, the third result indicates whether the maximum time delay of the valid voice frames within the first time period or the second time period is greater than a time delay threshold.

[0053] Optionally, in some embodiments, a latency threshold may be determined according to the depth of the first storage area, thereby effectively improving the accuracy and flexibility of audio processing. The depth of the first storage area can be used to represent the number of voice frames that can be cached simultaneously in the first storage area. By way of example, the latency threshold may be related to the depth (size) of the first storage area. For example, the depth (size) of the first storage area may be used as the latency threshold.

[0054] In the embodiments of the present disclosure, the depth of the first storage area may be dynamically changed. For example, the depth of the first storage area may be adaptively adjusted according to actual audio processing requirements.

[0055] In some embodiments, when reading the first voice frame each time, the first result and / or the second result and / or the third result may be determined. Then, based on the first result and / or the second result and / or the third result, it may be analyzed whether to play the first voice frame. When analyzing whether to play the first voice frame based on the first result and / or the second result and / or the third result, it can effectively support subsequent corresponding processing measures to avoid an increase in the playback latency of the voice frames after the first voice frame, and can also effectively improve the overall continuity of audio playback.

[0056] Step S303: Determine whether to play the first voice frame according to the first result and / or the second result and / or the third result.

[0057] After obtaining the first voice frame and determining the first result and / or the second result and / or the third result as described above, it may be determined whether to play the first voice frame according to the first result and / or the second result and / or the third result. By way of example, various results obtained as described above may be analyzed in an artificial intelligence-based manner to determine whether to play the first voice frame currently; or it may also be determined whether to play the first voice frame currently based on a configured manner and with reference to various results obtained as described above; or any other possible manner may also be used to implement determining whether to play the first voice frame according to the first result and / or the second result and / or the third result, and this is not limited.

[0058] In this embodiment, by obtaining a first speech frame from a first storage area, where the first storage area is used to cache at least some speech frames, the first speech frame belongs to at least some speech frames, and at least some speech frames belong to a plurality of speech frames obtained by decoding at least one data packet payload; determining whether each speech frame obtained from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determining whether the number of invalid speech frames obtained from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determining whether the maximum delay of valid speech frames within the first time period or the second time period is greater than a delay threshold to obtain a third result; and determining whether to play the first speech frame according to the first result and / or the second result and / or the third result, it is possible to effectively reduce the audio playback delay, improve the continuity of audio playback, and improve the audio quality.

[0059] In some embodiments of the present disclosure, before obtaining the first speech frame from the first storage area, at least one speech data packet may be received, the data packet payload may be extracted from the at least one speech data packet, and the data packet payload may be decoded to obtain a plurality of speech frames. Each speech frame has a corresponding caching order. According to the caching order, a caching position is allocated for the speech frame in the first storage area, and each speech frame is cached at the corresponding caching position in the first storage area. Thus, decoding of one or more received speech data packets is realized, the data packet payload of the speech data packet is extracted, and each decoded speech frame is cached in the first storage area according to the caching order, which can support that when reading each speech frame from the first storage area, there is no need to decode and sort again, thereby greatly improving the audio processing efficiency.

[0060] In some embodiments of the present disclosure, the caching order is related to the sequence number of the speech data packet to which the speech frame belongs and / or the transmission timestamp of the speech frame. By way of example, a smaller sequence number may be assigned to the speech data packet received first, and a larger sequence number may be assigned to the speech data packet received later. Thus, the caching order is determined in ascending order of the sequence number. The caching order of the speech frames of the speech data packet with a smaller sequence number is in the front, and the caching order of the speech frames of the speech data packet with a larger sequence number is in the back. For multiple speech frames belonging to the same speech data packet (since these multiple speech frames belong to the same speech data packet, they have the same sequence number), the caching order is allocated according to the sequence of the transmission timestamps. The caching order of the speech frame generated first is in the front, and the caching order of the speech frame generated later is in the back.

[0061] In some embodiments of the present disclosure, when the first result indicates that at least one speech frame obtained from the first storage area within the first time period is not a valid speech frame, it is determined to play the first speech frame. This effectively improves the timeliness and effectiveness of playing the first speech frame. Additionally, if an invalid speech frame or a second speech frame (for the definition and acquisition timing of the second speech frame, refer to the following embodiments) is obtained from the first storage area within the first time period, the first time period can be updated, and the audio playback processing within the next first time period can be initiated.

[0062] In some embodiments of the present disclosure, when the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is less than a preset number, it is determined to play the first speech frame. In this case, it indicates that the continuity of the speech frames is relatively good, so it can be determined to play the first speech frame, thereby accurately analyzing the timing of playing the first speech frame.

[0063] In some embodiments of the present disclosure, when the third result indicates that the maximum time delay of valid speech frames within the first time period is less than or equal to the time delay threshold, it is determined to play the first speech frame. In this case, it indicates that the time delay of the speech frames subsequent to the first speech frame is relatively good, so it can be determined to play the first speech frame, thereby accurately analyzing the timing of playing the first speech frame.

[0064] In some embodiments of the present disclosure, when the third result indicates that the maximum time delay of valid speech frames within the second time period is greater than the time delay threshold, it is determined to play the first speech frame. This effectively improves the timeliness and effectiveness of playing the first speech frame.

[0065] In some embodiments of the present disclosure, when including at least two of the following combinations, it can also be determined to play the first speech frame: the first result indicates that at least one speech frame obtained from the first storage area within the first time period is not a valid speech frame; the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is less than a preset number; the third result indicates that the maximum time delay of valid speech frames within the first time period is less than or equal to the time delay threshold; the third result indicates that the maximum time delay of valid speech frames within the second time period is greater than the time delay threshold. This greatly improves the timeliness, effectiveness, and flexibility of playing the first speech frame, and can effectively be applied to personalized audio processing scenarios.

[0066] Figure 5 It is a schematic flowchart of another audio processing method provided by the embodiments of the present disclosure.

[0067] As Figure 5 shown, this audio processing method includes the following steps:

[0068] Step S501: Obtain a first speech frame from a first storage area, where the first storage area is used to cache at least some speech frames, the first speech frame belongs to at least some speech frames, and at least some speech frames belong to multiple speech frames obtained by decoding the payloads of at least one data packet.

[0069] Step S502: Determine whether each speech frame obtained from the first storage area within a first time period is a valid speech frame to obtain a first result, and determine whether the maximum time delay of the valid speech frames within the first time period is greater than a time delay threshold to obtain a third result.

[0070] For the descriptions of steps S501 - S502, specific reference can be made to the above descriptions, and no limitation is imposed herein.

[0071] Step S503: When the first result indicates that each speech frame obtained from the first storage area within the first time period is a valid speech frame, and the third result indicates that the maximum time delay of the valid speech frames within the first time period is greater than the time delay threshold, determine not to play the first speech frame.

[0072] In some embodiments, if the first result indicates that each speech frame obtained from the first storage area within the first time period is a valid speech frame, and the third result indicates that the maximum time delay of the valid speech frames within the first time period is greater than the time delay threshold, this indicates that the speech frames cached in the first storage area are continuous, but the end - to - end communication time delay of some speech frames has increased significantly. At this time, it can be determined not to play the first speech frame and take corresponding countermeasures to reduce the end - to - end communication time delay of some speech frames.

[0073] Step S504: Discard the first speech frame.

[0074] Step S505: Obtain a second speech frame, where the first speech frame and the second speech frame are different.

[0075] In some embodiments, the first speech frame can be discarded and a second speech frame different from the first speech frame can be obtained.

[0076] In some embodiments, the next speech frame of the first speech frame can be obtained from the first storage area and used as the second speech frame. The next speech frame refers to the speech frame whose cache position in the first storage area is after the first speech frame and adjacent to the cache position of the first speech frame.

[0077] The above - mentioned discarding the first speech frame, obtaining the next speech frame of the first speech frame from the first storage area, and using the next speech frame as the second speech frame can be referred to as an acceleration algorithm. And since the acceleration algorithm can be executed once within the first time period, while effectively ensuring the audio quality, it can effectively reduce the end - to - end communication time delay of some speech frames.

[0078] Step S506, play the second voice frame.

[0079] Taking the first storage area as the jitter buffer for example, as Figure 6 shown, Figure 6 is a schematic diagram of the application of the acceleration algorithm in an embodiment of the present disclosure. The first voice frame can be, for example, Figure 6 the current voice frame in Figure 6 All the voice frames shown in are valid voice frames, and if the time delays of voice frames n+3, n+4, and n+5 are greater than the time delay threshold (the size (depth) of the jitter buffer), it can be indicated that within the first time period (for example, 1 second), valid voice frames can be continuously fetched from the jitter buffer for playback, and the time delays of some voice frames exceed the size (depth) of the jitter buffer. Then, the acceleration algorithm can be applied, and the following operation can be performed once within this first time period: discard the current voice frame n, fetch the next voice frame n+1 (an optional example of the second voice frame) for playback, and then enter the acceleration operation for the next first time period.

[0080] In this embodiment, by obtaining the first voice frame from the first storage area, where the first storage area is used to cache at least some voice frames, the first voice frame belongs to at least some voice frames, and at least some voice frames belong to multiple voice frames obtained by decoding at least one packet payload. When the first result indicates that all the voice frames obtained from the first storage area within the first time period are valid voice frames, and the third result indicates that the maximum time delay of the valid voice frames within the first time period is greater than the time delay threshold, it is determined not to play the first voice frame, the first voice frame is discarded, the second voice frame is obtained, and the second voice frame is played. It can effectively reduce the end-to-end communication time delay of some voice frames while ensuring audio continuity, thereby supporting shortening the overall audio playback time delay.

[0081] Figure 7 is a schematic flowchart of another audio processing method provided by an embodiment of the present disclosure.

[0082] As Figure 7 shown, this audio processing method includes the following steps:

[0083] Step S701, obtain a first voice frame from a first storage area, where the first storage area is used to cache at least some voice frames, the first voice frame belongs to at least some voice frames, and at least some voice frames belong to multiple voice frames obtained by decoding at least one packet payload.

[0084] Step S702, determine whether the number of invalid voice frames obtained from the first storage area within the second time period is less than a preset number to obtain a second result, and determine whether the maximum time delay of the valid voice frames within the second time period is greater than the time delay threshold to obtain a third result.

[0085] For the descriptions of steps S701 - S702, specific reference can be made to the above descriptions, and no limitations are imposed here.

[0086] Step S703: When the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is greater than or equal to a preset number, and the third result indicates that the maximum time delay of the valid speech frames within the second time period is less than or equal to the time delay threshold, it is determined not to play the first speech frame.

[0087] In some embodiments, if the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is greater than or equal to a preset number, and the third result indicates that the maximum time delay of the valid speech frames within the second time period is less than or equal to the time delay threshold, it indicates that there may be link packet loss or large jitter, resulting in some speech data packets not arriving at the receiving end in sequence. At this time, it can be determined not to play the first speech frame and corresponding countermeasures can be taken to effectively improve the continuity of audio playback while ensuring the end - to - end playback time delay.

[0088] Step S704: Obtain a second speech frame, where the first speech frame and the second speech frame are different.

[0089] In some embodiments, a second speech frame different from the first speech frame can be obtained.

[0090] In some embodiments, the previous speech frame of the first speech frame can be obtained from the first storage area and used as the second speech frame. The previous speech frame refers to the speech frame whose caching position in the first storage area is before the first speech frame and adjacent to the caching position of the first speech frame.

[0091] The above method of obtaining the previous speech frame of the first speech frame from the first storage area and using the previous speech frame as the second speech frame can be called a deceleration algorithm (in the deceleration algorithm, the first speech frame is not discarded and the first speech frame is also read again). And since the deceleration algorithm can be executed once within the second time period, it can effectively improve the continuity of audio playback while ensuring the end - to - end playback time delay.

[0092] Step S705: Play the second speech frame.

[0093] Taking the first storage area as an example of the jitter buffer, as Figure 8 shown, Figure 8 is a schematic diagram of the application of the deceleration algorithm in an embodiment of the present disclosure. The first speech frame can be, for example, Figure 8 the current speech frame in, NO_DATA represents an invalid speech frame, as Figure 8As shown, within a second time period (e.g., 0.2 seconds), there are several times (e.g., 5 times) when no valid speech frames can be obtained from the jitter buffer (e.g., the continuously obtained 5 frames are all invalid speech frames). And, if the delay of the valid speech frame (such as the (n + 4)-th speech frame in Figure 8 is not more than the size (depth) of the jitter buffer, it indicates that there may be link packet loss or large jitter, resulting in some speech data packets not arriving at the receiving end in order. Then, the deceleration algorithm can be applied, and the following operation can be performed once within this second time period: move the pointer position for reading the current speech frame (an optional example of the first speech frame) forward by one position (as Figure 8 shown, move forward from the position of reading the (n + 2)-th speech frame to the position of reading the (n + 1)-th speech frame), so that more speech frames can be waited for caching, and then enter the deceleration operation of the next second time period.

[0094] In this embodiment, by obtaining the first speech frame from the first storage area, where the first storage area is used to cache at least part of the speech frames, the first speech frame belongs to at least part of the speech frames, and at least part of the speech frames belong to multiple speech frames obtained by decoding at least one data packet payload. When the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is greater than or equal to the preset number, and the third result indicates that the maximum delay of the valid speech frames within the second time period is less than or equal to the delay threshold, it is determined not to play the first speech frame, and the second speech frame is obtained, where the first speech frame and the second speech frame are different, and the second speech frame is played. It can effectively improve the continuity of audio playback while ensuring the end-to-end playback delay.

[0095] To implement the above embodiment, the present disclosure also proposes an audio processing device. Figure 9 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present disclosure.

[0096] As Figure 9 shown, the audio processing device 90 includes:

[0097] An obtaining module 901, configured to obtain a first speech frame from a first storage area, where the first storage area is used to cache at least part of the speech frames, the first speech frame belongs to at least part of the speech frames, and at least part of the speech frames belong to multiple speech frames obtained by decoding at least one data packet payload.

[0098] The first determination module 902 is configured to determine whether each voice frame obtained from the first storage area within the first time period is a valid voice frame, to obtain a first result, and / or to determine whether the number of invalid voice frames obtained from the first storage area within the second time period is less than a preset number, to obtain a second result, and / or to determine whether the maximum time delay of the valid voice frames within the first time period or the second time period is greater than a time delay threshold, to obtain a third result.

[0099] The second determination module 903 is configured to determine whether to play the first voice frame according to the first result and / or the second result and / or the third result.

[0100] In some embodiments of the present disclosure, the second determination module 903 is configured to perform at least one of the following:

[0101] In the case where the first result indicates that at least one voice frame obtained from the first storage area within the first time period is not a valid voice frame, determine to play the first voice frame;

[0102] In the case where the second result indicates that the number of invalid voice frames obtained from the first storage area within the second time period is less than a preset number, determine to play the first voice frame;

[0103] In the case where the third result indicates that the maximum time delay of the valid voice frames within the first time period is less than or equal to the time delay threshold, determine to play the first voice frame;

[0104] In the case where the third result indicates that the maximum time delay of the valid voice frames within the second time period is greater than the time delay threshold, determine to play the first voice frame.

[0105] In some embodiments of the present disclosure, the second determination module 903 is configured to perform:

[0106] In the case where the first result indicates that each voice frame obtained from the first storage area within the first time period is a valid voice frame, and the third result indicates that the maximum time delay of the valid voice frames within the first time period is greater than the time delay threshold, determine not to play the first voice frame.

[0107] In some embodiments of the present disclosure, the apparatus further includes:

[0108] The first processing module is configured to discard the first voice frame and obtain a second voice frame after determining not to play the first voice frame, where the first voice frame and the second voice frame are different, and play the second voice frame.

[0109] In some embodiments of the present disclosure, the first processing module is further configured to obtain the next voice frame of the first voice frame from the first storage area and use the next voice frame as the second voice frame.

[0110] In some embodiments of the present disclosure, the second determination module 903 is configured to perform:

[0111] When the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is greater than or equal to a preset number, and the third result indicates that the maximum time delay of the valid speech frames within the second time period is less than or equal to the time delay threshold, it is determined not to play the first speech frame.

[0112] In some embodiments of the present disclosure, the apparatus further includes:

[0113] The second processing module is configured to, after determining not to play the first speech frame, obtain a second speech frame, where the first speech frame and the second speech frame are different, and play the second speech frame.

[0114] In some embodiments of the present disclosure, the second processing module is further configured to obtain the speech frame preceding the first speech frame from the first storage area, and use the preceding speech frame as the second speech frame.

[0115] In some embodiments of the present disclosure, the apparatus further includes:

[0116] The third determination module is configured to determine the time delay threshold according to the depth of the first storage area.

[0117] In some embodiments of the present disclosure, the method further includes:

[0118] The transceiver module is configured to receive at least one speech data packet;

[0119] The third processing module is configured to extract the packet payload from at least one speech data packet, decode the packet payload to obtain a plurality of speech frames, where each speech frame has a corresponding caching order, allocate a caching position for the speech frame in the first storage area according to the caching order, and cache each speech frame at the corresponding caching position in the first storage area.

[0120] In some embodiments of the present disclosure, the caching order is related to the sequence number of the speech data packet to which the speech frame belongs and / or the transmission timestamp of the speech frame.

[0121] It should be noted that the foregoing explanation of the embodiments of the audio processing method also applies to the audio processing apparatus of this embodiment, and will not be elaborated here.

[0122] In this embodiment, by obtaining a first speech frame from a first storage area, where the first storage area is used to cache at least some speech frames, the first speech frame belongs to at least some speech frames, and at least some speech frames belong to multiple speech frames obtained by decoding at least one data packet payload; determining whether each speech frame obtained from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determining whether the number of invalid speech frames obtained from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determining whether the maximum time delay of the valid speech frames within the first time period or the second time period is greater than a time delay threshold to obtain a third result; and determining whether to play the first speech frame according to the first result and / or the second result and / or the third result can effectively shorten the audio playback time delay, improve the continuity of audio playback, and improve the audio quality.

[0123] To implement the above embodiment, the present disclosure also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiment.

[0124] Figure 10 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 10 The illustrated electronic device 12 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure. The electronic device may be, for example, an electronic device or a terminal.

[0125] As Figure 10 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a memory 28, and a bus 18 connecting different system components (including the memory 28 and the processing unit 16).

[0126] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of the various bus architectures. By way of example, such architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.

[0127] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including both volatile and nonvolatile media, removable and non-removable media.

[0128] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Electronic device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 10 not shown, typically called a "hard disk drive").

[0129] Although Figure 10 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile magnetic disk (such as a "floppy disk"), and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as a Compact Disc Read Only Memory (CD-ROM), Digital Video Disc Read Only Memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present disclosure.

[0130] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in a memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present disclosure.

[0131] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a human body to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0132] The processing unit 16 executes various functional applications and data processing by running programs stored in the memory 28, such as implementing the methods mentioned in the foregoing embodiments.

[0133] To implement the above embodiments, the present disclosure also proposes a chip, including: The chip includes a processing circuit, and the processing circuit is configured to execute the method provided in the foregoing embodiments.

[0134] Figure 11 is a schematic structural diagram of the chip proposed in the embodiments of the present disclosure. Reference can be made to Figure 11 the schematic structural diagram of the chip 1100 shown, but not limited thereto.

[0135] The chip 1100 includes a processing circuit 1101, and the processing circuit 1101 is configured to execute any of the above methods.

[0136] In some embodiments, the chip 1100 further includes one or more interface circuits 1102. Optionally, the interface circuit 1102 is connected to the memory 1103. The interface circuit 1102 can be used to receive signals from the memory 1103 or other devices, and the interface circuit 1102 can be used to send signals to the memory 1103 or other devices. For example, the interface circuit 1102 can read the instructions stored in the memory 1103 and send the instructions to the processing circuit 1101.

[0137] In some embodiments, the interface circuit 1102 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1101 performs other steps.

[0138] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be used interchangeably.

[0139] In some embodiments, the chip 1100 further includes one or more memories 1103 for storing instructions. Optionally, all or part of the memory 1103 can be outside the chip 1100.

[0140] To implement the above embodiments, the present disclosure also proposes a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method proposed in the foregoing embodiments of the present disclosure.

[0141] To implement the above embodiments, the present disclosure also proposes a computer program product, when the instructions in the computer program product are executed by a processor, it executes the method proposed in the foregoing embodiments of the present disclosure.

[0142] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0143] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing the relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0144] The present disclosure anticipates embodiments that can provide users with the option to selectively block the use or access of personal information data. That is, the present disclosure anticipates that hardware and / or software can be provided to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0145] In the foregoing description of the various embodiments, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0146] Furthermore, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0147] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code that includes one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0148] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence of executable instructions for implementing logical functions, which can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0149] It should be understood that various parts of the present disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0150] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0151] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into one processing module, may exist separately physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0152] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. An audio processing method, characterized in that, The method includes the following steps: Obtain a first speech frame from a first storage area, where the first storage area is used to cache at least some speech frames, the first speech frame belongs to the at least some speech frames, and the at least some speech frames belong to multiple speech frames obtained by decoding at least one data packet payload; Determine whether each speech frame obtained from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determine whether the number of invalid speech frames obtained from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determine whether the maximum delay of valid speech frames within the first time period or the second time period is greater than a delay threshold to obtain a third result; and Determine whether to play the first speech frame according to the first result and / or the second result and / or the third result.

2. The method according to claim 1, wherein Determining to play the first speech frame according to the first result and / or the second result and / or the third result includes at least one of the following: Determine to play the first speech frame when the first result indicates that at least one speech frame obtained from the first storage area within the first time period is not a valid speech frame; Determine to play the first speech frame when the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is less than the preset number; Determine to play the first speech frame when the third result indicates that the maximum delay of valid speech frames within the first time period is less than or equal to the delay threshold; Determine to play the first speech frame when the third result indicates that the maximum delay of valid speech frames within the second time period is greater than the delay threshold.

3. The method according to claim 1, wherein Determining not to play the first speech frame according to the first result and / or the second result and / or the third result includes: Determine not to play the first speech frame when the first result indicates that each speech frame obtained from the first storage area within the first time period is a valid speech frame, and the third result indicates that the maximum delay of valid speech frames within the first time period is greater than the delay threshold.

4. The method according to claim 3, characterized in that, After determining not to play the first speech frame, the method further includes: Discard the first speech frame; Obtain a second speech frame, where the first speech frame and the second speech frame are different; Play the second speech frame.

5. The method according to claim 4, wherein The obtaining of the second speech frame includes: Obtain the next speech frame of the first speech frame from the first storage area and use the next speech frame as the second speech frame.

6. The method according to claim 1, wherein Determining not to play the first speech frame according to the first result and / or the second result and / or the third result includes: Determine not to play the first speech frame when the second result indicates that the number of invalid speech frames obtained from the first storage area within the second time period is greater than or equal to the preset number, and the third result indicates that the maximum delay of valid speech frames within the second time period is less than or equal to the delay threshold.

7. The method according to claim 6, wherein After determining not to play the first speech frame, the method further includes: Obtain a second speech frame, where the first speech frame and the second speech frame are different; Play the second speech frame.

8. The method according to claim 7, wherein The obtaining the second speech frame includes: Obtain the previous speech frame of the first speech frame from the first storage area, and use the previous speech frame as the second speech frame.

9. The method according to claim 1, characterized in that, The method further includes: Determine the delay threshold according to the depth of the first storage area.

10. The method according to any one of claims 1-9, characterized in that, Before obtaining the first speech frame from the first storage area, the method further includes: Receive at least one speech data packet; Extract the packet payload from the at least one speech data packet; Decode the packet payload to obtain the plurality of speech frames, where each speech frame has a corresponding caching order; Allocate a caching position in the first storage area for the speech frame according to the caching order; Cache each speech frame to the corresponding caching position in the first storage area.

11. The method according to claim 10, characterized in that, The caching order is related to the sequence number of the speech data packet to which the speech frame belongs and / or the transmission timestamp of the speech frame.

12. An audio processing device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a first speech frame from a first storage area, where the first storage area is used to cache at least part of speech frames, the first speech frame belongs to the at least part of speech frames, and the at least part of speech frames belong to a plurality of speech frames obtained by decoding at least one packet payload; A first determination module, configured to determine whether each speech frame obtained from the first storage area within a first time period is a valid speech frame to obtain a first result, and / or determine whether the number of invalid speech frames obtained from the first storage area within a second time period is less than a preset number to obtain a second result, and / or determine whether the maximum delay of valid speech frames within the first time period or the second time period is greater than a delay threshold to obtain a third result; and A second determination module, configured to determine whether to play the first speech frame according to the first result and / or the second result and / or the third result.

13. An electronic device, characterized in that, Includes: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-11.

14. A chip, characterized in that, The chip includes a processing circuit configured to execute the method according to any one of claims 1-11.

15. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when executed by a processor, are used to implement the method according to any one of claims 1-11.