Audio detection method and device, equipment and storage medium

By detecting the health of audio collected data and the device status, distinguishing the lag caused by network and device abnormalities, filtering out the lag caused by device abnormalities, and achieving accurate optimization of audio call quality.

CN120583078APending Publication Date: 2025-09-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410236622.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the process of audio and video calls, it is impossible to accurately distinguish between network abnormalities and uplink terminal equipment abnormalities, resulting in dirty data generated records, affecting call quality optimization.

Method used

By detecting the audio health information of the audio collected data of the uplink terminal, analyzing the operating status of the device, distinguishing the network abnormalities and the stuttering caused by the device abnormalities, filtering out the stuttering data caused by the device abnormalities, only the stuttering data caused by the network abnormalities are retained, and the storage status information of the audio packet cache area of ​​the downlink terminal is determined whether there is an audio packet to be played.

Benefits of technology

Accurately detect the hearing stuttering felt by the caller, more accurately reflect the real stuttering situation, and improve the quality of audio calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583078A_ABST
    Figure CN120583078A_ABST
Patent Text Reader

Abstract

The invention discloses an audio detection method and device, equipment and a storage medium. The embodiment of the invention can be applied to various scenes such as voice. Specifically, the method comprises the following steps: receiving an audio data packet generated based on audio acquisition data; receiving equipment state information corresponding to the audio data packet; the equipment state information is obtained by performing equipment operation state analysis on the uplink terminal based on audio health degree information corresponding to the audio acquisition data; and when it is indicated that the uplink terminal operates normally according to the equipment state information, an audio lagging index is obtained based on storage state information of the audio data packet cache region, and the storage state information is used for indicating whether the audio data packet to be played currently is stored in the audio data packet cache region. According to the technical scheme provided by the invention, the lagging dirty data caused by the abnormal state of the uplink terminal equipment can be filtered out, only the lagging data caused by the abnormal network is reserved, and the real lagging condition in the audio call process is reflected more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to an audio detection method, apparatus, device and storage medium. Background Art

[0002] With the development of the Internet, audio and video calls are becoming more and more popular in our lives. Call quality has a great impact on the call experience, and audio stuttering is a very important indicator of call quality. The level of the stuttering indicator can largely determine the quality of a call experience.

[0003] Currently, statistics on audio and video call jams contain a certain degree of dirty data. When an abnormal device status on the uplink terminal causes abnormal audio data, existing technologies may detect call jams and generate jam records. For example, if a video call on the uplink terminal is interrupted by a phone call, the uplink audio data may be abnormal. Even if there is no packet loss on the downlink terminal, the downlink call recipient may experience jams and other unpleasant listening experiences. However, jams in this case can be created on any device and have no guiding value for optimizing call quality. Summary of the Invention

[0004] The present application provides an audio detection method, apparatus, device, and storage medium. The method can accurately distinguish between jamming caused by network anomalies and jamming caused by uplink terminal device anomalies through the device status information of the uplink terminal, thereby filtering out jamming dirty data caused by uplink terminal device anomalies and retaining only jamming data caused by network anomalies. Furthermore, the method can determine whether there are currently audio data packets to be played through the storage status information of the audio data packet cache area of ​​the downlink terminal. The method can accurately detect the auditory jamming felt by the caller, more accurately reflect the actual jamming situation, and provide guidance for optimizing the quality of the audio call. The technical solution of the present application is as follows:

[0005] In one aspect, an audio detection method is provided, the method comprising:

[0006] receiving an audio data packet generated based on the audio acquisition data;

[0007] Receiving device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data;

[0008] When the device status information indicates that the uplink terminal is operating normally, an audio jam indicator is obtained based on storage status information of an audio data packet cache area, where the storage status information is used to indicate whether an audio data packet to be played is stored in the audio data packet cache area.

[0009] In another aspect, an audio detection device is provided, the device comprising:

[0010] An audio data packet receiving module, configured to receive an audio data packet generated based on the audio acquisition data;

[0011] A device status information receiving module, configured to receive device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data;

[0012] A jam detection module is used to obtain an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played.

[0013] On the other hand, an audio detection device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio detection method as described above.

[0014] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the audio detection method as described above.

[0015] In another aspect, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described audio detection method.

[0016] The present application provides an audio detection method, apparatus, device, and storage medium, which have the following technical effects:

[0017] By utilizing the technical solution provided by the present application, during an audio call, the audio health information corresponding to the audio collection data of the uplink terminal is detected, and based on the audio health information, the device operation status of the uplink terminal is analyzed to obtain device status information, thereby judging whether the uplink terminal has an audio health abnormality problem caused by an abnormal device status, and sending the audio data packet and device status information corresponding to the audio collection data to the downlink terminal, so that the downlink terminal obtains an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate whether the audio data packet cache area is normal. Whether the storage area contains audio data packets to be played by the downlink terminal, and the device status information of the uplink terminal can accurately distinguish between the jamming caused by network abnormalities and the uplink terminal device abnormalities, filter out the jamming dirty data caused by the abnormal status of the uplink terminal device, and only retain the jamming data caused by the network abnormalities. The storage status information of the audio data packet cache area of ​​the downlink terminal is used to determine whether there are currently audio data packets to be played, thereby determining whether there is playback jamming on the downlink terminal. This can accurately detect the auditory jamming felt by the caller, and more accurately reflect the actual jamming situation, so as to guide the quality optimization of the audio call and further improve the quality of the audio call. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0020] Figure 2 This is a flow chart of an audio detection method provided in an embodiment of the present application;

[0021] Figure 3 This is a flow chart of another audio detection method provided in an embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of a decision tree for device operating status provided in an embodiment of the present application;

[0023] Figure 5 This is a flow chart of a downlink terminal obtaining an audio jam indicator based on storage status information of an audio data packet buffer area provided by an embodiment of the present application;

[0024] Figure 6This is a flowchart of an audio call method provided by an embodiment of the present application;

[0025] Figure 7 This is a flow chart of another audio detection method provided in an embodiment of the present application;

[0026] Figure 8 A flowchart of another audio detection method provided in an embodiment of the present application;

[0027] Figure 9 This is a schematic diagram of the architecture of an audio detection system provided in an embodiment of the present application;

[0028] Figure 10 This is a block diagram of an audio detection device provided in an embodiment of the present application;

[0029] Figure 11 This is a block diagram of another audio detection device provided in an embodiment of the present application;

[0030] Figure 12 This is a structural diagram of an audio detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0033] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the following embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0034] The following explains the terms used in this application.

[0035] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0036] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0037] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.

[0038] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction.

[0039] PCM (Pulse Code Modulation) is the audio information after digital conversion. It is also the most primitive audio format. In essence, it is a series of digitally converted audio signals.

[0040] 3A: is a general term for the three core technologies of audio processing, including AEC (Acoustic Echo Cancelling), AGC (Automatic Gain Control) and ANS (Automatic Noise Suppression).

[0041] WebRTC (Web Real-Time Communications) is a real-time communication technology that allows web applications or sites to establish peer-to-peer connections between browsers without the help of an intermediary, enabling the transmission of video streams and / or audio streams or any other arbitrary data.

[0042] NetEQ: A core component of the WebRTC voice engine, it handles network jitter and packet loss, improving audio quality and buffering latency. NetEQ contains several buffers for temporarily storing audio data, primarily the JitterBuffer (PacketBuffer) and the SyncBuffer, which store pre-decoded audio data (i.e., audio packets) and post-decoding audio data, respectively.

[0043] A jitter buffer is a buffer used to handle network jitter, ensuring a smooth playback experience for audio or video data. When transmitting audio or video data over a network, due to network transmission delays and unpredictable jitter, packets may arrive at the receiving end at different speeds. This can cause packets to play at inconsistent speeds at the receiving end, affecting the quality and continuity of the audio or video. The jitter buffer caches received packets in a buffer, where they wait for a period of time for other packets to arrive.

[0044] Packetbuffer: A packet buffer in the jitter buffer, mainly used to store received audio data packets. When playback is required, the data packets are extracted from the buffer, decoded, and accelerated or decelerated as needed.

[0045] In the existing technology, the statistics of jams during audio and video calls contain a certain degree of dirty data. When the audio data is abnormal due to an abnormal device status on the uplink terminal, the existing technology may detect the call jam and generate a jam record. For example, if the video call of the uplink terminal is interrupted by a phone call, the uplink call audio data may be abnormal. At this time, even if the downlink terminal does not lose packets, the downlink call party will experience jams and other unpleasant listening experiences. However, the jams in this case can be constructed on any device and have no guiding value for optimizing call quality.

[0046] In order to solve this technical problem, the embodiment provided by the present application can detect the audio health information corresponding to the audio collection data of the uplink terminal during the audio call, and analyze the device operation status of the uplink terminal based on the audio health information to obtain device status information, thereby judging whether the uplink terminal has an audio health abnormality problem caused by an abnormal device status, and sending the audio data packet and device status information corresponding to the audio collection data to the downlink terminal, so that the downlink terminal can obtain an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate that the audio Whether the data packet cache area stores the audio data packet currently to be played by the downlink terminal, and the device status information of the uplink terminal can accurately distinguish between the jamming caused by network abnormalities and the uplink terminal device abnormalities, filter out the jamming dirty data caused by the abnormal status of the uplink terminal device, and only retain the jamming data caused by the network abnormalities, and judge whether there are currently audio data packets to be played through the storage status information of the audio data packet cache area of ​​the downlink terminal, thereby judging whether there is playback jamming on the downlink terminal. It can accurately detect the auditory jamming felt by the caller, and more accurately reflect the actual jamming situation, so as to guide the quality optimization of the audio call and further improve the quality of the audio call.

[0047] The solution provided in the embodiments of this application can be applied to various scenarios such as cloud technology, artificial intelligence, and smart transportation. Figure 1 , Figure 1is a schematic diagram of an application environment provided by an embodiment of the present application, which may include a terminal 10 and a server 20. The terminal 10 and the server 20 can be directly or indirectly connected via wired or wireless communication. The number of terminals 10 can be two or more, and similarly, the number of servers 20 can be one or more. In other words, there is no limit on the number of terminals 10 or servers 20. Schematically, the terminal 10 may include: an uplink terminal 11 and a downlink terminal 12 for audio calls. The uplink terminal 11 and the downlink terminal 12 can make an audio call through the server 20. During the audio call between the uplink terminal 11 and the downlink terminal 12, the uplink terminal 11 obtains the audio acquisition data in real time, performs a health check on the audio acquisition data, obtains audio health information, and then performs a device operation status analysis based on the audio health information to obtain the device status information of the uplink terminal 11, and sends the audio data packet generated based on the audio acquisition data and the device status information corresponding to the audio data packet to the downlink terminal 12 through the server 20; when the received device status information indicates that the uplink terminal 11 is operating normally, the downlink terminal 12 obtains an audio jam indicator based on the storage status information of the audio data packet cache area. The storage status information can be used to indicate whether the audio data packet cache area stores the audio data packet currently to be played by the downlink terminal 12. It should be noted that Figure 1 Just an example.

[0048] The terminal can be a physical device such as a smartphone, a computer (such as a desktop computer, a tablet computer, or a laptop computer), a digital assistant, an intelligent voice interaction device (such as a smart speaker), a smart wearable device, a smart home appliance, a vehicle-mounted terminal, or an aircraft. The terminal can be installed with a call application. The application involved in the embodiments of the present application can be a software client or a client such as a web page or a small program. The server is a background server corresponding to the software or web page or small program, and the specific type of the client is not limited. The operating system corresponding to the terminal can be Android system (Android system), iOS system (a mobile operating system developed by Apple), Linux system (an operating system), Microsoft Windows system (Microsoft Windows operating system), etc.

[0049] The server side can be the backend server corresponding to the call application installed on the terminal, which can provide the backend service functions of the call system. The server side can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can include a network communication unit, a processor, and memory, etc.

[0050] In some embodiments, the server 20 undertakes the main computing work and the terminal 10 undertakes the secondary computing work; or, the server 20 undertakes the secondary computing service and the terminal 10 undertakes the main computing work; or, the server 20 and the terminal 10 adopt a distributed computing architecture to perform collaborative computing.

[0051] The following describes a specific embodiment of an audio detection method provided by this application. Figure 2 It is a flowchart of an audio detection method provided in an embodiment of the present application. The present application provides method operation steps as described in the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment). Specifically, as Figure 2 As shown, the method may include:

[0052] S201: The uplink terminal obtains audio collection data in real time.

[0053] In the embodiments of this specification, the audio collection data may be sound collection data during an audio call. In a specific embodiment, the audio collection data may be data obtained by a call application of the uplink terminal collecting the speech content of the call party of the uplink terminal in real time. Specifically, the call application of the uplink terminal may be the application for conducting this audio call, and the call party may be the call user.

[0054] In the embodiments of this specification, the audio call can be a two-party call or a multi-party call, wherein uplink refers to the process of sending data from a terminal, and downlink refers to the process of receiving data to a terminal. Accordingly, the uplink terminal refers to the terminal that sends audio data packets, and the downlink terminal refers to the terminal that receives audio data packets. It can be understood that the uplink terminal and downlink terminal here are relative concepts. Schematically, taking a two-party call as an example, the first terminal and the second terminal conduct an audio call. The "first" and "second" here are only used to distinguish between the two terminals. When the call partner of the first terminal speaks, the first terminal is the uplink terminal and the second terminal is the downlink terminal. When the call partner of the second terminal speaks, the second terminal is the uplink terminal and the first terminal is the downlink terminal.

[0055] In an optional embodiment, the audio call may be a voice call or a video call. Optionally, the call scenarios corresponding to the audio call may include but are not limited to: voice chat, live chorus, voice connection in games, and other scenarios.

[0056] S202: The uplink terminal performs a health check on the audio collection data to obtain audio health information.

[0057] In an embodiment of the present specification, audio health information can be used to characterize the health of the audio collection status and audio playback status of the call application of the uplink terminal. Specifically, the audio health information may include: audio health indication information and audio unhealthy indication information. The audio health indication information can indicate that the audio collection status of the call application is normal, and the audio unhealthy indication information can indicate that the audio collection status of the call application is abnormal.

[0058] In a specific embodiment, the health of the audio collection data read within the preset collection detection period can be checked according to the preset collection detection period to obtain audio health information. The audio health information is used to characterize the health of the audio collection status of the call application of the uplink terminal within the preset collection detection period.

[0059] In a specific embodiment, the preset acquisition detection period can be set in combination with the generation period of the audio data packet. In an optional embodiment, the length of each preset acquisition detection period can be equal to the length of the generation period of the audio data packet during the call. Schematically, the length of the preset acquisition detection period can be 2ms.

[0060] In a specific embodiment, the uplink terminal performs health detection on the audio collection data, and the obtained audio health information may include:

[0061] S2021: The uplink terminal performs audio duration conversion on the audio collection data read within the preset collection detection period to obtain the audio duration corresponding to the preset collection detection period.

[0062] In a specific embodiment, the uplink terminal may perform audio duration conversion on the audio collection data read within the preset collection detection period based on the audio duration conversion formula to obtain the audio duration. Specifically, the audio duration conversion formula is as follows:

[0063] BytesToMs=Bytes×1000 / (sample_rate×channel×2),

[0064] Among them, Bytes represents the audio data collected during the current preset collection detection period, sample_rate represents the sampling rate set by the call application of the uplink terminal when starting collection, channel represents the number of channels set by the call application when starting collection, and BytesToMs represents the audio duration corresponding to the audio data collected during the current preset collection detection period.

[0065] In actual applications, under different collection parameters, the audio duration corresponding to each preset collection detection cycle will also change. The audio duration corresponding to multiple preset collection detection cycles during a call can be accumulated to ensure the accuracy of audio duration statistics.

[0066] S2022: The uplink terminal performs error analysis based on the audio duration and the length of the preset acquisition detection period to obtain audio health information.

[0067] In a specific embodiment, the audio health information may include: audio health indication information and audio unhealthy indication information. The uplink terminal performs error analysis based on the audio duration and the length of the preset acquisition detection period. The obtained audio health information may include:

[0068] In the case where the difference between the audio duration and the period length of the preset collection and detection period is less than the preset deviation threshold, the uplink terminal determines that the audio health information is: audio health indication information;

[0069] Alternatively, when the difference between the audio duration and the period length of the preset collection and detection period is greater than or equal to a preset deviation threshold, the uplink terminal determines that the audio health information is: audio unhealthy indication information.

[0070] It can be understood that under the normal state of audio collection, the audio duration corresponding to the audio collection data collected within each preset collection detection period should be equal to the length of each preset collection detection period. Therefore, when the deviation between the audio duration corresponding to the audio collection data collected within a preset collection detection period and the period duration is greater than the preset deviation threshold, the audio collection status of the call application of the uplink terminal is considered to be abnormal.

[0071] It can be seen from the above embodiments that according to the preset acquisition detection cycle, the amount of acquisition data read in the current cycle is converted into audio duration, and the audio acquisition health is judged based on the deviation between the read audio duration and the cycle length, which can improve the efficiency and accuracy of the audio health judgment.

[0072] S203: The uplink terminal performs device operation status analysis based on the audio health information to obtain device status information of the uplink terminal.

[0073] In the embodiment of this specification, the device status information may be used to indicate whether the uplink terminal has an abnormal device status.

[0074] In a specific embodiment, the device status information may include: normal device indication information and abnormal device indication information. The normal device indication information may be used to indicate that the uplink terminal is operating normally, and the abnormal device indication information may be used to indicate that the uplink terminal is operating abnormally.

[0075] In a specific embodiment, Figure 3 As shown, before the uplink terminal performs device operation status analysis based on the audio health information to obtain device status information of the uplink terminal, the method may further include:

[0076] S206: The uplink terminal performs a foreground / background switching detection on the call application, and obtains state switching information corresponding to the call application.

[0077] Specifically, the state switching information can indicate the switching between the foreground and background states of the call application. In actual applications, when the call application is in the background, some terminal devices will lower the scheduling priority of the call application, resulting in abnormal audio collection status of the call application.

[0078] S207: The uplink terminal performs audio collection interruption detection on the call application to obtain a collection interruption record.

[0079] Specifically, the collection interruption record can indicate whether the audio collection of the calling application in this call has been interrupted by other applications. In actual applications, in a terminal device, only one application can use the collection function of the terminal device at the same time. When other applications also start collection, the current audio collection of the calling application may be interrupted and various abnormalities may occur. For example, receiving a system call, WeChat call, etc. during the game connection process will cause the current audio collection of the calling application (that is, the game application that is connected to the microphone) to be interrupted.

[0080] In a specific embodiment, a collection interruption detection module may be registered in the system background of the uplink terminal. When the system senses a collection interruption, it notifies the call application of the collection interruption.

[0081] S208: The uplink terminal performs resource occupancy detection to obtain first resource occupancy information corresponding to the call application and second resource occupancy information corresponding to the uplink terminal.

[0082] In a specific embodiment, the first resource occupancy information can be used to characterize the resource occupancy of the call application, and the second resource occupancy information can be used to characterize the resource occupancy of the uplink terminal. Specifically, the first resource occupancy information may include: the CPU occupancy information of the call application, and the second resource occupancy information may include: the CPU occupancy information of the uplink terminal. Schematically, the CPU occupancy information of the call application may be the CPU occupancy rate of the call application, and the CPU occupancy information of the uplink terminal may be the CPU occupancy rate of the uplink terminal. In actual applications, if the CPU usage of the call application is normal, but the CPU usage rate of the uplink terminal exceeds the threshold, it is considered that the background of the uplink terminal is running a computationally complex background task, causing the CPU to be fully occupied (for example, other computing applications occupy the CPU), which may cause abnormalities in the audio collection of the call application.

[0083] Accordingly, the uplink terminal performs device operation status analysis based on the audio health information, and the device status information of the uplink terminal obtained may include:

[0084] S2031, when the audio collection status of the call application of the uplink terminal is abnormal according to the audio health information, the uplink terminal performs device operation status analysis based on the status switching information, the collection interruption record, the first resource occupancy information and the second resource occupancy information to obtain device status information.

[0085] Specifically, when the audio health information indicates that the audio collection status of the call application is abnormal, the uplink terminal can decide whether there is an audio health abnormality problem caused by the abnormal device status of the uplink terminal based on the status switching information, the collection interruption record, the first resource occupancy information and the second resource occupancy information.

[0086] It can be understood that the audio health information indicates that the audio collection status of the call application of the uplink terminal is normal, indicating that the current audio collection and playback status of the uplink terminal is in line with expectations. Even if the device status of the uplink terminal is abnormal, it has no effect on the current audio collection and playback status. Therefore, when the audio health information indicates that the audio collection status of the call application of the uplink terminal is normal, the device status analysis of the uplink terminal may not be performed; when the audio health information indicates that the audio collection status of the call application of the uplink terminal is abnormal, if the device status of the uplink terminal is also abnormal, it indicates that the abnormal audio collection status may be caused by the abnormal device status.

[0087] In a specific embodiment, Figure 4This is a schematic diagram of a decision tree for determining the operating status of a device provided in an embodiment of the present application. Based on the decision tree, the uplink terminal can analyze the device status of the background status switching information, the collection interruption record, the first resource occupancy information of the call application, and the second resource occupancy information of the uplink terminal to obtain the device status information of the uplink terminal to determine whether the audio health abnormality (abnormal audio collection status of the call application) is caused by the abnormal device status of the uplink terminal. Specifically, Figure 4 As shown:

[0088] S401: Determine whether the audio collection status of the call application is normal;

[0089] S402, when the audio collection state of the call application is abnormal, determining whether the audio collection of the call application is interrupted;

[0090] S403, when the audio collection is interrupted, it is considered that the device status of the uplink terminal is abnormal, indicating that the abnormal audio collection status of the call application is caused by the abnormal device status of the uplink terminal;

[0091] S404, if the audio collection is not interrupted, determining whether the call application is in the background state;

[0092] S405, when the audio collection is not interrupted and the call application is in the background state, it is considered that the device state of the uplink terminal is abnormal, indicating that the abnormal audio collection state of the call application is caused by the abnormal device state of the uplink terminal;

[0093] S406, when the call application is in the foreground state, determining whether the resource occupancy rate of the uplink terminal exceeds a threshold;

[0094] S407: If the audio collection is not interrupted, the call application is in the foreground, and the resource usage of the uplink terminal does not exceed the threshold, the device status of the uplink terminal is considered normal, indicating that the abnormal audio collection status of the call application is not caused by the abnormal device status of the uplink terminal;

[0095] S408, when the resource occupancy rate of the uplink terminal exceeds the threshold, determining whether the resource occupancy rate of the call application exceeds the threshold;

[0096] S409: If the audio collection is not interrupted, the call application is in the foreground state, the resource usage of the uplink terminal exceeds the threshold, and the resource usage of the call application exceeds the threshold, it is considered that the device status of the uplink terminal is abnormal, indicating that the abnormal audio collection status of the call application is caused by the abnormal device status of the uplink terminal;

[0097] S410: When the audio collection is not interrupted, the call application is in the foreground state, the resource occupancy rate of the uplink terminal exceeds the threshold and the resource occupancy rate of the call application does not exceed the threshold, it is considered that the device status of the uplink terminal is normal, indicating that the abnormal audio collection status of the call application is not caused by the abnormal device status of the uplink terminal.

[0098] It can be seen from the above embodiments that in the case of abnormal audio health, based on the detection results of collection interruption detection, foreground and background switching detection, and resource occupancy detection of the call application, it is decided whether there is an audio health abnormality problem caused by abnormal device status. It can accurately distinguish between the jamming caused by network abnormalities and the jamming caused by abnormal uplink terminal equipment, thereby filtering out the jamming dirty data caused by abnormal uplink terminal equipment status, and only retaining the jamming data caused by network abnormalities, which more accurately reflects the actual jamming situation and guides the quality optimization of audio calls.

[0099] S204: The uplink terminal sends an audio data packet generated based on the audio collection data and device status information corresponding to the audio data packet to the downlink terminal.

[0100] Specifically, the audio data collected by the call application of the uplink terminal is first processed by 3A to filter out echoes and noise in the collected data, and adjust the sound gain. Then, audio encoding is performed, and after adding a packet header, an audio data packet is generated and uploaded to the server.

[0101] In an embodiment of the present specification, during an audio call, the uplink terminal sends continuous frame audio data packets to the downlink terminal. Each audio data packet sent by the uplink terminal may carry device status information within a preset acquisition detection period of the corresponding audio acquisition data. In a specific embodiment, the device status information may be added to the header of the audio data packet. Schematically, the header format of the audio data packet may include: data packet type, timestamp, sequence number, audio encoding format, data packet size, audio data duration, device status information and other fields.

[0102] S205, when the device status information indicates that the uplink terminal is operating normally, the downlink terminal obtains an audio jam indicator based on the storage status information of the audio data packet cache area, where the storage status information is used to indicate whether the audio data packet cache area stores an audio data packet currently to be played by the downlink terminal.

[0103] Specifically, the downlink terminal can receive the audio data packets sent by the uplink terminal through the server, and store the audio data packets in the audio data packet buffer area, which is used to cache the audio data packets received by the downlink terminal and currently to be played.

[0104] In an embodiment of the present specification, the call application of the downlink terminal will periodically read the audio data packets to be played in the audio data packet cache area (the played audio data packets will be extracted from the cache area by the call application player), decode the audio data packets and perform acceleration and deceleration processing before transmitting them to the player for playback. Therefore, the storage status information of the audio data packet cache area represents the reception status of the audio data packets by the downlink terminal. When the audio data packet cache area is empty, it means that the downlink terminal has not received the audio data packets to be played, resulting in the player having no audio data packets to play, resulting in jams. The reason for the abnormal reception of data packets by the downlink terminal may be network packet loss or abnormal device status of the uplink terminal. Therefore, by periodically reading the storage status information of the audio data packet cache area, it is possible to know whether the audio data packet reception is normal, that is, whether the downlink terminal currently has audio data packets to be played, and then perform jam statistics.

[0105] In an embodiment of the present specification, the downlink terminal may first obtain the latest device status information currently received by the uplink terminal before reading the storage status information of the audio data packet cache area. In an optional embodiment, when the downlink terminal currently receives an audio data packet, the device status information corresponding to the currently received audio data packet may be used as the latest device status information; when the downlink terminal currently does not receive an audio data packet, the device status information corresponding to the most recently received historical audio data packet received by the downlink terminal may be used as the latest device status information.

[0106] Specifically, only when the latest device status information indicates that the uplink terminal is operating normally, the audio jam indicator is obtained based on the storage status information of the audio data packet cache area, which can accurately distinguish between jams caused by network abnormalities and uplink terminal device abnormalities, thereby filtering out the jam dirty data caused by uplink terminal device status abnormalities, and only retaining the jam data caused by network abnormalities. The storage status information of the audio data packet cache area of ​​the downlink terminal is used to determine whether there are currently audio data packets to be played, thereby determining whether the downlink terminal has playback jams. This can accurately detect the auditory jams felt by the caller and more accurately reflect the actual jam situation.

[0107] In an optional embodiment, before step S205, the above method may further include: reading device status information, and when the device status information indicates that the downlink terminal is operating abnormally, no longer performing a jam detection on the downlink terminal to avoid generating jam dirty data due to abnormal status of the uplink terminal device.

[0108] In the embodiment of this specification, the audio jam index can be used to characterize the jamming situation of the downlink terminal during the audio call. Specifically, the audio jam index can include but is not limited to: audio jam rate, audio jam duration, audio jam number, etc.

[0109] In a specific embodiment, Figure 5 As shown, the downlink terminal may obtain the audio jam indicator based on the storage status information of the audio data packet buffer area, which may include:

[0110] S501: The downlink terminal reads storage status information of the audio data packet buffer area according to a preset processing cycle; the preset processing cycle is a sub-cycle within a preset freeze detection cycle.

[0111] In an embodiment of the present specification, the preset processing cycle can be set in combination with the call application of the downlink terminal to set the data packet reading cycle of the audio data packet cache area. In actual applications, the call application of the downlink terminal extracts the audio data packets to be played from the audio data packet cache area according to the data packet reading cycle. Therefore, the downlink terminal needs to detect the storage status information of the audio data packet cache area before its own call application extracts the audio data packets to be played. In an optional embodiment, the period length of the preset processing cycle can be equal to the period length of the data packet reading cycle, and the starting time of the preset processing cycle is before the starting time of the data packet reading cycle.

[0112] In a specific embodiment, the preset jam detection cycle may include multiple sub-cycles, and the preset processing cycle may be a sub-cycle within the preset jam detection cycle. Schematically, the cycle length of the preset jam detection cycle may be 2s, and the cycle length of the preset processing cycle may be 20ms. Therefore, every 100 preset processing cycles can be regarded as a preset jam detection cycle.

[0113] Schematically, taking the communication between the uplink terminal and the downlink terminal using WebRTC as an example, the audio data packet cache area can be PacketBuffer. The call application of the downlink terminal can periodically read the audio data packets to be played from the PacketBuffer, and decode and accelerate and decelerate the audio data packets to be played to obtain decoded audio data, and then write the decoded audio data into the SyncBuffer. Finally, the player of the call application can read the data from the SyncBuffer for playback. Therefore, the storage status information of the PacketBuffer can indicate the status of the audio data packets received from the uplink. When the storage status information indicates that the PacketBuffer is empty, it means that the downlink terminal has not received the current audio data packet to be played, resulting in the player having no audio data packet to play, thereby causing lag.

[0114] S502: The downlink terminal performs audio freeze detection on a preset processing period based on the storage status information, and obtains an audio freeze record corresponding to the preset processing period.

[0115] In a specific embodiment, the downlink terminal performs audio jam detection on a preset processing period based on the storage status information, and obtaining the audio jam record corresponding to the preset processing period may include:

[0116] S5021, when the storage status information indicates that the audio data packet cache area does not store the audio data packet to be played by the downlink terminal, the downlink terminal uses the period length of the preset processing period as the audio freeze period corresponding to the preset processing period;

[0117] S5022: The downlink terminal generates an audio freeze record corresponding to the preset processing cycle according to the audio freeze duration corresponding to the preset processing cycle.

[0118] Specifically, the audio jam record is used to characterize the audio jam situation within a preset processing cycle. In a specific embodiment, when the storage status information corresponding to a certain processing cycle indicates that the audio data packet cache area does not store the audio data packet currently to be played by the downlink terminal, a jam is recorded, and the cycle length of the processing cycle is used as the audio jam duration of the jam. When the storage status information corresponding to multiple consecutive processing cycles all indicates that the audio data packet cache area does not store the audio data packet currently to be played by the downlink terminal, the continuous jam duration is recorded, and the total length of the multiple processing cycles is used as the continuous jam duration.

[0119] It can be seen from the above embodiments that when the latest device status information of the uplink terminal received indicates that the uplink terminal is operating normally, the downlink terminal reads the storage status information of the audio data packet cache area according to the preset processing cycle. When the storage status information indicates that the audio data packet cache area does not store the audio data packet currently to be played by the downlink terminal, an audio stutter record corresponding to the preset processing cycle is generated, thereby obtaining an audio stutter index, which can accurately detect the auditory stuttering felt by the caller and more accurately reflect the actual stuttering situation, so as to guide the quality optimization of the audio call.

[0120] S503: The downlink terminal determines audio scene type information.

[0121] Specifically, the audio scene type information may be the type information of the call scene in which the audio call is located. Schematically, the audio scene type information may include but is not limited to: voice chat, live chorus, voice connection in games, and other scenes.

[0122] In an optional embodiment, a scene type switch may occur during an audio call. For example, taking the call application used in the audio call as a live broadcast application, the call object of the audio call may be to switch the live broadcast application from voice chat mode to live chorus mode during the audio call.

[0123] S504: The downlink terminal performs audio jam analysis on the audio jam records within a preset jam detection period based on the preset jam influence factor corresponding to the audio scene type information to determine an audio jam index.

[0124] Specifically, the preset jamming impact factor can represent the degree of impact of audio jamming on the listening experience of the caller. In actual applications, the preset jamming impact factor can be pre-set based on the smoothness requirements of audio playback in different call scenarios. Indicatively, the numerical value corresponding to the preset jamming impact factor for the live chorus scenario should be greater than the numerical value corresponding to the preset jamming impact factor for the voice chat scenario.

[0125] In a specific embodiment, the audio jam record may include: audio jam duration and audio jam number. The downlink terminal performs audio jam analysis on the audio jam record within a preset jam detection period based on a preset jam influence factor corresponding to the audio scene type information. Determining the audio jam index may include:

[0126] S5041: The downlink terminal performs weighted processing on the number of audio jams within a preset jam detection period based on a preset jam impact factor to determine a first audio jam index.

[0127] Specifically, the first audio jam indicator can measure the audio jam situation within a preset jam detection period from the dimension of the number of jams.

[0128] In a specific embodiment, the downlink terminal may use a jam number conversion formula to determine the first audio jam index. Specifically, the jam number conversion formula may be expressed as:

[0129] y1=100×(1–nL), where L represents a preset jamming influence factor, n represents the number of audio jams within a preset jamming detection period, and y1 represents a first audio jamming indicator.

[0130] S5042: The downlink terminal performs continuous jam analysis on the preset jam detection period based on the audio jam duration within the preset jam detection period to determine a second audio jam indicator.

[0131] Specifically, the second audio jam indicator can measure the audio jam situation within a preset jam detection period from the dimension of jam duration.

[0132] In a specific embodiment, the downlink terminal may determine the second audio jam index using a jam duration conversion formula. Specifically, the jam duration conversion formula may be expressed as:

[0133] y2=100×[(1-m1 / T)×…×(1-mi / T)×…×(1-mp / T)], where T represents the length of a preset jam detection period, p represents the number of consecutive jams within the preset jam detection period, i=1, 2, …, p, mi represents the duration of the i-th consecutive jam, and y2 represents the second audio jam indicator.

[0134] S5043: The downlink terminal determines an audio jam indicator based on the first audio jam indicator and the second audio jam indicator.

[0135] In an optional embodiment, the downlink terminal may use the smaller value of the first audio jam indicator and the second audio jam indicator as the audio jam indicator corresponding to the preset jam detection period. In another optional embodiment, the downlink terminal may perform weighted fusion processing on the first audio jam indicator and the second audio jam indicator to obtain the audio jam indicator corresponding to the preset jam detection period.

[0136] It can be seen from the above embodiments that the downlink terminal performs audio jam analysis on the audio jam records within the preset jam detection period based on the preset jam influence factor corresponding to the audio scene type information, which can improve the flexibility of the jam analysis, and converts the audio jam records into indicators from the dimensions of the number of jams and the jam duration, respectively, to determine the audio jam index, which can more accurately measure the auditory jam felt by the caller and reflect the actual jam situation.

[0137] See also Figure 6 , Figure 6 : is a flowchart of an audio call method provided in an embodiment of the present application. Specifically, the audio call method may include:

[0138] S601: The uplink terminal collects the voice of the caller and obtains audio data.

[0139] S602: During the audio collection process, device status detection is performed. Specifically, if the audio health is abnormal, based on the results of call application collection interruption detection, foreground and background switching detection, and resource usage detection, it is determined whether the device status abnormality is causing the audio health abnormality, and device status information of the uplink terminal is generated.

[0140] S603: Performing pre-processing operations such as 3A processing on the audio data, filtering out echoes and noise in the collected data, and adjusting the sound gain;

[0141] S604: Perform audio encoding on the pre-processed audio collection data to obtain encoded audio data;

[0142] S605: Add a header to the encoded audio data to generate an audio data packet;

[0143] S606: Add device status information to the header of the audio data packet;

[0144] S607: Uplink the audio data packet to the server;

[0145] S608: The server transmits the audio data packet to the downstream terminal;

[0146] S609: The downlink terminal places the audio data packet into the audio data packet buffer;

[0147] S610: Performing jitter detection based on the device status information of the uplink terminal and the storage status of the audio data packet buffer to obtain an audio jitter indicator. Specifically, the audio data packet buffer here may be a PacketBuffer in a Jitter Buffer.

[0148] S611: Decoding the audio data packet to be played read from the audio data packet buffer to obtain decoded audio data;

[0149] S612: Accelerate and decelerate the decoded audio data to obtain audio data to be played;

[0150] S613: Send the audio data to be played to the player of the call application for playing.

[0151] The following describes another audio detection method provided by an embodiment of the present application using the uplink terminal as the execution subject. Figure 7 A flowchart of another audio detection method provided in an embodiment of the present application. It should be noted that this specification provides method operation steps as described in the embodiment or flowchart, but more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiment is only one way of executing the steps among many, and does not represent the only execution order. When the actual system or product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 7 As shown, the above method may include:

[0152] S701, real-time acquisition of audio data;

[0153] S702, performing a health check on the audio collection data to obtain audio health information;

[0154] S703, performing device operation status analysis based on the audio health information to obtain device status information of the uplink terminal;

[0155] S704, sending an audio data packet generated based on the audio acquisition data and device status information corresponding to the audio data packet, so that the downlink terminal obtains an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate whether the audio data packet cache area stores an audio data packet currently to be played by the downlink terminal.

[0156] In a specific embodiment, the health check of the audio collection data to obtain the audio health information may include:

[0157] S7021, performing audio duration conversion on the audio collection data read within the preset collection detection period to obtain the audio duration corresponding to the preset collection detection period;

[0158] S7022: Perform error analysis based on the audio duration and the length of a preset acquisition detection period to obtain audio health information.

[0159] In a specific embodiment, before performing the device operating status analysis based on the audio health information to obtain the device status information of the uplink terminal, the method may further include:

[0160] S705, performing foreground and background switching detection on the call application to obtain state switching information corresponding to the call application;

[0161] S706, performing audio collection interruption detection on the call application to obtain a collection interruption record;

[0162] S707, performing resource occupancy detection to obtain first resource occupancy information corresponding to the call application and second resource occupancy information corresponding to the uplink terminal;

[0163] Accordingly, the device operating status analysis based on the audio health information may include:

[0164] S7031, when the audio collection status of the call application of the uplink terminal is abnormal according to the audio health information, the device operation status is analyzed based on the status switching information, the collection interruption record, the first resource occupancy information and the second resource occupancy information to obtain the device status information.

[0165] The specific detailed steps of the embodiment of the audio detection method written from the uplink terminal side can be referred to the embodiment of the interactive side audio detection method, which will not be repeated here.

[0166] The following describes another audio detection method provided by an embodiment of the present application using a terminal as the execution subject. Figure 8A flowchart of another audio detection method provided in an embodiment of the present application. It should be noted that this specification provides method operation steps as described in the embodiment or flowchart, but more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiment is only one way of executing the steps among many, and does not represent the only execution order. When the actual system or product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 8 As shown, the above method may include:

[0167] S801, receiving an audio data packet generated based on audio acquisition data;

[0168] S802, receiving device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data;

[0169] S803, when the device status information indicates that the uplink terminal is operating normally, an audio jam indicator is obtained based on the storage status information of the audio data packet cache area, where the storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played.

[0170] In a specific embodiment, the device status information is obtained after analyzing the device operation status of the uplink terminal based on the status switching information of the call application, the collection interruption record of the call application, the first resource occupancy information corresponding to the call application, and the second resource occupancy information corresponding to the uplink terminal when the audio health information indicates that the audio collection status of the call application of the uplink terminal is abnormal.

[0171] In a specific embodiment, the state switching information is obtained after performing a foreground-background switching detection on the call application; the collection interruption record is obtained after performing an audio collection interruption detection on the call application; the first resource occupancy information is obtained after performing a resource occupancy detection on the call application; and the second resource occupancy information is obtained after performing a resource occupancy detection on the uplink terminal.

[0172] In a specific embodiment, the audio health information is obtained after error analysis between the audio duration corresponding to the preset acquisition detection period and the period length of the preset acquisition detection period; the audio duration is obtained after converting the audio acquisition data read within the preset acquisition detection period into an audio duration.

[0173] In a specific embodiment, obtaining the audio jam indicator based on the storage status information of the audio data packet cache area may include:

[0174] S8021, reading storage status information of the audio data packet cache area according to a preset processing cycle; the preset processing cycle is a sub-cycle within the preset jam detection cycle;

[0175] S8022: Perform audio jam detection for a preset processing period based on the storage status information to obtain an audio jam record corresponding to the preset processing period.

[0176] S8023, determining audio scene type information;

[0177] S8024: Based on a preset jamming influence factor corresponding to the audio scene type information, perform audio jamming analysis on the audio jamming records within a preset jamming detection period to determine an audio jamming index.

[0178] In a specific embodiment, the audio freeze detection is performed on the preset processing period based on the storage status information, and the audio freeze record corresponding to the preset processing period is obtained, which may include:

[0179] 1) when the storage status information indicates that the audio data packet buffer area does not store the audio data packet to be played by the downlink terminal, the period length of the preset processing period is used as the audio freeze period corresponding to the preset processing period;

[0180] 2) Generate an audio freeze record corresponding to the preset processing cycle according to the audio freeze duration corresponding to the preset processing cycle.

[0181] In a specific embodiment, the audio freeze record may include: audio freeze duration and audio freeze number. The preset freeze impact factor corresponding to the audio scene type information may be based on the preset freeze impact factor. The audio freeze analysis is performed on the audio freeze record within the preset freeze detection period. Determining the audio freeze index may include:

[0182] 1) performing weighted processing on the number of audio freezes within a preset freeze detection period based on a preset freeze impact factor to determine a first audio freeze index;

[0183] 2) Based on the audio jam duration within the preset jam detection period, performing continuous jam analysis on the preset jam detection period to determine a second audio jam indicator;

[0184] 3) Determine an audio jam indicator based on the first audio jam indicator and the second audio jam indicator.

[0185] The specific detailed steps of the embodiment of the audio detection method written from the downlink terminal side can be referred to the embodiment of the interactive side audio detection method, which will not be repeated here.

[0186] It can be seen from the technical solution provided by the above embodiments of the present application that during an audio call, by converting the amount of collected data read in the current cycle into audio duration according to the preset collection detection cycle, and judging the audio collection health according to the deviation between the read audio duration and the cycle length, the efficiency and accuracy of the audio health judgment can be improved. In the case of abnormal audio health, based on the detection results of collection interruption detection, foreground and background switching detection, and resource occupancy detection of the call application, it is decided whether there is an audio health abnormality caused by abnormal device status, and then the audio data packet and device status information corresponding to the audio collection data are sent to the downlink terminal, so that the downlink terminal can perform audio freeze detection when the device status information indicates that the uplink terminal is operating normally. The device status information of the uplink terminal can accurately distinguish between freezes caused by network abnormalities and uplink terminal device abnormalities. Filter out the dirty data caused by abnormal status of the uplink terminal device, and only retain the jam data caused by network anomalies, and read the storage status information of the audio data packet cache area according to the preset processing cycle. When the storage status information indicates that the audio data packet cache area does not store the audio data packet currently to be played by the downlink terminal, generate an audio jam record corresponding to the preset processing cycle, and perform audio jam analysis on the audio jam record within the preset jam detection cycle based on the preset jam influence factor corresponding to the audio scene type information, which can improve the flexibility and accuracy of the jam analysis, and convert the audio jam record into indicators from the dimensions of the number of jams and the jam duration, and determine the audio jam index, which can more accurately measure the hearing jam felt by the caller and reflect the actual jam situation, so as to guide the quality optimization of the audio call and further improve the quality of the audio call.

[0187] See also Figure 9 , Figure 9 Schematic diagram of the architecture of an audio detection system provided by an embodiment of the present application. Specifically, the audio detection system may include an uplink terminal 910 and a downlink terminal 920. The uplink terminal 910 may include a device detection module 911 and a supplementary packet header module 912. The downlink terminal 920 may include a freeze detection module 921.

[0188] In a specific embodiment, the device detection module 911 may include a health detection unit 912, a collection interruption detection unit 913, an application foreground and background detection unit 914, a resource usage detection unit 915 and a detection and analysis unit 916. Specifically, the health detection unit 912 can be used to perform health detection on audio collection data to obtain audio health information; the collection interruption detection unit 913 can be used to perform audio collection interruption detection on the call application to obtain collection interruption records; the application foreground and background detection unit 914 can be used to perform foreground and background switching detection on the call application to obtain state switching information corresponding to the call application; the resource usage detection unit 915 can be used to perform resource occupancy detection to obtain first resource occupancy information corresponding to the call application and second resource occupancy information corresponding to the uplink terminal; the detection and analysis unit 916 can be used to perform device operation status analysis based on the state switching information, collection interruption records, first resource occupancy information and second resource occupancy information to obtain device status information when the audio health information indicates that the audio collection status of the call application of the uplink terminal is abnormal. That is, based on the detection results of the previous units, it is determined whether there is an audio health abnormality problem caused by the abnormal device status of the uplink terminal.

[0189] In a specific embodiment, the header supplementing module 912 can be used to add device status information to the header of the audio data packet in real time.

[0190] In a specific embodiment, the jam detection module 921 may include: a storage status reading unit 922, a device status reading unit 923 and a jam statistics unit 924. Specifically, the storage status reading unit 922 can be used to detect the storage status information of the audio data packet cache area of ​​the downlink terminal, and the storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played; the device status reading unit 923 can be used to read the latest device status information of the uplink terminal received by the downlink terminal; the jam statistics unit 924 can be used to perform jam statistics based on the latest device status information of the uplink terminal received and the storage status information of the audio data packet cache area to obtain an audio jam index.

[0191] The embodiment of the present application provides an audio detection device with an uplink terminal as the execution subject, such as Figure 10 As shown, the audio detection device may include:

[0192] The audio data acquisition module 1010 is used to acquire audio data in real time;

[0193] Health detection module 1020, used to perform health detection on the audio collection data to obtain audio health information;

[0194] The device operation status analysis module 1030 is used to analyze the device operation status based on the audio health information to obtain the device status information of the uplink terminal;

[0195] The sending module 1040 is used to send audio data packets generated based on audio acquisition data and device status information corresponding to the audio data packets, so that the downlink terminal can obtain an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played by the downlink terminal.

[0196] In a specific embodiment, the health detection module 1020 may include:

[0197] The audio duration conversion unit is used to convert the audio collection data read within the preset collection detection period into an audio duration to obtain the audio duration corresponding to the preset collection detection period;

[0198] The error analysis unit is used to perform error analysis based on the audio duration and the period length of the preset acquisition and detection period to obtain audio health information.

[0199] In a specific embodiment, the above device may further include:

[0200] The front-end and back-end switching detection module is used to detect the front-end and back-end switching of the call application and obtain the state switching information corresponding to the call application;

[0201] The interruption detection module is used to detect the interruption of audio collection in the call application and obtain the collection interruption record;

[0202] A resource occupancy detection module is used to perform resource occupancy detection and obtain first resource occupancy information corresponding to the call application and second resource occupancy information corresponding to the uplink terminal;

[0203] Accordingly, the device operation status analysis module 1030 may include:

[0204] The status decision unit is used to analyze the device operation status based on the status switching information, the collection interruption record, the first resource occupancy information and the second resource occupancy information to obtain the device status information when the audio collection status of the call application of the uplink terminal is abnormal according to the audio health information.

[0205] It should be noted that the device in the device embodiment and the method embodiment are based on the same inventive concept.

[0206] The embodiment of the present application provides an audio detection device as an execution subject of the following terminal, such as Figure 11As shown, the audio detection device may include:

[0207] The audio data packet receiving module 1110 is configured to receive an audio data packet generated based on the audio acquisition data;

[0208] The device status information receiving module 1120 is used to receive device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data;

[0209] The jam detection module 1130 is used to obtain an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally. The storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played.

[0210] In a specific embodiment, the device status information is obtained after analyzing the device operation status of the uplink terminal based on the status switching information of the call application, the collection interruption record of the call application, the first resource occupancy information corresponding to the call application, and the second resource occupancy information corresponding to the uplink terminal when the audio health information indicates that the audio collection status of the call application of the uplink terminal is abnormal.

[0211] In a specific embodiment, the state switching information is obtained after performing a foreground-background switching detection on the call application; the collection interruption record is obtained after performing an audio collection interruption detection on the call application; the first resource occupancy information is obtained after performing a resource occupancy detection on the call application; and the second resource occupancy information is obtained after performing a resource occupancy detection on the uplink terminal.

[0212] In a specific embodiment, the audio health information is obtained after error analysis between the audio duration corresponding to the preset acquisition detection period and the period length of the preset acquisition detection period; the audio duration is obtained after converting the audio acquisition data read within the preset acquisition detection period into an audio duration.

[0213] In a specific embodiment, the jam detection module 1120 may include:

[0214] A storage status information reading unit, configured to read storage status information of an audio data packet buffer area according to a preset processing cycle; the preset processing cycle being a sub-cycle within a preset freeze detection cycle;

[0215] An audio jam detection unit, configured to perform audio jam detection on a preset processing cycle based on the storage status information, and obtain an audio jam record corresponding to the preset processing cycle;

[0216] A scene type determination unit, configured to determine audio scene type information;

[0217] The audio jam analysis unit is used to perform audio jam analysis on audio jam records within a preset jam detection period based on a preset jam influence factor corresponding to the audio scene type information to determine an audio jam index.

[0218] In a specific embodiment, the audio jam detection unit may include:

[0219] A freeze duration determining unit, configured to, when the storage status information indicates that the audio data packet cache area does not store the audio data packet currently to be played by the downlink terminal, use the period duration of the preset processing period as the audio freeze duration corresponding to the preset processing period;

[0220] The audio jam record generating unit is used to generate an audio jam record corresponding to a preset processing cycle according to the audio jam duration corresponding to the preset processing cycle.

[0221] In a specific embodiment, the audio jam record may include: audio jam duration and audio jam number, and the audio jam analysis unit may include:

[0222] A first indicator determination unit is configured to perform weighted processing on the number of audio freezes within a preset freeze detection period based on a preset freeze impact factor to determine a first audio freeze indicator;

[0223] A second indicator determination unit is configured to perform continuous jam analysis on the preset jam detection period based on the audio jam duration within the preset jam detection period to determine a second audio jam indicator;

[0224] The third indicator determination unit is configured to determine an audio jam indicator based on the first audio jam indicator and the second audio jam indicator.

[0225] It should be noted that the device in the device embodiment and the method embodiment are based on the same inventive concept.

[0226] An embodiment of the present application provides an audio detection device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio detection method provided in the above method embodiment.

[0227] Furthermore, Figure 12 The figure shows a hardware structure diagram of an audio detection device for implementing the audio detection method provided in the embodiment of the present application. The audio detection device may participate in or include the audio detection device provided in the embodiment of the present application. Figure 12As shown, the audio detection device 120 may include one or more ( Figure 12 The computer system includes a processor 1202 (shown as 1202a, 1202b, ..., 1202n) (the processor 1202 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1204 for storing data, and a transmission device 1206 for communication functions. In addition, the computer system may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. It will be understood by those skilled in the art that Figure 12 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 12 More or fewer components than shown, or with Figure 12 Different configurations shown.

[0228] It should be noted that the one or more processors 1202 and / or other data processing circuits may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any one of the other components in the audio detection device 120 (or mobile device). As described in the embodiments of the present application, the data processing circuitry acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0229] The memory 1204 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the audio detection method described in the embodiments of the present application. The processor 1202 executes various functional applications and data processing by running the software programs and modules stored in the memory 1204, thereby implementing the above-mentioned audio detection method. The memory 1204 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1204 may further include a memory remotely located relative to the processor 1202, and these remote memories may be connected to the audio detection device 120 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0230] Transmission device 1206 is used for receiving or sending data via a network.Above-mentioned network specific instance can comprise the wireless network that the communication supplier of audio detection equipment 120 provides.In an example, transmission device 1206 comprises a network adapter (Network Interface Controller, NIC), thereby it can be connected to other network devices by base station and can communicate with the Internet.In one embodiment, transmission device 1206 can be radio frequency (Radio Frequency, RF) module, and it is used for communicating with the Internet by wireless mode.

[0231] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the audio detection device 120 (or mobile device).

[0232] An embodiment of the present application also provides a computer-readable storage medium, which can be set in an audio detection device to store at least one instruction or at least one program related to the audio detection method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the audio detection method provided by the above-mentioned method embodiment.

[0233] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0234] Embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio detection method provided in the method embodiment.

[0235] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0236] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and apparatus embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.

[0237] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0238] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0239] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. An audio detection method, characterized in that: The method comprises: receiving an audio data packet generated based on the audio acquisition data; Receiving device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data; When the device status information indicates that the uplink terminal is operating normally, an audio jam indicator is obtained based on storage status information of an audio data packet cache area, where the storage status information is used to indicate whether an audio data packet to be played is stored in the audio data packet cache area.

2. The method according to claim 1, characterized in that The step of obtaining the audio jam indicator based on the storage status information of the audio data packet cache area includes: Reading storage status information of the audio data packet buffer area according to a preset processing cycle; the preset processing cycle is a sub-cycle within a preset freeze detection cycle; Based on the storage status information, performing audio jam detection on the preset processing period to obtain an audio jam record corresponding to the preset processing period; Determining audio scene type information; Based on a preset jamming influence factor corresponding to the audio scene type information, an audio jamming analysis is performed on the audio jamming records within the preset jamming detection period to determine the audio jamming index.

3. The method according to claim 2, characterized in that The performing audio jam detection on the preset processing period based on the storage status information to obtain the audio jam record corresponding to the preset processing period includes: When the storage status information indicates that no audio data packet to be played is stored in the audio data packet cache area, the duration of the preset processing cycle is used as the audio freeze duration corresponding to the preset processing cycle; An audio freeze record corresponding to the preset processing cycle is generated according to the audio freeze duration corresponding to the preset processing cycle.

4. The method according to claim 2, characterized in that The audio jam record includes: audio jam duration and audio jam number; the preset jam impact factor corresponding to the audio scene type information is based on the audio jam analysis of the audio jam record within the preset jam detection period; and determining the audio jam index includes: Performing weighted processing on the number of audio freezes within the preset freeze detection period based on the preset freeze impact factor to determine a first audio freeze index; Based on the audio freeze duration within the preset freeze detection period, performing continuous freeze analysis on the preset freeze detection period to determine a second audio freeze indicator; The audio jam indicator is determined based on the first audio jam indicator and the second audio jam indicator.

5. The method according to any one of claims 1 to 4, characterized in that: The device status information is obtained after analyzing the device operation status of the uplink terminal based on the status switching information of the call application, the collection interruption record of the call application, the first resource occupancy information corresponding to the call application, and the second resource occupancy information corresponding to the uplink terminal when the audio health information indicates that the audio collection status of the call application of the uplink terminal is abnormal.

6. The method according to claim 5, characterized in that The state switching information is obtained after the call application performs a foreground and background switching detection; the collection interruption record is obtained after the call application performs an audio collection interruption detection; the first resource occupancy information is obtained after the call application performs a resource occupancy detection; The second resource occupancy information is obtained after performing resource occupancy detection on the uplink terminal.

7. The method according to claim 1, characterized in that The audio health information is obtained after error analysis between the audio duration corresponding to the preset acquisition detection period and the period length of the preset acquisition detection period; the audio duration is obtained after audio duration conversion of the audio acquisition data read within the preset acquisition detection period.

8. An audio detection device, characterized in that: The device comprises: An audio data packet receiving module, configured to receive an audio data packet generated based on the audio acquisition data; A device status information receiving module, configured to receive device status information corresponding to the audio data packet; the device status information is obtained by analyzing the device operating status of the uplink terminal based on the audio health information corresponding to the audio collection data; A jam detection module is used to obtain an audio jam indicator based on the storage status information of the audio data packet cache area when the device status information indicates that the uplink terminal is operating normally, and the storage status information is used to indicate whether the audio data packet cache area stores the audio data packet currently to be played.

9. An audio detection device, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the audio detection method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product includes at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the audio detection method according to any one of claims 1 to 7.