Audio detection method, terminal equipment and computer readable storage medium

By dividing the audio stream collected by the microphone into multiple audio data frames and detecting their working state, the problems of high detection cost and low adaptability in the prior art are solved, and a higher precision microphone working state detection is achieved.

CN120034816APending Publication Date: 2025-05-23SHENZHEN ADDX INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195269.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When detecting the working state of a microphone, the prior art is susceptible to the limitations of additional hardware support and strict environmental conditions, resulting in high detection cost and low adaptability, and the situation where the microphone is blocked cannot be accurately detected.

Method used

By dividing the audio stream collected by the microphone into multiple audio data frames, and detecting the working state of the microphone based on these frames, the impact of abnormal data on the detection results is reduced, thereby improving detection accuracy.

Benefits of technology

This method can more accurately capture details in the audio stream, reduce the impact of abnormal data, improve the accuracy of microphone operating state detection, and reduce detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034816A_ABST
    Figure CN120034816A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of audio processing, and particularly relates to an audio detection method, terminal equipment and a computer readable storage medium. The method comprises the following steps: acquiring an audio stream collected by a microphone to be detected; dividing the audio stream into a plurality of audio data frames; and detecting the working state of the microphone to be detected according to the plurality of audio data frames to obtain a first detection result. Through the method, details of different audio segments in the audio stream can be captured, and the influence of abnormal data of a certain audio segment on the whole detection result can be reduced, so that the detection precision of the working state of the microphone is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of audio processing technology, and in particular, relates to an audio detection method, a terminal device, and a computer-readable storage medium. Background Art

[0002] With the popularity of mobile smart terminals, microphones, as an important audio input device, play an important role in scenarios such as voice calls, voice recognition, and environmental monitoring. However, in actual applications, microphones are easily blocked by fingers or other objects, resulting in the microphone being unable to collect audio data normally, thus affecting the accuracy of voice calls or voice recognition.

[0003] At present, the detection method for the working status of the microphone usually needs to set detection conditions, such as adding additional hardware support or strictly controlling the current sound pickup environment, which not only increases the detection cost, but also cannot accurately detect the working status of the microphone when the detection conditions are not met, and the adaptability of the detection method is low. Summary of the invention

[0004] The embodiments of the present application provide an audio detection method, a terminal device, and a computer-readable storage medium, which can improve the detection accuracy of the working status of a microphone.

[0005] In a first aspect, an embodiment of the present application provides an audio detection method, comprising:

[0006] Get the audio stream collected by the microphone to be detected;

[0007] Dividing the audio stream into a plurality of audio data frames;

[0008] The working state of the microphone to be detected is detected according to the multiple audio data frames to obtain a first detection result.

[0009] In an embodiment of the present application, by dividing the audio stream collected by the microphone into multiple audio data frames, and then detecting the working status of the microphone based on the multiple audio data frames, it is equivalent to detecting according to different audio segments in an audio stream. This not only can capture the details of different audio segments in the audio stream, but also can reduce the impact of abnormal data of a certain audio segment on the overall detection result, thereby effectively improving the detection accuracy of the working status of the microphone.

[0010] In a possible implementation manner of the first aspect, dividing the audio stream into multiple audio data frames includes:

[0011] Get the preset frame length;

[0012] The audio stream is divided into a plurality of audio data frames according to the preset frame length.

[0013] It is understandable that the preset frame length is smaller than the duration of the audio stream. The smaller the preset frame length, the smaller the granularity of the division, the fewer audio points included in each audio data frame, the higher the detection accuracy, and the larger the data processing volume; the larger the preset frame length, the larger the granularity of the division, the more audio points included in each audio data frame, the lower the detection accuracy, and the smaller the data processing volume.

[0014] In a possible implementation manner of the first aspect, detecting the blocking state of the to-be-detected microphone according to the multiple audio data frames includes:

[0015] Detecting the working state of the microphone to be detected corresponding to each audio data frame in N consecutive audio data frames, and obtaining a second detection result of each of the N consecutive audio data frames; wherein N is less than or equal to the number of audio data frames divided by the audio stream;

[0016] The working state of the microphone to be detected is determined according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result.

[0017] In a possible implementation manner of the first aspect, detecting the working state of the microphone to be detected corresponding to each audio data frame in the N consecutive audio data frames to obtain a second detection result of each of the N consecutive audio data frames includes:

[0018] For any one of the N consecutive audio data frames, counting the number of audio points in the audio data frame whose signal values ​​are less than a first preset value to obtain a first number;

[0019] If the first number is greater than a second preset value, the second detection result of the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocked state; wherein the second preset value is less than or equal to the total number of audio points in one of the audio data frames.

[0020] In a possible implementation manner of the first aspect, determining the working state of the microphone to be detected according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result includes:

[0021] Counting the number of audio data frames that meet a preset condition in N consecutive audio data frames to obtain a second number; wherein the preset condition is that the second detection result corresponding to the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocking state;

[0022] If the second number is greater than a third preset value, the first detection result indicates that the working state of the microphone to be detected is a blocking state; wherein the third preset value is less than or equal to N.

[0023] In an embodiment of the present application, the final detection result is determined by detecting N consecutive audio data frames. Since the consecutive audio data frames are connected in time, they can more accurately reflect the sound pickup conditions of the microphone compared to discontinuous audio data frames, thereby helping to improve the accuracy of audio detection.

[0024] In a possible implementation manner of the first aspect, the method further includes:

[0025] If the first detection result indicates that the working state of the microphone to be detected is a blocked state, a prompt message is issued to prompt the user that the microphone to be detected is blocked.

[0026] In this way, when the microphone is blocked, the user can be notified in time.

[0027] In a possible implementation manner of the first aspect, obtaining the preset frame length includes:

[0028] Obtaining the sampling frequency of the microphone to be detected;

[0029] The preset frame length is determined according to the sampling frequency of the microphone to be detected.

[0030] In this implementation, the preset frame length is determined according to the sampling frequency of the microphone, so that the division of the audio stream is more consistent with the sampling frequency of the microphone, which helps to improve the detection accuracy.

[0031] In a possible implementation manner of the first aspect, dividing the audio stream into a plurality of audio data frames according to the preset frame length includes:

[0032] Performing calibration processing on the audio stream to obtain the calibrated audio stream;

[0033] The calibrated audio stream is divided into a plurality of audio data frames according to the preset frame length. When the detection device is not connected to a microphone, the detection device may also detect an audio signal that is not zero due to its own detection error, which will affect subsequent detection results. In the above implementation, the calibration process can effectively filter out the influence caused by the detection error of the detection device, thereby helping to improve the accuracy of audio detection.

[0034] In a second aspect, an embodiment of the present application provides an audio detection device, including:

[0035] An acquisition unit, used to acquire an audio stream collected by the microphone to be detected;

[0036] A division unit, used for dividing the audio stream into a plurality of audio data frames;

[0037] The detection unit is used to detect the working state of the microphone to be detected according to the multiple audio data frames to obtain a first detection result.

[0038] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, an audio detection method as described in any one of the first aspects above is implemented.

[0039] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the audio detection method as described in any one of the above-mentioned first aspects is implemented.

[0040] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the audio detection method described in any one of the above-mentioned first aspects.

[0041] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0043] Figure 1 is a flowchart of an audio detection method provided in an embodiment of the present application;

[0044] Figure 2 is a schematic diagram of a detection system provided in an embodiment of the present application;

[0045] Figure 3 It is an audio schematic diagram provided in an embodiment of the present application;

[0046] Figure 4 is a structural block diagram of an audio detection device provided in an embodiment of the present application;

[0047] Figure 5 It is a schematic diagram of the structure of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0049] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0050] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0051] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0052] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0053] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0054] With the popularity of mobile smart terminals, microphones, as an important audio input device, play an important role in scenarios such as voice calls, voice recognition, and environmental monitoring. However, in actual applications, microphones are easily blocked by fingers or other objects, resulting in the microphone being unable to collect audio data normally, thus affecting the accuracy of voice calls or voice recognition.

[0055] In some related technologies, the working state of the microphone can be determined by key events. For example, the response to the case event is monitored; if the response is monitored, the microphone is determined to be working normally; if the response is not monitored, the microphone is determined to be blocked. In this way, when the key event is not synchronized with the actual blocking state, the working state of the microphone cannot be accurately determined. In addition, this method requires additional hardware support, which increases the detection cost.

[0056] Based on this, an embodiment of the present application provides an audio detection method. In the embodiment of the present application, by dividing the audio stream collected by the microphone into multiple audio data frames, and then detecting the working state of the microphone according to the multiple audio data frames, it is equivalent to detecting according to different audio segments in an audio stream, which can not only capture the details of different audio segments in the audio stream, but also reduce the influence of abnormal data of a certain audio segment on the overall detection result, thereby effectively improving the detection accuracy of the working state of the microphone.

[0057] See also Figure 1 , is a flowchart of an audio detection method provided in an embodiment of the present application. As an example but not a limitation, the method may include the following steps:

[0058] S101, obtaining an audio stream collected by a microphone to be detected.

[0059] For example, see Figure 2 , is a schematic diagram of a detection system provided in an embodiment of the present application. As an example and not a limitation, Figure 2 As shown, the detection system may include a microphone 21 and a detection device 22. The microphone 21 and the detection device 22 may be electrically connected or communicatively connected.

[0060] During the detection process, the detection device 22 executes the audio detection method in the embodiment of the present application, that is, obtains the audio stream collected by the microphone 21, and then detects the working state of the microphone 21 according to the audio stream.

[0061] It is understandable that in Figure 2 In the example, the microphone 21 can be used as the microphone to be detected.

[0062] An audio stream refers to a continuous data series of sound, which contains the smallest meaningful frame set of a given audio data format.

[0063] Optionally, the detection device 22 may obtain an audio stream of a preset duration. For example, the detection device 22 detects the working state of the microphone once according to the audio stream of 10 seconds each time it obtains the audio stream of 10 seconds.

[0064] Optionally, the detection device 22 may obtain an audio stream of a preset duration every preset period. For example, the detection device 22 obtains a 5s audio stream every 10s, and each time a 5s audio stream is obtained, the microphone working state is detected according to the 5s audio stream.

[0065] It should be noted that, in actual applications, the preset cycle and preset duration can be set according to detection requirements.

[0066] Optionally, the preset duration may also be set according to the sampling frequency of the microphone.

[0067] For example, the preset time length is greater than or equal to the sampling period corresponding to the sampling frequency. For example, assuming that the sampling frequency of the microphone is 48 Hz and the corresponding sampling period is 1 / 48 s, the preset time length is greater than or equal to 1 / 48 s.

[0068] For another example, the preset duration is a preset multiple of the sampling period corresponding to the sampling frequency. For example, assuming that the sampling frequency of the microphone is 48 Hz and the corresponding sampling period is 1 / 48 s, the preset duration is 8 times the sampling period, that is, 1 / 6 s.

[0069] S102: Divide the audio stream into a plurality of audio data frames.

[0070] It is understandable that the duration corresponding to the audio data frame is shorter than the duration corresponding to the audio stream. In other words, S102 is equivalent to dividing the audio stream into more fine-grained units in time.

[0071] In one embodiment, S102 may include:

[0072] Get the preset frame length;

[0073] The audio stream is divided into a plurality of audio data frames according to the preset frame length.

[0074] The frame length refers to the length of an audio data frame. An audio data frame includes a frame header (also called the start time or start time), a data portion, and a frame tail (also called the end time or end time). The time from the frame header to the frame tail is the frame length.

[0075] In one implementation, the preset frame length may be preset based on experience or detection requirements.

[0076] However, it is understandable that the preset frame length is smaller than the duration of the audio stream. The smaller the preset frame length, the smaller the granularity of the division, the fewer audio points included in each audio data frame, the higher the detection accuracy, and the larger the data processing volume; the larger the preset frame length, the larger the granularity of the division, the more audio points included in each audio data frame, the lower the detection accuracy, and the smaller the data processing volume.

[0077] In another implementation, the method of obtaining the preset frame length may include:

[0078] Obtaining the sampling frequency of the microphone to be detected;

[0079] The preset frame length is determined according to the sampling frequency of the microphone to be detected.

[0080] Optionally, the preset frame length may be greater than a sampling period corresponding to the sampling frequency and less than a duration corresponding to the audio stream.

[0081] Exemplarily, assuming that the duration corresponding to the audio stream is 1s, the sampling frequency of the microphone is 48Hz, and the corresponding sampling period is 1 / 48s, which is about 200ms. The preset frame length can be greater than 200ms and less than 1s. For example, the preset frame length is 300ms. In this case, the preset frame length is not an integer multiple of the sampling period, which is equivalent to splitting the audio data of one sampling period into different audio data frames. For another example, the preset frame length is 400ms. In this case, the preset frame length is an integer multiple of the sampling period, which is equivalent to including the audio data of the same sampling period in the same audio data frame.

[0082] In this implementation, the preset frame length is determined according to the sampling frequency of the microphone, so that the division of the audio stream is more consistent with the sampling frequency of the microphone, which helps to improve the detection accuracy.

[0083] Optionally, the audio data frames may be divided in such a manner that, starting from the start time of the audio stream, each time a preset frame length is reached, an audio data frame is divided.

[0084] For example, if the duration of an audio stream is 1 second and the preset frame length is 200 ms, the audio stream may be divided into five audio data frames.

[0085] In one case, the duration of the audio stream may not be an integral multiple of the preset frame length. In this case, the last remaining audio data may be discarded, or the last remaining audio data may be used as an audio data frame.

[0086] For example, if the duration of the audio stream is 1 second and the preset frame length is 300 ms, the first 900 ms of the audio stream can be divided into 3 audio data frames, and the remaining 100 ms of audio data can be discarded or used as one audio data frame.

[0087] Based on the above example, preferably, the preset duration and preset frame length of the audio stream are both integer multiples of the sampling period, and the preset frame length is smaller than the preset duration. In this case, it is convenient to divide the complete audio data frame.

[0088] In one implementation, the method of dividing the audio data frame may include:

[0089] Performing calibration processing on the audio stream to obtain the calibrated audio stream;

[0090] The calibrated audio stream is divided into a plurality of audio data frames according to the preset frame length.

[0091] In the embodiment of the present application, the calibration process is used to adjust the detection error of the detection device.

[0092] Optionally, the calibration processing method may include: obtaining an initial signal value of audio data monitored by the detection device when the microphone is not connected; subtracting the initial signal value from the signal value of each audio point in the audio stream to obtain a calibrated audio stream.

[0093] Of course, it is understandable that after the audio data frames are divided, each audio data frame may be calibrated separately.

[0094] When the detection device is not connected to a microphone, due to its own detection error, the detection device may also detect an audio signal that is not 0, which will affect subsequent detection results. In the above implementation, the influence caused by the detection error of the detection device can be effectively filtered out through calibration processing, thereby helping to improve the accuracy of audio detection.

[0095] S103: Detect the working state of the microphone to be detected according to the multiple audio data frames to obtain a first detection result.

[0096] In one embodiment, S103 may include:

[0097] Detect the working state of the microphone to be detected corresponding to each audio data frame to obtain a second detection result of each audio data frame; count the number of second detection results that meet the preset conditions; if the ratio of the number of second detection results that meet the preset conditions to the number of second detection results that do not meet the preset conditions is greater than a preset ratio, then the first detection result indicates that the working state of the microphone to be detected is a blocked state. Wherein, the preset condition is that the second detection result corresponding to the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocked state.

[0098] For example, assuming that the preset ratio is 1, the audio stream is divided into 5 audio data frames, among which the second detection results of 3 audio data frames meet the preset conditions, and the second detection results of 2 audio data frames do not meet the preset conditions, and the ratio of the two is 3 / 2>1, then the first detection result indicates that the working state of the microphone to be detected is in a blocked state.

[0099] In another embodiment, S103 may include:

[0100] Detecting the working state of the microphone to be detected corresponding to each audio data frame in N consecutive audio data frames, and obtaining a second detection result of each of the N consecutive audio data frames; wherein N is less than or equal to the number of audio data frames divided by the audio stream;

[0101] The working state of the microphone to be detected is determined according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result.

[0102] Exemplarily, the audio stream is divided into 5 audio data frames, and the first detection result is determined according to the second detection results of each of 3 consecutive audio data frames.

[0103] Optionally, the first detection result may be determined based on the second detection results of any N consecutive audio data frames in the multiple audio data frames. Specifically: the first detection result is determined based on the second detection results of the 1st to 3rd audio data frames; or, the first detection result is determined based on the second detection results of the 2nd to 4th audio data frames; or, the first detection result is determined based on the second detection results of the 3rd to 5th audio data frames.

[0104] Optionally, the first detection result may be determined based on every N consecutive audio data frames in the multiple audio data frames. Specifically: first determine the third detection result based on the second detection results of the 1st to 3rd audio data frames, determine the third detection result based on the second detection results of the 2nd to 4th audio data frames, and determine the third detection result based on the second detection results of the 3rd to 5th audio data frames; then determine the first detection result based on all the third detection results.

[0105] In one implementation, a method of obtaining the second detection result may include:

[0106] For any one of the N consecutive audio data frames, counting the number of audio points in the audio data frame whose signal values ​​are less than a first preset value to obtain a first number;

[0107] If the first number is greater than a second preset value, the second detection result of the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocked state; wherein the second preset value is less than or equal to the total number of audio points in one of the audio data frames.

[0108] For example, an audio data frame includes 10 audio points, among which the signal values ​​of 8 audio points are less than the first preset value, the signal values ​​of 2 audio points are greater than the first preset value, and the first number 8 is greater than the second preset value 5. Then the second detection result of the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is in a blocked state.

[0109] See also Figure 3 , is an audio schematic diagram provided in an embodiment of the present application. As an example and not a limitation, Figure 3 (a) in the figure shows a schematic diagram of the audio collected by the microphone normally. Figure 3 (b) in the figure shows an audio diagram when the microphone is blocked.

[0110] like Figure 3 As shown, in normal working state, the microphone can capture audio signals, and even if the surrounding environment is very quiet, the microphone can capture noise data in the environment, etc. When the microphone is blocked, the microphone cannot pick up any sound, including environmental noise.

[0111] Based on this, optionally, the first preset value may be 0.

[0112] Optionally, an implementation manner of determining the first detection result according to the second detection result may include:

[0113] Counting the number of audio data frames that meet a preset condition in N consecutive audio data frames to obtain a second number; wherein the preset condition is that the second detection result corresponding to the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocking state;

[0114] If the second number is greater than a third preset value, the first detection result indicates that the working state of the microphone to be detected is a blocking state; wherein the third preset value is less than or equal to N.

[0115] In one example, the audio stream is divided into 5 audio data frames, and when N=3, the first detection result is determined based on any N consecutive audio data frames. For example, for the 1st to 3rd audio data frames, where the 1st and 3rd audio data frames meet the preset condition, that is, the number 2 of the audio data frames that meet the preset condition in the N consecutive audio data frames is greater than the third preset value 1.5, then the first detection result indicates that the working state of the microphone to be detected is a blocked state.

[0116] In another example, the audio stream is divided into 5 audio data frames, and when N=3, the first detection result is determined according to every N consecutive audio data frames in the multiple audio data frames. Specifically: for the 1st to 3rd audio data frames, wherein the 1st and 3rd audio data frames meet the preset conditions, the third detection results corresponding to the 1st to 3rd audio data frames indicate that the working state of the microphone to be detected is a blocking state; for the 2nd to 4th audio data frames, wherein the 3rd and 4th audio data frames meet the preset conditions, the third detection results corresponding to the 2nd to 4th audio data frames indicate that the working state of the microphone to be detected is a blocking state; for the 3rd to 5th audio data frames, wherein the 3rd and 4th audio data frames meet the preset conditions, the third detection results corresponding to the 2nd to 4th audio data frames indicate that the working state of the microphone to be detected is a blocking state. If the number 3 of the third detection results that meet the preset conditions is greater than the number 0 of the third detection results that do not meet the preset conditions, then the first detection result indicates that the working state of the microphone to be detected is a blocking state.

[0117] In an embodiment of the present application, the final detection result is determined by detecting N consecutive audio data frames. Since the consecutive audio data frames are connected in time, they can more accurately reflect the sound pickup conditions of the microphone compared to discontinuous audio data frames, thereby helping to improve the accuracy of audio detection.

[0118] In one embodiment, the method further comprises:

[0119] If the first detection result indicates that the working state of the microphone to be detected is a blocked state, a prompt message is issued to prompt the user that the microphone to be detected is blocked.

[0120] Optionally, the prompt information may be an alarm sound, a text prompt or an alarm light, etc.

[0121] Optionally, if the first detection result indicates that the working state of the microphone to be detected is a blocking state, the detection device may automatically generate a fault report, wherein the fault report may include fault time and fault data (such as the audio stream corresponding to the blocking state).

[0122] In one example, an audio stream is obtained from a microphone to be detected; the audio stream is divided into multiple audio data frames of 20 ms; the second detection results of 5 consecutive audio data frames are determined; for each of the 5 consecutive audio data frames, if the signal values ​​of the audio points in the audio data frame are all 0, then the audio data frame meets the preset conditions; if the second detection results of the 5 consecutive audio data frames all meet the preset conditions, it is determined that the microphone to be detected is blocked.

[0123] In another example, an audio stream is obtained from the microphone to be detected; the audio stream is divided into multiple audio data frames of 30ms; the second detection results of three consecutive audio data frames are determined; for each of the three consecutive audio data frames, if the signal values ​​of the audio points in the audio data frame are all 0, then the audio data frame meets the preset conditions; if the second detection results of the three consecutive audio data frames all meet the preset conditions, it is determined that the microphone to be detected is blocked.

[0124] In another example, an audio stream is obtained from the microphone to be detected; the audio stream is divided into multiple audio data frames of 40ms; the second detection results of 8 consecutive audio data frames are determined; for each of the 8 consecutive audio data frames, if the signal values ​​of the audio points in the audio data frame are all 0, then the audio data frame meets the preset conditions; if the second detection results of the 8 consecutive audio data frames all meet the preset conditions, it is determined that the microphone to be detected is blocked.

[0125] In the embodiment of the present application, by dividing the audio stream collected by the microphone into multiple audio data frames, and then detecting the working state of the microphone based on the multiple audio data frames, it is equivalent to detecting based on different audio segments in an audio stream, which can not only capture the details of different audio segments in the audio stream, but also reduce the impact of abnormal data of a certain audio segment on the overall detection result, thereby effectively improving the detection accuracy of the working state of the microphone. In addition, determining the final detection result based on the detection results of N consecutive audio data frames can more accurately reflect the sound pickup situation of the microphone compared to discontinuous audio data frames, thereby helping to improve the accuracy of audio detection.

[0126] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0127] Corresponding to the audio detection method described in the above embodiment, Figure 4 This is a structural block diagram of the audio detection device provided in an embodiment of the present application. For the sake of convenience of explanation, only the parts related to the embodiment of the present application are shown.

[0128] Reference Figure 4 , the device 4 comprises:

[0129] The acquisition unit 41 is used to acquire the audio stream collected by the microphone to be detected.

[0130] The dividing unit 42 is configured to divide the audio stream into a plurality of audio data frames.

[0131] The detection unit 43 is used to detect the working state of the microphone to be detected according to the multiple audio data frames to obtain a first detection result.

[0132] Optionally, the dividing unit 42 is further configured to:

[0133] Get the preset frame length;

[0134] The audio stream is divided into a plurality of audio data frames according to the preset frame length.

[0135] Optionally, the detection unit 43 is further used for:

[0136] Detecting the working state of the microphone to be detected corresponding to each audio data frame in N consecutive audio data frames, and obtaining a second detection result of each of the N consecutive audio data frames; wherein N is less than or equal to the number of audio data frames divided by the audio stream;

[0137] The working state of the microphone to be detected is determined according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result.

[0138] Optionally, the detection unit 43 is further used for:

[0139] For any one of the N consecutive audio data frames, counting the number of audio points in the audio data frame whose signal values ​​are less than a first preset value to obtain a first number;

[0140] If the first number is greater than a second preset value, the second detection result of the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocked state; wherein the second preset value is less than or equal to the total number of audio points in one of the audio data frames.

[0141] Optionally, the detection unit 43 is further used for:

[0142] Counting the number of audio data frames that meet a preset condition in N consecutive audio data frames to obtain a second number; wherein the preset condition is that the second detection result corresponding to the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocking state;

[0143] If the second number is greater than a third preset value, the first detection result indicates that the working state of the microphone to be detected is a blocking state; wherein the third preset value is less than or equal to N.

[0144] Optionally, the detection unit 43 is further used for:

[0145] If the first detection result indicates that the working state of the microphone to be detected is a blocked state, a prompt message is issued to prompt the user that the microphone to be detected is blocked.

[0146] Optionally, the dividing unit 42 is further configured to:

[0147] Obtaining the sampling frequency of the microphone to be detected;

[0148] The preset frame length is determined according to the sampling frequency of the microphone to be detected.

[0149] Optionally, the dividing unit 42 is further configured to:

[0150] Performing calibration processing on the audio stream to obtain the calibrated audio stream;

[0151] The calibrated audio stream is divided into a plurality of audio data frames according to the preset frame length.

[0152] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0153] in addition, Figure 4 The audio detection device shown may be a software unit, a hardware unit, or a combination of software and hardware, which is built into an existing terminal device, or may be integrated into the terminal device as an independent accessory, or may exist as an independent terminal device.

[0154] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0155] Figure 5 Schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 5As shown, the terminal device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one is shown in the figure) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, and when the processor 50 executes the computer program 52, the steps in any of the above-mentioned audio detection method embodiments are implemented.

[0156] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 5 It is only an example of the terminal device 5 and does not constitute a limitation on the terminal device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0157] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0158] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as a hard disk or memory of the terminal device 5. In other embodiments, the memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the terminal device 5. The memory 51 is used to store an operating system, an application program, a boot loader, data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0159] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0160] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0161] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0162] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0163] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0164] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0165] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An audio detection method, characterized in that: include: Get the audio stream collected by the microphone to be detected; Dividing the audio stream into a plurality of audio data frames; The working state of the microphone to be detected is detected according to the multiple audio data frames to obtain a first detection result.

2. The audio detection method according to claim 1, characterized in that: The step of dividing the audio stream into a plurality of audio data frames comprises: Get the preset frame length; The audio stream is divided into a plurality of audio data frames according to the preset frame length.

3. The audio detection method according to claim 1, characterized in that: The detecting the blocking state of the microphone to be detected according to the multiple audio data frames includes: Detecting the working state of the microphone to be detected corresponding to each audio data frame in N consecutive audio data frames, and obtaining a second detection result of each of the N consecutive audio data frames; wherein N is less than or equal to the number of audio data frames divided by the audio stream; The working state of the microphone to be detected is determined according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result.

4. The audio detection method according to claim 3, characterized in that: The detecting the working state of the microphone to be detected corresponding to each audio data frame in the N consecutive audio data frames to obtain the second detection result of each of the N consecutive audio data frames includes: For any one of the N consecutive audio data frames, counting the number of audio points in the audio data frame whose signal values ​​are less than a first preset value to obtain a first number; If the first number is greater than a second preset value, the second detection result of the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocked state; wherein the second preset value is less than or equal to the total number of audio points in one of the audio data frames.

5. The audio detection method according to claim 3, characterized in that: The step of determining the working state of the microphone to be detected according to the second detection results of each of the N consecutive audio data frames to obtain the first detection result includes: Counting the number of audio data frames that meet a preset condition in N consecutive audio data frames to obtain a second number; wherein the preset condition is that the second detection result corresponding to the audio data frame indicates that the working state of the microphone to be detected corresponding to the audio data frame is a blocking state; If the second number is greater than a third preset value, the first detection result indicates that the working state of the microphone to be detected is a blocking state; wherein the third preset value is less than or equal to N.

6. The audio detection method according to any one of claims 1 to 5, characterized in that: The method further comprises: If the first detection result indicates that the working state of the microphone to be detected is a blocked state, a prompt message is issued to prompt the user that the microphone to be detected is blocked.

7. The audio detection method according to claim 2, characterized in that: The obtaining of the preset frame length comprises: Obtaining the sampling frequency of the microphone to be detected; The preset frame length is determined according to the sampling frequency of the microphone to be detected.

8. The audio detection method according to claim 2, characterized in that: The step of dividing the audio stream into a plurality of audio data frames according to the preset frame length includes: Performing calibration processing on the audio stream to obtain the calibrated audio stream; The calibrated audio stream is divided into a plurality of audio data frames according to the preset frame length.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.