Audio signal detection method and apparatus, electronic device, and readable storage medium

By using a speech probability recognition model to detect the probability of speech signals in the ultrasonic frequency band in the context of speech recognition in electronic devices, the problem of undetected manipulation of electronic devices is solved, improving security and accuracy.

CN115938359BActive Publication Date: 2026-07-28GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2022-10-24
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

In the existing technology, electronic devices are operated through voice control commands that are not easily heard by the human ear, which reduces the safety of use.

Method used

In the context of voice recognition on electronic devices, a pre-acquired voice probability recognition model is used to determine whether the target probability of the voice signal is included in the ultrasonic frequency band of the audio signal to be recognized. When the probability is greater than or equal to a specified threshold, a prompt message is displayed to warn the user of the existence of audio signal interference.

Benefits of technology

It improves the security of electronic devices in voice recognition scenarios, accurately detects and alerts to audio signal interference, and prevents unauthorized manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938359B_ABST
    Figure CN115938359B_ABST
Patent Text Reader

Abstract

The application discloses an audio signal detection method and device, electronic equipment and a readable storage medium, which are applied to electronic equipment. The method comprises the following steps: in a voice recognition scene of the electronic equipment, determining a target probability that a voice signal is included in an ultrasonic frequency band of a to-be-recognized audio signal based on a pre-acquired voice probability recognition model; and if the target probability is greater than or equal to a specified threshold, displaying prompt information to prompt a user that there is audio signal interference. When the target probability that the voice signal is included in the ultrasonic frequency band is greater than or equal to the specified threshold, the user is prompted that there is audio signal interference, so that the detection of the audio signal interference in the voice recognition scene of the electronic equipment is realized, and the safety of the electronic equipment used in the voice recognition scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and more specifically, to an audio signal detection method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] Currently, with the rapid development of electronic information technology, voice control is increasingly appearing in people's lives. Electronic devices can acquire voice control commands issued by users and execute the corresponding operations. However, some devices manipulate electronic devices through voice commands that are not easily heard by the human ear, thereby controlling the electronic devices without the user's knowledge and reducing the security of using electronic devices. Summary of the Invention

[0003] This application proposes an audio signal detection method, apparatus, electronic device, and readable storage medium to improve the above-mentioned deficiencies.

[0004] In a first aspect, embodiments of this application provide an audio signal detection method applied to an electronic device. The method includes: in a speech recognition scenario of the electronic device, determining the target probability that the ultrasonic frequency band of the audio signal to be recognized includes a speech signal based on a pre-acquired speech probability recognition model; if the target probability is greater than or equal to a specified threshold, displaying a prompt message to alert the user that there is audio signal interference.

[0005] Secondly, this application also provides an audio signal detection device applied to the main control module of an electronic device. The device includes: an acquisition unit, configured to determine the target probability of a voice signal being included in the ultrasonic frequency band of the audio signal to be recognized, based on a pre-acquired voice probability recognition model, in the voice recognition scenario of the electronic device; and a prompting unit, configured to display a prompting message if the target probability is greater than or equal to a specified threshold, so as to prompt the user that there is audio signal interference.

[0006] Thirdly, embodiments of this application also provide an electronic device, including:

[0007] One or more processors; memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the method described in the first aspect.

[0008] The audio signal detection method, apparatus, electronic device, and readable storage medium provided in this application, in the context of speech recognition in electronic devices, can determine the target probability that the ultrasonic frequency band of the audio signal to be recognized includes a speech signal based on a pre-acquired speech probability recognition model. When the target probability is greater than or equal to a specified threshold, the user is alerted to the presence of audio signal interference. This application realizes the detection of audio signal interference in the speech recognition scenario of the electronic device, thereby alerting the user when audio signal interference is present, improving the security of using electronic devices in speech recognition scenarios. Furthermore, since the target probability represents the probability that the ultrasonic frequency band includes a speech signal, and the ultrasonic frequency band generally does not contain a speech signal (speech signals are usually artificially modulated into the ultrasonic frequency band as a type of interference audio signal), determining the presence of audio signal interference based on the target probability has high accuracy.

[0009] Other features and advantages of the embodiments of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the embodiments of this application. The objects and other advantages of the embodiments of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A structural block diagram of a microphone is shown;

[0012] Figure 2 A schematic diagram of a modulation method is shown;

[0013] Figure 3 The following diagram illustrates an application scenario of the audio signal detection method provided in this application embodiment;

[0014] Figure 4 A flowchart of the audio signal detection method provided in an embodiment of this application is shown;

[0015] Figure 5 A flowchart of an audio signal detection method according to another embodiment of this application is shown;

[0016] Figure 6 It shows Figure 5 A diagram illustrating one embodiment of step S210;

[0017] Figure 7 It shows Figure 5 Another embodiment of step S210 is shown in the figure;

[0018] Figure 8 It shows Figure 5 Another embodiment of step S210 is shown in the figure;

[0019] Figure 9 A flowchart of an audio signal detection method according to another embodiment of this application is shown;

[0020] Figure 10 A flowchart of an audio signal detection method according to another embodiment of this application is shown;

[0021] Figure 11 A structural block diagram of the audio signal detection device provided in an embodiment of this application is shown;

[0022] Figure 12 This paper shows a structural block diagram of an electronic device provided in an embodiment of the present application;

[0023] Figure 13 This paper shows a structural block diagram of a computer-readable storage medium provided in an embodiment of this application;

[0024] Figure 14 A structural block diagram of a computer program product provided in an embodiment of this application is shown. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.

[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] Currently, with the rapid development of electronic information technology, voice control is increasingly appearing in people's lives. Electronic devices can acquire voice control commands issued by people and then execute the corresponding operations.

[0028] Currently, users can control electronic devices by emitting voice signals. Specifically, the voice signal can include voice control commands, and the electronic device can include an audio acquisition device, such as a microphone or microphone array. The electronic device can acquire the user's voice signal through this audio acquisition device, then determine the corresponding voice control command based on the voice signal, and perform the corresponding operation based on the voice control command. For example, a user can emit the voice signal "play music." The electronic device can acquire the voice signal "play music" through the audio acquisition device, and then determine the voice control command as "play music," thereby controlling the electronic device to play music, such as launching a music playback-related application and playing music.

[0029] The audio acquisition device can be a microphone. Please refer to [link / reference]. Figure 1 , Figure 1 A structural block diagram of a microphone 100 is shown. Specifically, the microphone 100 may include an audio-to-electrical conversion module 110, a power pump 120, a low dropout regulator (LDO) 130, a low noise amplifier (LNA) 140, non-volatile memory (NVM) 150, a buffer circuit 160, and a reference circuit (Vref) 170. The audio-to-electrical conversion module 110 converts analog audio signals into electrical signals; the power pump 120 provides power to various modules in the electronic device; the low dropout regulator 130 stabilizes the voltage and prevents excessive voltage fluctuations; the low noise amplifier 140 amplifies the signal with a high signal-to-noise ratio; the non-volatile memory 150 stores data such as gain calibration parameters; the buffer circuit 160 provides output buffering or impedance transformation; and the reference circuit 170 provides a reference signal.

[0030] Furthermore, there are also some smart microphones, in Figure 1 Based on the microphone 100 shown, it also has digital interfaces such as a serial peripheral interface, through which internal parameters can be changed. For example, specific acoustic and electrical parameters such as the microphone's gain, sensitivity, and frequency response can be changed.

[0031] Generally, the frequency of voice signals emitted by users is in the range of 20Hz to 20kHz, which is also the range of frequencies that the human ear can normally hear. Therefore, when a user emits a voice signal to an electronic device, the human ear can hear the voice signal controlling the electronic device. However, the inventors discovered in their research that some devices operate electronic devices through voice control commands that are not easily heard by the human ear.

[0032] Specifically, some devices can use voice signals as baseband signals, modulating them onto carrier signals in other frequency bands. The voice signals include voice control commands for controlling electronic devices. By sending the modulated carrier signal to the electronic device, the device receives and demodulates the signal to obtain the voice signal contained within it. The device then operates the electronic device based on the voice control commands included in the voice signal. Since the frequency of this carrier signal is outside the range of human hearing, it is difficult for the human ear to hear during the transmission of the modulated carrier signal to the electronic device.

[0033] The modulation can include analog modulation and digital modulation. Analog modulation can include phase modulation (PM), frequency modulation (FM), and amplitude modulation (AM), etc.; digital modulation can include phase-shift keying (PSK), quadrature amplitude modulation (QAM), and frequency-shift keying (FSK), etc. For an example, please refer to [link to example]. Figure 2 , Figure 2 A schematic diagram is shown illustrating how a baseband signal is modulated onto a carrier wave using amplitude modulation (AM). Specifically, Figure 2 The horizontal axis represents time, and the unit can be microseconds (µs). Figure 2 Curve 101 represents a baseband signal, i.e., a low-frequency audio signal. Curve 102 represents a carrier signal, i.e., a high-frequency audio signal. Curve 103 represents an amplitude-modulated (AM) wave signal obtained by modulating the baseband signal onto the carrier signal using AM modulation.

[0034] The aforementioned carrier signal can be an ultrasonic signal, which is a signal with a frequency higher than 20kHz. The human ear's hearing frequency range is typically between 20 and 20kHz; therefore, ultrasonic signals cannot be directly heard by humans. However, audio acquisition devices included in electronic devices, such as microphones, can generally receive audio signals with frequencies above 20kHz, meaning they can acquire ultrasonic signals. Therefore, a lower-frequency speech signal can be modulated onto a high-frequency carrier to obtain a modulated signal, where the high-frequency carrier is the ultrasonic signal. This modulated signal can then be sent to the electronic device. After the electronic device acquires the modulated signal through the audio acquisition device, it can demodulate it to obtain the lower-frequency speech signal included in the modulated signal, thereby enabling control of the electronic device based on this speech signal.

[0035] The above analysis shows that the modulated signal is a high-frequency signal, such as an ultrasonic signal. Ultrasonic signals are not easily perceived by users. Therefore, if a low-frequency voice signal is modulated onto a high-frequency carrier, the modulated signal can be used to control electronic devices without the user's awareness, reducing the security of the electronic devices. This ultrasonic signal carrying a voice signal is generally called an ultrasonic attack signal or a dolphin-like attack signal.

[0036] Therefore, in order to overcome the above-mentioned defects, this application provides an audio signal detection method, apparatus, electronic device and readable storage medium. The method can prompt the user that there is audio signal interference when the target probability of including voice signal in the ultrasonic frequency band is greater than or equal to a specified threshold, thereby realizing the detection of audio signal interference in the voice recognition scenario of the electronic device and improving the security of the electronic device in the voice recognition scenario.

[0037] Please see Figure 3 , Figure 3 The illustration shows an application scenario diagram of an audio signal detection method provided in an embodiment of this application, namely audio signal detection scenario 300. Specifically, the audio signal detection scenario 300 includes an electronic device 310 and an audio interference signal 320.

[0038] In some implementations, the electronic device 310 can acquire audio signals from the surrounding environment in a specified scenario, and then acquire voice signals included in the audio signals. The specified scenario can characterize a scenario in which the electronic device 310 is running a specified application, which may include applications that require the acquisition of voice signals, such as call applications, recording applications, voice control applications, etc.

[0039] Furthermore, the specified scenario may include a voice recognition scenario. When the electronic device 310 is running a voice control application, it can be considered to be in a voice recognition scenario. At this time, the electronic device 310 can acquire voice signals, such as collecting audio signals from the surrounding environment, and then acquire voice signals based on the audio signals. It can then determine the control command corresponding to the voice signal and perform corresponding control on the electronic device 310 based on the control command. For example, a user can issue a voice signal of "play music." The electronic device can acquire the voice signal of "play music," and then determine the control command as "play music" based on "play music," thereby controlling the electronic device to play music, such as launching a music playback-related application and playing music.

[0040] Furthermore, the electronic device 310 can also control other devices that have pre-established a connection with it based on the acquired audio signal. These other devices can be smart devices, such as smart sockets, smart speakers, smart curtains, and smart lamps. Specifically, these other devices can establish a connection with the electronic device 310 through at least one communication method, such as Wi-Fi, Zigbee, or Bluetooth. For example, if the electronic device 310 acquires a voice signal saying "open the curtains," it can control the smart curtains to open based on that voice signal.

[0041] An audio interference signal 320 may exist in the environment surrounding the electronic device 310. This audio interference signal 320 can be a modulated signal carrying a voice signal, such as an ultrasonic signal. Since the frequency of ultrasonic signals exceeds the range of human hearing, this ultrasonic signal carrying a voice signal can be detected by the electronic device 310 without the user's notice, and subsequently used to control the electronic device 310 based on the voice signal, thereby reducing the safety of using the electronic device. The audio interference signal 320 can be emitted by an ultrasonic device.

[0042] Therefore, the electronic device 310 can detect the presence of audio interference signal 320. If an audio interference signal is detected, a prompt message is displayed to alert the user that audio signal interference exists. The specific detection method can be found in the description of subsequent embodiments.

[0043] Please see Figure 4 , Figure 4 An audio signal detection method provided in this application embodiment is illustrated. This method can be applied to the electronic device 300 in the foregoing embodiments. Specifically, the method includes steps S110 and S130.

[0044] Step S110: In the voice recognition scenario of the electronic device, based on the pre-acquired voice probability recognition model, determine the target probability that the ultrasonic frequency band of the audio signal to be recognized includes the voice signal.

[0045] In some implementations, a voice recognition scenario for an electronic device can be used to characterize the current scenario in which the electronic device triggers its voice recognition function. This triggering of the voice recognition function involves the electronic device acquiring an audio signal, extracting the voice signal from that audio signal, determining the control command within the voice signal, and then executing the operation corresponding to that control command. For example, if the application currently running on the electronic device is a voice control application, it can be determined that the electronic device is in a voice recognition scenario. For a detailed method for determining whether an electronic device is in a voice recognition scenario, please refer to the description of the following embodiments.

[0046] The above analysis shows that when electronic devices are in a voice recognition scenario, they can acquire audio signals. If the audio signal includes a voice signal, the electronic device can be controlled based on the control commands within that voice signal. However, if the audio signal also includes an ultrasonic signal carrying a voice signal, the electronic device can be controlled without the user's notice. For example, the electronic device could be controlled to perform operations corresponding to the voice signal, or other devices pre-connected to it could be controlled to perform operations corresponding to the voice signal, thus reducing the security of the electronic device. Therefore, in a voice recognition scenario, it is crucial to detect audio signal interference. If interference is detected, the user should be alerted to improve the security of the electronic device. Specifically, the presence of audio signal interference can indicate the presence of an ultrasonic signal carrying a voice signal, i.e., an ultrasonic attack or a dolphin-like attack.

[0047] In some implementations, the electronic device may first acquire the audio signal to be identified. This audio signal can be obtained by an audio acquisition device configured on the electronic device, which collects audio from the surrounding environment. The audio signal to be identified can include multiple frequency bands, such as the frequency band corresponding to the speech signal (20Hz to 20kHz); it can also include frequency bands lower than the speech signal (below 20Hz, also known as the infrasound band); and it can also include frequency bands higher than the speech signal (above 20kHz, also known as the ultrasonic band). Furthermore, the environment in which the electronic device is located may contain various audio signals. For example, the audio signal may include a voice signal, which may contain operating instructions for controlling electronic devices; the audio signal may also include an ultrasonic signal, which may carry a baseband signal, which may be a voice signal; the audio signal may also include other ultrasonic signals generated by special devices in the environment, which do not carry a baseband signal, such as the wave signals generated by some smart home devices when performing distance measurement through ultrasonic modules; the audio signal may also include infrasound signals generated by special devices in the environment, such as infrasound signals emitted by the vibration or shaking of large equipment.

[0048] The above analysis shows that audio signal interference indicates the presence of ultrasonic signals carrying speech signals. Therefore, it can be determined whether the ultrasonic frequency band of the acquired audio signal to be identified includes a speech signal. If it does, audio signal interference is determined. Specifically, the target probability that the ultrasonic frequency band of the audio signal to be identified includes a speech signal can be determined to judge whether audio signal interference exists. If the target probability is high, it indicates a high probability that the ultrasonic frequency band of the audio signal to be identified includes a speech signal, meaning there is a significant possibility of audio signal interference. If the target probability is low, it indicates a low probability that the ultrasonic frequency band of the audio signal to be identified includes a speech signal, meaning there is a significant possibility that there is no audio signal interference.

[0049] In some implementations, a target probability can be determined using a speech probability recognition model. This model determines the probability that an audio signal contains a speech signal. It's easy to understand that the speech probability recognition model outputs only one probability, which characterizes the probability that the audio signal contains a speech signal. A higher probability indicates a higher probability of the audio signal containing a speech signal; a lower probability indicates a lower probability of the audio signal containing a speech signal. For example, if the audio signal contains a user-generated speech signal, such as "turn up the volume," the probability determined by the speech probability recognition model is a large value, such as 90% or 100%; if the audio signal does not contain a user-generated speech signal, the probability determined by the speech probability recognition model is a small value, such as 5% or 0%.

[0050] Since speech signals are unique sound signals, they possess distinctive characteristics such as unique voiceprints, articulation intervals, and frequency ranges. These characteristics can be collected and organized as target features. The more closely the features of the acquired audio signal match the target features, the more likely it is to be considered to contain a speech signal, and thus a higher output probability value can be assigned. Currently, many mature algorithms or algorithm libraries are readily available for direct use. By fine-tuning these algorithms or libraries using the collected and organized speech signal features, the speech probability recognition model can be obtained; further details will not be elaborated upon here.

[0051] Furthermore, as the foregoing analysis shows, the audio signal to be identified can include multiple frequency bands, while the target probability is only the probability that the ultrasonic frequency band of the corresponding audio signal to be identified includes a speech signal. Therefore, for some implementation methods, the method of determining the target probability through a speech probability recognition model can be to first directly acquire the audio signal in the environment as the first audio signal, and acquire the audio signal in the environment as the second signal through a first filtering process, wherein the first filtering process is used to filter out the audio signal in the ultrasonic frequency band, that is, the difference between the frequency band corresponding to the first signal and the frequency band corresponding to the second signal is the ultrasonic frequency band. Then, through the speech probability recognition model, the first probability that a speech signal exists in the first signal and the second probability that a speech signal exists in the second signal are obtained respectively, and the difference between the first probability and the second probability is taken as the target probability. For example, if the first probability is 90% and the second probability is 0%, then the difference between the first probability and the second probability is 90%, that is, the target probability is 90%. A detailed description of this method can be found in the following embodiments.

[0052] In other implementations, the method for determining the target probability using a speech probability recognition model can also involve first obtaining an audio signal from the environment as a third audio signal through a second filtering process, wherein the second filtering process filters out audio signals other than the ultrasonic frequency band. Therefore, the frequency band corresponding to the third audio signal obtained after the second filtering process is the ultrasonic frequency band. Then, based on the pre-acquired speech probability recognition model, the target probability that the third audio signal includes a speech signal is determined, which is the target probability that the ultrasonic frequency band of the audio signal to be recognized includes a speech signal. A detailed description of this method can be found in the following embodiments.

[0053] Step S130: If the target probability is greater than or equal to a specified threshold, display a prompt message to alert the user that there is audio signal interference.

[0054] In some implementations, after obtaining the target probability through the aforementioned steps, the relationship between the target probability and a specified threshold can be determined. Therefore, a specified threshold can be preset. The obtained target probability is compared with the specified threshold. If the target probability is greater than or equal to the specified threshold, it can be considered that the ultrasonic frequency band of the audio signal to be identified contains a high probability of including a speech signal, and audio signal interference is determined to exist. If the target probability is less than the specified threshold, it can be considered that the ultrasonic frequency band of the audio signal to be identified contains a low probability of including a speech signal, and audio signal interference is determined to exist.

[0055] It is easy to understand that the higher the specified threshold is set, the higher the target probability required to determine the presence of audio signal interference, and thus the higher the accuracy of determining the presence of audio signal interference. Accuracy refers to correctly identifying the presence of audio signal interference. Conversely, the lower the specified threshold is set, the lower the target probability required to determine the presence of audio signal interference, and the less likely it is to result in false negatives. False negatives refer to incorrectly classifying the presence of audio signal interference as absent. It should be noted that this application does not specifically limit the specified threshold and can be flexibly set as needed.

[0056] Furthermore, if the target probability is greater than or equal to a specified threshold, audio signal interference is determined to exist. At this time, a prompt message can be displayed on the electronic device to alert the user of the audio signal interference. For example, the prompt message can be displayed through a display module configured on the electronic device, such as through the device's display screen. This prompt message can be text, such as "Attention: Audio signal interference detected"; it can also be text that flashes at specified time intervals, such as flashing once every 0.5 seconds, to increase its appeal to the user and make it easier for them to notice the prompt message. In other implementations, the prompt message can also be an audio message, such as being played through the electronic device's speaker.

[0057] Optionally, when audio signal interference is detected, the electronic device can also automatically stop running applications that require audio signal acquisition in the current speech recognition scenario. For example, if the application currently running on the electronic device that requires audio signal acquisition is a voice control application, then when audio signal interference is detected, the voice control application can be stopped. Stopping operation includes several aspects, specifically including stopping the execution of corresponding control operations through the voice control application and stopping the acquisition of ambient audio signals based on the voice control application.

[0058] Optionally, after stopping the application that requires audio signal acquisition in the current speech recognition scenario, the electronic device can also display a confirmation message to instruct the user to enter a confirmation command. Based on the user's confirmation command, the electronic device can then rerun the application that requires audio signal acquisition in the current speech recognition scenario, such as rerunning the voice control application. For example, the electronic device can display "Rerun voice control application?", and the user can enter a confirmation command to instruct the electronic device to rerun the voice control application.

[0059] The audio signal detection method, apparatus, electronic device, and readable storage medium provided in this application, in the context of speech recognition in electronic devices, can determine the target probability that the ultrasonic frequency band of the audio signal to be recognized includes a speech signal based on a pre-acquired speech probability recognition model. When the target probability is greater than or equal to a specified threshold, the user is alerted to the presence of audio signal interference. This application realizes the detection of audio signal interference in the speech recognition scenario of the electronic device, thereby alerting the user when audio signal interference is present, improving the security of using electronic devices in speech recognition scenarios. Furthermore, since the target probability represents the probability that the ultrasonic frequency band includes a speech signal, and the ultrasonic frequency band generally does not contain a speech signal (speech signals are usually artificially modulated into the ultrasonic frequency band as a type of interference audio signal), determining the presence of audio signal interference based on the target probability has high accuracy.

[0060] Please see Figure 5 , Figure 5 An audio signal detection method provided in this application embodiment is illustrated, which can be applied to the electronic device 300 in the foregoing embodiments. Specifically, the method includes steps S210 to S250.

[0061] Step S210: In the speech recognition scenario of the electronic device, based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal and determine the second probability that the second audio signal includes a speech signal.

[0062] As can be seen from the analysis of the foregoing embodiments, a first audio signal and a second audio signal can be collected by an electronic device, wherein the first audio signal is an audio signal collected by the electronic device, and the second audio signal is an audio signal obtained by the electronic device after being collected and processed by a first filtering process, wherein the first filtering process is used to filter out audio signals in the ultrasonic frequency band.

[0063] In some implementations, an audio acquisition device may be provided in the electronic device to acquire a first audio signal. Furthermore, the audio acquisition device may also have a first filtering function, which is used to perform the first filtering process, thus enabling the acquisition of a second audio signal via the audio acquisition device with the first filtering function enabled.

[0064] In other implementations, the electronic device may also include multiple audio acquisition devices, such as a first audio acquisition device and a second audio acquisition device. The first filtering function in the first audio acquisition device is turned off, while the first filtering function in the second audio acquisition device is turned on. This allows the acquisition of a first audio signal based on the first audio acquisition device with the first filtering function turned off, and the acquisition of a second signal based on the second audio acquisition device with the first filtering function turned on.

[0065] Furthermore, a first probability that the first audio signal contains a speech signal and a second probability that the second audio signal contains a speech signal can be determined using a pre-acquired speech probability recognition model. For details, please refer to [link to relevant documentation]. Figure 6 , Figure 6 A diagram illustrating one embodiment of step S210 is shown, wherein Figure 6 This includes steps S211 to S213.

[0066] Step S211: Obtain the audio signal of the first ambient sound as the first audio signal.

[0067] In some implementations, the electronic device may include an audio acquisition device, such as a microphone or microphone array, which can acquire audio from the environment, for example, the audio signal at the location of the electronic device.

[0068] Furthermore, the audio acquisition unit may also have a first filtering function, which can be used to perform the first filtering process. Since the first filtering process is used to filter out audio signals in the ultrasonic frequency band, the first filtering function can be implemented using a first filter, for example, a low-pass filter (LPF). The low-pass filter LPF can exhibit the characteristic of allowing input signals below the cutoff frequency to pass through, while blocking input signals above the cutoff frequency. Since the first filtering process is to filter out signals in the ultrasonic frequency band, the cutoff frequency of the low-pass filter LPF can be set at 20kHz, allowing audio signals with frequencies below 20kHz to pass through, while blocking audio signals above 20kHz.

[0069] In other embodiments, the first filter can also be a band-pass filter (BPF). A BPF can exhibit the characteristic of allowing input signals within its bandpass frequency range to pass through, while blocking input signals outside its bandpass frequency range. Therefore, in order to retain signals within the frequency range of the speech signal as much as possible while performing the first filtering process, the bandpass frequency range of the BPF can be set to 20Hz to 20kHz, or slightly less than 20Hz to 20kHz. In this case, the BPF allows audio signals within the 20Hz to 20kHz or slightly less than 20Hz to 20kHz frequency range to pass through, while blocking audio signals outside this range.

[0070] Furthermore, the audio acquisition device can also be equipped with a control module, which can control the first filter to be turned on or off. When the first filter is turned on, the audio acquisition device can perform the first filtering function, and when the control module controls the first filter to be turned off, the audio acquisition device can not perform the first filtering function.

[0071] In some implementations, the first audio signal can be an audio signal acquired without undergoing the first filtering process; that is, the first audio signal may include at least one of the infrasound frequency band, the speech signal frequency band, or the ultrasonic frequency band. Therefore, the first filter in the audio acquisition device can be turned off, thereby disabling the first filtering function. Then, with the first filtering function of the audio acquisition device turned off, the audio signal of the first ambient sound acquired by the audio acquisition device is obtained as the first audio signal.

[0072] Step S212: After a preset time delay, acquire the audio signal of the second ambient sound, and perform the first filtering process on the audio signal of the second ambient sound to obtain the second audio signal.

[0073] In some implementations, the second audio signal can be the acquired audio signal that has undergone the first filtering process, while the first audio signal is the acquired audio signal that has not undergone the first filtering process. Therefore, it can be seen that when acquiring the first audio signal through the audio acquisition device, the first filter needs to be turned off, thus disabling the first filtering function; while when acquiring the second audio signal through the audio acquisition device, the first filter needs to be turned on, thus enabling the first filtering function. Therefore, it is known that the audio acquisition device cannot simultaneously acquire both the first and second audio signals. Therefore, after acquiring the first audio signal, a preset time delay can be established before acquiring the second ambient sound audio signal acquired by the audio acquisition device, and the first filtering process can be performed on the second ambient sound audio signal to obtain the second audio signal.

[0074] Specifically, the first filter can be activated by the control module shown in the preceding steps to achieve the first filtering function. Then, with the first filtering function of the audio collector activated, the audio signal of the second ambient sound collected by the audio collector and processed by the first filtering is acquired as the second audio signal. For example, if the collected second ambient sound audio signal includes infrasound, speech signal, and ultrasonic frequency bands, after the first filtering process, the resulting second audio signal may include both infrasound and speech signal frequency bands.

[0075] In some implementations, the audio acquisition device may include one, which can acquire the audio signal of a first ambient sound when the first filter is not turned on, as the first audio signal; and then, when the first filter is turned on, acquire the audio signal of a second ambient sound, and perform the first filtering process on the audio signal of the second ambient sound to obtain the second audio signal.

[0076] In other embodiments, the audio acquisition device may include multiple devices, such as two, specifically including a first audio acquisition device and a second audio acquisition device. The first audio acquisition device may be controlled to not turn on the first filter and to acquire the first audio signal; the second audio acquisition device may be controlled to turn on the first filter and to acquire the second audio signal.

[0077] Specifically, the target probability needs to be the difference between the first probability corresponding to the first audio signal and the second probability corresponding to the second audio signal. This requires that the first ambient sound of the first audio signal and the second ambient sound of the second audio signal be the same ambient sound. The same ambient sound means that the audio signals present in the environment are identical. For example, if the ambient sound present in the environment at time A1 is B1, the ambient sound present in the environment at time A2 is B1, and the ambient sounds present in the environment at time A3 are both B1 and B2, it can be determined that the ambient sounds present in the environments corresponding to times A1 and A2 are the same, while the ambient sound present in the environment at time A3 is different from the ambient sounds present in the environments at times A1 and A2.

[0078] The preset time length can approximate the difference between the moment when the first audio signal is acquired by the audio acquisition device and the moment when the second audio signal is obtained after passing through the audio acquisition device and undergoing the first filtering process. Therefore, the preset time length can be set to a small value, ensuring that the difference between the moment when the first audio signal is acquired by the audio acquisition device and the moment when the second audio signal is obtained after passing through the audio acquisition device and undergoing the first filtering process is small. For example, since the pronunciation time of a single word or letter generally exceeds 10ms, the preset time length can be set to 10ms. Even if the preset time length is less than the pronunciation time of a single word or letter, the second audio signal is acquired after a 10ms delay. This ensures that the moment when the first audio signal is acquired by the audio acquisition device and the moment when the second audio signal is obtained after passing through the audio acquisition device and undergoing the first filtering process are approximately 10ms. Therefore, it can be ensured that the first ambient sound from which the first audio signal is acquired and the second ambient sound from which the second audio signal is acquired are the same ambient sound, thereby improving the reliability of determining the target probability based on the first probability and the second probability.

[0079] Step S213: Based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal and the second probability that the second audio signal includes a speech signal.

[0080] In some implementations, after obtaining the first audio signal and the second audio signal through the aforementioned steps, a first probability that the first audio signal includes a speech signal and a second probability that the second audio signal includes a speech signal can be determined based on a speech probability recognition model. Specifically, the obtained first audio signal can be used as an input signal to the speech probability recognition model, and the probability value output by the obtained speech probability recognition model can be used as the first probability. Since the first audio signal has not undergone the first filtering process, the first audio signal may include at least one frequency band among the infrasound frequency band, the speech signal frequency band, and the ultrasonic frequency band. Therefore, the first probability is the probability that the at least one frequency band includes a speech signal. As an example, if the first audio signal includes both the speech signal frequency band and the ultrasonic frequency band, then the first probability is the probability that the speech signal frequency band and the ultrasonic frequency band include a speech signal.

[0081] Furthermore, the acquired second audio signal can be used as an input signal into the speech probability recognition model, and the probability value output by the acquired speech probability recognition model can be used as the second probability. Since the second audio signal has undergone the first filtering process, the first audio signal does not include the ultrasonic frequency band. Therefore, the second probability can be the probability that at least one frequency band, including the infrasound frequency band and the speech signal frequency band, includes the speech signal. For example, if the second audio signal includes the speech signal frequency band, then the second probability is the probability that the speech signal frequency band includes the speech signal.

[0082] For other implementations, please refer to Figure 7 , Figure 7 A diagram illustrating another embodiment of step S210 is shown, wherein... Figure 7 This includes steps S214 to S217.

[0083] Step S214: Obtain the audio signal of the first ambient sound as the first audio signal.

[0084] Step S215: Based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal.

[0085] Step S214 has been described in detail in the foregoing embodiments and will not be repeated here. The method for determining the first probability based on the pre-acquired speech probability recognition model can also be found in the foregoing embodiments and will not be repeated here.

[0086] Step S216: If the first probability satisfies the probability condition, after a preset time delay, acquire the audio signal of the second ambient sound, and perform the first filtering process on the audio signal of the second ambient sound to obtain the second audio signal.

[0087] In some implementations, since the target probability is characterized by the difference between the first probability and the second probability, if the first probability is already small, it indicates that the probability that the ultrasonic frequency band of the audio signal to be identified includes a voice signal is also small. Therefore, it is not necessary to perform subsequent determination of the second probability and the determination of the target probability based on the first and second probabilities. Instead, it can be directly determined that there is no audio signal interference, thereby optimizing the process of the signal detection method and reducing the consumption of operating resources.

[0088] Specifically, subsequent operations such as acquiring the second audio signal can only be performed when the first probability meets the probability condition. In some implementations, the first probability meeting the probability condition may include a first probability greater than or equal to a specified probability. That is, if the first probability is greater than or equal to the specified probability, after a preset time delay, the audio signal of the second ambient sound is acquired, and the first filtering process is performed on the audio signal of the second ambient sound to obtain the second audio signal. The specified probability can be preset, for example, 10% or 20%.

[0089] It's easy to understand that the smaller the specified probability is set, the greater the chance that the first probability is greater than or equal to the specified probability. This means the first probability is more likely to meet the probability condition, leading to a higher probability of obtaining the second probability, determining the target probability, and then using the target probability to determine if audio signal interference exists. Conversely, the larger the specified probability is set, the smaller the chance that the first probability is greater than or equal to the specified probability. This means the first probability is less likely to meet the probability condition, thus saving system resources to a greater extent. It should be noted that this application does not limit the specific value of the specified probability; it can be flexibly set as needed.

[0090] Step S217: Based on the pre-acquired speech probability recognition model, determine the second probability that the second audio signal includes a speech signal.

[0091] The method for determining the second probability based on the pre-acquired speech probability recognition model can be found in the description of the aforementioned implementation, and will not be repeated here.

[0092] For some other implementation methods, please refer to Figure 8 , Figure 8 A diagram illustrating another embodiment of step S210 is shown, wherein... Figure 7 This includes steps S218 to S220.

[0093] Step S218: When the first filtering function of the first audio collector is turned off, acquire the audio signal of the first ambient sound collected by the first audio collector as the first audio signal, wherein the first filtering function is used to perform the first filtering process.

[0094] Step S219: When the first filtering function of the second audio collector is enabled, acquire the first ambient sound audio signal collected by the second audio collector and processed by filtering, and use it as the second audio signal.

[0095] In some implementations, the electronic device may also include multiple audio acquisition units, such as a first audio acquisition unit and a second audio acquisition unit. A first audio signal is acquired through the first audio acquisition unit, and a second audio signal is acquired through the second audio acquisition unit, thus ensuring that both the first and second audio signals can be acquired simultaneously. Both the first and second audio acquisition units may have a first filtering function, which performs the first filtering process. Therefore, the first audio signal can be acquired through the first audio acquisition unit with the first filtering function disabled, and the second audio signal can be acquired through the second audio acquisition unit with the first filtering function enabled.

[0096] Specifically, the method for acquiring the first audio signal through the first audio acquisition device and the method for acquiring the second audio signal through the second audio acquisition device are similar to the methods for acquiring the first and second audio signals through the audio acquisition device in the previous embodiments, and will not be described again here.

[0097] Step S220: Based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal and the second probability that the second audio signal includes a speech signal.

[0098] Step S220 has been described in detail in the foregoing embodiments and will not be repeated here.

[0099] Step S230: Obtain the difference between the first probability and the second probability as the target probability.

[0100] In some implementations, the first probability is used to characterize the probability that the first audio signal includes a speech signal, while the second probability is used to characterize the probability that the second audio signal includes a speech signal. Since the first audio signal has not undergone the first filtering process, and the second audio signal has undergone the first filtering process, the difference between the frequency band corresponding to the first audio signal and the frequency band corresponding to the second audio signal is the frequency band filtered by the first filtering process, which is the ultrasonic frequency band. Therefore, the difference between the first probability and the second probability can be used to characterize the target probability. For example, if the first probability is 90% and the second probability is 0%, it means that the speech signal has a high probability of existing in the ultrasonic frequency band filtered out by the first filtering process. Therefore, the difference between the first probability and the second probability, 90% - 0% = 90%, can characterize the target probability.

[0101] Step S250: If the target probability is greater than or equal to a specified threshold, display a prompt message to alert the user that there is audio signal interference.

[0102] Step S250 has been described in detail in the foregoing embodiments and will not be repeated here.

[0103] The audio signal detection method, apparatus, electronic device, and readable storage medium provided in this application can acquire first audio information and second audio information respectively in the context of speech recognition in electronic devices. Then, based on a pre-acquired speech probability recognition model, a first probability and a second probability are determined respectively. A target probability is determined based on the first and second probabilities. When the target probability is greater than or equal to a specified threshold, the user is alerted to the presence of audio signal interference. The target probability determined using the first and second probabilities has high accuracy, and the determination of the existence of audio signal interference based on this target probability has high reliability.

[0104] Please see Figure 9 , Figure 9 An audio signal detection method provided in this application embodiment is illustrated, which can be applied to the electronic device 300 in the foregoing embodiments. Specifically, the method includes steps S310 to S350.

[0105] Step S310: In the voice recognition scenario of the electronic device, acquire the audio signal to be recognized collected by the electronic device.

[0106] Step S330: Extract a third audio signal from the audio signal to be identified based on the second filtering process, wherein the second filtering process is used to filter out audio signals outside the ultrasonic frequency band.

[0107] In some implementations, in the voice recognition scenario of the electronic device, an audio signal to be recognized can be acquired by the electronic device. This audio signal to be recognized may include audio signals in different frequency bands, such as at least one of the infrasound band, the speech signal band, and the ultrasonic band. Furthermore, this audio signal to be recognized may exist in the environment, for example, as ambient sound.

[0108] Furthermore, since it is necessary to determine whether the ultrasonic frequency band in the audio signal to be identified includes a speech signal, the ultrasonic frequency band in the audio signal to be identified can be extracted to obtain a third audio signal, that is, the third audio signal only includes the ultrasonic frequency band. Then, based on the pre-acquired speech probability recognition module, the third probability that the third audio signal includes a speech signal is determined.

[0109] As an example, an electronic device may include a second filter that performs a second filtering process to filter out audio signals outside the ultrasonic frequency band. For example, the second filter may be a high-pass filter (HPF). The HPF can exhibit characteristics of blocking input signals below its cutoff frequency while allowing input signals above its cutoff frequency to pass through. Since the second filtering process is to filter out audio signals outside the ultrasonic frequency band, the cutoff frequency of the HPF can be set to 20kHz, meaning the HPF can block audio signals with frequencies below 20kHz while allowing audio signals with frequencies above 20kHz to pass through.

[0110] Step S350: Based on the pre-acquired speech probability recognition model, determine the third probability that the third audio signal includes the speech signal, as the target probability.

[0111] The method for determining the third probability that a third audio signal includes a speech signal based on a pre-acquired speech probability recognition model can be found in the foregoing embodiments of the method for determining the first probability that a first audio signal includes a speech signal and the second probability that a second audio signal includes a speech signal based on a pre-acquired speech probability recognition model, which will not be elaborated here.

[0112] Step S370: If the target probability is greater than or equal to a specified threshold, display a prompt message to alert the user that there is audio signal interference.

[0113] Step S370 has been described in detail in the foregoing embodiments and will not be repeated here.

[0114] The audio signal detection method, apparatus, electronic device, and readable storage medium provided in this application, in the context of speech recognition in electronic devices, directly acquire a third audio signal through a second filtering process, and then determine a third probability corresponding to the third audio signal based on a pre-acquired speech probability recognition model as a target probability. When the target probability is greater than or equal to a specified threshold, the user is alerted to the presence of audio signal interference. Determining the target probability through the third probability is simple to implement.

[0115] Please see Figure 10 , Figure 10 An audio signal detection method provided in this application embodiment is illustrated, which can be applied to the electronic device 300 in the foregoing embodiments. Specifically, the method includes steps S410 to S470.

[0116] Step S410: Obtain the program type of the currently running audio signal acquisition program.

[0117] Step S430: If the program type is a specified type, then the electronic device is determined to be in a voice recognition scenario.

[0118] It's easy to understand that electronic devices can run various applications. Some applications can call the audio signal acquisition module within the electronic device to capture audio signals and perform their corresponding functions. For example, if the application is a call application, the electronic device can call the audio signal acquisition module to capture audio signals and thus make a call; another example is that the application can be a recording application, in which case the electronic device can call the audio signal acquisition module to capture audio signals and thus make a call; yet another example is that the application can be a control application, in which case the electronic device can call the audio signal acquisition module to capture audio signals and thus control the electronic device, such as increasing the playback volume.

[0119] Furthermore, the aforementioned applications that require audio signal acquisition can correspond to different types. For example, call applications and recording applications do not need to control electronic devices based on the acquired audio signals, so these applications can be non-control types; while control applications need to control electronic devices based on the acquired audio signals, so these applications can be control types.

[0120] One example is that the control application of this type of control can be triggered by the user. For example, the user can input a command to start the control application on the electronic device, such as clicking the icon of the control application to start the control application; or the user can send a pre-agreed message to trigger the electronic device to start the control application. For example, the specified message can be "Hi mobile phone", then when the electronic device receives "Hi mobile phone", it can start the control application.

[0121] Since electronic devices are typically running control applications, if audio information interference occurs (i.e., they are subjected to ultrasonic attacks), they may be controlled by the voice signals carried within the ultrasonic attacks, thus reducing the security of the electronic devices. Therefore, it is possible to obtain the program type of the currently running audio signal acquisition program. If the program type is a specified type, it is determined that the electronic device is in a voice recognition scenario.

[0122] Step S450: In the voice recognition scenario of the electronic device, based on the pre-acquired voice probability recognition model, determine the target probability that the ultrasonic frequency band of the audio signal to be recognized includes the voice signal.

[0123] Step S470: If the target probability is greater than or equal to a specified threshold, display a prompt message to alert the user that there is audio signal interference.

[0124] Steps S450 and S470 have been described in detail in the foregoing embodiments and will not be repeated here.

[0125] Please see Figure 11 The diagram shows a structural block diagram of an audio signal detection device 1100 provided in an embodiment of this application. The device, applied to an electronic device, may include: an acquisition unit 1110 and a prompting unit 1120.

[0126] The acquisition unit 1110 is used to determine the target probability that the ultrasonic frequency band of the audio signal to be recognized includes a speech signal in the speech recognition scenario of the electronic device, based on a pre-acquired speech probability recognition model.

[0127] Furthermore, the acquisition unit 1110 is also used to determine a first probability that the first audio signal includes a speech signal and a second probability that the second audio signal includes a speech signal based on a pre-acquired speech probability recognition model; and to acquire the difference between the first probability and the second probability as the target probability.

[0128] Furthermore, the acquisition unit 1110 is also used to acquire the audio signal of the first ambient sound as the first audio signal; after a preset time delay, acquire the audio signal of the second ambient sound collected by the audio collector, and perform the first filtering process on the audio signal of the second ambient sound to obtain the second audio signal; based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal and determine the second probability that the second audio signal includes a speech signal.

[0129] Furthermore, the acquisition unit 1110 is also used to acquire the audio signal of the first ambient sound as the first audio signal; based on the pre-acquired speech probability recognition model, determine the first probability that the first audio signal includes a speech signal; if the first probability satisfies the probability condition, after a preset time delay, acquire the audio signal of the second ambient sound collected by the audio collector, and perform the first filtering process on the audio signal of the second ambient sound to obtain the second audio signal; based on the pre-acquired speech probability recognition model, determine the second probability that the second audio signal includes a speech signal.

[0130] Furthermore, the acquisition unit 1110 is also configured to acquire the audio signal of the first ambient sound collected by the audio collector as the first audio signal when the first filtering function of the audio collector is turned off; after a preset time delay, turn on the first filtering function; and when the first filtering function of the audio collector is turned on, acquire the audio signal of the second ambient sound collected by the audio collector and processed by the first filtering as the second audio signal.

[0131] Furthermore, the acquisition unit 1110 is also configured to acquire the audio signal of the first ambient sound collected by the first audio collector as the first audio signal when the first filtering function of the first audio collector is turned off, wherein the first filtering function is used to perform the first filtering process; when the first filtering function of the second audio collector is turned on, acquire the audio signal of the first ambient sound collected by the second audio collector and after filtering process as the second audio signal; and determine the first probability that the first audio signal includes a speech signal and the second probability that the second audio signal includes a speech signal based on a pre-acquired speech probability recognition model.

[0132] Furthermore, the acquisition unit 1110 is also used to acquire the audio signal to be identified collected by the electronic device; extract a third audio signal from the audio signal to be identified based on the second filtering process, wherein the second filtering process is used to filter out audio signals other than the ultrasonic frequency band; and determine the third probability that the third audio signal includes a speech signal as the target probability based on the pre-acquired speech probability recognition model.

[0133] Furthermore, the acquisition unit 1110 is also used to acquire the program type of the currently running audio signal acquisition program; if the program type is a specified type, then it is determined that the electronic device is in a voice recognition scenario.

[0134] The prompting unit 1120 is used to display a prompt message if the target probability is greater than or equal to a specified threshold, so as to prompt the user that there is audio signal interference.

[0135] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0136] In the several embodiments provided in this application, the coupling between the units can be electrical, mechanical or other forms of coupling.

[0137] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0138] Please see Figure 12 , Figure 12 This diagram illustrates a structural block diagram of an electronic device 1200 according to an embodiment of this application. The electronic device 1200 can be a smartphone, tablet computer, server, or other electronic device capable of running applications. The electronic device 1200 in this application may include one or more of the following components: a processor 1210 and a memory 1220, wherein the one or more processors are used in the methods described above.

[0139] The processor 1210 may include one or more processing cores. Optionally, the processor 1210 may be implemented in at least one of the following hardware forms: Microcontroller Unit (MCU), Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 1210 connects to various parts within the electronic device 1200 using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1220, and by calling data stored in the memory 1220, it performs various functions of the electronic device 1200 and processes data.

[0140] The memory 1220 may include Double Data Rate Synchronous Dynamic Random Access Memory (DDR) and Static Random Access Memory (SRAM), etc. Optionally, the memory 1220 may also include external storage, such as a hard disk drive (HDD), solid-state drive (SSD), USB flash drive, flash memory card, etc. The memory 1220 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1220 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the electronic device 1200 during use.

[0141] Please refer to Figure 13This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 1300 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0142] The computer-readable storage medium 1300 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1300 has storage space for program code 1310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1310 may be compressed, for example, in a suitable form.

[0143] Please refer to Figure 14 The diagram illustrates a structural block diagram of a computer program product 1400 provided in an embodiment of this application. The computer program product 1400 includes a computer program / instructions 1410, which, when executed by a processor, implements the steps of the aforementioned method.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An audio signal detection method, characterized in that, Applied to electronic devices, including: In the voice recognition scenario of the electronic device, based on the pre-acquired voice probability recognition model, a first probability that the first audio signal includes a voice signal and a second probability that the second audio signal includes a voice signal are determined. The first audio signal is an audio signal collected by the electronic device, and the second audio signal is an audio signal collected by the electronic device and obtained after a first filtering process. The first filtering process is used to filter out audio signals in the ultrasonic frequency band. The difference between the first probability and the second probability is obtained as the target probability; If the target probability is greater than or equal to a specified threshold, a prompt message will be displayed to alert the user that there is audio signal interference.

2. The method according to claim 1, characterized in that, The method of determining a first probability that a first audio signal contains a speech signal and a second probability that a second audio signal contains a speech signal, based on a pre-acquired speech probability recognition model, includes: Acquire the audio signal of the first ambient sound and use it as the first audio signal; After a preset delay, the audio signal of the second ambient sound is acquired, and the first filtering process is performed on the audio signal of the second ambient sound to obtain the second audio signal. Based on a pre-acquired speech probability recognition model, a first probability that the first audio signal includes a speech signal and a second probability that the second audio signal includes a speech signal are determined.

3. The method according to claim 1, characterized in that, The method of determining a first probability that a first audio signal contains a speech signal and a second probability that a second audio signal contains a speech signal, based on a pre-acquired speech probability recognition model, includes: Acquire the audio signal of the first ambient sound and use it as the first audio signal; Based on a pre-acquired speech probability recognition model, a first probability is determined that the first audio signal includes a speech signal; If the first probability satisfies the probability condition, after a preset time delay, the audio signal of the second ambient sound is acquired, and the first filtering process is performed on the audio signal of the second ambient sound to obtain the second audio signal. Based on a pre-acquired speech probability recognition model, a second probability is determined that the second audio signal includes a speech signal.

4. The method according to claim 2 or 3, characterized in that, The electronic device includes an audio acquisition unit, which has a first filtering function, and the first filtering function is used to perform the first filtering process. The acquisition of the first ambient sound audio signal, as the first audio signal, includes: When the first filtering function of the audio collector is turned off, the audio signal of the first ambient sound collected by the audio collector is acquired and used as the first audio signal; After the preset delay time, the audio signal of the second ambient sound is acquired, and the first filtering process is performed on the audio signal of the second ambient sound to obtain the second audio signal, including: After a preset delay, the first filtering function is activated; When the first filtering function of the audio collector is enabled, the audio signal of the second ambient sound collected by the audio collector and processed by the first filtering is acquired and used as the second audio signal.

5. The method according to claim 1, characterized in that, The electronic device includes a first audio acquisition unit and a second audio acquisition unit; the step of determining a first probability that the first audio signal includes a speech signal and a second probability that the second audio signal includes a speech signal based on a pre-acquired speech probability recognition model includes: When the first filtering function of the first audio collector is turned off, the audio signal of the first ambient sound collected by the first audio collector is acquired as the first audio signal, wherein the first filtering function is used to perform the first filtering process; When the first filtering function of the second audio collector is enabled, the audio signal of the first ambient sound collected by the second audio collector and filtered is acquired as the second audio signal; based on the pre-acquired speech probability recognition model, the first probability that the first audio signal includes a speech signal and the second probability that the second audio signal includes a speech signal are determined.

6. The method according to claim 1, characterized in that, Before determining the first probability that a first audio signal includes a speech signal and the second probability that a second audio signal includes a speech signal in the speech recognition scenario of the electronic device, based on a pre-acquired speech probability recognition model, the method further includes: Obtain the program type of the currently running audio signal acquisition program; If the program type is a specified type, then the electronic device is determined to be in a voice recognition scenario.

7. An audio signal detection device, characterized in that, The device, used as a main control module for electronic devices, includes: The acquisition unit is configured to, in the speech recognition scenario of the electronic device, determine a first probability that a first audio signal includes a speech signal and a second probability that a second audio signal includes a speech signal based on a pre-acquired speech probability recognition model, wherein the first audio signal is an audio signal collected by the electronic device, and the second audio signal is an audio signal collected by the electronic device and obtained after a first filtering process, wherein the first filtering process is used to filter out audio signals in the ultrasonic frequency band; and acquire the difference between the first probability and the second probability as a target probability. The prompting unit is used to display a prompt message if the target probability is greater than or equal to a specified threshold, so as to prompt the user that there is audio signal interference.

8. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-6.