Voice detection method and device in wake-up-free scene and electronic equipment
By combining primary and secondary voice activity detection in wake-free scenarios, the system determines whether target voice data exists within a preset time period, thus solving the problem of false VAD triggering and achieving efficient voice interaction and accurate voice recognition.
Patent Information
- Application Number
- CN202111439706.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-11-30
AI Technical Summary
In wake-free scenarios, Voice Activity Detection (VAD) is easily affected by environmental noise, resulting in a high probability of false triggering.
After the primary speech activity detection outputs the first speech start signal, it further determines whether target speech data exists within a preset duration. Combined with secondary speech activity detection, preset rules are used to determine whether valid speech data exists, thereby reducing the impact of noise.
It effectively reduces the probability of false triggering of VAD, improves user experience, simplifies voice interaction operations, and improves the accuracy of voice recognition.
Smart Images

Figure CN114121042B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic technology, specifically relating to a voice detection method, device, and electronic device in a wake-up-free scenario. Background Technology
[0002] With the development of electronic devices, most electronic devices can now incorporate voice interaction functionality. Typically, electronic devices use Voice Activity Detection (VAD) to detect voice activity in real-time recorded data. To improve user experience, many electronic devices currently employ a wake-up-free mode for voice interaction. In this way, once the VAD detects sound, it outputs voice data for subsequent speech recognition processing to achieve voice interaction. However, due to the influence of environmental noise, the probability of false triggering of the VAD is relatively high. Summary of the Invention
[0003] The purpose of this application is to provide a voice detection method, device, and electronic device in a wake-up-free scenario, which can solve the problem that the probability of VAD false triggering is high due to the influence of environmental noise.
[0004] In a first aspect, embodiments of this application provide a speech detection method in a wake-up-free scenario, including:
[0005] Obtain the audio data to be recognized;
[0006] Perform main speech activity detection on the audio data to be identified;
[0007] When the main speech activity detection outputs a first speech start signal, it is determined whether the audio data to be identified contains target speech data within a preset duration.
[0008] Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected.
[0009] Secondly, embodiments of this application provide a voice detection device for a wake-up-free scenario, characterized in that it includes:
[0010] The acquisition module is used to acquire the audio data to be recognized;
[0011] The detection module is used to detect the main speech activity in the audio data to be identified;
[0012] The determination module is used to determine whether the target speech data exists in the audio data to be identified within a preset duration when the main speech activity detection outputs a first speech start signal;
[0013] Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected.
[0014] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0015] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0016] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0017] This application embodiment acquires audio data to be identified; performs main voice activity detection on the audio data to be identified; and, when the main voice activity detection outputs a first voice start signal, determines whether target voice data exists in the audio data to be identified within a preset duration. Specifically, when target voice data exists in the audio data to be identified within the preset duration, a voice signal is detected. Because the existence of target voice data within a preset duration is further determined after the first voice start signal is output based on the main voice activity detection, environmental noise can be filtered to a certain extent, reducing the probability of false VAD triggering. Furthermore, this application embodiment achieves voice interaction without voice wake-up, simplifying the operation of voice interaction. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the speech detection method in a wake-up-free scenario provided in the embodiments of this application;
[0020] Figure 2 This is a structural diagram of the voice detection device in a wake-up-free scenario provided in the embodiments of this application;
[0021] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of this application;
[0022] Figure 4 This is a structural diagram of another electronic device provided in an embodiment of this application. Detailed Implementation
[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The speech detection method in the wake-free scenario provided by the embodiments of this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0026] See Figure 1 , Figure 1 This is a flowchart of a speech detection method in a wake-up-free scenario provided in an embodiment of this application, such as... Figure 1 As shown, it includes the following steps:
[0027] Step 101: Obtain the audio data to be recognized;
[0028] In this embodiment, the audio data to be identified can be real-time recording data, or audio data sent by other devices or stored locally. For example, an electronic device with a recording function can record in real time when the electronic device is powered on or when it is in operation. The electronic device can process the recorded data (e.g., noise reduction and echo cancellation) to obtain the recorded data, i.e., the aforementioned audio data. The operating state can be understood as being powered on and not in standby mode; for example, for an electronic device with a lock screen, it can be understood as the electronic device being in an unlocked state.
[0029] Step 102: Detect main speech activity in the audio data to be identified;
[0030] In this embodiment, taking audio recording as an example, during the recording process, main voice activity detection can be initiated to detect the recording data in real time, or the main detection process can be initiated after the recording ends to detect the recording data in real time using main voice activity detection. When the main voice activity detection detects the presence of sound, a first voice start signal can be output. It should be understood that in this embodiment, the first voice start signal can be understood as a voice start signal output after voice is detected, or as a voice start signal output after sound is detected.
[0031] Step 103: When the main speech activity detection outputs the first speech start signal, determine whether the audio data to be identified contains target speech data within a preset duration.
[0032] In this embodiment, the start time of the preset duration is located after the time when the first voice start signal is output. Specifically, if target voice data exists within the preset duration, it is determined that the audio data contains valid sound, i.e., a voice signal is detected; otherwise, it is determined that the audio data does not contain valid sound.
[0033] Optionally, if it is determined that target speech data exists within a preset duration in the audio data to be recognized, then valid target speech data can be extracted and output to perform related operations. For example, speech recognition can be performed on the output target speech data, and corresponding voice interaction functions can be executed based on the speech recognition results.
[0034] It should be understood that the aforementioned preset duration does not include the moment when the first voice start signal is output. Optionally, the start time of the preset duration being after the moment when the first voice start signal is output can be understood as the start time of the preset duration being later than the moment when the first voice start signal is output, and the interval between the start time of the preset duration and the moment when the first voice start signal is output being a target duration. For example, taking a certain time unit as the processing time unit, when the first voice start signal is detected and output in time unit 1, the start time of time unit 2 is determined as the start time of the preset duration, wherein time unit 2 and time unit 1 are two adjacent time units, and time unit 2 is located after time unit 1.
[0035] It should be noted that the preset duration can be set according to actual needs. For example, it can be a user-defined threshold (e.g., 500ms), or the preset duration can be intelligently adjusted. Determining that the target audio data exists within the preset duration can be understood as the electronic device determining, assuming, or believing that valid target audio data exists within the preset duration. Since the first speech start signal is output based on the speech activity detection, and then the presence of target audio data within the preset duration is further determined, environmental noise can be filtered to a certain extent, reducing the probability of VAD false triggering. VAD false triggering can be understood as the VAD falsely triggering the output of audio data, i.e., outputting abnormal audio data (e.g., noise data) to perform related operations. For example, in scenarios where the dialogue status is displayed in real time, this can significantly reduce false triggering of dialogue display, improving the user experience.
[0036] This application embodiment acquires audio data to be identified; performs main voice activity detection on the audio data to be identified; and, when the main voice activity detection outputs a first voice start signal, determines whether target voice data exists in the audio data to be identified within a preset duration. Specifically, when target voice data exists in the audio data to be identified within the preset duration, a voice signal is detected. Because the existence of target voice data within a preset duration is further determined after the first voice start signal is output based on the main voice activity detection, environmental noise can be filtered to a certain extent, reducing the probability of false VAD triggering. Furthermore, this application embodiment achieves voice interaction without voice wake-up, simplifying the operation of voice interaction.
[0037] It should be understood that the method for determining whether the audio data to be identified contains target speech data within a preset duration can be set according to actual needs. For example, in some embodiments, when the main speech activity detection outputs a first speech start signal, determining whether the audio data to be identified contains target speech data within a preset duration includes:
[0038] When the primary speech activity detection outputs a first speech start signal, the secondary speech activity detection is triggered to detect the audio data.
[0039] Determine whether target voice data exists within the preset duration according to preset rules;
[0040] The preset rule includes: if a preset condition is met, determining that no target speech data exists within the preset duration; otherwise, determining that target speech data exists within the preset duration. The preset condition includes at least one of the following:
[0041] Within the preset duration, the main speech activity detection outputs a speech end signal or the auxiliary speech activity detection outputs a speech end signal;
[0042] Within the preset duration, the primary speech activity detection does not output a speech end signal, and within the preset duration, the secondary speech activity detection does not output a speech end signal or a second speech start signal.
[0043] In this embodiment, whenever the primary speech activity detection outputs a first speech start signal, a secondary speech activity detection is triggered to determine the validity of the sound. At this time, the primary speech activity detection and the secondary speech activity detection simultaneously detect audio data within a preset duration, and judge the detection results based on preset rules to determine whether there is valid sound or target speech data within the preset duration.
[0044] Optionally, during the detection process, once the main speech activity detection or the auxiliary speech activity detection outputs a speech end signal, it indicates that the duration of the sound is short and it is an invalid sound. In other words, the electronic device can assume, consider, or confirm that the sound detected based on the first speech start signal is noise (i.e., determine that there is no target speech data within the preset duration), thereby not triggering the main VAD to output speech data. When the preset duration ends, if neither the primary speech activity detection nor the secondary speech activity detection outputs a speech end signal, it can be determined whether the secondary speech activity detection outputs a second speech start signal within the preset duration. If the secondary speech activity detection does not output a second speech start signal, it can be understood that the duration of the sound is short and therefore invalid. In other words, the electronic device can assume, consider, or confirm that the sound detected based on the first speech start signal is noise (i.e., it is determined that there is no target speech data within the preset duration), and thus will not trigger the primary VAD to output speech data. If the secondary speech activity detection outputs a second speech start signal, it can be understood that the duration of the sound is long and therefore valid. In other words, the electronic device can assume, consider, or confirm that the sound detected based on the first speech start signal is normal speech (i.e., it is determined that there is target speech data within the preset duration), and thus trigger the primary VAD to output speech data.
[0045] It should be noted that in some embodiments, the existence of valid sound within the preset duration can be determined solely based on the voice end signals output by the main voice activity detection and the auxiliary voice activity detection. For example, once the main voice activity detection or the auxiliary voice activity detection outputs a voice end signal, it can be determined that there is no target voice data within the preset duration; otherwise (i.e., neither the main voice activity detection nor the auxiliary voice activity detection outputs a voice end signal), it is determined that there is target voice data within the preset duration.
[0046] In this embodiment of the application, by using voice activity detection to determine whether target voice data exists within a preset time period, the impact of noise generation can be reduced, thereby reducing the frequency of false triggering of VAD.
[0047] Optionally, in some embodiments, where the above-described voice interaction is used to control an electronic device to perform corresponding voice interaction functions, the method may further include, after determining that a voice signal has been detected:
[0048] The target speech data is subjected to speech recognition to obtain the speech recognition result;
[0049] The corresponding voice interaction function is executed based on the speech recognition result.
[0050] In this embodiment, the aforementioned voice interaction function may include performing corresponding control operations, such as playing a specified multimedia file or adjusting the volume. The voice interaction function may also include performing voice dialogue, such as answering questions asked by the user.
[0051] Optionally, the execution information corresponding to the aforementioned voice interaction function can be stored in advance. For example, in some embodiments, multiple execution information can be stored in a preset database in advance, and the recognition result can be compared with the execution information stored in the preset database to determine the operation to be executed based on the comparison result.
[0052] Optionally, in some embodiments, the speech recognition result includes voiceprint recognition result and speech content recognition result, and the execution of the corresponding voice interaction function based on the speech recognition result includes:
[0053] The voiceprint recognition result is matched with preset voiceprint information;
[0054] If the voiceprint recognition result matches the preset voiceprint information, the corresponding voice interaction function is executed based on the voice content recognition result.
[0055] In this embodiment, when recognizing target voice data, voiceprint recognition can be added to obtain voiceprint recognition results. When the voiceprint recognition results match preset voiceprint information, the corresponding voice interaction function can be executed. Since voice interaction can be achieved directly without wake-up, when multiple people use the electronic device together, it is easy to accidentally trigger the electronic device to execute the voice interaction function. In order to improve the accuracy of voice recognition and reduce the probability of accidental triggering of the voice interaction function, it can only respond to the voice content that matches the pre-stored voiceprint information, that is, only respond to the voice commands of the specified user.
[0056] It should be understood that in this embodiment, the user can first record the corresponding voice in the corresponding settings interface, so that the electronic device can recognize and store the corresponding voiceprint information. In this way, the probability of false triggering of the voice interaction function can be reduced by comparing the voiceprint information during subsequent voice interaction.
[0057] Optionally, in some embodiments, the method further includes:
[0058] Obtain the sound information of the current ambient sound;
[0059] The preset duration is determined based on the sound information of the current ambient sound.
[0060] In this embodiment, the sound information of the current ambient sound can be understood as ambient noise information, which can specifically include the type of ambient noise, such as whether it is steady-state noise or non-steady-state noise. Steady-state noise can be understood as continuous noise, and non-steady-state noise can be understood as disordered noise. When the noise type is determined to be steady-state noise, the duration of the preset time can be appropriately increased; when it is non-steady-state noise, the duration of the preset time can be appropriately shortened.
[0061] Optionally, in some embodiments, the method further includes:
[0062] Get the historical frequency of false triggers;
[0063] The preset duration is determined based on the historical frequency of false triggers.
[0064] The aforementioned historical false trigger frequency can be understood as the number of times the VAD is falsely triggered within a certain historical period. For example, it can be the number of VAD false triggers corresponding to the N first voice start signals output by the main voice activity detection. The value of N can be set according to actual needs and is not further limited here. Optionally, in some embodiments, when the false trigger frequency is higher, the corresponding preset duration can be set to be shorter, and when the false trigger frequency is lower, the corresponding preset duration can be set to be longer.
[0065] It should be noted that in some implementations, the preset duration can be adjusted simultaneously based on the current ambient sound information and the historical frequency of false triggers. Because the preset duration can be adjusted based on the current ambient sound information and / or the historical frequency of false triggers, it can better adapt to detection in different environments, thereby further reducing the probability of VAD false triggers.
[0066] Furthermore, in some embodiments, the method further includes:
[0067] Determine the operating mode of the electronic device;
[0068] The electronic device operates in two modes: a first mode and a second mode. Determining whether target voice data exists within the preset duration according to preset rules includes:
[0069] When the electronic device is in the first working mode, it is determined whether target voice data exists within the preset duration according to preset rules.
[0070] In this embodiment, the first working mode can be understood as a wake-up-free mode, and the second working mode can be understood as a wake-up mode. When the second working mode is used, voice wake-up is required first (e.g., by voice input of the name of the electronic device), and voice input can be performed after wake-up.
[0071] It should be understood that the above-mentioned operating modes can be set through a designated interactive interface, or by voice or gestures, without further limitation. Because users can set the operating modes of electronic devices, they can be configured based on different application environments, such as using wake-up mode in public places and wake-free mode in private places. This improves the flexibility of using electronic products.
[0072] It should be noted that the speech detection method in the wake-free scenario provided in this application embodiment can be executed by a speech detection device in the wake-free scenario, or by a control module in the speech detection device in the wake-free scenario for executing the speech detection method in the wake-free scenario. This application embodiment uses the execution of the speech detection method in the wake-free scenario by the speech detection device in the wake-free scenario as an example to illustrate the speech detection method in the wake-free scenario provided in this application embodiment.
[0073] See Figure 2 , Figure 2 This is a structural diagram of the voice detection device in a wake-up-free scenario provided in the embodiments of this application, as shown below. Figure 2 As shown, the voice detection device 200 for wake-up-free scenarios includes:
[0074] The acquisition module 201 is used to acquire the audio data to be recognized;
[0075] Detection module 202 is used to detect the main speech activity of the audio data to be identified;
[0076] The determination module 203 is used to determine whether the target voice data exists in the audio data to be identified within a preset duration when the main voice activity detection outputs a first voice start signal;
[0077] Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected.
[0078] Optionally, the determining module 203 includes:
[0079] The triggering unit is used to trigger the auxiliary speech activity detection to detect audio data when the main speech activity detection outputs a first speech start signal;
[0080] The determining unit is used to determine whether target speech data exists within the preset duration according to preset rules;
[0081] The preset rule includes: if a preset condition is met, determining that no target speech data exists within the preset duration; otherwise, determining that target speech data exists within the preset duration. The preset condition includes at least one of the following:
[0082] Within the preset duration, the main speech activity detection outputs a speech end signal or the auxiliary speech activity detection outputs a speech end signal;
[0083] Within the preset duration, the primary speech activity detection does not output a speech end signal, and within the preset duration, the secondary speech activity detection does not output a speech end signal or a second speech start signal.
[0084] Optionally, the voice control device 200 further includes:
[0085] The recognition module is used to perform speech recognition on the target speech data when the audio data to be recognized contains target speech data within a preset duration, and to obtain a speech recognition result.
[0086] The execution module is used to execute the corresponding voice interaction function based on the speech recognition result.
[0087] Optionally, the speech recognition result includes voiceprint recognition result and speech content recognition result, and the execution module is specifically used to: match the voiceprint recognition result with preset voiceprint information; and, if the voiceprint recognition result matches the preset voiceprint information, execute the corresponding voice interaction function based on the speech content recognition result.
[0088] Optionally, the voice control device 200 further includes an adjustment module, wherein,
[0089] The acquisition module is also used to acquire the sound information of the current ambient sound;
[0090] The adjustment module is used to determine the preset duration based on the sound information of the current ambient sound.
[0091] Optionally, the voice control device 200 further includes an adjustment module, wherein,
[0092] The acquisition module is also used to acquire the historical frequency of false triggers;
[0093] The adjustment module is used to determine the preset duration based on the historical false trigger frequency.
[0094] The voice detection device in the wake-free scenario in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0095] The voice detection device in the wake-free scenario in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0096] The speech detection device for wake-up-free scenarios provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.
[0097] Optionally, this application embodiment also provides an electronic device, including a processor 310, a memory 309, and a program or instructions stored in the memory 309 and executable on the processor 310. When the program or instructions are executed by the processor 310, they implement the various processes of the above-described speech detection method embodiment in the wake-free scenario and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0098] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0099] Figure 4 A schematic diagram of the hardware structure of an electronic device for implementing the various embodiments of this application.
[0100] The electronic device 400 includes, but is not limited to, components such as: radio frequency unit 401, network module 402, audio output unit 403, input unit 404, sensor 405, display unit 406, user input unit 407, interface unit 408, memory 409, and processor 410.
[0101] Those skilled in the art will understand that the electronic device 400 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 410 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0102] The processor 410 is used to perform the following operations:
[0103] Obtain the audio data to be recognized;
[0104] Perform main speech activity detection on the audio data to be identified;
[0105] When the main speech activity detection outputs a first speech start signal, it is determined whether the audio data to be identified contains target speech data within a preset duration.
[0106] Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected.
[0107] Optionally, the processor 410 is specifically configured to perform the following operations:
[0108] When the primary speech activity detection outputs a first speech start signal, the secondary speech activity detection is triggered to detect the audio data.
[0109] Determine whether target voice data exists within the preset duration according to preset rules;
[0110] The preset rule includes: if a preset condition is met, determining that no target speech data exists within the preset duration; otherwise, determining that target speech data exists within the preset duration. The preset condition includes at least one of the following:
[0111] Within the preset duration, the main speech activity detection outputs a speech end signal or the auxiliary speech activity detection outputs a speech end signal;
[0112] Within the preset duration, the primary speech activity detection does not output a speech end signal, and within the preset duration, the secondary speech activity detection does not output a speech end signal or a second speech start signal.
[0113] Optionally, the processor 410 is also configured to perform the following operations:
[0114] When the audio data to be recognized contains target speech data within a preset duration, speech recognition is performed on the target speech data to obtain a speech recognition result.
[0115] The corresponding voice interaction function is executed based on the speech recognition result.
[0116] Optionally, the speech recognition result includes voiceprint recognition result and speech content recognition result, and the processor 410 is specifically used to perform the following operations:
[0117] The voiceprint recognition result is matched with preset voiceprint information;
[0118] If the voiceprint recognition result matches the preset voiceprint information, the corresponding voice interaction function is executed based on the voice content recognition result.
[0119] Optionally, the processor 410 is also configured to perform the following operations:
[0120] Obtain the sound information of the current ambient sound;
[0121] The preset duration is determined based on the sound information of the current ambient sound.
[0122] Optionally, the processor 410 is also configured to perform the following operations:
[0123] Get the historical frequency of false triggers;
[0124] The preset duration is determined based on the historical frequency of false triggers.
[0125] In this embodiment, after outputting the first voice start signal based on the main voice activity detection, it further determines whether target voice data exists within a preset time period. This can filter out environmental noise to a certain extent and reduce the probability of VAD false triggering. Furthermore, this embodiment achieves voice interaction without voice wake-up, simplifying the voice interaction operation.
[0126] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described speech detection method embodiments in the wake-free scenario and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0127] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0128] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described speech detection method embodiment in the wake-up-free scenario, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0129] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0132] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A speech detection method in a wake-up-free scenario, characterized in that, include: Obtain the audio data to be recognized; Perform main speech activity detection on the audio data to be identified; When the main speech activity detection outputs a first speech start signal, it is determined whether the audio data to be identified exists within a preset duration; wherein, when the first speech start signal is detected and output in time unit 1, the start time of time unit 2 is determined as the start time of the preset duration, time unit 2 and time unit 1 are two adjacent time units, and time unit 2 is located after time unit 1. Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected; Wherein, when the main speech activity detection outputs a first speech start signal, determining whether the target speech data exists in the audio data to be identified within a preset duration includes: When the primary speech activity detection outputs a first speech start signal, the secondary speech activity detection is triggered to detect the audio data. Determine whether target voice data exists within the preset duration according to preset rules; The preset rules include: if preset conditions are met, determining that no target speech data exists within the preset duration; otherwise, determining that target speech data exists within the preset duration. The preset conditions include: Within the preset duration, the primary speech activity detection does not output a speech end signal, and within the preset duration, the secondary speech activity detection does not output a speech end signal or a second speech start signal.
2. The method according to claim 1, characterized in that, After determining that a voice signal has been detected, the method further includes: The target speech data is subjected to speech recognition to obtain the speech recognition result; The corresponding voice interaction function is executed based on the speech recognition result.
3. The method according to claim 2, characterized in that, The speech recognition result includes voiceprint recognition result and speech content recognition result, and the execution of the corresponding voice interaction function based on the speech recognition result includes: The voiceprint recognition result is matched with preset voiceprint information; If the voiceprint recognition result matches the preset voiceprint information, the corresponding voice interaction function is executed based on the voice content recognition result.
4. The method according to claim 1, characterized in that, The method further includes: Obtain the sound information of the current ambient sound; The preset duration is determined based on the sound information of the current ambient sound.
5. The method according to claim 1, characterized in that, The method further includes: Get the historical frequency of false triggers; The preset duration is determined based on the historical frequency of false triggers.
6. A voice detection device for wake-up-free scenarios, characterized in that, include: The acquisition module is used to acquire the audio data to be recognized; The detection module is used to detect the main speech activity in the audio data to be identified; The determination module is used to determine whether the target voice data exists in the audio data to be identified within a preset duration when the main voice activity detection outputs a first voice start signal; wherein, when the first voice start signal is detected and output in time unit 1, the start time of time unit 2 is determined as the start time of the preset duration, time unit 2 and time unit 1 are two adjacent time units, and time unit 2 is located after time unit 1. Specifically, when the audio data to be identified contains target speech data within a preset duration, it is determined that a speech signal has been detected; The determining module includes: The triggering unit is used to trigger the auxiliary speech activity detection to detect audio data when the main speech activity detection outputs a first speech start signal; The determining unit is used to determine whether target speech data exists within the preset duration according to preset rules; The preset rules include: if preset conditions are met, determining that no target speech data exists within the preset duration; otherwise, determining that target speech data exists within the preset duration. The preset conditions include: Within the preset duration, the primary speech activity detection does not output a speech end signal, and within the preset duration, the secondary speech activity detection does not output a speech end signal or a second speech start signal.
7. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the voice detection method in the wakeless scenario as described in any one of claims 1 to 5.
8. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps of the speech detection method in the wake-free scenario as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Voice equipment and control method, device and equipment thereof
CN111816192A
Voice interaction method and device and voice chip module
CN113689847A