Voice wake-up method, apparatus, storage medium, and electronic device
By integrating speech segments into voice wake-up detection, the problem of wake-up words being incorrectly segmented at slow speech speeds is solved, thus improving the wake-up success rate and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-09-28
- Publication Date
- 2026-05-12
AI Technical Summary
In slow speech scenarios, the success rate of voice wake-up in existing technologies is low, and the wake-up word is easily split into multiple speech segments, leading to wake-up failure.
After determining the first speech segment through speech activity endpoint detection, the time difference between its speech activity endpoint and the next speech activity start point is calculated. When the time difference is less than a preset value, the next speech segment is obtained and merged with the first speech segment to form a fused speech segment for wake-up detection.
It improves the success rate of voice wake-up in slow speech scenarios, avoids wake-up failure caused by mis-splitting of wake words, and improves the accuracy and efficiency of wake-up detection.
Smart Images

Figure CN115831109B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, specifically to a voice wake-up method, device, storage medium, and electronic device. Background Technology
[0002] With the continuous development of artificial intelligence technology, various AI-based products are constantly appearing in people's lives and work, bringing great convenience to their daily lives. Among them, smart assistants and smart speakers based on voice recognition technology are currently the most widely used.
[0003] Taking smart assistants as an example, when a user needs to use a smart assistant, they need to say the corresponding wake word to activate it. However, in some cases, such as when the user speaks slowly, saying the wake word may not be enough to activate the smart assistant, resulting in a low success rate. Summary of the Invention
[0004] This application provides a voice wake-up method, device, storage medium, and electronic device, which can improve the success rate of voice wake-up.
[0005] The voice wake-up method provided in this application includes:
[0006] Acquire voice information and perform voice activity endpoint detection on the voice information;
[0007] When the first speech segment is determined based on the speech activity endpoint detection result, wake-up detection is performed on the first speech segment;
[0008] When the wake-up detection result is wake-up failure, calculate the time difference between the end point of the speech activity of the first speech segment and the start point of the next speech activity;
[0009] When the time difference is less than a preset value, the second voice segment corresponding to the starting point of the next voice activity is obtained;
[0010] The first speech segment and the second speech segment are fused to obtain a fused speech segment, and wake-up detection is performed based on the fused speech segment.
[0011] The voice wake-up device provided in this application includes:
[0012] The first acquisition module is used to acquire voice information and perform voice activity endpoint detection on the voice information;
[0013] The first detection module is used to perform wake-up detection on the first speech segment when the first speech segment is determined based on the speech activity endpoint detection result;
[0014] The calculation module is used to calculate the time difference between the end point of the speech activity of the first speech segment and the start point of the next speech activity when the wake-up detection result is wake-up failure;
[0015] The second acquisition module is used to acquire the second voice segment corresponding to the starting point of the next voice activity when the time difference is less than a preset value.
[0016] The second detection module is used to fuse the first speech segment and the second speech segment to obtain a fused speech segment, and to perform wake-up detection based on the fused speech segment.
[0017] The storage medium provided in this application stores a computer program that, when loaded by a processor, executes the steps in the voice wake-up method provided in this application.
[0018] The electronic device provided in this application includes a processor and a memory. The memory stores a computer program, and the processor loads the computer program to execute the steps in the voice wake-up method provided in this application.
[0019] The computer program product provided in this application includes a computer program / instruction, which, when executed by a processor, implements the steps in the voice wake-up method provided in this application.
[0020] In this application, voice information is acquired, and voice activity endpoint detection is performed on the voice information. When a first voice segment is determined based on the voice activity endpoint detection result, wake-up detection is performed on the first voice segment. When the wake-up detection result is wake-up failure, the time difference between the voice activity endpoint of the first voice segment and the next voice activity start point is calculated. When the time difference is less than a preset value, a second voice segment corresponding to the next voice activity start point is acquired. The first voice segment and the second voice segment are fused to obtain a fused voice segment, and wake-up detection is performed based on the fused voice segment. Compared with related technologies, this application sets reasonable intervals between voice segments, so that even intermittently spoken wake words can be accurately recognized. This method can greatly improve the wake-up success rate in slow speech scenarios. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of a scenario for the voice wake-up system provided in an embodiment of this application.
[0023] Figure 2 This is a flowchart illustrating the voice wake-up method provided in the embodiments of this application.
[0024] Figure 3 This is a schematic diagram of voice activity endpoint detection provided in an embodiment of this application.
[0025] Figure 4 This is another schematic diagram illustrating the voice activity endpoint detection provided in the embodiments of this application.
[0026] Figure 5 This is another flowchart illustrating the voice wake-up method provided in this application.
[0027] Figure 6 This is a structural block diagram of the voice wake-up device provided in the embodiments of this application.
[0028] Figure 7 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0029] It should be noted that the principles of this application are illustrated by example in a suitable computing environment. The following description is based on the specific embodiments of this application exemplified, and should not be considered as limiting other specific embodiments not detailed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0030] The relational terms such as "first" and "second" used in the following embodiments of this application are only used to distinguish one object or operation from another, and are not intended to limit the actual order of these objects or operations. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0031] Artificial intelligence (AI) is the theory, methods, technology, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0032] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include machine learning (ML), with deep learning (DL) being a relatively new research direction within ML. It has been introduced into machine learning to bring it closer to its original goal: artificial intelligence. Currently, deep learning is mainly applied in fields such as computer vision and natural language processing.
[0033] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0034] To improve the success rate of voice wake-up, this application provides a voice wake-up method apparatus, a storage medium, and an electronic device. The voice wake-up method can be executed by an electronic device.
[0035] Please refer to Figure 1 This application also provides a voice wake-up system, such as Figure 1 The diagram illustrates a usage scenario of the voice wake-up system provided in this application. The system includes an electronic device 100. For example, the electronic device acquires voice information and performs voice activity endpoint detection on the voice information. When a first voice segment is determined based on the voice activity endpoint detection result, wake-up detection is performed on the first voice segment. When the wake-up detection result indicates a wake-up failure, the time difference between the voice activity endpoint of the first voice segment and the next voice activity start point is calculated. When the time difference is less than a preset value, a second voice segment corresponding to the next voice activity start point is acquired. The first and second voice segments are fused to obtain a fused voice segment, and wake-up detection is performed based on the fused voice segment. The electronic device 100 can send the fused voice segment to a server 200 loaded with a wake-up detection model for wake-up detection to obtain a detection result. After recognizing the recognition result, the server 200 returns the recognition result to the electronic device 100, enabling the electronic device 100 to perform voice wake-up based on the recognition result.
[0036] Electronic device 100 can be any device equipped with a processor and having voice processing capabilities, such as smartphones, tablets, PDAs, laptops, smart speakers and other mobile electronic devices with processors, or desktop computers, televisions, servers and other fixed electronic devices with processors.
[0037] It should be noted that, Figure 1 The schematic diagram of the voice wake-up system shown is merely an example. The voice wake-up system and scenario described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of voice wake-up systems and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0038] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating the voice wake-up method provided in an embodiment of this application. Figure 2 As shown, the flow of the voice wake-up method provided in this application embodiment can be as follows:
[0039] In S301, voice information is acquired and voice activity endpoint detection is performed on the voice information.
[0040] In related technologies, voice wake-up of electronic devices generally relies on endpoint detection technology to identify the wake word, and then wakes the electronic device based on the recognition result. Endpoint detection, also known as Voice Activity Detection (VAD), aims to distinguish between speech and non-speech regions. Specifically, endpoint detection accurately locates the start and end points of speech in noisy speech, removes silent parts, and finds the truly effective content of the speech. Voice activity endpoint detection is one of the key technologies for the correctness of voice wake-up; the accuracy of the endpoints directly determines the success or failure of wake-up to a certain extent. If endpoint detection is accurate, it can improve the wake-up effect and reduce the wake-up time. Endpoint detection can be roughly divided into two categories: feature-based methods and model-based methods. Feature-based methods refer to finding features that can distinguish between speech and noise, and judging speech segments and noise segments according to certain rules. The features used mainly include energy, fundamental frequency, zero-crossing rate, entropy, etc. Model-based methods refer to modeling noise and speech separately and using classification methods to detect endpoints. Among them, energy-based VAD is the simplest method; as long as the energy is less than a certain threshold, we consider it as silence or noise. This method is logically simple, computationally efficient, and easily applied to real-time systems. At normal speech speeds, the energy-based VAD can correctly detect endpoints. However, at slow speech speeds, due to discontinuous pronunciation and long periods of silence between wake-up words, traditional energy-based VADs will also classify these pauses as silent segments and cut them off, resulting in the wake-up word being split into two parts. This leads to inaccurate endpoint detection and poor wake-up performance in subsequent steps. For example, saying "Hello Xiaofei" at normal speech speed can wake up the device. However, if there is a pause between "Hello" and "Xiaofei," the electronic device will recognize "Hello" and "Xiaofei" as two separate speech segments, neither of which can independently wake up the device, resulting in a lower wake-up success rate in slow speech scenarios. To address these issues, this application proposes a voice wake-up method to improve the success rate of voice wake-up in slow speech scenarios. The voice wake-up method provided in this application will be described in detail below.
[0041] First, when the electronic device is in standby mode, it will intermittently acquire ambient voice. When the energy of the ambient voice is always below a set threshold, the environment can be considered to be in a silent state, and the electronic device will continue to acquire ambient voice intermittently. When the energy of the ambient voice is detected to be greater than the threshold, the electronic device will be triggered to continuously acquire ambient voice and perform wake word analysis on the acquired ambient voice.
[0042] In some embodiments, voice activity endpoint detection of voice information includes:
[0043] The speech information is divided into multiple consecutive audio frames;
[0044] Energy detection is performed on each audio frame;
[0045] Voice activity endpoint detection is performed on the voice information based on the energy of each audio frame.
[0046] Below, we can first introduce the specific methods for endpoint detection. For example... Figure 3 The diagram illustrates a speech activity endpoint detection method. For an audio stream, the energy value of each audio frame can be determined. An energy threshold can then be set. If, within a first number of audio frames starting from a given frame, the energy value of a second number of audio frames exceeds the energy threshold, that frame is considered the start of a speech activity. Here, the first number is greater than the second number. For example, if, within ten audio frames starting from the fifth frame, the energy value of eight frames exceeds the energy threshold, the fifth frame is considered the start of a speech activity. If, within a third consecutive number of frames preceding a given frame, the energy value is less than the energy threshold, that frame is considered the end of a speech activity. For instance, if the energy values of the ten audio frames preceding the 100th frame (frames 91 to 100) are all less than the energy threshold, the 100th frame is considered the end of a speech activity. Frames 91 to 100 can be referred to as silence frames. The speech data between the start and end points of a speech activity, i.e., the audio frames 5 to 100 mentioned above, are the speech segments for which wake-word recognition needs to be performed. The start point of a speech activity and its nearest subsequent end point can be called a speech activity endpoint combination.
[0047] When the user's input wake-up voice has a normal speaking speed, that is, when the user speaks a complete wake-up word in a speech segment between a combination of speech activity endpoints, such as saying "Hello Xiaofei", the wake-up voice can be determined based on the speech segment between the aforementioned combination of speech activity endpoints. Then, the wake-up voice is input into the wake-up detection model for feature extraction and detection to obtain the detection result. When the detection result meets the wake-up condition, the electronic device is woken up.
[0048] When the user's wake-up voice input is discontinuous, meaning there are pauses or silences between wake-up words—for example, pausing after saying "hello" before saying "Xiaofei," or pausing after saying "you" before saying "hello," and then pausing again before saying "Xiaofei"—then the next voice activity starting point can be detected. For example... Figure 4The figure shows a schematic diagram of the determination of voice activity endpoints in this application. As shown, in this embodiment, a first voice activity start point and a first voice activity end point can be determined in the audio stream corresponding to the acquired voice information, and a first voice segment can be determined accordingly. If the wake-up detection of the first voice segment fails, voice activity endpoint detection can be further performed to obtain a second voice activity start point and a second voice activity end point. Specifically, determining at least two consecutive combinations of voice activity endpoints in the voice information can be based on the energy value of each audio frame in the voice information according to the aforementioned voice activity endpoint detection method, that is, determining at least two consecutive combinations of voice activity endpoints in the voice information.
[0049] Specifically, voice activity endpoint detection is performed on the voice information based on the energy of each audio frame, including:
[0050] The energy of each audio frame is compared with the preset energy value one by one;
[0051] When a second number of audio frames with energy greater than the preset energy value are detected in a first number of audio frames after the first target audio frame, the first target audio frame is determined to be the starting point of the voice activity, and the first number is greater than or equal to the second number.
[0052] When the energy of the third consecutive number of audio frames preceding the second target audio frame is less than or equal to the preset energy value, the second target audio frame is determined to be the end point of the voice activity.
[0053] That is, this can be done according to Figure 3 The example speech activity endpoint detection method performs speech activity endpoint detection sequentially from the first frame of the speech information, that is, comparing the energy of each audio frame with a preset energy value. Here, the preset energy value can be the aforementioned energy threshold. Based on the comparison result of each audio frame with the energy threshold, and in conjunction with the aforementioned speech endpoint detection method, the speech activity start point can be detected. Once a speech activity start point is detected, the comparison of audio frame energy with the energy threshold can continue frame by frame until the speech activity end point corresponding to the speech activity start point is detected.
[0054] In some embodiments, energy detection is performed on each audio frame, including:
[0055] Obtain the number of sampling points and the amplitude of each sampling point for each audio frame;
[0056] The energy of each audio frame is calculated based on the number of sampling points in each audio frame and the amplitude of each sampling point.
[0057] In this embodiment, the energy value of each audio frame can specifically be the short-time energy of the audio frame. The short-time energy of the audio frame can be calculated based on the number of sampling points of the audio frame and the amplitude of each sampling point. This embodiment provides the following formula for calculating the short-time energy of the audio frame:
[0058]
[0059] Where n is the number of frames, and N is the frame length, i.e., the number of sampling points. N can be 160. It is the square of the amplitude value of the m-th sampling point of the n-th audio frame.
[0060] In S302, when the first speech segment is determined based on the speech activity endpoint detection result, wake-up detection is performed on the first speech segment.
[0061] As mentioned earlier, in related technologies, wake-up detection is generally performed on a single speech segment corresponding to a set of endpoints obtained from endpoint detection. That is, in related technologies, once a first speech segment is determined based on the speech activity endpoint detection results, only this first speech segment is used for wake-up detection. Specifically, this first speech segment can be input into a wake-up detection model to obtain the output wake-up detection result. Regardless of the wake-up detection model's detection result, it will output the wake-up detection result and flush its state. Thus, when a user speaks a wake-up word slowly, the interval between wake-up words is longer than the duration corresponding to the aforementioned silence frames (91 to 100 frames), causing the wake-up word to be split into two or more speech segments. Due to the wake-up detection model's mechanism of flushing its state after detection, wake-up detection is performed on each of the split speech segments individually. However, the detection result of individual wake-up detection on each of the split speech segments fails to meet the wake-up conditions, leading to wake-up failure.
[0062] In S303, when the wake-up detection result is wake-up failure, the time difference between the end point of the first speech segment and the start point of the next speech segment is calculated.
[0063] In this embodiment, when the wake-up detection result for the first speech segment is a wake-up failure, the wake-up detection model is not directly reset. Instead, speech activity endpoint detection continues to be performed on the speech information to determine the next speech activity start point. It can be understood that the next speech activity start point is the most recent speech activity start point after the end point of the first speech activity segment in terms of temporal sequence.
[0064] Once the start of the next speech activity is detected, the time difference between the end of the first speech activity segment and the start of the next speech activity can be calculated.
[0065] In some embodiments, the time difference between two speech segments can also be determined by calculating the number of audio frames between two adjacent speech segments. It is understood that the time difference between two adjacent speech segments must be greater than the time length corresponding to the aforementioned silence frames. For example, if the time length corresponding to the aforementioned 10 silence frames is 0.1 seconds, then the time difference here must be greater than 0.1 seconds.
[0066] In S304, when the time difference is less than a preset value, the second speech segment corresponding to the starting point of the next speech activity is obtained.
[0067] After calculating the time difference between adjacent speech segments, this time difference can be further compared with a preset value. This preset value can be a time interval, such as 1 second. In this embodiment, considering the user's slow speaking speed, when the user speaks the wake-up word slowly, there may be interruptions or silences between the spoken wake-up words. Therefore, speech segments spoken within a 1-second interval can also be considered part of the user's wake-up word segment. Thus, the wake-up voice spoken by the user can be determined based on the combination of these multiple speech segments, rather than based on a single speech segment corresponding to a single speech endpoint combination.
[0068] Therefore, when the time difference is determined to be less than a preset value, speech activity endpoint detection can continue to be performed on the speech information to determine the speech activity endpoint after the aforementioned next speech activity start point, that is, the speech activity end point corresponding to the next speech activity start point. Further, the speech segment between these two speech activity endpoints can be obtained, which can be referred to here as the second speech segment.
[0069] If the time difference between two speech segments is greater than the aforementioned preset value, such as 1 second, it can be determined that the user is speaking two speech segments, and it is not that the speech was mistakenly split into two segments due to slow speech speed. In this case, wake-up detection can be performed on each speech segment one by one.
[0070] In S305, the first speech segment and the second speech segment are fused to obtain a fused speech segment, and wake-up detection is performed based on the fused speech segment.
[0071] After obtaining the second speech segment, the first and second speech segments can be further fused to obtain a fused speech segment, also known as a wake-up speech. Then, wake-up detection is performed on the fused speech segment.
[0072] In some embodiments, the speech used for wake-up detection can be understood as wake-up speech. When wake-up detection is performed on the first speech segment, the wake-up speech is the first speech segment. If wake-up detection on the first speech segment fails, and a second speech segment is obtained and fused with the first speech segment to obtain a fused speech segment, the wake-up speech can be updated using the fused speech segment, and then wake-up detection can be performed on the wake-up speech again. If wake-up detection still fails, the next speech activity start point can be obtained, and the time difference between it and the speech activity end point of the second speech activity segment can be calculated. If the time difference is still less than the aforementioned preset value, a third speech activity segment can be obtained, and the third speech activity segment can be further fused with the aforementioned fused segment to update the wake-up speech, and then wake-up detection can continue to be performed on the wake-up speech.
[0073] In some embodiments, when the time difference is less than a preset value, the wake-up voice is determined based on the voice segment corresponding to the corresponding voice activity endpoint combination, including:
[0074] When two consecutive combinations of speech activity endpoints are identified in the speech information and the time difference is less than a preset value, the two speech segments corresponding to the two consecutive combinations of speech activity endpoints are determined.
[0075] The wake-up voice is obtained by splicing two audio segments in chronological order.
[0076] In this embodiment of the application, when the number of consecutive voice activity endpoint combinations determined in the voice information is two, i.e., as follows Figure 4 The example shown illustrates that when two sets of speech activity endpoints are identified in the speech information, i.e., when two speech segments are identified, these two speech segments can be spliced together in chronological order to obtain the wake-up speech.
[0077] Specifically, for example, if the first voice segment is "Hello" and the second voice segment is "Xiaofei", then splicing these two voice segments together will produce a new voice segment "Hello Xiaofei" as the wake-up voice. Then, wake-up detection can be performed on the wake-up voice "Hello Xiaofei".
[0078] In some embodiments, when the time difference is less than a preset value, the wake-up voice is determined based on the voice segment corresponding to the corresponding voice activity endpoint combination, including:
[0079] When multiple consecutive combinations of voice activity endpoints are identified in the voice information, and the time difference between any two adjacent combinations of voice activity endpoints is less than a preset value, the voice segment corresponding to each combination of voice activity endpoints is determined, resulting in multiple voice segments.
[0080] The wake-up voice is obtained by splicing multiple voice segments in chronological order.
[0081] In this embodiment of the application, when multiple sets (greater than or equal to three sets) of voice activity endpoint combinations are determined in the voice information, it is necessary to determine the time difference between each adjacent voice segment one by one. That is, to determine whether the time difference between the beginning and end of each pair of adjacent voice activity endpoint combinations is less than the aforementioned preset value. When the time difference between all adjacent voice segments is less than the aforementioned preset value, the multiple voice segments corresponding to the multiple sets of voice activity endpoint combinations can be spliced together in chronological order to obtain the wake-up voice.
[0082] Specifically, for example, when four sets of voice activity endpoint combinations are identified from the voice information, i.e., four voice segments corresponding to "you," "hello," "small," and "fly," the time difference between all adjacent voice segments can be determined one by one. When the time difference between all adjacent voice segments is less than a preset value, the multiple voice segments can be concatenated in chronological order to form the wake-up voice "Hello Xiaofei." It is understandable that the multiple sets of voice activity endpoint combinations identified from the voice information, i.e., the multiple voice segments, may also include information other than the wake-up word, such as "play music." For example, if the concatenated wake-up voice is "Hello Xiaofei, play music," when the wake-up detection system detects "Hello Xiaofei" during the wake-up voice time-based wake-up detection process, it can determine that the detection result meets the wake-up conditions, and the electronic device can be directly woken up.
[0083] In some embodiments, when multiple consecutive combinations of voice activity endpoints are identified in the voice information, the consecutive voice segments can be spliced together to determine the wake-up voice if only the time difference between consecutive voice segments is less than a preset value. Specifically, for example, if three combinations of voice activity endpoints are identified in the voice information, and the corresponding voice segments include "you," "hello," "Xiaofei," and "play music," then if the time difference between adjacent voice segments of the first three voice segments is less than a preset value, the first three voice segments can be spliced together to form the wake-up voice "hello Xiaofei."
[0084] As mentioned above, the step of determining whether the time difference between adjacent speech segments is less than a preset value can also be achieved by determining whether the number of audio frames between adjacent speech segments is less than a preset value, such as determining whether the number of audio frames between adjacent speech segments is less than 100 frames.
[0085] The method provided in this application can avoid the problem of failure to wake up the user in slow speech scenarios when the wake-up voice is divided into multiple speech segments and then the wake-up detection is performed on each individual speech segment. This can improve the accuracy of wake-up detection and thus improve the wake-up success rate of electronic devices.
[0086] In this application, after determining the wake-up speech in a slow-speed scenario using the method described above, wake-up detection can be performed based on the wake-up speech, and then the wake-up target can be woken up by voice based on the wake-up detection result. Specifically, the wake-up target can be the electronic device itself, or other objects that need to be woken up.
[0087] In some embodiments, wake-up detection of the first speech segment includes:
[0088] The first speech segment is input into the wake-up detection model for wake-up detection, and the wake-up detection result is obtained.
[0089] The first and second speech segments are fused to obtain a fused speech segment, and wake-up detection is performed based on the fused speech segment, including:
[0090] The second speech segment is input into the wake-up detection model and used together with the first speech segment for wake-up detection.
[0091] Specifically, in this embodiment, wake-up detection of speech segments can be implemented using a wake-up detection model. The recognized speech segment (wake-up speech) is input into the wake-up detection model for wake-up detection. Upon receiving the wake-up speech, the wake-up detection model first extracts features from the speech, then sends the extracted features to a multilayer perceptron (MLP) for feature mapping to obtain a detection result. This detection result can be a probability value, referred to as the wake-up probability, which indicates the probability that the wake-up speech meets the wake-up conditions. When the wake-up probability is greater than a preset value, wake-up can be determined as successful, and the target object can then be woken up via voice.
[0092] In some embodiments, the voice wake-up method provided in this application may further include:
[0093] When the wake-up detection result is a successful wake-up, the object to be woken up is woken up by voice.
[0094] The wake-up detection model is used to reset its state.
[0095] In this embodiment, when the wake-up detection model detects a wake-up failure, the model state is not reset temporarily. It can wait for the input of the next speech segment. After the next speech segment is input, it can be merged with the previous speech segment into a single speech segment for wake-up detection. If the currently obtained recognition result meets the wake-up conditions, the wake-up object is activated and the wake-up detection model is initialized.
[0096] Specifically, for example, when three speech segments—"Hello," "Xiaofei," and "Play Music"—are identified in the speech information, and the time difference between any two adjacent speech segments is less than a preset value, the speech segment corresponding to "Hello" is first identified and input into the wake-up detection model for detection. At this point, the wake-up detection result does not meet the wake-up condition, and the wake-up model does not reset its state but only returns the wake-up result. Then, when it is detected that the time difference between the "Xiaofei" speech segment and the "Hello" speech segment is less than the preset value, the "Xiaofei" speech segment can be input into the wake-up detection model for wake-up detection. At this point, the detected result "Hello Xiaofei" meets the wake-up condition, and the electronic device can be directly woken up without needing to input the "Play Music" speech segment into the wake-up detection model for further recognition. Furthermore, the wake-up detection model state can be reset, and then the "Play Music" speech segment can be input into the wake-up detection model for a new round of speech recognition.
[0097] In some embodiments, the voice wake-up method provided in this application further includes:
[0098] When the time difference is greater than or equal to a preset value, the wake-up detection model is reset.
[0099] Specifically, when calculating the time difference between two speech segments, if the calculated time difference is not less than the aforementioned preset value, the wake-up detection model's waiting state is terminated, and its state is directly reset so that it can enter the next round of wake-up detection.
[0100] The solution provided in this embodiment can avoid the problem of slow wake-up in slow speech scenarios. It can wake up the electronic device immediately after recognizing a wake-up word that meets the wake-up conditions, thereby improving wake-up efficiency.
[0101] The following will provide a detailed description of the specific implementation process of the technical solution provided in this application using a concrete example.
[0102] like Figure 5 The diagram shown is another flowchart illustrating the voice wake-up method provided in this application.
[0103] In S401, the initial state of endpoint detection is set to "no voice activity start point found" and "no voice activity end point found".
[0104] First, set the initial state of VAD. The initial state is that no voice activity start point (i.e., isVADStart = false) and no voice activity end point (i.e., isVADEnd = false).
[0105] In S402, an audio frame is buffered, short-time energy is calculated, and the comparison result of the short-time energy of each audio frame with the energy threshold value is accumulated.
[0106] During endpoint detection, the electronic device continuously acquires audio frames and then performs short-time energy detection on each audio frame. After detecting the short-time energy of the audio frame, the short-time energy of the audio frame is compared with the energy threshold value Eth, and the comparison result is recorded.
[0107] In S403, it is determined whether the conditions for the start of voice activity are met.
[0108] For each acquired audio frame, it can be checked whether a speech activity start point (VADStart) has been detected. If not detected, the process returns to S402 to continue acquiring the next audio frame. If a speech activity start point is detected, the process proceeds to S404 to detect the speech activity end point.
[0109] In S404, it is determined whether the conditions for the end of the voice activity are met.
[0110] After detecting the start point of the voice activity, the end point of the voice activity, i.e., VADEnd, can be further detected. If no VADEnd is detected, proceed to step 405, where each audio frame is sent to the wake-up detection model for wake-up detection. If a VADEnd is detected, proceed to step 406, where it is determined whether the current detection result meets the wake-up conditions.
[0111] In S405, audio frames are sent to the wake-up detection model for wake-up detection.
[0112] That is, after detecting the start point of the speech activity and before detecting the end point of the speech activity, every audio frame acquired during this period needs to be sent to the wake-up detection model for speech recognition. In other words, the audio segments between the start and end points of the identified speech activity need to be sent to the wake-up detection model for wake-up detection to obtain the detection results.
[0113] In S406, it is determined whether the detection result meets the wake-up condition.
[0114] After the audio frame is fed into the wake-up detection model for wake-up detection and the detection result is obtained, the result can be further judged to determine whether the wake-up condition is met. If the result meets the wake-up condition, the electronic device is directly woken up, and the process jumps to S412 to reset all system states, i.e., a flush operation is performed. Then, a new round of wake-up detection is performed. If the current detection result does not meet the wake-up condition, the audio frame is continued to be acquired and endpoint detection is performed.
[0115] In S407, an audio frame is buffered, short-time energy is calculated, and the comparison result of the short-time energy of each audio frame with the energy threshold value is accumulated.
[0116] This step serves the same purpose as S402, and will not be elaborated upon here. It is simply a further operation after a set of voice activity endpoints has been identified.
[0117] In S408, it is determined whether the conditions for the start of voice activity are met.
[0118] This step is the same as step S403, and will not be repeated here.
[0119] In S409, it is determined whether the number of frames between the start of the speech activity and the end of the previous speech activity is greater than N.
[0120] Here, N is a preset value, i.e., a preset number of audio frames. If the number of audio frames between the start of a new speech activity and the end of the previous speech activity is less than N frames, it indicates that the interval between the new speech segment and the previous speech segment is small, which is likely due to the slow speech rate causing them to be divided into different speech segments. In this case, the audio frames can continue to be sent to the wake-up detection model for joint recognition with the previous speech segment. If the interval of audio frames is greater than N frames, it means that a new speech task has started, and at this time, we can proceed to step 412 for flushing to start a new round of speech recognition.
[0121] In S410, audio frames are sent to the wake-up detection model for wake-up detection.
[0122] This step is the same as step S405, and will not be repeated here.
[0123] In S411, it is determined whether the conditions for the end of the voice activity are met.
[0124] This step is the same as for non-S404 cases, and will not be repeated here.
[0125] In S412, all states are reset.
[0126] Specifically, when the previous round of VAD task is determined to have ended—for example, when the wake-up detection result meets the wake-up conditions, or when the interval between the new VADStart and the previous VADEnd is greater than the preset N frames—the previous round of VAD task is considered to have ended. At this point, all states can be reset, and a new round of wake-up recognition task can begin again from S401.
[0127] For slow-speed wake-up word data, in the related art, VAD will cut off the pause in the middle of the wake-up word as a silent segment, resulting in a wake-up word being split into multiple segments and inaccurate VAD endpoints. For the slow-speed problem, "Hello" and "Xiaofei" are actually two parts of the same wake-up word. They should not be decoded independently. Instead, after decoding "Hello", the audio of "Xiaofei" should be sent continuously, and the two should be decoded together to produce a wake-up result. Between the two pieces of audio cut by VAD, this solution adjusts the flush position. Only when more than the specified number of frames N have passed after VadEnd, will a flush occur at the next VadStart, resetting all states. Otherwise, no flush will be performed, and the audio will continue to be decoded in the existing state, and the number of frames in the intermediate interval will be given in the final decoded string for subsequent correction of the end point. When the pause between "Hello" and "Xiaofei" does not exceed the specified number of frames N, no flush will occur, but the decoding of "Xiaofei" will continue. The wake-up word audio will not be decoded independently but decoded together, and finally a wake-up message will be produced. When the silence between two wake-up words exceeds the specified number of frames N, a flush will occur, resetting all states to prepare for the next wake-up. If the length of the silence between two wake-up words also does not exceed the specified number of frames N, but because the previous wake-up word has been recognized, the decoding module will be automatically reset after the wake-up word is produced to prepare for the next wake-up.
[0128] Among them, in actual use, the electronic device generally produces a wake-up result during the flush operation. In the embodiment of this application, when multiple speech segments are recognized in the speech information, the flush operation node will be at the position N frames after the last VadEnd. Then there will be a problem that the user still needs to wait for the time corresponding to N audio frames after speaking the complete wake-up speech to be woken up. That is, there will be a problem that a long time needs to be waited after speaking the wake-up word before the wake-up result is suddenly produced, affecting the user experience. In response to this, this application also provides a design, that is, appropriately increasing the length of the silent frames used to judge VADEnd. Then design to produce the wake-up result at VADEnd. In this way, the problem that the wake-up speech is split into multiple speech segments in the slow-speed scenario can be solved by increasing the silent frames, and the problem of producing the wake-up result after a long time difference due to setting an overly long time difference can also be avoided, that is, the problem of waiting for a long time after speaking the wake-up word before the wake-up result is suddenly produced as described above.
[0129] In addition, after increasing the length of the silent frames, although the wake-up result can be produced in time, the consequence is that the audio sent to the subsequent module has an increased length of silent frames at slow speed. Taking "Hello Xiaofei" as an example, the audio sent to the subsequent wake-up detection model may become "Hello + silent segment + Xiaofei". To improve the wake-up rate, an entry of "Hello + silent segment + Xiaofei" can be added to the wake-up word table.
[0130] As described above, the voice wake-up method provided in this application acquires voice information and identifies at least two consecutive combinations of voice activity endpoints within the voice information. Each combination of voice activity endpoints includes a voice activity start point and a voice activity end point. The method calculates the time difference between the voice activity end point of the preceding combination and the voice activity start point of the following combination. When the time difference is less than a preset value, a wake-up voice is determined based on the corresponding voice segment of the voice activity endpoint combination. The wake-up voice is then used to wake up the target object. Compared to related technologies, this application, by setting reasonable intervals between voice segments, enables accurate recognition even for intermittently spoken wake-up words. This method can significantly improve the wake-up success rate in slow-speech scenarios.
[0131] Please refer to Figure 6 To better implement the voice wake-up method provided in this application, this application further provides a voice wake-up device 500, such as... Figure 6 As shown, the voice wake-up device 500 includes:
[0132] The first acquisition module 510 is used to acquire voice information and perform voice activity endpoint detection on the voice information;
[0133] The first detection module 520 is used to perform wake-up detection on the first speech segment when the first speech segment is determined based on the speech activity endpoint detection result;
[0134] The calculation module 530 is used to calculate the time difference between the end point of the speech activity of the first speech segment and the start point of the next speech activity when the wake-up detection result is wake-up failure;
[0135] The second acquisition module 540 is used to acquire the second voice segment corresponding to the starting point of the next voice activity when the time difference is less than a preset value.
[0136] The second detection module 550 is used to fuse the first speech segment and the second speech segment to obtain a fused speech segment, and to perform wake-up detection based on the fused speech segment.
[0137] Optionally, in some embodiments, the first acquisition module includes:
[0138] The segmentation submodule is used to divide the voice information into multiple consecutive audio frames;
[0139] The first detection submodule is used to perform energy detection on each audio frame;
[0140] The second detection submodule is used to detect speech activity endpoints based on the energy of each audio frame.
[0141] Optionally, in some embodiments, the second detection submodule includes:
[0142] The comparison unit is used to compare the energy of each audio frame with the preset energy value one by one;
[0143] The first determining unit is used to determine the first target audio frame as the starting point of speech activity when, after detecting the first target audio frame, there is a second number of audio frames with energy greater than a preset energy value among a first number of audio frames, and the first number is greater than or equal to the second number.
[0144] The second determining unit is used to determine the second target audio frame as the end point of the speech activity when the energy of the third consecutive number of audio frames preceding the second target audio frame is less than or equal to a preset energy value.
[0145] Optionally, in some embodiments, the first detection submodule includes:
[0146] The acquisition unit is used to acquire the number of sampling points in each audio frame and the amplitude of each sampling point;
[0147] The calculation unit is used to calculate the energy of each audio frame based on the number of sampling points and the amplitude of each sampling point.
[0148] Optionally, in some embodiments, the first detection module includes:
[0149] The input submodule is used to input the first speech segment into the wake-up detection model for wake-up detection and obtain the wake-up detection result.
[0150] The second detection module is also used for:
[0151] The second speech segment is input into the wake-up detection model and used together with the first speech segment for wake-up detection.
[0152] Optionally, in some embodiments, the voice wake-up device provided in this application further includes:
[0153] The first reset module is used to reset the state of the wake-up detection model when the time difference is greater than or equal to a preset value.
[0154] Optionally, in some embodiments, the voice wake-up device provided in this application further includes:
[0155] The wake-up module is used to wake up the object to be woken up when the wake-up detection result is successful.
[0156] The second reset module is used to reset the state of the wake-up detection model.
[0157] It should be noted that the voice wake-up device 500 provided in this application embodiment belongs to the same concept as the voice wake-up method in the above embodiment. Its specific implementation process can be found in the above related embodiments, and will not be repeated here.
[0158] As described above, the voice wake-up device provided in this application acquires voice information through a first acquisition module 510 and performs voice activity endpoint detection on the voice information; when a first voice segment is determined based on the voice activity endpoint detection result, a first detection module 520 performs wake-up detection on the first voice segment; when the wake-up detection result is a wake-up failure, a calculation module 530 calculates the time difference between the voice activity endpoint of the first voice segment and the next voice activity start point; when the time difference is less than a preset value, a second acquisition module 540 acquires the second voice segment corresponding to the next voice activity start point; and a second detection module 550 fuses the first and second voice segments to obtain a fused voice segment, and performs wake-up detection based on the fused voice segment. Compared with related technologies, this application, by setting reasonable intervals between voice segments, enables the accurate recognition of intermittently spoken wake-up words, which can greatly improve the wake-up success rate in slow speech scenarios.
[0159] This application also provides an electronic device, including a memory and a processor, wherein the processor executes the steps in the voice wake-up method provided in this embodiment by calling a computer program stored in the memory.
[0160] Please refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of the electronic device 100 provided in the embodiments of this application.
[0161] The electronic device 100 may include components such as a network interface 110, a memory 120, a processor 130, and a screen assembly. Those skilled in the art will understand that... Figure 7 The structure of the electronic device 100 shown does not constitute a limitation on the electronic device 100, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0162] Network interface 110 can be used for network connections between devices.
[0163] The memory 120 can be used to store computer programs and data. The computer program stored in the memory 120 contains executable code. The computer program can be divided into various functional modules. The processor 130 executes various functional applications and data processing by running the computer program stored in the memory 120.
[0164] The processor 130 is the control center of the electronic device 100. It connects various parts of the electronic device 100 through various interfaces and lines. By running or executing computer programs stored in the memory 120 and calling data stored in the memory 120, it performs various functions of the electronic device 100 and processes data, thereby controlling the electronic device 100 as a whole.
[0165] In this embodiment, the processor 130 in the electronic device 100 loads executable code corresponding to one or more computer programs into the memory 120 according to the following instructions, and the processor 130 executes the steps in the voice wake-up method provided in this application, such as:
[0166] Acquire speech information and perform speech activity endpoint detection on the speech information; when a first speech segment is determined based on the speech activity endpoint detection result, wake-up detection is performed on the first speech segment; when the wake-up detection result is wake-up failure, the time difference between the speech activity endpoint of the first speech segment and the start point of the next speech activity is calculated; when the time difference is less than a preset value, the second speech segment corresponding to the start point of the next speech activity is acquired; the first speech segment and the second speech segment are fused to obtain a fused speech segment, and wake-up detection is performed based on the fused speech segment.
[0167] It should be noted that the electronic device 100 provided in this application embodiment belongs to the same concept as the voice wake-up method in the above embodiment. Its specific implementation process can be found in the above related embodiments, and will not be repeated here.
[0168] This application also provides a computer-readable storage medium storing a computer program thereon. When the computer program stored thereon is executed on the processor of the electronic device provided in the embodiments of this application, the processor of the electronic device performs any of the steps in the above-described voice wake-up method suitable for electronic devices. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0169] The above provides a detailed description of a voice wake-up method, apparatus, storage medium, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A voice wake-up method, characterized in that, The method includes: Acquire voice information and perform voice activity endpoint detection on the voice information; When the first speech segment is determined based on the speech activity endpoint detection result, wake-up detection is performed on the first speech segment; When the wake-up detection result is wake-up failure, calculate the time difference between the end point of the speech activity of the first speech segment and the start point of the next speech activity; When the time difference is less than a preset value, the second voice segment corresponding to the starting point of the next voice activity is obtained; The first speech segment and the second speech segment are fused to obtain a fused speech segment, and wake-up detection is performed based on the fused speech segment.
2. The method according to claim 1, characterized in that, The step of detecting voice activity endpoints in the voice information includes: The voice information is divided into multiple consecutive audio frames; Energy detection is performed on each audio frame; Voice activity endpoint detection is performed on the voice information based on the energy of each audio frame.
3. The method according to claim 2, characterized in that, The step of detecting speech activity endpoints based on the energy of each audio frame includes: The energy of each audio frame is compared with the preset energy value one by one; When a second number of audio frames with energy greater than the preset energy value are detected in a first number of audio frames after the first target audio frame, the first target audio frame is determined to be the starting point of the voice activity, and the first number is greater than or equal to the second number. When the energy of the third consecutive number of audio frames preceding the second target audio frame is less than or equal to the preset energy value, the second target audio frame is determined to be the end point of the voice activity.
4. The method according to claim 2, characterized in that, The energy detection for each audio frame includes: Obtain the number of sampling points and the amplitude of each sampling point for each audio frame; The energy of each audio frame is calculated based on the number of sampling points in each audio frame and the amplitude of each sampling point.
5. The method according to any one of claims 1 to 4, characterized in that, The wake-up detection of the first speech segment includes: The first speech segment is input into the wake-up detection model for wake-up detection, and the wake-up detection result is obtained. The step of fusing the first speech segment and the second speech segment to obtain a fused speech segment, and performing wake-up detection based on the fused speech segment, includes: The second speech segment is input into the wake-up detection model and used together with the first speech segment for wake-up detection.
6. The method according to claim 5, characterized in that, The method further includes: When the time difference is greater than or equal to the preset value, the wake-up detection model is reset.
7. The method according to claim 6, characterized in that, The method further includes: When the wake-up detection result is a successful wake-up, the object to be woken up is woken up. The wake-up detection model is then reset.
8. A voice wake-up device, characterized in that, The device includes: The first acquisition module is used to acquire voice information and perform voice activity endpoint detection on the voice information; The first detection module is used to perform wake-up detection on the first speech segment when the first speech segment is determined based on the speech activity endpoint detection result; The calculation module is used to calculate the time difference between the end point of the speech activity of the first speech segment and the start point of the next speech activity when the wake-up detection result is wake-up failure; The second acquisition module is used to acquire the second voice segment corresponding to the starting point of the next voice activity when the time difference is less than a preset value. The second detection module is used to fuse the first speech segment and the second speech segment to obtain a fused speech segment, and to perform wake-up detection based on the fused speech segment.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is loaded by the processor, it performs the steps of the voice wake-up method as described in any one of claims 1-7.
10. An electronic device comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor loads the computer program to perform the steps in the voice wake-up method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the voice wake-up method according to any one of claims 1 to 7.