Voice interaction method and device, model training method and device, equipment and product
By training a wake-up recognition model based on audio reception distance, the terminal can recognize the user's close-range voice interaction intention without a specific wake-up word, solving the problem of low wake-up efficiency in the prior art and achieving more efficient voice interaction.
Patent Information
- Application Number
- CN202510215930.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing voice interaction technology requires users to wake up voice assistants through specific wake-up words, which is inefficient and does not conform to the voice communication habits between natural persons.
By training the wake-up recognition model based on sample audio signals at different audio reception distances in advance, the terminal can wake-up recognition model based on the audio characteristics of the audio signal collected by the microphone, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up.
Users can trigger voice interaction through close speech without using specific wake-up words, simplifying the wake-up process and improving the efficiency of voice interaction.
Smart Images

Figure CN119993149A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of voice interaction technology, and in particular to a voice interaction, model training method, device, equipment and product. Background Art
[0002] Voice interaction is a technology in which a user sends commands to a terminal through natural language, and the terminal responds to the commands. In related technologies, a user can wake up a voice interaction assistant through a specific wake-up word to perform voice interaction functions. Summary of the invention
[0003] The present application provides a method, device, equipment and product for voice interaction and model training. The technical solution is as follows:
[0004] On the one hand, an embodiment of the present application provides a voice interaction method, the method comprising:
[0005] Extract features of the audio signal collected by the microphone to obtain audio features;
[0006] Inputting the audio feature into a wake-up recognition model to obtain a wake-up recognition result output by the wake-up recognition model, wherein the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, where the audio receiving distance is the distance from the sound source to the microphone;
[0007] When the wake-up recognition result indicates wake-up, the voice interaction function is woken up.
[0008] On the other hand, an embodiment of the present application provides a model training method, the method comprising:
[0009] Acquire a sample audio signal, wherein the sample audio signal corresponds to an audio receiving distance, and the audio receiving distance is the distance from the sound source to the microphone;
[0010] Extracting features from the sample audio signal to obtain sample audio features;
[0011] Inputting the sample audio feature into a wake-up recognition model to obtain a sample wake-up recognition result output by the wake-up recognition model, wherein the sample wake-up recognition result includes wake-up and non-wake-up;
[0012] The wake-up recognition model is trained based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal.
[0013] On the other hand, an embodiment of the present application provides a voice interaction device, the device comprising:
[0014] A feature extraction module is used to extract features from the audio signal collected by the microphone to obtain audio features;
[0015] an inference module, configured to input the audio feature into a wake-up recognition model to obtain a wake-up recognition result output by the wake-up recognition model, wherein the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, wherein the audio receiving distance is the distance from the sound source to the microphone;
[0016] The voice interaction module is used to wake up the voice interaction function when the wake-up recognition result indicates wake-up.
[0017] On the other hand, an embodiment of the present application provides a model training device, the device comprising:
[0018] A sample acquisition module is used to acquire a sample audio signal, wherein the sample audio signal corresponds to an audio receiving distance, and the audio receiving distance is the distance from the sound source to the microphone;
[0019] A feature extraction module, used to extract features from the sample audio signal to obtain sample audio features;
[0020] A training module is used to input the sample audio features into a wake-up recognition model to obtain a sample wake-up recognition result output by the wake-up recognition model, wherein the sample wake-up recognition result includes wake-up and non-wake-up; based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal, the wake-up recognition model is trained.
[0021] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to implement the voice interaction method as described in the above aspects, or the model training method.
[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer instruction, and the at least one computer instruction is used to be executed by a processor to implement the voice interaction method as described in the above aspects, or the model training method.
[0023] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the voice interaction method as described in the above aspects, or the model training method.
[0024] In the embodiment of the present application, the wake-up recognition model is pre-trained based on sample audio signals at different audio receiving distances. After the wake-up recognition model is subsequently deployed to the terminal, the terminal can perform wake-up recognition based on the audio features of the audio signal collected by the microphone through the wake-up recognition model, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up. With the close-range audio recognition capability of the wake-up recognition model, the user can trigger voice interaction with the terminal by speaking close to the terminal without using a specific wake-up word, which simplifies the wake-up process of the voice interaction function and improves the efficiency of voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flow chart of a voice interaction method provided by an exemplary embodiment of the present application is shown;
[0026] Figure 2 A schematic diagram of an implementation of a wake-up recognition model deployment method provided by an exemplary embodiment of the present application is shown;
[0027] Figure 3 is a schematic diagram of an implementation of an audio feature extraction process shown in an exemplary embodiment of the present application;
[0028] Figure 4 is a structural diagram of a wake-up recognition model provided by an exemplary embodiment of the present application;
[0029] Figure 5 A flow chart of a model training method provided by an exemplary embodiment of the present application is shown;
[0030] Figure 6 is a schematic diagram of a breath signal and a short-term speech signal shown in an exemplary embodiment of the present application;
[0031] Figure 7 is a flow chart of a sample audio signal acquisition process shown in an exemplary embodiment of the present application;
[0032] Figure 8 A structural block diagram of a voice interaction device provided by another exemplary embodiment of the present application is shown;
[0033] Fig. 9 A structural block diagram of a model training device provided by another exemplary embodiment of the present application is shown;
[0034] Fig.10 A schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0035] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0036] The term "multiple" as used herein refers to two or more than two. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.
[0037] In the related art, when a user needs to interact with the voice assistant of the terminal, he first needs to say a specific wake-up word (such as "Hi, Xiaobu") to wake up the voice interaction function. Among them, the specific wake-up word can be a system preset wake-up word, or a user-defined wake-up word. The low-power microphone of the terminal continuously collects audio signals and detects whether the collected audio signals contain specific wake-up words. If included, the voice interaction function is awakened, and voice interaction is performed based on the continuously collected audio signals.
[0038] However, before using the voice interaction function each time, you need to wake up with a specific wake-up word. On the one hand, the voice interaction efficiency is low and does not conform to the voice communication habits between natural persons; on the other hand, the degree of pronunciation standard will directly affect the wake-up success rate.
[0039] In an embodiment of the present application, a wake-up recognition model is pre-trained based on sample audio signals at different audio receiving distances, and the wake-up recognition model can identify the audio receiving distance of the audio signal, and then determine whether the audio signal is a close-range audio signal for waking up the voice interaction function. After the wake-up recognition model is deployed to the terminal, the terminal can perform wake-up recognition through the wake-up recognition model based on the audio features of the audio signal collected by the microphone, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up. After applying the wake-up recognition model, when the user needs to interact with the terminal by voice, he only needs to speak close to the terminal without using a specific wake-up word, which simplifies the wake-up process of the voice interaction function and improves the efficiency of voice interaction.
[0040] The voice interaction method provided in the embodiment of the present application can be applied to a terminal with a voice interaction function, and the terminal can be a smart phone, a tablet computer, a wearable device, a smart TV, a smart speaker, a smart home gateway, a personal computer, etc., which is not limited in the embodiment of the present application. In addition, the terminal is provided with a microphone for collecting external audio signals. Among them, the microphone can be a low-power microphone to achieve continuous audio collection at extremely low power consumption.
[0041] For the convenience of description, the following embodiments are described by taking the voice interaction method applied to a terminal as an example.
[0042] Please refer to Figure 1, which shows a flow chart of a voice interaction method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method used in a terminal as an example, and the method may include the following steps:
[0043] Step 101: extract features from the audio signal collected by the microphone to obtain audio features.
[0044] When the voice interaction function is turned on, the terminal continuously collects audio signals through the microphone and extracts features of the collected audio signals in real time to obtain audio features. The terminal extracts features according to the feature requirements of the wake-up recognition model so that the extracted audio features meet the model input requirements of the wake-up recognition model.
[0045] Step 102: input the audio feature into the wake-up recognition model to obtain the wake-up recognition result output by the wake-up recognition model, where the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, where the audio receiving distance is the distance from the sound source to the microphone.
[0046] In the embodiment of the present application, the wake-up recognition model is trained based on sample audio signals at different audio reception distances. The sample label of the sample audio signal is wake-up or not wake-up. Therefore, the wake-up recognition model can output a wake-up recognition result indicating wake-up or not wake-up based on the audio feature based on the audio reception distance.
[0047] Optionally, the output of the wake-up recognition model may be represented by 0 and 1, wherein 0 represents wake-up and 1 represents no wake-up.
[0048] Among them, when the wake-up recognition result output by the wake-up recognition model indicates wake-up, it indicates that the audio signal is recognized as a close-range audio signal, that is, the audio receiving distance of the audio signal is within the audio receiving distance range for triggering the voice interaction function (the distance range between the sound source and the microphone when waking up the voice interaction function). Accordingly, the audio signal is most likely collected when the user needs to interact with the terminal by voice and speaks close to the terminal microphone.
[0049] When the wake-up recognition result output by the wake-up recognition model indicates no wake-up, it indicates that the recognized audio signal is not a close-range audio signal, that is, the audio receiving distance of the audio signal is outside the audio receiving distance range that triggers the voice interaction function. Accordingly, the audio signal is most likely not collected when the user speaks close to the terminal microphone, but a collected far-end audio signal.
[0050] In a possible design, the wake-up recognition model is built on the sensorhub platform. The sensorhub platform uses AON (Always ON) low-power technology to enable the model to be in standby mode for a long time. The sensorhub platform can be set between the AP (Application Processor) and various MEMS (Micro-Electro-Mechanical System).
[0051] Indicatively, Figure 2 As shown, the sensorhub platform 22 is set between the microphone 21 and the AP 23. The microphone 21 receives the external audio signal at a low power consumption moment and writes it into the buffer. The sensorhub platform 22 obtains the buffered audio signal through the sensor driver 221 and performs feature extraction, thereby performing wake-up recognition on the extracted features through the wake-up recognition model 222. Among them, the wake-up recognition model 222 runs independently on the sensorhub platform 22, using the computing power of the NPU 24 in the inference framework on the sensorhub side.
[0052] Step 103: When the wake-up recognition result indicates wake-up, wake up the voice interaction function.
[0053] When the wake-up recognition result indicates wake-up, the terminal wakes up the voice interaction function. After the voice interaction function is awakened, the terminal continuously collects audio through a microphone and performs voice recognition on the collected audio (for example, through ASR technology), thereby performing voice interaction with the user based on the voice recognition result.
[0054] In some embodiments, when the wake-up function is implemented through the sensorhub platform, the voice interaction after the voice interaction function is awakened can be implemented by the AP. Figure 2 As shown, when it is recognized that wake-up is required, the sensorhub platform 22 reports the wake-up instruction to the AP 23, and the AP 23 wakes up the voice interaction function and performs subsequent voice interaction. Compared with the AP side call, deploying the model on the sensorhub platform can achieve ultra-low power consumption and always-on effect.
[0055] Since the solution provided in the embodiment of the present application is adopted, the user can directly issue voice commands without having to say a specific wake-up word first. Therefore, in order to ensure the integrity of voice command recognition, the audio signal used for wake-up recognition will also be used as input for voice recognition.
[0056] Optionally, when the wake-up recognition result of the audio signal at time T is wake-up, the terminal takes time Tt as the speech recognition starting point, that is, performs speech recognition on the audio collected after time Tt. Wherein, t is usually set to a shorter time length, such as 0.5 seconds, 1 second, or 1.5 seconds.
[0057] It should be noted that when the solution provided in the embodiment of the present application is adopted, the terminal can still wake up the voice interaction function when a specific wake-up word is recognized. Among them, recognizing the specific wake-up word and model recognition can be implemented in two different processes.
[0058] In summary, in the embodiment of the present application, the wake-up recognition model is pre-trained based on sample audio signals at different audio receiving distances. After the wake-up recognition model is subsequently deployed to the terminal, the terminal can perform wake-up recognition based on the audio features of the audio signal collected by the microphone through the wake-up recognition model, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up. With the help of the close-range audio recognition capability of the wake-up recognition model, the user can trigger voice interaction with the terminal by speaking close to the terminal without using a specific wake-up word, which simplifies the wake-up process of the voice interaction function and improves the efficiency of voice interaction.
[0059] Audio feature extraction process
[0060] Since the audio is continuous, the terminal can extract the collected audio signal through a sliding window and perform feature extraction on the extracted audio signal. In a possible implementation, the above step 101 may include the following sub-steps:
[0061] Step 101A: extract the audio signal collected by the microphone through a sliding window to obtain an audio signal segment.
[0062] In a possible implementation, the terminal slides on the collected audio signal using a sliding window of fixed length according to a fixed sliding step size to obtain audio signal segments within the sliding window at different sliding positions, where the length of the audio signal segment is less than or equal to the length of the sliding window.
[0063] In an illustrative example, the length of the sliding window is 1 second, that is, the sliding window is used to extract a 1-second audio signal segment from a continuous audio signal. Furthermore, when the sliding step length of the sliding window is 0.2 seconds (for illustrative purposes only, the sliding can be performed in frames), there is an intersection of 0.8 seconds of audio signals between sliding windows at adjacent sliding positions.
[0064] In a possible scenario, in the initial audio signal acquisition stage, since the length of the acquired audio signal may be less than the length of the sliding window, in order to ensure the normal extraction of subsequent audio features (extracting audio features of sufficient length) and improve recognition efficiency (it is possible to acquire an audio signal for wake-up in the initial acquisition stage), if the length of the audio signal does not reach the length of the sliding window, the terminal amplifies the signal within the sliding window to obtain an amplified signal, and the length of the amplified signal is the length of the sliding window. The terminal subsequently determines the amplified signal as an audio signal segment.
[0065] When the length of the audio signal reaches the length of the sliding window, the signal within the sliding window is extracted as the audio signal segment.
[0066] Regarding the signal amplification method, in a possible implementation manner, the terminal performs signal interpolation processing on the audio signal based on the length of the audio signal and the length of the sliding window to obtain an amplified signal.
[0067] When performing signal interpolation processing, the terminal selects a specified audio signal frame in the audio signal, and determines an interpolation audio signal frame through an interpolation algorithm, thereby inserting the interpolation audio signal frame between the specified frame audio signals. The interpolation algorithm can be a linear interpolation algorithm or a nonlinear interpolation algorithm.
[0068] Of course, in other possible implementations, the terminal may use other time domain signal stretching technologies to amplify the audio signal, which is not limited in this embodiment of the present application.
[0069] Illustratively, when the duration of the signal in the sliding window is 0.5 seconds, the terminal may amplify the signal in the sliding window to 1 second by signal amplification.
[0070] In a possible implementation, the terminal also needs to pre-process the collected audio signal to reduce the impact of the audio loudness on subsequent recognition. For different microphone settings, the terminal pre-processes the audio signal in different ways.
[0071] In some embodiments, when the microphone is a dual microphone (for example, two microphones respectively arranged at the top and bottom of the terminal), the terminal performs differential processing on the dual-channel audio signals collected by the dual microphones to obtain an audio signal. Subsequently, the audio signal obtained by the differential processing is subjected to segment extraction.
[0072] In other embodiments, when the microphone is a single-channel microphone, the terminal performs normalization processing on the single-channel audio signal collected by the single microphone to obtain an audio signal. Subsequently, a segment is extracted from the audio signal obtained by the normalization processing. Among them, normalization processing on the single-channel audio signal refers to normalization of the loudness, that is, the loudness of the audio signal obtained after the normalization processing is in the loudness range of 0-1.
[0073] Step 101B: extract features from the audio signal segment to obtain audio features of the audio signal segment.
[0074] For each extracted audio signal segment, the terminal performs feature extraction on the audio signal segment to obtain audio features of the audio signal segment. In a possible implementation, the audio feature extraction process may include the following sub-steps:
[0075] Sub-step 1: Perform short-time Fourier transform on the audio signal segment to obtain the signal time-frequency spectrum.
[0076] For an audio signal segment containing signal time domain amplitude information, the terminal first performs a short-time Fourier transform (STFT) on the audio signal segment to decompose the signal into time and frequency to obtain a signal time-frequency spectrum. The audio signal segment before the STFT transform contains the relationship between time and the time domain signal amplitude, and the signal time-frequency spectrum obtained by the STFT transform contains the relationship between time and frequency, as well as the amplitude of the signal at a specific time and frequency.
[0077] Indicatively, Figure 3 As shown, the terminal performs STFT transformation on the audio signal segment 31 to obtain a signal time-frequency spectrum 32.
[0078] Sub-step 2: filtering the signal time-frequency spectrum through a triangular filter to obtain the audio features of the audio signal segment.
[0079] After completing the STFT transformation, the terminal further performs triangular filtering on the signal time-frequency spectrum through a triangular filter, and determines the triangular filtering result as the audio feature of the audio signal segment.
[0080] Optionally, the triangular filter used for triangular filtering is a Mel filter bank, which is usually composed of multiple triangular filters and is used to convert the frequency axis from a linear scale to a Mel scale. This filter bank can better simulate the nonlinear perception of the human ear to frequency. Accordingly, the triangular filtering result obtained by performing triangular filtering on the signal time-frequency spectrum is a Mel spectrum.
[0081] The process of filtering the signal time-frequency spectrum is to perform a dot multiplication operation between the signal time-frequency spectrum and the triangular filter.
[0082] Indicatively, Figure 3 As shown, the terminal further performs triangular filtering on the signal time-frequency spectrum 32 through a triangular filter 33 to obtain an audio feature 34 in the form of a Mel spectrum.
[0083] Of course, in addition to extracting audio features in the above manner, the terminal may also extract features of other dimensions of the audio signal segment, which is not limited in this embodiment of the present application.
[0084] In some embodiments, after the audio features in the form of time-frequency spectrum are obtained in the above manner, a wake-up recognition model based on a convolutional neural network can be used for wake-up recognition. The wake-up recognition model based on a convolutional neural network can be composed of a convolutional layer, a pooling layer, and a fully connected layer.
[0085] Indicatively, Figure 4 As shown, after the audio feature 41 in the form of a signal time-frequency spectrum is input into the wake-up recognition model 42, it undergoes convolution, pooling, convolution and full connection processing in sequence, and finally outputs the wake-up recognition result 43 through the output layer.
[0086] Of course, the wake-up recognition model can also adopt other network structures, such as recurrent neural network (RNN), long short-term memory network (LSTM), etc., and the embodiments of the present application are not limited to this.
[0087] Training process of wake-up recognition model
[0088] The wake-up recognition model deployed on the terminal side can be pre-trained by a computer device based on a sample audio signal. The sample audio signal corresponds to the respective audio receiving distance and the respective wake-up tag, which is used to indicate whether the voice interaction function is awakened based on the sample audio signal. The computer device can be a server, workstation, personal computer or other device with a model training function.
[0089] Please refer to Figure 5 , which shows a flow chart of a model training method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method used in a computer device as an example, and the method may include the following steps:
[0090] Step 501: Acquire a sample audio signal. The sample audio signal corresponds to an audio receiving distance, which is the distance from the sound source to the microphone.
[0091] In a possible implementation, the computer device first obtains sample audio signals at multiple audio receiving distances, including sample audio signals within the audio receiving distance range for triggering the voice interaction function and sample audio signals outside the audio receiving distance range for triggering the voice interaction function.
[0092] For example, when the audio receiving distance range for triggering the voice interaction function is 0-10cm, the samples used to train the wake-up recognition model include sample audio signals with an audio receiving distance less than or equal to 10cm, and sample audio signals with an audio receiving distance greater than 10cm.
[0093] In some embodiments, the sample audio signal is obtained by artificial recording (such as by a real person reading any text), and / or is generated by performing data enhancement on the artificially recorded audio.
[0094] Furthermore, in order to improve the recognition robustness of the wake-up recognition model, the sample audio signals cover natural persons of different ages, genders, speaking habits, timbres, and languages.
[0095] Step 502: extract features from the sample audio signal to obtain sample audio features.
[0096] The computer device extracts features according to the feature requirements of the wake-up recognition model, so that the extracted sample audio features meet the model input requirements of the wake-up recognition model.
[0097] Step 503: Input the sample audio feature into the wake-up recognition model to obtain the sample wake-up recognition result output by the wake-up recognition model, where the sample wake-up recognition result includes wake-up and non-wake-up.
[0098] During the training process, the computer device inputs the sample audio features into the wake-up recognition model to be trained, and obtains the sample wake-up recognition results corresponding to the sample audio features. Wherein, when the sample wake-up recognition results represent wake-up, it indicates that the model predicts that the voice interaction function needs to be awakened when the sample audio signal corresponding to the sample audio feature is received; when the sample wake-up recognition results represent no wake-up, it indicates that the model predicts that the voice interaction function does not need to be awakened when the sample audio signal corresponding to the sample audio feature is received.
[0099] Optionally, the output of the wake-up recognition model may be represented by 0 and 1, wherein 0 represents wake-up and 1 represents no wake-up.
[0100] Step 504: training a wake-up recognition model based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal.
[0101] In the embodiment of the present application, each sample audio signal is provided with a wake-up tag, which is a truth value tag indicating whether the sample audio signal needs to be woken up, that is, the wake-up tag of the sample audio signal is wake-up or not wake-up. Optionally, the wake-up tag indicating wake-up can be represented by 1, and the wake-up tag indicating not wake-up can be represented by 0.
[0102] In some embodiments, the wake-up tag corresponding to the sample audio signal is related to the audio collection distance corresponding to the sample audio signal and the audio collection distance range that triggers the wake-up voice interaction function.
[0103] In one possible implementation, for the same sample audio signal, the computer device determines the prediction loss of the model for the sample audio signal based on the difference between the sample wake-up recognition result corresponding to the sample audio signal and the wake-up label, thereby training the wake-up recognition model based on the prediction losses of different sample audio signals until the training end conditions are met.
[0104] Among them, the prediction loss can be a cross entropy loss; a gradient descent back propagation algorithm can be used when training the wake-up recognition model; the training end condition can include an upper limit on the number of iterations, a loss convergence condition, etc., which is not limited in the embodiments of the present application.
[0105] In summary, in the embodiment of the present application, the wake-up recognition model is pre-trained based on sample audio signals at different audio receiving distances. After the wake-up recognition model is subsequently deployed to the terminal, the terminal can perform wake-up recognition based on the audio features of the audio signal collected by the microphone through the wake-up recognition model, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up. With the help of the close-range audio recognition capability of the wake-up recognition model, the user can trigger voice interaction with the terminal by speaking close to the terminal without using a specific wake-up word, which simplifies the wake-up process of the voice interaction function and improves the efficiency of voice interaction.
[0106] Wake-up label settings
[0107] Regarding the method of setting the wake-up tag of the sample audio signal, in a possible implementation, the computer device sets the wake-up tag for the sample audio signal based on the relationship between the audio collection distance corresponding to the sample audio signal and the distance threshold.
[0108] The distance threshold is the maximum audio collection distance of the audio signal for waking up the voice interaction function.
[0109] When the audio reception distance corresponding to the sample audio signal is less than or equal to the distance threshold, the computer device determines the wake-up tag corresponding to the sample audio signal as wake-up; when the audio reception distance corresponding to the sample audio signal is greater than the distance threshold, the computer device determines the wake-up tag corresponding to the sample audio signal as not wake-up.
[0110] In some embodiments, the terminal may provide multiple distance thresholds for the user to choose from. For example, the user may set the distance threshold for waking up the voice interaction function to 5 cm, or 10 cm.
[0111] Accordingly, in order to adapt to different user needs, the computer device can pre-set different wake-up tags for the same sample audio signal based on different distance thresholds (for example, for a sample audio signal with an audio collection distance of 7 cm, when the distance threshold is 10 cm, the wake-up tag of the sample audio signal is wake-up, and when the distance threshold is 5 cm, the wake-up tag of the sample audio signal is not wake-up), and train to obtain a wake-up recognition model suitable for different distance thresholds. The subsequent terminal can deploy a wake-up recognition model suitable for the distance threshold set by the user.
[0112] The sample audio signal is in the form of
[0113] In order to improve the accuracy of the timing of waking up the voice interaction function, in a possible implementation, the sample audio signal includes a breath signal and a short-term voice signal. The breath signal is the breath sound first collected when approaching the microphone of the terminal, and the short-term voice signal is a short-term voice (usually the beginning of the voice command) emitted after approaching the microphone.
[0114] Using the sample setting method of "breath signal + short-time voice signal", the trained wake-up recognition model can recognize audio signals that are close to the terminal microphone (breath signal) and truly have voice interaction intentions (short-time voice signals).
[0115] Schematically, when you put your mouth close to the microphone and issue a voice command, the dual microphones collect the following Figure 6 The upper and lower audio signals are shown. Among them, the audio signal located in the first area 61 is a breath signal, and the audio signal located in the second area 62 is a short-term speech signal.
[0116] In one possible implementation, for an artificially recorded audio signal, the breath signal position and the short-time speech signal position in the artificially recorded audio signal are marked by manual labeling and / or feature recognition, so that a sample audio signal serving as a training sample is extracted from the artificially recorded audio signal based on the marked breath signal position and the short-time speech signal position.
[0117] The process of extracting sample audio features
[0118] Similar to the audio feature extraction process in the application process, in a possible implementation, the above step 502 may include the following sub-steps:
[0119] Step 502A: extract the sample audio signal through a sliding window to obtain a sample audio signal segment.
[0120] In a possible implementation, a computer device slides a sample audio signal using a sliding window of fixed length at a fixed sliding step size to obtain sample audio signal segments within the sliding window at different sliding positions, wherein the length of the sample audio signal segment is less than or equal to the length of the sliding window.
[0121] In an illustrative example, the length of the sliding window is 1 second, that is, the sliding window is used to extract a 1-second sample audio signal segment from a continuous sample audio signal. Furthermore, when the sliding step length of the sliding window is 0.2 seconds (for illustrative purposes only, the sliding can be performed in frames), there is an intersection of 0.8 seconds of audio signals between sliding windows at adjacent sliding positions.
[0122] Since the duration of the sample audio signal may be shorter than the duration of the sliding window (related to factors such as the speaker's speaking speed, speaking habits, and spoken content), in order to ensure the normal extraction of subsequent audio features (extracting audio features of sufficient length) and improve recognition efficiency, when the length of the sample audio signal does not reach the length of the sliding window, the computer device amplifies the signal within the sliding window to obtain a sample amplified signal, and the length of the sample amplified signal is the length of the sliding window. The subsequent terminal determines the sample amplified signal as an audio signal segment.
[0123] When the length of the sample audio signal reaches the length of the sliding window, the signal within the sliding window is extracted as the sample audio signal segment.
[0124] Regarding the signal amplification method, in a possible implementation, the computer device performs signal interpolation processing on the sample audio signal based on the length of the sample audio signal and the length of the sliding window to obtain a sample amplified signal.
[0125] When performing signal interpolation processing, the computer device selects a specified audio signal frame from the sample audio signal, and determines an interpolation audio signal frame through an interpolation algorithm, thereby inserting the interpolation audio signal frame between the specified frame audio signals. The interpolation algorithm can be a linear interpolation algorithm or a nonlinear interpolation algorithm.
[0126] Of course, in other possible implementations, the computer device may use other time domain signal stretching technologies to amplify the audio signal, which is not limited in this embodiment of the present application.
[0127] Illustratively, when the duration of the sample audio signal is 0.5 seconds and the duration of the sliding window is 1 second, the computer device may amplify the sample audio signal to 1 second by signal amplification.
[0128] In a possible implementation, the computer device also needs to pre-process the sample audio signal to reduce the impact of audio loudness on subsequent recognition. In some embodiments, when the sample audio signal is a two-way audio signal collected by a two-way microphone, the computer device performs differential processing on the two-way sample audio signal. Subsequently, the sample audio signal obtained by the differential processing is subjected to segment extraction.
[0129] In other embodiments, when the sample audio signal is a single-channel audio signal collected by a single-channel microphone, the computer device performs normalization processing on the single-channel audio signal. Subsequently, the sample audio signal obtained by the normalization processing is subjected to segment extraction. The normalization processing on the single-channel audio signal refers to normalizing the loudness, that is, the loudness of the sample audio signal obtained after the normalization processing is within the loudness range of 0-1.
[0130] Step 502B: extract features from the sample audio signal segment to obtain sample audio features of the sample audio signal segment.
[0131] For each sample audio signal segment extracted, the terminal performs feature extraction on the sample audio signal segment to obtain the sample audio feature of the sample audio signal segment. It should be noted that when the duration of the sample audio signal is greater than the duration of the sliding window, at least two sample audio signal segments can be extracted from the sample audio signal. The sample audio signal segments extracted from the same sample audio signal correspond to the same audio acquisition distance.
[0132] In a possible implementation, during the sample audio feature extraction process, the computer device performs a short-time Fourier transform on the sample audio signal segment to obtain a sample signal time spectrum, and performs filtering processing on the signal sample time spectrum through a triangular filter to obtain a sample audio feature of the sample audio signal segment. The specific extraction process of the sample audio feature can refer to the extraction process of the audio feature in the application stage, and this embodiment will not be described in detail here.
[0133] The process of obtaining sample audio signals
[0134] Since the quality of model training is closely related to the quality and quantity of training samples, in order to improve the quality of model training, the computer device can use data enhancement technology to obtain sample audio signals. The process of obtaining sample audio signals is as follows: Figure 7 shown.
[0135] Step 701: Acquire a first sample audio signal, where the first sample audio signal is a user audio signal collected at a preset audio receiving distance.
[0136] In one possible implementation, a computer device obtains a user audio signal (manually reading any text) collected at a preset audio receiving distance, and marks the breath signal position and the short-time speech signal position in the user audio signal, thereby extracting a first sample audio signal from the user audio signal based on the marked position.
[0137] Step 702: Perform data enhancement processing on the first sample audio signal to obtain a second sample audio signal.
[0138] Since it is impossible to collect audio of all people manually, in order to increase the number of training samples, the computer device performs data enhancement on the basis of the first sample audio signal to obtain the second sample audio signal.
[0139] The data enhancement method may include at least one of the following:
[0140] Data enhancement method 1: performing frequency band adjustment processing on the first sample audio signal to obtain a second sample audio signal, wherein the frequency band adjustment processing includes at least one of increasing the signal frequency band and reducing the signal frequency band.
[0141] In order to simulate the intonations of different groups of people, the computer device obtains a second sample audio signal having a different intonation from the first sample audio signal by changing the frequency band of the first sample audio signal, thereby enriching the intonation covered by the training sample.
[0142] In a possible implementation, the computer device may increase and / or decrease the signal frequency band of the first sample audio signal (increase and / or decrease as a whole) according to the unit frequency band adjustment amount (e.g., 10 Hz). The increase amount and decrease amount of the signal frequency band are integer multiples of the unit frequency band adjustment amount. Accordingly, the computer device may obtain at least one second sample audio signal having a higher pitch than the first sample audio signal, and / or at least one second sample audio signal having a lower pitch than the first sample audio signal.
[0143] It should be noted that, when performing the frequency band adjustment process, it is necessary to ensure that the signal frequency band of the second sample audio signal is within the human vocalization frequency band (eg, 85 Hz-1100 Hz).
[0144] Data enhancement method 2: performing speech inversion processing on the first sample audio signal to obtain a second sample audio signal, wherein the speech inversion processing is used to invert the front and back relationship of the speech signal.
[0145] Typically, the artificially recorded first sample audio signal contains semantics, and the computer device can perform speech inversion on the first sample audio signal, that is, invert the front and back signals of the first sample audio signal, to construct a training sample (that is, the second sample audio signal) that does not contain semantics but contains speech features.
[0146] In one possible implementation, since the normally collected audio signals all have the breath signal appearing first and then the speech signal, in order to ensure the rationality of the generated second sample audio signal, the computer device maintains the order of the breath signal in the first sample audio signal and only performs speech inversion processing on the short-term speech signal in the first sample audio signal.
[0147] It should be noted that when the above two data enhancement methods are used at the same time, the computer device can perform speech inversion processing on the audio signal obtained after the frequency band adjustment processing to obtain a second sample audio signal.
[0148] Of course, in addition to the above data enhancement methods, the computer device may also use other methods to augment the training sample data, which will not be described in detail in this embodiment.
[0149] Step 703: determine the first sample audio signal and the second sample audio signal as sample audio signals.
[0150] Furthermore, the computer device determines the first sample audio signal obtained by artificial recording and the generated second sample audio signal together as the sample audio signal, wherein the second sample audio signal generated based on the first sample audio signal corresponds to the same audio collection distance.
[0151] In this embodiment, in addition to obtaining the first sample audio signal through manual recording, the computer device generates a second sample audio signal based on the first sample audio signal through frequency band adjustment and / or voice inversion, thereby increasing the number of training samples, helping to reduce the cost of collecting training samples and improving the training quality of the model.
[0152] Noise processing of sample audio signal
[0153] In real voice wake-up scenarios, there is usually environmental noise, which may affect recognition robustness. Therefore, in order to improve the robustness of the wake-up recognition model, the computer device can perform noise processing on the sample audio signal (the first sample audio signal and / or the second sample audio signal), so as to use the sample audio signal containing noise for model training.
[0154] In a possible implementation, after generating the second sample audio signal through data enhancement, the computer device performs noise processing on the first sample audio signal and the second sample audio signal to obtain the noisy first sample audio signal and the second sample audio signal, thereby determining the noisy first sample audio signal and the second sample audio signal as sample audio signals.
[0155] Regarding the specific manner of noise addition processing, the computer device adopts noise superposition technology to add preset noise to the first and second sample audio signals (energy superposition).
[0156] In order to improve the diversity of the superimposed noise, the computer device may perform noise scaling on the preset noise and then add the scaled noise to the first and second sample audio signals.
[0157] In a possible implementation, the computer device determines a noise power spectrum of a preset noise and a signal power spectrum of the first sample audio signal and the second sample audio signal. The computer device performs scaling processing on the noise power spectrum and superimposes the scaled noise power spectrum on the signal power spectrum to obtain the first sample audio signal and the second sample audio signal after noise addition.
[0158] Optionally, the computer device may be set with a variety of preset noises to simulate different noise environments, such as white noise, traffic noise, noise in a building, etc., which is not limited in the present embodiment of the application.
[0159] Furthermore, in order to simulate more noise levels, the computer device may scale the noise power spectrum according to a plurality of scaling ratios to obtain scaled noise of different powers. For example, the scaling ratio may include 0.7, 0.8, 0.9, 1.1, 1.2, 1.3, etc., which is not limited in this embodiment.
[0160] It should be noted that the first sample audio signal can be artificially recorded in different noise environments to further improve the model's ability to resist noise.
[0161] In this embodiment, the computer device adds noise to the sample audio signal through the noise superposition technology, so as to train the wake-up recognition model with the sample audio signal containing noise, improve the model's ability to resist noise, and improve the robustness of the model.
[0162] Adaptive fine-tuning of the model
[0163] In the above embodiments, the wake-up recognition models are all trained based on big data and have good universality. However, since some users may have special pronunciation or voice interaction habits, the models trained based on big data may not perform well on such users.
[0164] In order to further improve the quality of the wake-up recognition model, the terminal or computer device can fine-tune the wake-up recognition model based on whether the voice interaction function on the terminal side is correctly awakened.
[0165] In a possible implementation, during the application process, when the voice interaction is successful, the terminal generates a positive sample based on the audio signal and the wake-up recognition result. When a voice interaction cancellation operation is received, the terminal generates a negative sample based on the audio signal and the wake-up recognition result. The terminal then fine-tunes the wake-up recognition model based on the positive and negative samples.
[0166] Among them, successful voice interaction means that after the voice interaction function is awakened based on the wake-up recognition result output by the model, a voice interaction instruction is further received, which indicates that the wake-up recognition result output by the model indicating awakening is correct; the voice interaction cancellation operation means that after the voice interaction function is awakened based on the wake-up recognition result output by the model, a voice interaction cancellation instruction is further received, and the voice interaction cancellation instruction can be a voice instruction (for example, the user immediately expresses that no interaction is needed after the voice interaction function is awakened) or a touch instruction.
[0167] In addition, the generated positive samples include the audio signal that correctly triggers the awakening voice interaction function and the awakening label "awakening", and the generated negative samples include the audio signal that incorrectly triggers the awakening voice interaction function and the awakening label "not awakening".
[0168] Optionally, fine-tuning the wake-up recognition model based on positive and negative samples can be executed locally by the terminal (the terminal has strong computing power), or reported by the terminal to the server, executed by the server, and the fine-tuned model parameters are sent to the terminal.
[0169] In a possible design, the computer device receives positive samples and negative samples reported by the terminal, and fine-tunes the wake-up recognition model based on the positive samples and the negative samples.
[0170] In this embodiment, the terminal generates corresponding positive and negative samples based on the success of voice interaction after awakening the voice interaction function and the voice interaction cancellation operation, and uses the positive and negative samples to fine-tune the model parameters of the wake-up recognition model, so that the fine-tuned model is more adapted to the current user's voice wake-up habits, which helps to improve the accuracy of voice wake-up.
[0171] Please refer to Figure 8 , which shows a structural block diagram of a voice interaction device provided by an exemplary embodiment of the present application. The device includes:
[0172] The feature extraction module 801 is used to extract features from the audio signal collected by the microphone to obtain audio features;
[0173] An inference module 802 is used to input the audio feature into a wake-up recognition model to obtain a wake-up recognition result output by the wake-up recognition model, wherein the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, wherein the audio receiving distance is the distance from the sound source to the microphone;
[0174] The voice interaction module 803 is used to wake up the voice interaction function when the wake-up recognition result indicates wake-up.
[0175] Optionally, the feature extraction module 801 is used to:
[0176] Extracting the audio signal collected by the microphone through a sliding window to obtain an audio signal segment;
[0177] Perform feature extraction on the audio signal segment to obtain the audio feature of the audio signal segment.
[0178] Optionally, the feature extraction module 801 is used to:
[0179] When the length of the audio signal does not reach the length of the sliding window, amplify the signal within the sliding window to obtain an amplified signal, wherein the length of the amplified signal is the length of the sliding window; and determine the amplified signal as the audio signal segment;
[0180] When the length of the audio signal reaches the length of the sliding window, the signal within the sliding window is extracted as the audio signal segment.
[0181] Optionally, the feature extraction module 801 is used to:
[0182] Based on the length of the audio signal and the length of the sliding window, signal interpolation processing is performed on the audio signal to obtain the amplified signal.
[0183] Optionally, the feature extraction module 801 is used to:
[0184] Performing short-time Fourier transform on the audio signal segment to obtain a signal time-frequency spectrum;
[0185] The signal time-frequency spectrum is filtered by a triangular filter to obtain the audio feature of the audio signal segment.
[0186] Optionally, the device further comprises a preprocessing module, which is used to:
[0187] In the case where the microphone is a dual microphone, performing differential processing on the dual-channel audio signals collected by the dual microphones to obtain the audio signal;
[0188] In the case that the microphone is a single microphone, a single-channel audio signal collected by the single microphone is normalized to obtain the audio signal.
[0189] Optionally, the device further comprises:
[0190] A local sample generation module, configured to generate a positive sample based on the audio signal and the wake-up recognition result when the voice interaction is successful; and to generate a negative sample based on the audio signal and the wake-up recognition result when a voice interaction cancellation operation is received;
[0191] A fine-tuning module is used to fine-tune the wake-up recognition model based on the positive sample and the negative sample.
[0192] Please refer to Fig. 9 , which shows a structural block diagram of a model training device provided by an exemplary embodiment of the present application. The device includes:
[0193] The sample acquisition module 901 is used to acquire a sample audio signal, wherein the sample audio signal corresponds to an audio receiving distance, and the audio receiving distance is the distance from the sound source to the microphone;
[0194] A feature extraction module 902 is used to extract features from the sample audio signal to obtain sample audio features;
[0195] The training module 903 is used to input the sample audio feature into the wake-up recognition model to obtain the sample wake-up recognition result output by the wake-up recognition model, and the sample wake-up recognition result includes wake-up and non-awakening; based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal, the wake-up recognition model is trained.
[0196] Optionally, the feature extraction module 902 is used to:
[0197] Extracting the sample audio signal through a sliding window to obtain a sample audio signal segment;
[0198] Feature extraction is performed on the sample audio signal segment to obtain the sample audio feature of the sample audio signal segment.
[0199] Optionally, the feature extraction module 902 is used to:
[0200] In the case that the length of the sample audio signal is greater than or equal to the length of the sliding window, a signal within the sliding window is extracted as the sample audio signal segment.
[0201] When the length of the sample audio signal is less than the length of the sliding window, amplify the signal within the sliding window to obtain a sample amplified signal, the length of the sample amplified signal is the length of the sliding window; and determine the sample amplified signal as the sample audio signal segment.
[0202] Optionally, the feature extraction module 902 is used to:
[0203] Based on the length of the sample audio signal and the length of the sliding window, signal interpolation processing is performed on the sample audio signal to obtain the sample amplified signal.
[0204] Optionally, the feature extraction module 902 is used to:
[0205] Performing short-time Fourier transform on the sample audio signal segment to obtain a sample signal time-frequency spectrum;
[0206] The sample signal time-frequency spectrum is filtered by a triangular filter to obtain the sample audio feature of the sample audio signal segment.
[0207] Optionally, the sample audio signal includes a breath signal and a short-term speech signal.
[0208] Optionally, the sample acquisition module 901 is used to:
[0209] Acquire a first sample audio signal, where the first sample audio signal is a user audio signal collected at a preset audio receiving distance;
[0210] Performing data enhancement processing on the first sample audio signal to obtain a second sample audio signal;
[0211] The first sample audio signal and the second sample audio signal are determined as the sample audio signals.
[0212] Optionally, the sample acquisition module 901 is used to:
[0213] Performing frequency band adjustment processing on the first sample audio signal to obtain the second sample audio signal, wherein the frequency band adjustment processing includes at least one of increasing the signal frequency band and decreasing the signal frequency band;
[0214] Performing speech inversion processing on the first sample audio signal to obtain the second sample audio signal, wherein the speech inversion processing is used to invert the front-to-back relationship of the speech signal.
[0215] Optionally, the device further includes a noise adding module, configured to:
[0216] Performing noise processing on the first sample audio signal and the second sample audio signal to obtain the first sample audio signal and the second sample audio signal after noise is added;
[0217] The sample acquisition module 901 is used to:
[0218] The first sample audio signal and the second sample audio signal after the noise is added are determined as the sample audio signals.
[0219] Optionally, the noise adding module is used to:
[0220] Determine a noise power spectrum of a preset noise, and a signal power spectrum of the first sample audio signal and the second sample audio signal;
[0221] The noise power spectrum is scaled, and the scaled noise power spectrum is superimposed on the signal power spectrum to obtain the first sample audio signal and the second sample audio signal after noise is added.
[0222] Optionally, the device further comprises:
[0223] a tag determination module, configured to determine the wake-up tag corresponding to the sample audio signal as wake-up when the audio receiving distance corresponding to the sample audio signal is less than or equal to a distance threshold;
[0224] When the audio receiving distance corresponding to the sample audio signal is greater than the distance threshold, the wake-up tag corresponding to the sample audio signal is determined as not to wake up.
[0225] Optionally, the device further comprises a fine-tuning module, which is used to:
[0226] Receiving positive samples and negative samples reported by the terminal, wherein the positive samples are training samples generated by the terminal based on the audio signal and the wake-up recognition result when the voice interaction is successful, and the negative samples are training samples generated by the terminal based on the audio signal and the wake-up recognition result when a voice interaction cancellation operation is received;
[0227] The wake-up recognition model is fine-tuned based on the positive sample and the negative sample.
[0228] In summary, in the embodiment of the present application, the wake-up recognition model is pre-trained based on sample audio signals at different audio receiving distances. After the wake-up recognition model is subsequently deployed to the terminal, the terminal can perform wake-up recognition based on the audio features of the audio signal collected by the microphone through the wake-up recognition model, and wake up the voice interaction function when the obtained wake-up recognition result indicates wake-up. With the help of the close-range audio recognition capability of the wake-up recognition model, the user can trigger voice interaction with the terminal by speaking close to the terminal without using a specific wake-up word, which simplifies the wake-up process of the voice interaction function and improves the efficiency of voice interaction.
[0229] It should be noted that: the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0230] See also Fig.10 , Fig.10 10 is a schematic diagram of a computer device provided by an exemplary embodiment of the present application. The computer device can be implemented as a terminal or a computer device in the above embodiment. The computer device may include one or more of the following components: a processor 1010 and a memory 1020.
[0231] Optionally, the processor 1010 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1020, and calling data stored in the memory 1020. Optionally, the processor 1010 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA).
[0232] The processor 1010 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), and a baseband chip. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the touch screen; the NPU is used to implement artificial intelligence (AI) functions; and the baseband chip is used to process wireless communications. It is understandable that the above-mentioned baseband chip may not be integrated into the processor 1010, but may be implemented by a single chip.
[0233] The memory 1020 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 1020 includes a non-transitory computer-readable storage medium. The memory 1020 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 1020 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the electronic device, etc.
[0234] In addition, those skilled in the art will understand that the structure of the terminal shown in the above drawings does not constitute a limitation on the terminal, and the terminal may include more (such as a microphone, a speaker, a power supply component, a display component, a sensor component) or fewer components than shown in the figure, or a combination of certain components, or different component arrangements.
[0235] An embodiment of the present application provides a computer-readable storage medium, which stores at least one computer instruction, and the at least one computer instruction is used to be executed by a processor to implement the voice interaction as described in the above embodiment, or the model training method.
[0236] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the voice interaction as described in the above embodiments, or a model training method.
[0237] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented with hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.
[0238] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A voice interaction method, characterized in that: The method comprises: Extract features of the audio signal collected by the microphone to obtain audio features; Inputting the audio feature into a wake-up recognition model to obtain a wake-up recognition result output by the wake-up recognition model, wherein the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, where the audio receiving distance is the distance from the sound source to the microphone; When the wake-up recognition result indicates wake-up, the voice interaction function is woken up.
2. The method according to claim 1, characterized in that The feature extraction of the audio signal collected by the microphone to obtain the audio feature includes: Extracting the audio signal collected by the microphone through a sliding window to obtain an audio signal segment; Perform feature extraction on the audio signal segment to obtain the audio feature of the audio signal segment.
3. The method according to claim 2, characterized in that The step of extracting the audio signal collected by the microphone through a sliding window to obtain an audio signal segment includes: When the length of the audio signal does not reach the length of the sliding window, amplify the signal within the sliding window to obtain an amplified signal, wherein the length of the amplified signal is the length of the sliding window; and determine the amplified signal as the audio signal segment; When the length of the audio signal reaches the length of the sliding window, the signal within the sliding window is extracted as the audio signal segment.
4. The method according to claim 3, characterized in that: The amplifying the signal within the sliding window to obtain an amplified signal includes: Based on the length of the audio signal and the length of the sliding window, signal interpolation processing is performed on the audio signal to obtain the amplified signal.
5. The method according to claim 2, characterized in that: The extracting features from the audio signal segment to obtain the audio features of the audio signal segment includes: Performing short-time Fourier transform on the audio signal segment to obtain a signal time-frequency spectrum; The signal time-frequency spectrum is filtered by a triangular filter to obtain the audio feature of the audio signal segment.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: In the case where the microphone is a dual microphone, performing differential processing on the dual-channel audio signals collected by the dual microphones to obtain the audio signal; In the case that the microphone is a single microphone, a single-channel audio signal collected by the single microphone is normalized to obtain the audio signal.
7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: In the case where the voice interaction is successful, generating a positive sample based on the audio signal and the wake-up recognition result; In case of receiving a voice interaction cancellation operation, generating a negative sample based on the audio signal and the wake-up recognition result; The wake-up recognition model is fine-tuned based on the positive sample and the negative sample.
8. A model training method, characterized in that: The method comprises: Acquire a sample audio signal, wherein the sample audio signal corresponds to an audio receiving distance, and the audio receiving distance is the distance from the sound source to the microphone; Extracting features from the sample audio signal to obtain sample audio features; Inputting the sample audio feature into a wake-up recognition model to obtain a sample wake-up recognition result output by the wake-up recognition model, wherein the sample wake-up recognition result includes wake-up and non-wake-up; The wake-up recognition model is trained based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal.
9. The method according to claim 8, characterized in that The extracting features of the sample audio signal to obtain sample audio features includes: Extracting the sample audio signal through a sliding window to obtain a sample audio signal segment; Feature extraction is performed on the sample audio signal segment to obtain the sample audio feature of the sample audio signal segment.
10. The method according to claim 9, characterized in that The step of extracting the sample audio signal through a sliding window to obtain a sample audio signal segment includes: In the case that the length of the sample audio signal is greater than or equal to the length of the sliding window, a signal within the sliding window is extracted as the sample audio signal segment. When the length of the sample audio signal is less than the length of the sliding window, amplify the signal within the sliding window to obtain a sample amplified signal, the length of the sample amplified signal is the length of the sliding window; and determine the sample amplified signal as the sample audio signal segment.
11. The method according to claim 10, characterized in that The amplifying the signal within the sliding window to obtain the sample amplified signal includes: Based on the length of the sample audio signal and the length of the sliding window, signal interpolation processing is performed on the sample audio signal to obtain the sample amplified signal.
12. The method according to claim 9, characterized in that The extracting features from the sample audio signal segment to obtain the sample audio features of the sample audio signal segment includes: Performing short-time Fourier transform on the sample audio signal segment to obtain a sample signal time-frequency spectrum; The sample signal time-frequency spectrum is filtered by a triangular filter to obtain the sample audio feature of the sample audio signal segment.
13. The method according to any one of claims 8 to 12, characterized in that: The sample audio signal includes a breath signal and a short-term speech signal.
14. The method according to any one of claims 8 to 12, characterized in that: The step of obtaining a sample audio signal comprises: Acquire a first sample audio signal, where the first sample audio signal is a user audio signal collected at a preset audio receiving distance; Performing data enhancement processing on the first sample audio signal to obtain a second sample audio signal; The first sample audio signal and the second sample audio signal are determined as the sample audio signals.
15. The method according to claim 14, characterized in that The performing data enhancement processing on the first sample audio signal to obtain a second sample audio signal includes at least one of the following: Performing frequency band adjustment processing on the first sample audio signal to obtain the second sample audio signal, wherein the frequency band adjustment processing includes at least one of increasing the signal frequency band and decreasing the signal frequency band; Performing speech inversion processing on the first sample audio signal to obtain the second sample audio signal, wherein the speech inversion processing is used to invert the front-to-back relationship of the speech signal.
16. The method according to claim 14, characterized in that The method further comprises: Performing noise processing on the first sample audio signal and the second sample audio signal to obtain the first sample audio signal and the second sample audio signal after noise is added; The step of determining the first sample audio signal and the second sample audio signal as the sample audio signals comprises: The first sample audio signal and the second sample audio signal after the noise is added are determined as the sample audio signals.
17. The method according to claim 16, characterized in that The performing noise processing on the first sample audio signal and the second sample audio signal to obtain the first sample audio signal and the second sample audio signal after noise addition, comprises: Determine a noise power spectrum of a preset noise, and a signal power spectrum of the first sample audio signal and the second sample audio signal; The noise power spectrum is scaled, and the scaled noise power spectrum is superimposed on the signal power spectrum to obtain the first sample audio signal and the second sample audio signal after noise is added.
18. The method according to any one of claims 8 to 12, characterized in that: The method further comprises: When the audio receiving distance corresponding to the sample audio signal is less than or equal to the distance threshold, determining the wake-up tag corresponding to the sample audio signal as wake-up; When the audio receiving distance corresponding to the sample audio signal is greater than the distance threshold, the wake-up tag corresponding to the sample audio signal is determined as not to wake up.
19. The method according to any one of claims 8 to 12, characterized in that: The method further comprises: Receiving positive samples and negative samples reported by the terminal, wherein the positive samples are training samples generated by the terminal based on the audio signal and the wake-up recognition result when the voice interaction is successful, and the negative samples are training samples generated by the terminal based on the audio signal and the wake-up recognition result when a voice interaction cancellation operation is received; The wake-up recognition model is fine-tuned based on the positive sample and the negative sample.
20. A voice interaction device, characterized in that: The device comprises: A feature extraction module is used to extract features from the audio signal collected by the microphone to obtain audio features; an inference module, configured to input the audio feature into a wake-up recognition model to obtain a wake-up recognition result output by the wake-up recognition model, wherein the wake-up recognition result includes wake-up and non-wake-up, and the wake-up recognition model is trained based on sample audio signals at different audio receiving distances, wherein the audio receiving distance is the distance from the sound source to the microphone; The voice interaction module is used to wake up the voice interaction function when the wake-up recognition result indicates wake-up.
21. A model training device, characterized in that: The device comprises: A sample acquisition module is used to acquire a sample audio signal, wherein the sample audio signal corresponds to an audio receiving distance, and the audio receiving distance is the distance from the sound source to the microphone; A feature extraction module, used to extract features from the sample audio signal to obtain sample audio features; A training module is used to input the sample audio features into a wake-up recognition model to obtain a sample wake-up recognition result output by the wake-up recognition model, wherein the sample wake-up recognition result includes wake-up and non-wake-up; based on the sample wake-up recognition result and the wake-up label corresponding to the sample audio signal, the wake-up recognition model is trained.
22. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to implement the voice interaction method as described in any one of claims 1 to 7, or the model training method as described in any one of claims 8 to 19.
23. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer instruction, and the at least one computer instruction is used to be executed by a processor to implement the voice interaction method as described in any one of claims 1 to 7, or the model training method as described in any one of claims 8 to 19.
24. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; the processor reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the voice interaction method as described in any one of claims 1 to 7, or the model training method as described in any one of claims 8 to 19.
Citation Information
Patent Citations
Sound source positioning method and device, readable storage medium and electronic equipment
CN111161757A
Distributed identification in networked system
CN111448549A
Speech recognition method and device, and storage medium
CN114464184A
Voice wake-up method, acoustic model training method and related device
CN115223555A
Voice processing method and device, terminal equipment, server equipment and storage medium
CN115762504A