Audio playing control method and device, and electronic device

By automatically adjusting the audio playback status of electronic devices through voice wake-up and type recognition models, the hassle of users manually adjusting the audio playback status is eliminated, improving ease of operation and detection accuracy, and ensuring call quality.

CN115579002BActive Publication Date: 2026-04-21SOUNDAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUNDAI TECH CO LTD
Filing Date
2022-08-18
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, after a user answers a call from a device that can make calls, they need to manually adjust the audio playback status of the electronic device, which is cumbersome and easily affected by interference signals, resulting in poor call quality.

Method used

The confidence level of the wake word is detected by a voice wake-up model, and the wake word category is identified by a type recognition model. The audio playback status of the electronic device is automatically adjusted, including obtaining the target audio adjustment method and performing the corresponding operation.

Benefits of technology

It enables automatic adjustment of the audio playback status of electronic devices when users answer calls from callable devices, improving ease of operation and detection accuracy, avoiding erroneous adjustments, and ensuring call quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115579002B_ABST
    Figure CN115579002B_ABST
Patent Text Reader

Abstract

This invention provides an audio playback control method, apparatus, and electronic device, relating to the field of voice processing technology. The audio playback control method includes: acquiring a sampled signal and inputting it into a voice wake-up model to obtain the confidence level of the sampled signal containing a wake-up word, the voice wake-up model being used for wake-up word detection; if the confidence level is greater than a preset confidence level, extracting a signal segment containing the wake-up word from the sampled signal and inputting this segment into a type recognition model to obtain the category of the wake-up word output by the type recognition model, the type recognition model being used for wake-up word type recognition; if the category is a call intro wake-up word and the electronic device is detected to be playing audio, adjusting the current audio playback state of the electronic device. The technical solution of this invention can automatically adjust the audio playback state of an electronic device when a user connects to a callable device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to an audio playback control method, apparatus, and electronic device. Background Technology

[0002] With the development of artificial intelligence technology, electronic devices with audio playback functions are becoming increasingly intelligent, giving rise to devices such as smart robots, smart TVs, and smart speakers that can interact via voice. During use, users may need to answer calls from mobile phones, smartwatches, or other devices capable of making calls. This may require adjusting the audio playback settings of the electronic device to suit the call environment. Currently, this adjustment is primarily done manually by the user, which is rather cumbersome. Summary of the Invention

[0003] This invention provides an audio playback control method, device, and electronic device to solve the defect in the prior art where users need to manually adjust the audio playback status of the electronic device after answering a callable device, and to realize the automatic adjustment of the audio playback status of the electronic device when the user answers a callable device.

[0004] This invention provides an audio playback control method, comprising:

[0005] A sampling signal is acquired and input into a voice wake-up model to obtain the confidence level of the sampling signal containing a wake-up word, which is output by the voice wake-up model. The voice wake-up model is used to detect wake-up words.

[0006] When the confidence level is greater than a preset confidence level, a signal segment containing the wake word is extracted from the sampled signal, and the signal segment containing the wake word is input into the type recognition model to obtain the category of the wake word output by the type recognition model. The type recognition model is used to identify the type of wake word.

[0007] When the category is a call start-up wake word and the electronic device is detected to be playing audio, the current audio playback state of the electronic device is adjusted.

[0008] According to an audio playback control method provided by the present invention, adjusting the current audio playback state of the electronic device includes:

[0009] Obtain the target audio adjustment method;

[0010] The electronic device's current audio playback state is adjusted based on the target audio adjustment method.

[0011] According to an audio playback control method provided by the present invention, the method for obtaining the target audio adjustment includes:

[0012] Get the current time information;

[0013] Based on the current time information, the audio adjustment method is matched from the audio adjustment method information database to obtain the target audio adjustment method;

[0014] The audio adjustment method information database stores the correspondence between time information and audio adjustment methods, and the correspondence is determined based on the configuration operation in the audio adjustment method configuration interface.

[0015] According to an audio playback control method provided by the present invention, the step of extracting a signal segment containing a wake-up word from the sampled signal when the confidence level is greater than a preset confidence level includes:

[0016] If the confidence level is determined to be greater than the preset confidence level, the current detection position in the sampled signal is determined as the end position of the wake-up word.

[0017] Based on the position of the end of the wake word, a signal segment containing the wake word is extracted from the sampled signal.

[0018] According to an audio playback control method provided by the present invention, the step of extracting a signal segment containing the wake word from the sampled signal based on the position of the wake word's end point includes:

[0019] From the sampled signal, a signal segment with a preset time period preceding the end position of the wake-up word is extracted to obtain a signal segment containing the wake-up word; or...

[0020] From the sampled signal, a signal segment containing the wake word is obtained by extracting a preset number of audio frames before the end position of the wake word.

[0021] According to an audio playback control method provided by the present invention, after adjusting the current audio playback state of the electronic device, the method further includes:

[0022] In the case where the category is a call end-of-call wake-up word, the audio playback state of the electronic device is restored to the previous audio playback state.

[0023] According to an audio playback control method provided by the present invention, the voice wake-up model is trained based on the following steps:

[0024] Obtain sample wake-up words and non-sample wake-up words, wherein the sample wake-up words include sample device wake-up words, sample call opening wake-up words, and sample call closing wake-up words;

[0025] The initial voice wake-up model is trained based on the sample wake-up words, the non-sample wake-up words, and their corresponding label information to obtain the voice wake-up model.

[0026] According to an audio playback control method provided by the present invention, the type recognition model is trained based on the following steps:

[0027] Obtain sample wake-up words, which include sample device wake-up words, sample call opening wake-up words, and sample call closing wake-up words;

[0028] The initial type recognition model is trained based on the wake-up words of the sample devices, the wake-up words at the beginning of the sample calls, the wake-up words at the end of the sample calls, and their corresponding tag information to obtain the type recognition model.

[0029] The present invention also provides an audio playback control device, comprising:

[0030] An acquisition module is used to acquire a sampling signal and input the sampling signal into a voice wake-up model to obtain the confidence level of the sampling signal containing a wake-up word output by the voice wake-up model. The voice wake-up model is used to perform wake-up word detection.

[0031] The interception module is used to intercept a signal segment containing a wake-up word from the sampled signal when the confidence level is greater than a preset confidence level, and input the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model. The type recognition model is used to identify the type of wake-up word.

[0032] The adjustment module is used to adjust the current audio playback state of the electronic device when the category is a call opening wake-up word and the electronic device is detected to be playing audio.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio playback control method as described in any of the preceding claims.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the audio playback control method as described in any of the preceding claims.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the audio playback control method as described in any of the preceding claims.

[0036] The audio playback control method, device, and electronic device provided by this invention detect wake-up words in sampled signals using a voice wake-up model. The confidence level of the sampled signal containing a wake-up word is obtained. If this confidence level is greater than a preset confidence level, it can be determined that the sampled signal contains a wake-up word. Then, a signal segment containing the wake-up word is extracted from the sampled signal and input into a type recognition model. The type recognition model identifies the wake-up word type of the signal segment containing the wake-up word, obtaining the wake-up word category. If the category is a call-starting wake-up word and the electronic device is detected to be playing audio, the current audio playback state of the electronic device is adjusted. In this way, after a user answers a callable device, the electronic device is woken up based on the call-starting wake-up word, causing the electronic device to adjust its audio playback state, thus achieving automatic adjustment of the electronic device's audio playback state when a callable device is connected. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the audio playback control method provided by the present invention;

[0039] Figure 2 This is a flowchart illustrating the method for extracting a signal segment containing a wake-up word from a sampled signal provided by the present invention.

[0040] Figure 3 This is a flowchart illustrating the method for adjusting the current audio playback state of an electronic device provided by the present invention;

[0041] Figure 4 This is a schematic diagram of the structure of the audio playback control device provided by the present invention;

[0042] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] As electronic devices become increasingly intelligent, voice-interactive devices such as smart robots, smart TVs, smart speakers, and smart projectors have emerged. When these devices are outputting audio, users might need to answer calls from other devices like mobile phones or smartwatches. To avoid interrupting the call, users need to manually mute the audio or lower the volume, which is cumbersome. Furthermore, users may not always realize they need to mute the audio or adjust the volume, affecting call quality. Alternatively, in certain scenarios, it might be necessary to increase the volume of electronic devices to create a noisy environment. For example, if someone doesn't want to answer a call at a certain time but has to, increasing the volume can create an unpleasant atmosphere to politely decline the call.

[0045] In related technologies, it's possible to determine if someone is using a mobile phone for a call by detecting the phone signal. If so, the volume of the electronic device is lowered; otherwise, the volume is not adjusted. However, this method is susceptible to interference from other electronic devices, such as signals from other electronic devices, which can lead to erroneous volume adjustments.

[0046] Based on this, embodiments of the present invention provide an audio playback control method, which can detect wake words in the sampled signal using a voice wake-up model to obtain the confidence level of the wake word in the sampled signal. If the confidence level is greater than a preset confidence level, a signal segment containing the wake word is extracted from the sampled signal and input into a type recognition model. The type recognition model identifies the wake word type of the signal segment containing the wake word to obtain the wake word category. If the category is a call intro wake word and the electronic device is detected to be playing audio, the current audio playback state of the electronic device is adjusted.

[0047] The following is combined Figures 1-3 The audio playback control method of the present invention will be described.

[0048] Figure 1 An exemplary flowchart of an audio playback control method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the audio playback control method may include the following steps 110 to 130.

[0049] Step 110: Acquire the sampling signal and input the sampling signal into the voice wake-up model to obtain the confidence level of the wake-up word contained in the sampling signal output by the voice wake-up model. The voice wake-up model is used to perform wake-up word detection.

[0050] Voice-interactive electronic devices need to be woken up from sleep mode to enter working mode before they can process user commands. For example, an electronic device can be woken up using a voice wake-up function. In the voice wake-up function, the wake word is the key word that activates the electronic device. A voice wake-up model can be used to detect the wake word. If a wake word that can wake up the electronic device is detected, then the electronic device is activated.

[0051] Electronic devices can collect audio signals from the surrounding environment, obtain sampled signals, and input the sampled signals into a voice wake-up model. The voice wake-up model is then used to detect wake words and obtain the confidence level that the sampled signals output by the voice wake-up model contain wake words.

[0052] In one example embodiment, after the sampled signal is input into the voice wake-up model, the voice wake-up model can first extract audio features from the sampled signal, such as frequency domain feature extraction, to obtain audio feature data. Then, starting from the first frame of audio feature data, the confidence of each frame of audio feature data matching the wake-up word is calculated sequentially and accumulated to obtain the confidence that the sampled signal contains the wake-up word. When this confidence is greater than a preset confidence, it is determined that the wake-up word has been detected. For example, for a 5-second sampled signal, if the 2nd to 4th second sampled signal segment contains the wake-up word, during the process of inputting the sampled signal into the voice wake-up model for wake-up word detection, the confidence of each frame is accumulated starting from the first frame of audio feature data. The initial 2-second sampled signal segment does not contain the wake-up word, and the confidence that the sampled signal containing the wake-up word accumulated in these 2 seconds is relatively small or even close to 0. Starting from the 2nd second, the confidence that the sampled signal contains the wake-up word increases rapidly until the 4th second, when the confidence exceeds the preset confidence, at which point it can be determined that the sampled signal contains the wake-up word.

[0053] In this embodiment of the invention, conversation opening phrases such as "Hello," "Hello, who is this?" and "Hello, who is this?" can be used as a type of wake-up word, along with the existing device wake-up words, to serve as the wake-up words for the electronic device. Correspondingly, samples containing conversation opening phrases can be added to the training samples of the existing voice wake-up model to train the model; that is, sample device wake-up words and sample conversation opening phrase wake-up words are used together as sample wake-up words to train the voice wake-up model. Thus, by simply adding sample conversation opening phrase wake-up words to the training samples based on the existing training mechanism of the voice wake-up model, the voice wake-up model can be trained, making the method for obtaining the voice wake-up model simple and convenient.

[0054] Specifically, the voice wake-up model can be trained through the following steps: acquiring sample wake-up words and non-sample wake-up words, where sample wake-up words include sample device wake-up words and sample call initiation wake-up words; training the initial voice wake-up model based on the sample wake-up words, non-sample wake-up words, and their corresponding label information to obtain the voice wake-up model. This voice wake-up model can then be used to detect wake-up words in the input signal, specifically device wake-up words and call initiation wake-up words.

[0055] For example, the device wake-up word can be defined as "Xiao Ai, Xiao Ai". Pronunciation data of this wake-up word from different speakers in different pronunciation scenarios can be collected to obtain sample device wake-up words. Similarly, pronunciation data of different speakers' opening phrases such as "Hello," "Hello, who," and "Hello, who" in different pronunciation scenarios can be collected to obtain sample call opening wake-up words. Both the sample device wake-up word and the sample call opening wake-up word can be labeled with the same information as sample wake-up words. Audio data that does not contain a device wake-up word or call opening phrase can also be collected to obtain non-sample wake-up words, such as "Very good," "Nice to meet you," "It's a pleasure to meet you," etc.

[0056] For example, the initial voice wake-up model can be a Deep Neural Network (DNN) model, a Long Short-Term Memory (LSTM) model, or a Convolutional Neural Network (CNN) model, or any hybrid structure of these basic neural network models. This invention does not impose any special limitations on this.

[0057] Step 120: When the confidence level is greater than the preset confidence level, extract the signal segment containing the wake word from the sampled signal, and input the signal segment containing the wake word into the type recognition model to obtain the category of the wake word output by the type recognition model. The type recognition model is used to identify the type of wake word.

[0058] After obtaining the confidence level of the sampled signal output by the voice wake-up model to contain a wake-up word, this confidence level is compared with a preset confidence level. If the confidence level is greater than the preset confidence level, a wake-up word is detected. At this point, a signal segment containing the wake-up word is extracted from the sampled signal and input into the type recognition model. The type recognition model identifies the wake-up word type of this signal segment to determine whether the detected wake-up word is a device wake-up word or a call opening wake-up word. For example, the type recognition model can first extract audio features from the signal segment and then identify the wake-up word type based on the extracted audio feature data.

[0059] For example, the type recognition model can be trained based on the following steps: obtaining sample wake words, which include sample device wake words and sample call intro wake words; training the initial type recognition model based on the sample device wake words, sample call intro wake words and their corresponding label information to obtain the type recognition model.

[0060] In one example embodiment, the device wake-up word is defined as "Xiao Ai, Xiao Ai". A large audio dataset of different pronunciations of the device wake-up word can be recorded as a sample device wake-up word. A large audio dataset of different pronunciations of various call opening phrases such as "Hello", "Hello, who", and "Hello, who" can be recorded as a sample call opening phrase wake-up word. Then, different label information is labeled on the sample device wake-up word and the sample call opening phrase wake-up word. Based on the sample device wake-up word, the sample call opening phrase wake-up word and their corresponding label information, the initial type recognition model is trained to obtain the type recognition model.

[0061] In another example embodiment, an initial type recognition model can be trained using a pre-trained voice wake-up model to obtain a type recognition model. Specifically, training samples used during the training of the voice wake-up model can be used, and these training samples can be labeled with sample device wake-up words and sample call intro wake-up words. The labeled training samples are then input into the voice wake-up model, which detects wake-up words and obtains the confidence level of the training samples output by the voice wake-up model containing wake-up words. If this confidence level is greater than a preset confidence level, samples containing wake-up words are extracted from the training samples, and these extracted samples are input into the initial type recognition model to train the initial type recognition model.

[0062] For example, the initial type recognition model can be a Deep Neural Network (DNN) model, a Long Short-Term Memory (LSTM) model, or a Convolutional Neural Network (CNN) model, or any hybrid structure of these basic neural network models. This invention does not impose any special limitations on this.

[0063] For example, Figure 2 This diagram illustrates a method for extracting a signal segment containing a wake word from a sampled signal according to an embodiment of the present invention. (Refer to...) Figure 2 As shown, the method may include the following steps 210 to 220.

[0064] Step 210: Determine the current detection position in the sampled signal as the end position of the wake-up word.

[0065] For example, for a 5-second sampled signal, during the wake-up word detection process, the voice wake-up model accumulates the confidence score of each frame starting from the first frame of audio feature data. When the 4th second is detected, if the confidence score of the sampled signal output by the voice wake-up model containing the wake-up word is greater than the preset confidence score, then the position of the 4th second in the sampled signal can be determined as the end position of the wake-up word.

[0066] Step 220: Based on the position of the end of the wake word, extract the signal segment containing the wake word from the sampled signal.

[0067] After determining the end position of the wake-up word, in one example embodiment, a signal segment containing the wake-up word can be extracted from the sampled signal for a preset time period preceding the end position of the wake-up word. For example, the 4th second position in the sampled signal can be determined as the end position of the wake-up word. The preset time period can be determined based on the length of the wake-up word. For example, if it takes about 2 seconds to say a wake-up word, the preset time period can be set to at least 2 seconds. If the preset time period is 2 seconds, then the sampled signal for a period of 2 to 4 seconds can be extracted as the signal segment containing the wake-up word.

[0068] In another example embodiment, a signal segment containing the wake-up word can be obtained by extracting a preset number of audio frames before the end of the wake-up word from the sampled signal. For example, when the 20th audio frame is detected, if the confidence level of the wake-up word in the sampled signal output by the voice wake-up model is greater than a preset confidence level, then the position of the 20th frame is determined as the end of the wake-up word. The preset number of audio frames can be determined based on the length of the wake-up word. For example, if a wake-up word occupies 10 audio frames, then the preset number of audio frames can be set to 10. In this case, a signal segment containing the wake-up word can be extracted from the 10th to 20th audio frames in the sampled signal.

[0069] Step 120 is illustrated using the example of extracting a signal segment containing a wake-up word from the sampled signal. In another example embodiment, an audio feature data segment containing a wake-up word can also be extracted from the audio feature data extracted by the voice wake-up model, and this audio feature data segment can be input into the type recognition model. The type recognition model identifies the type of wake-up word based on this audio feature data segment. In this way, the type recognition model can directly utilize the audio feature extraction results of the voice wake-up model without having to extract audio features again, thus improving data processing efficiency.

[0070] Step 130: If the category is a call start-up wake word and the electronic device is detected to be playing audio, adjust the current audio playback status of the electronic device.

[0071] When the wake-up word output by the type recognition model falls under the category of a call intro wake-up word, the electronic device can determine that the user is currently answering a call and in the middle of a conversation. If the device is detected to be playing audio, it adjusts the audio playback settings accordingly. If audio playback is not detected, the device maintains its current state. In this way, when the user answers a call while the electronic device is playing audio, it can automatically adjust the audio playback to meet the user's call environment needs, such as lowering the volume to ensure call quality.

[0072] For example, an electronic device can adjust its current audio playback state according to a preset audio adjustment method, or it can adjust its current audio playback state according to a user-configured audio adjustment method.

[0073] For example, Figure 3 An exemplary flowchart illustrates a method for adjusting the current audio playback state of an electronic device according to an embodiment of the present invention. (Refer to...) Figure 3 As shown, the method may include the following steps 310 to 320.

[0074] Step 310: Obtain the target audio adjustment method.

[0075] Target audio adjustment methods may include increasing the audio playback volume, changing the audio playback content, pausing audio playback, turning off the audio, or decreasing the audio playback volume.

[0076] In one example embodiment, an electronic device may preset an audio adjustment method, such as reducing the audio playback volume. When the wake-up word output by the type recognition model is a call opening wake-up word and the electronic device is detected to be playing audio, the electronic device may obtain the preset audio adjustment method and determine the audio adjustment method as the target audio adjustment method.

[0077] In another example embodiment, the electronic device may also provide multiple optional audio adjustment methods for the user to select. For instance, the electronic device can provide an audio adjustment method configuration interface, displaying optional audio adjustment methods, from which the user can select one to complete the audio adjustment configuration. When the wake-up word output by the type recognition model is a call intro wake-up word and the electronic device is detected to be playing audio, the electronic device can obtain the user-selected audio adjustment method and identify it as the target audio adjustment method.

[0078] In another example embodiment, the electronic device can provide a user with an audio adjustment mode configuration interface. This interface includes audio adjustment mode configuration controls, allowing users to set different audio adjustment modes for different time periods, thus making the audio adjustment more flexible. Specifically, obtaining the target audio adjustment mode may include: obtaining current time information; matching audio adjustment modes from an audio adjustment mode information database based on the current time information to obtain the target audio adjustment mode; wherein the audio adjustment mode information database stores the correspondence between time information and audio adjustment modes, and this correspondence is determined based on the configuration operations in the audio adjustment mode configuration interface.

[0079] For example, if a user needs a quiet environment to ensure call quality, they can configure the electronic device's audio adjustment methods to change the audio playback content, pause audio playback, mute audio, or reduce audio volume. This way, when the user answers a call on an available device while the electronic device is outputting audio, the device can change the audio playback content, pause audio playback, mute audio, or reduce audio volume to maintain call quality. Changing the audio playback content could, for example, replace the currently playing audio with soothing background music.

[0080] For example, if a user may not want to answer a certain call during a certain time period but has to, in this scenario, the user can configure the electronic device to increase the audio playback volume during that time period. In this way, when the user answers the call during that time period, the electronic device can increase the volume to create an environment that makes it inconvenient to talk, thus politely refusing to talk.

[0081] Step 320: Adjust the current audio playback status of the electronic device based on the target audio adjustment method.

[0082] The audio playback control method provided in this invention uses a voice wake-up model to detect wake-up words in a sampled signal, obtaining the confidence level of the sampled signal containing a wake-up word. If this confidence level is greater than a preset confidence level, it can be determined that the sampled signal contains a wake-up word. Then, a signal segment containing the wake-up word is extracted from the sampled signal and input into a type recognition model. The type recognition model identifies the wake-up word type of the signal segment containing the wake-up word, obtaining the wake-up word category. If the category is a call-starting wake-up word and the electronic device is detected to be playing audio, the current audio playback state of the electronic device is adjusted. In this way, after a user answers a callable device, the electronic device is woken up based on the call-starting wake-up word, causing the electronic device to adjust its audio playback state. This achieves automatic adjustment of the electronic device's audio playback state when a call is connected, making it more convenient to use. Furthermore, compared to methods that detect call connection via mobile phone signal, this method avoids false detections caused by interference from other signals, improving the accuracy of call connection detection.

[0083] based on Figure 1 In one example embodiment of the audio playback control method corresponding to the embodiments, after the electronic device obtains the sampling signal, it can first extract audio features from the sampling signal to obtain audio feature data, and then input the audio feature data into a voice wake-up model. The voice wake-up model performs wake-up word detection based on the audio feature data. Correspondingly, an audio feature data segment containing the wake-up word can be extracted from the audio feature data and then input into a type recognition model. The type recognition model identifies the wake-up word type based on the audio feature data segment. In this way, audio feature extraction of the sampling signal only needs to be performed once, improving data processing efficiency.

[0084] based on Figure 1 In one example embodiment of the audio playback control method corresponding to the embodiments, after adjusting the current audio playback state of the electronic device, the audio playback control method may further include: if the wake-up word output by the type recognition model is a call end-of-call wake-up word, restoring the audio playback state of the electronic device to the state before adjustment. The call end-of-call wake-up word may include, for example, "goodbye" or "bye-bye".

[0085] Accordingly, the voice wake-up model can be trained based on the following steps: acquiring sample wake-up words and non-sample wake-up words, where the sample wake-up words include sample device wake-up words, sample call opening wake-up words, and sample call closing wake-up words; training the initial voice wake-up model based on the sample wake-up words, non-sample wake-up words, and their corresponding label information to obtain the voice wake-up model. The type recognition model can be trained based on the following steps: acquiring sample wake-up words, including sample device wake-up words, sample call opening wake-up words, and sample call closing wake-up words; training the initial type recognition model based on the sample device wake-up words, sample call opening wake-up words, sample call closing wake-up words, and their corresponding label information to obtain the type recognition model.

[0086] By restoring the electronic device's audio playback state to its previous state when the wake-up word output by the type recognition model is a call-ending wake-up word, the system can automatically restore the electronic device's audio playback state after the user finishes answering a call. This could include resuming paused audio or restoring the volume to its previous level, thus reducing user intervention.

[0087] based on Figure 1 The audio playback control method in the corresponding embodiment, taking the device wake-up word as "Xiao Ai, Xiao Ai", the call opening wake-up words as "Hello", "Hello, who is this?" and "Hello, who is this?", the call ending wake-up words as "Goodbye" and "Bye-bye", and the target audio adjustment method as reducing volume, as an example, assuming the electronic device is currently playing audio, such as music or video, the electronic device will continuously collect the surrounding environment's voice signals to obtain sampling signals. If the user answers an incoming call and emits the voice signal "Hello, who is this?" during the call, the electronic device can obtain the sampling signal of that voice signal and output it. The sampled signal is fed into a voice wake-up model, which detects wake-up words. If the confidence level of the voice wake-up model is greater than the preset confidence level when the last frame of audio of the "position" is detected, the electronic device will extract the "Hello, who is this?" signal segment and input it into a type recognition model. The type recognition model will identify the wake-up word type of the signal segment and recognize that the wake-up word contained in the signal segment is the wake-up word at the beginning of the call. At this time, the electronic device will reduce the volume, for example, it can reduce it to a preset volume value or reduce the volume by a preset percentage, to provide the user with a relatively quiet call environment and ensure call quality.

[0088] Before the call ends, for example, if the user sends a voice signal saying "Let's reschedule for another day, goodbye," the electronic device can obtain a sampled signal of this voice signal and input it into the voice wake-up model. The voice wake-up model performs wake-up word detection on the sampled signal. When it detects the last frame of audio for "see you," the confidence level output by the voice wake-up model is greater than the preset confidence level. At this point, the electronic device extracts the "goodbye" signal segment and inputs it into the type recognition model. The type recognition model can identify that the wake-up word contained in this signal segment is the call-ending wake-up word. At this time, the electronic device can restore the audio playback volume to the volume before answering the call, realizing automatic restoration of the audio playback state. For example, the electronic device can record the current audio playback volume value when the type recognition model recognizes the wake-up word as the call-starting wake-up word, and restore the volume to the recorded volume value when the type recognition model recognizes the call-ending wake-up word.

[0089] Regardless of whether the electronic device is playing audio, as long as the type recognition model identifies the wake word as the device wake word, the electronic device will be woken up and enter the working state of receiving voice commands.

[0090] The audio playback control method provided in this invention can utilize a type recognition model to further identify the wake-up word recognized by the voice wake-up model. Different control methods are applied to the electronic device based on different type recognition results. Only when the wake-up word is a device wake-up word can the electronic device be truly woken up, enabling it to enter a working state that receives voice commands. Call opening and closing wake-up words are only used to activate the audio playback control function of the electronic device, adjusting its audio playback state, but not to enable it to enter a working state that receives voice commands. For example, the voice prompt "Xiao Ai, Xiao Ai, what's the weather like today?" can be identified as containing the device wake-up word "Xiao Ai, Xiao Ai," waking the electronic device and allowing it to receive the voice query command "what's the weather like today?" to provide the user with weather information. However, the voice prompt "Hello, what's the weather like today?" can be identified as containing the call opening wake-up word "Hello." If the electronic device is currently in audio playback mode, its current audio playback state will be adjusted; otherwise, it will maintain its current working state and will not receive the voice command "what's the weather like today?". For the voice prompt "Goodbye. How's the weather today?", if the system recognizes the call-ending wake-up word "Goodbye", it can restore the electronic device to its previous audio playback state and will not receive the "How's the weather today?" voice command. Thus, by adding a type recognition model after the voice wake-up model, and using this model to further identify the type of wake-up word detected by the voice wake-up model, and then controlling the electronic device differently based on different type recognition results, it can avoid false wake-ups of the electronic device when at least one of the call-opening or call-ending wake-up words is introduced.

[0091] The audio playback control device provided by the present invention is described below. The audio playback control device described below can be referred to in correspondence with the audio playback control method described above.

[0092] Figure 4 An exemplary schematic diagram of the audio playback control device provided by the present invention is shown, with reference to... Figure 4As shown, the audio playback control device 400 may include an acquisition module 410, a segmentation module 420, and an adjustment module 430. The acquisition module 410 is used to acquire a sampled signal and input it into a voice wake-up model to obtain the confidence level of the sampled signal containing the wake-up word, which is used by the voice wake-up model for wake-up word detection. The segmentation module 420 is used to extract a signal segment containing the wake-up word from the sampled signal when the confidence level is greater than a preset confidence level, and input the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model, which is used for wake-up word type recognition. The adjustment module 430 is used to adjust the current audio playback state of the electronic device when the category is a call intro wake-up word and the electronic device is detected to be playing audio.

[0093] In one exemplary embodiment of the present invention, the adjustment module 430 may include an acquisition unit and an adjustment unit. The acquisition unit may be used to acquire a target audio adjustment method; the adjustment unit may be used to adjust the current audio playback state of the electronic device based on the target audio adjustment method. The target audio adjustment method may include increasing the audio playback volume, changing the audio playback content, pausing audio playback, turning off audio, or decreasing the audio playback volume, etc.

[0094] In one exemplary embodiment of the present invention, the acquisition unit may include an acquisition subunit and a matching subunit. The acquisition subunit may be used to acquire current time information; the matching subunit may be used to match audio adjustment methods from an audio adjustment method information database based on the current time information to obtain a target audio adjustment method; wherein the audio adjustment method information database stores the correspondence between time information and audio adjustment methods, and this correspondence is determined based on configuration operations in the audio adjustment method configuration interface.

[0095] In one example embodiment of the present invention, the interception module 420 may include a determining unit and an interception unit. The determining unit may be configured to determine the current detection position in the sampled signal as the wake-up word end-point position if the confidence level is determined to be greater than a preset confidence level; the interception unit may be configured to intercept a signal segment containing the wake-up word from the sampled signal based on the wake-up word end-point position.

[0096] In one exemplary embodiment of the present invention, the interception unit may include a first interception subunit or a second interception subunit. The first interception subunit may be used to intercept a signal segment from the sampled signal before the end position of the wake-up word, resulting in a signal segment containing the wake-up word; the second interception subunit may be used to intercept a signal segment from the sampled signal before the end position of the wake-up word, resulting in a signal segment containing the wake-up word.

[0097] In one example embodiment of the present invention, the audio playback control device 400 may further include a recovery module, which can be used to restore the audio playback state of the electronic device to the audio playback state before adjustment when the category is a call end-of-call wake-up word.

[0098] In one exemplary embodiment of the present invention, the audio playback control device 400 may further include a first training module, which can be used to train a voice wake-up model. Exemplarily, the first training module may include a first sample acquisition unit and a first training unit. The first sample acquisition unit can be used to acquire sample wake-up words and non-sample wake-up words, wherein the sample wake-up words include sample device wake-up words, sample call initiation wake-up words, and sample call end wake-up words; the first training unit can be used to train an initial voice wake-up model based on the sample wake-up words, non-sample wake-up words, and their corresponding label information to obtain a voice wake-up model.

[0099] In one exemplary embodiment of the present invention, the audio playback control device 400 may further include a second training module, which can be used to train a type recognition model. Exemplarily, the second training module may include a second sample acquisition unit and a second training unit. The second sample acquisition unit can be used to acquire sample wake-up words, including sample device wake-up words, sample call intro wake-up words, and sample call outtro wake-up words; the second training unit can be used to train an initial type recognition model based on the sample device wake-up words, sample call intro wake-up words, sample call outtro wake-up words, and their corresponding tag information to obtain a type recognition model.

[0100] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device 500 may include a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call logical instructions in the memory 530 to execute the steps of the audio playback control method described above. The audio playback control method may include: acquiring a sampled signal and inputting the sampled signal into a voice wake-up model to obtain the confidence level of the wake-up word contained in the sampled signal output by the voice wake-up model, which is used for wake-up word detection; if the confidence level is greater than a preset confidence level, extracting a signal segment containing the wake-up word from the sampled signal and inputting the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model, which is used for wake-up word type recognition; if the category is a call opening wake-up word and the electronic device is detected to be in the state of playing audio, adjusting the current audio playback state of the electronic device.

[0101] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the audio playback control method provided in the above-described method embodiments. The audio playback control method may include: acquiring a sampling signal and inputting the sampling signal into a voice wake-up model to obtain the confidence level of the sampled signal output by the voice wake-up model containing a wake-up word, the voice wake-up model being used for wake-up word detection; when the confidence level is greater than a preset confidence level, extracting a signal segment containing a wake-up word from the sampled signal and inputting the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model, the type recognition model being used for wake-up word type recognition; when the category is a call opening phrase wake-up word and the electronic device is detected to be playing audio, adjusting the current audio playback state of the electronic device.

[0103] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the steps of the audio playback control method provided in the above-described method embodiments. The audio playback control method may include: acquiring a sampled signal and inputting the sampled signal into a voice wake-up model to obtain the confidence level of the sampled signal output by the voice wake-up model containing a wake-up word, the voice wake-up model being used for wake-up word detection; when the confidence level is greater than a preset confidence level, extracting a signal segment containing a wake-up word from the sampled signal and inputting the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model, the type recognition model being used for wake-up word type recognition; when the category is a call opening phrase wake-up word and the electronic device is detected to be playing audio, adjusting the current audio playback state of the electronic device.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio playback control method for an electronic device capable of voice interaction, characterized in that, include: A sampling signal is acquired through an electronic device and input into a voice wake-up model to obtain the confidence level of the sample signal containing a wake-up word output by the voice wake-up model. The voice wake-up model is used to detect wake-up words, where the wake-up word is the activation keyword of the electronic device, and the electronic device has audio playback and voice interaction functions. When the confidence level is greater than a preset confidence level, a signal segment containing the wake word is extracted from the sampled signal, and the signal segment containing the wake word is input into the type recognition model to obtain the category of the wake word output by the type recognition model. The type recognition model is used to identify the type of wake word. When the category is a call start-up wake word and the electronic device is detected to be playing audio, the current audio playback state of the electronic device is adjusted.

2. The audio playback control method according to claim 1, characterized in that, Adjusting the current audio playback state of the electronic device includes: Obtain the target audio adjustment method; The electronic device's current audio playback state is adjusted based on the target audio adjustment method.

3. The audio playback control method according to claim 2, characterized in that, The methods for obtaining and adjusting the target audio include: Get the current time information; Based on the current time information, the audio adjustment method is matched from the audio adjustment method information database to obtain the target audio adjustment method; The audio adjustment method information database stores the correspondence between time information and audio adjustment methods, and the correspondence is determined based on the configuration operation in the audio adjustment method configuration interface.

4. The audio playback control method according to claim 1, characterized in that, When the confidence level is greater than a preset confidence level, the step of extracting a signal segment containing the wake-up word from the sampled signal includes: If the confidence level is determined to be greater than the preset confidence level, the current detection position in the sampled signal is determined as the end position of the wake-up word. Based on the position of the end of the wake word, a signal segment containing the wake word is extracted from the sampled signal.

5. The audio playback control method according to claim 4, characterized in that, The step of extracting a signal segment containing the wake word from the sampled signal based on the end position of the wake word includes: From the sampled signal, a signal segment with a preset time period preceding the end position of the wake-up word is extracted to obtain a signal segment containing the wake-up word; or... From the sampled signal, a signal segment containing the wake word is obtained by extracting a preset number of audio frames before the end position of the wake word.

6. The audio playback control method according to claim 1, characterized in that, After adjusting the current audio playback state of the electronic device, the method further includes: In the case where the category is a call end-of-call wake-up word, the audio playback state of the electronic device is restored to the previous audio playback state.

7. The audio playback control method according to claim 6, characterized in that, The voice wake-up model is trained based on the following steps: Obtain sample wake-up words and non-sample wake-up words, wherein the sample wake-up words include sample device wake-up words, sample call opening wake-up words, and sample call closing wake-up words; The initial voice wake-up model is trained based on the sample wake-up words, the non-sample wake-up words, and their corresponding label information to obtain the voice wake-up model.

8. The audio playback control method according to claim 6, characterized in that, The type recognition model is trained based on the following steps: Obtain sample wake-up words, which include sample device wake-up words, sample call intro wake-up words, and sample call outtro wake-up words; The initial type recognition model is trained based on the wake-up words of the sample devices, the wake-up words at the beginning of the sample calls, the wake-up words at the end of the sample calls, and their corresponding tag information to obtain the type recognition model.

9. An audio playback control device for an electronic device capable of voice interaction, characterized in that, include: The acquisition module is used to acquire a sampling signal through an electronic device and input the sampling signal into a voice wake-up model to obtain the confidence level of the sampling signal containing a wake-up word output by the voice wake-up model. The voice wake-up model is used to detect the wake-up word, which is the activation keyword of the electronic device. The electronic device has audio playback and voice interaction functions. The interception module is used to intercept a signal segment containing a wake-up word from the sampled signal when the confidence level is greater than a preset confidence level, and input the signal segment containing the wake-up word into a type recognition model to obtain the category of the wake-up word output by the type recognition model. The type recognition model is used to identify the type of wake-up word. The adjustment module is used to adjust the current audio playback state of the electronic device when the category is a call start-up wake word and the electronic device is detected to be playing audio.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio playback control method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio playback control method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice recognition method and device, and electronic equipment

    CN108694940A

  • Vehicle sound equipment control method and device, electronic equipment and storage medium

    CN109147820A