Voice equipment control method and device, equipment and storage medium
By dynamically adjusting the wake-up threshold and processing strategy in voice devices based on the energy distribution and intensity characteristics of ambient sound signals, the problems of false triggering and decreased wake-up rate of voice wake-up technology in different environments are solved, achieving a more efficient voice interaction experience.
Patent Information
- Application Number
- CN202510914680.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-16
AI Technical Summary
Existing voice wake-up technologies are prone to false triggering or decreased wake-up rates in different environments, affecting the user's interactive experience.
By acquiring the ambient sound signal of the environment where the voice device is located, determining its energy distribution characteristics and sound intensity characteristics, and dynamically adjusting the wake-up threshold and processing strategy based on these characteristics, the performance of voice wake-up is optimized.
A balance is achieved between the wake-up rate and false wake-up rate of voice devices in different environments, improving the user's interactive experience.
Smart Images

Figure CN120656452A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice wake-up technology, and in particular to a method, apparatus, device, and storage medium for controlling a voice device. Background Art
[0002] Voice wake-up technology is widely used in various devices so that users can control the devices through voice commands.
[0003] However, in related technologies, the voice wake-up solution adopts a fixed wake-up threshold and processing strategy. Regardless of how the surrounding noise changes, the speaker always recognizes the wake-up command with the same sensitivity, which makes it easy for voice wake-up to be falsely triggered in noisy environments. When using voice wake-up in a quiet environment, it will break the quiet atmosphere of the environment and reduce the interactive experience of voice wake-up. Summary of the Invention
[0004] The purpose of this application is to provide a voice device control method, device, equipment and storage medium, which can take into account the wake-up rate and false wake-up rate in different environments and provide a stable and efficient voice interaction experience.
[0005] The present invention provides a method for controlling a voice device, including: Obtain the ambient sound signal of the environment where the voice device is located; Determining energy distribution characteristics and sound intensity characteristics of the ambient sound signal; determining a target arousal threshold and a target processing strategy according to the energy distribution characteristics and the sound intensity characteristics; The voice device is controlled to operate according to the target wake-up threshold and the target processing strategy.
[0006] In some embodiments, determining the energy distribution characteristics and sound intensity characteristics of the ambient sound signal includes: Performing frequency domain analysis on the ambient sound signal to obtain a frequency domain signal; Determining the energy proportion of the frequency domain signal in each frequency band to obtain the energy distribution characteristics; The sum of the energy of the frequency domain signal in each frequency band is determined to obtain the sound intensity feature.
[0007] In some embodiments, determining a target arousal threshold and a target processing strategy based on the energy distribution characteristics and the sound intensity characteristics includes: When the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency range corresponding to the energy distribution feature is less than the reference energy proportion, the first target awakening threshold and the first target processing strategy are adopted; otherwise, the second target awakening threshold and the second target processing strategy are adopted; The first target wake-up threshold is less than the reference wake-up threshold, and the first target processing strategy is to extract the Mel-frequency cepstral coefficient features in the ambient sound signal. The second target wake-up threshold is greater than the reference wake-up threshold, and the second target processing strategy is to perform noise reduction on the ambient sound signal and then extract the Mel-frequency cepstral coefficient features in the ambient sound signal.
[0008] In some embodiments, extracting Mel-frequency cepstral coefficient features from the ambient sound signal includes: Performing pre-emphasis processing on the ambient sound signal to obtain a pre-emphasized sound signal; Performing frame and window processing on the pre-emphasized sound signal to obtain multiple sound segment signals; Frequency domain analysis, Mel filtering, logarithmic operation and discrete cosine transform are performed on the sound clip signal in sequence to obtain Mel frequency cepstral coefficient features in the ambient sound signal.
[0009] In some embodiments, when the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency segment corresponding to the energy distribution feature is less than the reference energy proportion, the first target wake-up threshold linearly decreases from the reference wake-up threshold, whereas the second target wake-up threshold linearly increases from the reference wake-up threshold.
[0010] In some embodiments, controlling the voice device to operate according to the target wakeup threshold and the target processing strategy includes: processing the ambient sound signal into an intermediate sound signal based on the target processing strategy; identifying the intermediate sound signal at the target arousal threshold to obtain a target sound signal; The target sound signal is matched with a reference sound signal. If the match is successful, the voice device is controlled to respond to the target sound signal; the reference sound signal is a sound signal corresponding to the wake-up word.
[0011] In some embodiments, the voice device control method further includes: Obtain historical arousal thresholds, historical processing strategies, and historical sound signals; According to the historical wake-up threshold, the historical processing strategy and the historical sound signal, a correspondence between the wake-up threshold, the processing strategy and the sound signal is configured.
[0012] The present application also provides a voice control device, including: The first module is used to obtain the ambient sound signal of the environment where the voice device is located; The second module is used to determine the energy distribution characteristics and sound intensity characteristics of the ambient sound signal; A third module is configured to determine a target arousal threshold and a target processing strategy based on the energy distribution characteristics and the sound intensity characteristics; The fourth module is used to control the voice device to operate according to the target wake-up threshold and the target processing strategy.
[0013] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned voice device control method when executing the computer program.
[0014] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned voice device control method is implemented.
[0015] The beneficial effects of the present application are as follows: based on the energy distribution characteristics and sound intensity characteristics of the ambient sound signal in the environment where the voice device is located, the wake-up threshold and the processing strategy are adjusted in a linked manner, and then the voice device is controlled to operate according to the target wake-up threshold and the target processing strategy. The voice device can be controlled to dynamically adjust the wake-up threshold and the processing strategy according to the noise parameters of the environment, ensuring that the false awakening rate and the false awakening rate are in a relatively balanced state, rather than reducing the false awakening rate by sacrificing the wake-up rate, thus ensuring that the false awakening rate is at a low level and the wake-up rate is at a high level, so that the voice device is usually in a state of high wake-up performance in the space where it is located, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a diagram of the application environment of the voice device control method provided in an embodiment of the present application.
[0017] Figure 2 This is a flowchart of the voice device control method provided in an embodiment of the present application.
[0018] Figure 3 It is a flowchart of the specific method of step S202 provided in an embodiment of the present application.
[0019] Figure 4 It is a flowchart of the specific method of step S204 provided in an embodiment of the present application.
[0020] Figure 5 This is an optional structural diagram of the voice device control device provided in an embodiment of the present application.
[0021] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0023] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps illustrated may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. Terms such as "first" and "second" in the specification, claims, and drawings are used to distinguish similar items and are not intended to describe a specific sequence or precedence.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0025] The voice device control method provided in the embodiments of the present application can be executed by a computer device, which can be a terminal device or a server. Terminal devices include, but are not limited to, mobile phones, computers, smart home appliances, vehicle-mounted terminals, aircraft, etc. The server can be a standalone physical server, a server cluster composed of multiple physical servers, a distributed system, or a cloud server.
[0026] In addition, the information, data, and signals involved in the embodiments of this application are authorized by the relevant objects or fully authorized by all parties, and the collection, use, and processing of relevant data comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0027] Figure 1 This is an application environment diagram of the voice device control method provided in the embodiment of the present application. Figure 1 , the voice device control method is applied to a voice device control system. The voice device control system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster composed of multiple servers. The terminal 110 is used to send the ambient sound signal of the environment where the voice device is located to the server 120. The server 120 is used to obtain the ambient sound signal of the environment where the voice device is located, determine the energy distribution characteristics and sound intensity characteristics of the ambient sound signal, determine the target wake-up threshold and the target processing strategy according to the energy distribution characteristics and the sound intensity characteristics, and control the voice device to operate according to the target wake-up threshold and the target processing strategy.
[0028] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the voice device control method provided in the embodiments of the present application can also be applied to other scenarios. For example, the above-mentioned voice device control method can be directly applied to terminal 110, which is used to obtain the ambient sound signal of the environment where the voice device is located, determine the energy distribution characteristics and sound intensity characteristics of the ambient sound signal, determine the target wake-up threshold and target processing strategy based on the energy distribution characteristics and sound intensity characteristics, and control the voice device to operate according to the target wake-up threshold and target processing strategy.
[0029] To facilitate understanding of the voice device control method provided in the embodiment of the present application, the application scenario of the voice device control method is exemplarily introduced below, taking the execution subject as the terminal 110 as an example.
[0030] Figure 2 This is a flow chart of a method for controlling a voice device provided by an embodiment of the present application. Figure 2 In some embodiments, the method includes but is not limited to steps S201 to S204.
[0031] Step S201: Acquire an ambient sound signal of the environment where the voice device is located.
[0032] Ambient sound signals refer to mixed audio data containing background noise and voice commands collected by a microphone array, such as the background noise emitted by a TV when a user is watching TV. Specifically, they can be implemented using a multi-channel acoustic sensor to reflect the real-time acoustic state of the environment in which the device is located.
[0033] As some examples, one or more microphone arrays can be deployed in the environment where the voice device is located, and the microphones can collect ambient sound signals in the current environment in real time. The execution entity can obtain the ambient sound signals collected by the microphone array for processing, such as sound feature extraction.
[0034] Step S202: determining the energy distribution characteristics and sound intensity characteristics of the ambient sound signal.
[0035] Energy distribution refers to the energy distribution of ambient sound signals across different frequency bands. This can be achieved through frequency domain analysis and calculation of the energy ratios within each frequency band. This is used to identify noise types and spectral characteristics. Sound intensity refers to the total energy of ambient sound signals within a specific frequency band. This can be achieved through integration to calculate the cumulative energy value within the target frequency band. This is used to quantify the intensity level of ambient noise.
[0036] As some examples, the ambient sound signal can be subjected to frequency domain analysis, the time domain ambient sound signal can be converted into a frequency domain signal, and the energy distribution characteristics representing the noise spectrum characteristics and the sound intensity characteristics reflecting the intensity of the ambient noise can be obtained by calculation for the obtained frequency domain signal.
[0037] Step S203: determining a target awakening threshold and a target processing strategy according to the energy distribution characteristics and the sound intensity characteristics.
[0038] The target wakeup threshold refers to a dynamically adjusted voice trigger sensitivity parameter. It can be linearly adjusted based on the relationship between energy distribution characteristics, sound intensity characteristics, and the wakeup threshold to balance the false trigger rate and missed detection rate. The target processing strategy refers to an audio processing flow optimized for noise characteristics. This can be achieved through various combinations of Mel coefficient extraction or noise reduction preprocessing to improve the effectiveness of speech feature extraction.
[0039] As some examples, the target environmental state corresponding to the energy distribution characteristics and sound intensity characteristics can be determined based on the correspondence between the preset energy distribution, sound intensity and environmental state, and then the wake-up threshold and processing strategy corresponding to the target environmental state can be determined as the target wake-up threshold and target processing strategy based on the correspondence between the environmental state, the wake-up threshold and the processing strategy. For example, if it is determined based on the correspondence between the energy distribution, sound intensity and environmental state that the environmental state corresponding to the currently extracted energy distribution characteristics and sound intensity characteristics is a quiet environment, the wake-up threshold is automatically lowered and a direct feature extraction strategy is adopted, so that the voice device responds to soft commands with higher sensitivity in a quiet environment. If it is determined based on the correspondence between the energy distribution, sound intensity and environmental state that the environmental state corresponding to the currently extracted energy distribution characteristics and sound intensity characteristics is a noisy environment, the wake-up threshold is automatically increased and noise reduction preprocessing is started to effectively suppress the interference of background noise on voice recognition.
[0040] Step S204: Control the voice device to operate according to the target wake-up threshold and target processing strategy.
[0041] The executive entity controls the voice device to operate according to the most recently determined target wake-up threshold and target processing strategy. During operation, the executive entity continuously obtains the ambient sound wave signals in the environment where the voice device is located and converts them into digital audio data. The energy distribution characteristics that characterize the noise spectrum characteristics and the sound intensity characteristics that reflect the intensity of the ambient noise are calculated through frequency domain analysis. According to the correspondence between energy distribution, sound intensity and environmental state, the target environmental state corresponding to the energy distribution characteristics and sound intensity characteristics is determined. Then, according to the correspondence between the environmental state, the wake-up threshold and the processing strategy, the wake-up threshold and processing strategy corresponding to the target environmental state are determined as the target wake-up threshold and target processing strategy. The working mode is dynamically switched according to the target wake-up threshold and target processing strategy determined in real time to achieve environmentally adaptive voice interaction control. Therefore, the wake-up threshold and processing strategy are adjusted in linkage based on the energy distribution characteristics and sound intensity characteristics of the ambient sound signal in the environment where the voice device is located, and then the voice device is controlled to operate according to the target wake-up threshold and target processing strategy. The voice device can be controlled to dynamically adjust the wake-up threshold and processing strategy according to the noise parameters of the environment, ensuring that the false awakening rate and the false awakening rate are in a relatively balanced state, rather than reducing the false awakening rate by sacrificing the wake-up rate. This ensures that the false awakening rate is at a low level and the wake-up rate is at a high level, so that the voice device is usually in a state of high wake-up performance in the space where it is located, thereby improving the user experience.
[0042] Figure 3 This is a flowchart of the specific method of step S202 provided in the embodiment of the present application. Figure 3 In some embodiments, the method includes but is not limited to steps S301 to S303.
[0043] Step S301 : Perform frequency domain analysis on the ambient sound signal to obtain a frequency domain signal.
[0044] Step S302: determine the energy proportion of the frequency domain signal in each frequency band to obtain energy distribution characteristics.
[0045] Step S303: Determine the sum of the energy of the frequency domain signal in each frequency band to obtain the sound intensity feature.
[0046] Frequency domain analysis refers to converting time domain sound signals into frequency domain signals. It can be achieved by using the fast Fourier transform algorithm, which decomposes the time domain waveform into a superposition of different frequency components.
[0047] After the ambient sound signal is collected, it is first converted into a frequency domain signal through a fast Fourier transform. The frequency domain signal is divided into multiple preset frequency bands, for example, 0-4kHz is divided into the human voice band, and 4kHz-8kHz is divided into the high-frequency noise band. The energy value of each frequency band is obtained by calculating the square sum of all frequency components within the band. The ratio of the energy value of each frequency band to the total energy forms the energy distribution feature. At the same time, the cumulative sum of the energy values of all frequency bands forms the sound intensity feature. Using these two features, the spectral distribution pattern and intensity level of the ambient noise can be accurately identified.
[0048] In some embodiments, step S203 specifically includes: when the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency band corresponding to the energy distribution feature is less than the reference energy proportion, adopting the first target awakening threshold and the first target processing strategy; otherwise, adopting the second target awakening threshold and the second target processing strategy. The first target awakening threshold is less than the reference awakening threshold, and the first target processing strategy is to extract the Mel-frequency cepstral coefficient feature from the ambient sound signal; the second target awakening threshold is greater than the reference awakening threshold, and the second target processing strategy is to extract the Mel-frequency cepstral coefficient feature from the ambient sound signal after denoising the ambient sound signal.
[0049] The reference sound intensity refers to a preset noise intensity threshold, which can be determined by the statistical average of historical environmental noise data or the user's parameter configuration operation, and is used to determine whether the current environment is in a low-noise state. The reference energy ratio refers to a preset energy distribution threshold, which can be determined by experimental data on the energy ratio of typical noise frequency bands, and is used to identify whether there is noise interference. It can be understood that when the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy ratio in the target frequency band corresponding to the energy distribution feature is less than the reference energy ratio, it is a quiet environment, such as a bedroom late at night. When the sound intensity corresponding to the sound intensity feature is not less than the reference sound intensity and the energy ratio in the target frequency band corresponding to the energy distribution feature is not less than the reference energy ratio, it is a noisy environment, such as a living room or outdoor place.
[0050] When the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency band corresponding to the energy distribution feature is less than the reference energy proportion, a wake-up threshold less than the reference wake-up threshold is selected as the target wake-up threshold, that is, the first target wake-up threshold. Before the first target wake-up threshold is used to identify the ambient sound signal, the first target processing strategy is used to process the ambient sound signal, that is, to extract the Mel-frequency cepstral coefficient features in the ambient sound signal, so as to avoid the user having to repeat the wake-up command due to the wake-up threshold being too high and the user's voice being too low to wake up. Conversely, a wake-up threshold greater than the reference wake-up threshold is selected as the target wake-up threshold, that is, the second target wake-up threshold. Before the second target wake-up threshold is used to identify the ambient sound signal, the second target processing strategy is used to process the ambient sound signal, that is, to first perform noise reduction processing on the ambient sound signal and then extract the Mel-frequency cepstral coefficient features in the ambient sound signal, thereby suppressing the interference of noise on voice wake-up. For example, in a conference room scenario, if the ambient sound intensity is detected to be less than 60 decibels and the energy proportion of the 300-3400Hz frequency band is less than 40%, the first target wake-up threshold and the first target processing strategy are adopted to make the voice device easier to be awakened by a gentle voice. If the sound intensity is detected to be greater than 60 decibels or the energy proportion of this frequency band is greater than 40%, the second target wake-up threshold and the second target processing strategy are adopted to prevent the voice device from being mistakenly triggered by background conversations. In this way, the wake-up threshold can be lowered in a low-noise environment and the human voice characteristics in the ambient sound signal can be emphasized through the corresponding processing strategy, avoiding the need for users to repeat the wake-up command due to the wake-up threshold being too high and the failure to wake up due to the user's voice being too low. In a high-noise environment, the wake-up threshold is raised and noise reduction processing is added to effectively suppress the interference of noise on voice wake-up.
[0051] In some embodiments, noise reduction of an ambient sound signal can be performed by inputting the ambient sound signal into a pre-trained noise reduction network. The noise reduction network processes the input ambient sound signal and extracts and transforms the input data through weights and biases in each layer of neurons in the noise reduction network, gradually removing noise interference and restoring speech features to separate the noise and speech components, outputting a relatively clear speech component. The noise reduction network can be a network based on a convolutional neural network or a recurrent neural network as the main structure, trained using a large amount of noisy speech data in a noisy environment.
[0052] In some embodiments, extracting Mel-frequency cepstral coefficient features in the ambient sound signal includes: pre-emphasis processing on the ambient sound signal to obtain a pre-emphasized sound signal; frame-dividing and windowing processing on the pre-emphasized sound signal to obtain multiple sound segment signals; performing frequency domain analysis, Mel filtering, logarithmic operation and discrete cosine transform on the sound segment signal in sequence to obtain Mel-frequency cepstral coefficient features in the ambient sound signal.
[0053] Pre-emphasis processing refers to increasing the energy of the high-frequency part of the ambient sound signal through a filter, which can be achieved by using a first-order high-pass filter. For example, through the transfer function H(z) = 1-αz -1 Signal processing is performed, where the value of α can range from 0.9 to 0.97. This step can compensate for the attenuation of high-frequency components during signal transmission and enhance the formant structure of the speech signal. Frame windowing involves segmenting the continuous ambient sound signal into short time segments of fixed length and applying a window function to each frame. This can be achieved using a Hamming window or a Hanning window, for example, with a frame length set to 25 milliseconds and a frame shift set to 10 milliseconds. This step can reduce spectral leakage and make the signal smoother in the time and frequency domains. Mel filtering involves smoothing the spectrum using a set of triangular filters. This can be achieved using a Mel-scale filter bank that covers the human hearing range, for example, with 20 to 40 filters. This step simulates the human ear's nonlinear perception of sound frequencies and highlights the key frequency band characteristics of the speech signal. Discrete cosine transform converts the logarithmic energy of the Mel spectrum into a set of cepstral coefficients. This can be achieved by removing high-order coefficients and retaining low-order coefficients, for example, retaining the first 12 to 16 coefficients. This step can decouple the vocal tract excitation source and the vocal tract response and extract key features reflecting the speech content.
[0054] After acquiring the ambient sound signal, pre-emphasis is first used to enhance the high-frequency components, enabling subsequent processing to more effectively capture the consonant information in the speech. Subsequently, frame-based windowing is used to segment the signal into short time segments, and a window function is used to suppress sudden changes at frame edges, resulting in more accurate spectral information during frequency domain transformation. The signal after frequency domain transformation is then weighted using a Mel filter bank, converting the linear frequency scale to a Mel scale that aligns with auditory perception. Logarithmic operations are then used to compress the dynamic range and enhance robustness to noise. Finally, a discrete cosine transform converts the Mel spectrum energy into cepstral coefficients, eliminating correlations between dimensions and producing a low-dimensional, highly discriminative feature vector. This series of processing steps extracts key features highly correlated with the wake-up word from complex ambient sounds, providing a reliable basis for subsequent wake-up threshold adjustment and signal matching. Consequently, by enhancing the high-frequency components of speech through pre-emphasis and suppressing frame edge effects through windowing, the Mel-filtered spectrum more accurately reflects the characteristics of speech. Furthermore, the nonlinear scale design of the Mel filter bank better aligns with human hearing and effectively suppresses interfering noise in non-speech frequency bands.
[0055] In some embodiments, when the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency band corresponding to the energy distribution feature is less than the reference energy proportion, the first target wake-up threshold linearly decreases from the reference wake-up threshold, and conversely, the second target wake-up threshold linearly increases from the reference wake-up threshold.
[0056] When the environment in which the voice device is located is maintained as a quiet environment, that is, the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy ratio within the target frequency range corresponding to the energy distribution feature is less than the reference energy ratio, the execution subject dynamically adjusts the first target wake-up threshold so that the first target wake-up threshold decreases linearly from the reference wake-up threshold, thereby linearly increasing the sensitivity of the voice device to weak ambient sound signals. Conversely, when the environment in which the voice device is located is maintained as a noisy environment, that is, the sound intensity corresponding to the sound intensity feature is not less than the reference sound intensity and / or the energy ratio within the target frequency range corresponding to the energy distribution feature is not less than the reference energy ratio, the execution subject dynamically adjusts the second target wake-up threshold so that the second target wake-up threshold increases linearly from the reference wake-up threshold, thereby linearly decreasing the sensitivity of the voice device to weak ambient sound signals. For example, in a quiet scene such as a library, the first target wake-up threshold can continue to decrease over time until the user whispers the wake-up word to trigger a response. In a noisy scene such as a shopping mall, the second target wake-up threshold dynamically increases as the background noise increases and is triggered only when the energy of the voice command significantly exceeds the noise. Therefore, the wake-up threshold can be adaptively changed with environmental parameters through a linear increase and decrease mechanism, which not only avoids breaking the environmental atmosphere in quiet scenes, but also reduces false triggering caused by noise interference.
[0057] Figure 4 This is a flowchart of the specific method of step S204 provided in the embodiment of the present application. Figure 4 In some embodiments, the method includes but is not limited to steps S401 to S403.
[0058] Step S401 : Processing the ambient sound signal into an intermediate sound signal based on a target processing strategy.
[0059] Step S402 : identifying an intermediate sound signal under a target wake-up threshold to obtain a target sound signal.
[0060] Step S403: Match the target sound signal with the reference sound signal. If the match is successful, the voice device is controlled to respond to the target sound signal.
[0061] The reference sound signal is the sound signal corresponding to the wake-up word. It is understood that the reference sound signal refers to a pre-stored standard voice feature template, which can be implemented using a Mel-frequency cepstral coefficient feature library, and is used to perform pattern matching with the target sound signal corresponding to the ambient voice signal.
[0062] After determining the target wakeup threshold and target processing strategy, the ambient sound signal is first preprocessed using the target processing strategy. For example, in noisy environments, a noise reduction filter is applied before extracting the Mel-frequency cepstral coefficient features from the ambient sound signal; in quiet environments, the Mel-frequency cepstral coefficient features are directly extracted from the ambient sound signal. The resulting intermediate sound signal is then input into the speech recognition module, where the execution agent dynamically identifies the intermediate sound signal based on the target wakeup threshold. When the sound intensity of the intermediate sound signal exceeds the sound intensity corresponding to the target wakeup threshold, the execution agent controls the speech recognition module to identify the voiceprint features in the intermediate sound signal and obtain the target sound signal. Otherwise, the intermediate sound signal is not responded to. After obtaining the target sound signal, the target sound signal is matched with a reference sound signal. If an acoustic pattern matching the preset wakeup word characteristics is detected in the target sound signal, the execution agent matches the target sound signal with the reference sound signal and calculates similarity against the reference sound signal database. If a match is successful, the voice device is triggered to respond to the target sound signal and execute the corresponding control command. Therefore, by dynamically adjusting the processing strategy and wake-up threshold, the device sensitivity is reduced in a quiet environment to avoid false triggering, and the signal processing strength is increased in a noisy environment to ensure effective wake-up, thus achieving environmentally adaptive voice interaction control and improving the environmental adaptability of voice devices.
[0063] In some embodiments, the voice device control method further includes: obtaining historical wake-up thresholds, historical processing strategies, and historical sound signals; and configuring the correspondence between the wake-up thresholds, processing strategies, and sound signals based on the historical wake-up thresholds, historical processing strategies, and historical sound signals.
[0064] The historical wake-up threshold refers to the wake-up threshold parameter actually used by the voice device during past operation. This information can be obtained from device logs or database records and is used to reflect sensitivity requirements in different environments. The historical processing strategy refers to the specific combination of steps used by the voice device to reduce noise and extract features from sound signals in different scenarios. For example, the Mel filter parameters or the type of noise reduction algorithm can be retrieved from historical operation records. The historical sound signal refers to the ambient sound signal collected and stored by the device during past operation. This signal can be continuously collected and cached using the microphone.
[0065] In actual applications, by analyzing the historical wake-up threshold and the relevant characteristics of the ambient sound signal in the corresponding time period, the optimal wake-up threshold setting under different noise intensities or spectral distributions can be identified. At the same time, combined with the execution effect of the historical processing strategy, such as the change in speech recognition accuracy after noise reduction, the association rules between the processing strategy and the sound signal characteristics can be established. For example, the correspondence between the wake-up threshold, processing strategy and sound signal can be learned based on machine learning methods, and the current ambient sound signal can be configured based on the learning results. Therefore, when the voice device detects that the current ambient sound characteristics match the historical data, it can automatically call the pre-configured wake-up threshold and processing strategy combination to achieve dynamic parameter adaptation.
[0066] See also Figure 5 The present application also provides a voice device control device that can implement the above-mentioned voice device control method. The device includes: The first module 501 is used to obtain the ambient sound signal of the environment where the voice device is located; The second module 502 is used to determine the energy distribution characteristics and sound intensity characteristics of the ambient sound signal; The third module 503 is used to determine the target awakening threshold and target processing strategy according to the energy distribution characteristics and the sound intensity characteristics; The fourth module 504 is used to control the voice device to operate according to the target wake-up threshold and the target processing strategy.
[0067] The specific implementation of the voice device control apparatus is substantially the same as the specific embodiment of the above-mentioned voice device control method, and will not be described in detail here.
[0068] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment.
[0069] Refer to the following Figure 6 hereinafter, an electronic device 600 according to this embodiment of the present disclosure is described. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0070] like Figure 6 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting various system components (including storage unit 620 and processing unit 610), a display unit 640, and the like.
[0071] The storage unit stores program code, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present disclosure described in the above voice device control method section of this specification.
[0072] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0073] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0074] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0075] The electronic device 600 can also communicate with one or more external devices 600' (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 650. Furthermore, the electronic device 600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 via the bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0076] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned voice device control method is implemented.
[0077] The voice device control method, apparatus, equipment and storage medium provided in the embodiments of the present application perform linkage adjustment of the wake-up threshold and processing strategy based on the energy distribution characteristics and sound intensity characteristics of the ambient sound signal of the environment in which the voice device is located, and then control the voice device to operate according to the target wake-up threshold and target processing strategy. The voice device can be controlled to dynamically adjust the wake-up threshold and processing strategy according to the noise parameters of the environment in which it is located, ensuring that the false awakening rate and the false awakening rate are in a relatively balanced state, rather than reducing the false awakening rate by sacrificing the wake-up rate. This ensures that the false awakening rate is at a low level and the wake-up rate is at a high level, so that the voice device is usually in a state of high wake-up performance in the space in which it is located, thereby improving the user experience.
[0078] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above-mentioned method according to the embodiments of the present disclosure.
[0079] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0080] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0081] Those skilled in the art will appreciate that the modules described above can be distributed in the device according to the description of the embodiment, or can be modified accordingly to be used in one or more devices that are different from the embodiment. The modules of the above embodiment can be combined into one module or further divided into multiple submodules.
[0082] While the exemplary embodiments of the present disclosure have been specifically illustrated and described above, it should be understood that the present disclosure is not limited to the detailed structures, configurations, or implementations described herein; rather, the present disclosure is intended to encompass various modifications and equivalent configurations within the spirit and scope of the appended claims.
Claims
1. A method for controlling a voice device, characterized in that: include: Obtain the ambient sound signal of the environment where the voice device is located; Determining energy distribution characteristics and sound intensity characteristics of the ambient sound signal; determining a target arousal threshold and a target processing strategy according to the energy distribution characteristics and the sound intensity characteristics; The voice device is controlled to operate according to the target wake-up threshold and the target processing strategy.
2. The voice device control method according to claim 1, characterized in that: The determining of the energy distribution characteristics and the sound intensity characteristics of the ambient sound signal includes: Performing frequency domain analysis on the ambient sound signal to obtain a frequency domain signal; Determining the energy proportion of the frequency domain signal in each frequency band to obtain the energy distribution characteristics; The sum of the energy of the frequency domain signal in each frequency band is determined to obtain the sound intensity feature.
3. The voice device control method according to claim 1, characterized in that: The determining of a target awakening threshold and a target processing strategy according to the energy distribution characteristics and the sound intensity characteristics includes: When the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency range corresponding to the energy distribution feature is less than the reference energy proportion, the first target awakening threshold and the first target processing strategy are adopted; otherwise, the second target awakening threshold and the second target processing strategy are adopted; The first target wake-up threshold is less than the reference wake-up threshold, and the first target processing strategy is to extract the Mel-frequency cepstral coefficient features in the ambient sound signal. The second target wake-up threshold is greater than the reference wake-up threshold, and the second target processing strategy is to perform noise reduction on the ambient sound signal and then extract the Mel-frequency cepstral coefficient features in the ambient sound signal.
4. The method for controlling a voice device according to claim 3, wherein: The extracting of Mel-frequency cepstral coefficient features from the ambient sound signal includes: Performing pre-emphasis processing on the ambient sound signal to obtain a pre-emphasized sound signal; Performing frame and window processing on the pre-emphasized sound signal to obtain multiple sound segment signals; Frequency domain analysis, Mel filtering, logarithmic operation and discrete cosine transform are performed on the sound clip signal in sequence to obtain Mel frequency cepstral coefficient features in the ambient sound signal.
5. The method for controlling a voice device according to claim 3, wherein: When the sound intensity corresponding to the sound intensity feature is less than the reference sound intensity and the energy proportion in the target frequency segment corresponding to the energy distribution feature is less than the reference energy proportion, the first target wake-up threshold linearly decreases from the reference wake-up threshold, whereas the second target wake-up threshold linearly increases from the reference wake-up threshold.
6. The voice device control method according to claim 1, characterized in that: The controlling the voice device to operate according to the target wake-up threshold and the target processing strategy includes: processing the ambient sound signal into an intermediate sound signal based on the target processing strategy; identifying the intermediate sound signal at the target arousal threshold to obtain a target sound signal; The target sound signal is matched with a reference sound signal. If the match is successful, the voice device is controlled to respond to the target sound signal; the reference sound signal is a sound signal corresponding to the wake-up word.
7. The method for controlling a voice device according to claim 1, wherein: Also includes: Obtain historical arousal thresholds, historical processing strategies, and historical sound signals; According to the historical wake-up threshold, the historical processing strategy and the historical sound signal, a correspondence between the wake-up threshold, the processing strategy and the sound signal is configured.
8. A voice equipment control device, characterized in that: include: The first module is used to obtain the ambient sound signal of the environment where the voice device is located; The second module is used to determine the energy distribution characteristics and sound intensity characteristics of the ambient sound signal; A third module is configured to determine a target arousal threshold and a target processing strategy based on the energy distribution characteristics and the sound intensity characteristics; The fourth module is used to control the voice device to operate according to the target wake-up threshold and the target processing strategy.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the voice device control method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the voice device control method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Noise suppression method and apparatus
CN105931647A
Voice wake-up method and system, storage medium and electronic device
CN110808030A
Awakening method and device of intelligent equipment, equipment and medium
CN111128155A
Voiceprint awakening method, device and equipment and storage medium
CN111223490A
Voice equipment control method and device, equipment and medium
CN119724168A