Intelligent electric appliance voice recognition control method and system
By performing noise reduction processing and dialect network model recognition on the voice data of smart appliances, the problem of poor speech recognition is solved, and the control accuracy and user experience of smart appliances are improved.
Patent Information
- Application Number
- CN202510746957.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the robustness of speech recognition is poor, and is affected by factors such as environmental noise, accents of different users and mood fluctuations, resulting in inaccurate control of smart appliances and reducing user experience.
The voice data is obtained through the sound pickup module of the smart appliance for noise reduction processing, and the noise reduction is reduced by using the Meer frequency cepspectral coefficient, wavelet energy distribution and noise state transition probability. The user's intention is identified in combination with the dialect network model to ensure that the speech recognition results comply with the semantic rules and then execute control instructions.
It improves the voice control efficiency of smart appliances, improves user experience, and solves the problem of inaccurate recognition caused by user differentiated voice and noise.
Smart Images

Figure CN120472900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech recognition control method and system for intelligent electrical appliances. Background Art
[0002] With the rapid development of modern technology and the widespread application of smart homes and smart devices, people have gained a deeper understanding of emerging technologies such as artificial intelligence, big data, and smart cities, and are beginning to pursue smarter and more convenient lifestyles. Speech recognition is a core technology for smart home control, enhancing the intelligence of homes through accurate and efficient speech recognition. As the core of smart home control systems, the performance of model data processing directly affects the overall experience of smart home systems.
[0003] In actual applications, speech recognition technology is challenged by various changing conditions, such as environmental noise, different accents of different users, and pronunciation changes caused by emotional fluctuations of the speaker. However, the speech recognition robustness of related technologies is very poor. These factors will affect the accuracy of speech recognition, and in turn affect the control of smart appliances, reducing the user experience of smart appliances. Summary of the Invention
[0004] The present invention aims to provide a method and system for voice recognition control of intelligent electrical appliances to address the deficiencies in the prior art. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0005] An embodiment of the present invention provides a method for voice recognition control of an intelligent electrical appliance, the method comprising: Acquire the voice data input by the user at the current time through the sound pickup module of the smart appliance, and perform noise reduction processing on the voice data; Determine the corresponding initial speech recognition result and speech feature data based on the speech data after noise reduction processing; Determining the dialect network model corresponding to the user through the initial speech recognition result and the speech feature data; Inputting the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result; Determining whether the semantics of the target speech recognition result conforms to semantic rules; If the semantics of the target speech recognition result conforms to the semantic rules, the control instruction corresponding to the target text is used to control the execution of the intelligent appliance.
[0006] In an optional embodiment, the performing noise reduction processing on the voice data includes: Acquire historical voice data of a historical time before the current time from a database of the smart appliance; Determining Mel-frequency cepstral coefficients, wavelet energy distribution data, and noise state transition probability corresponding to the historical speech data; Noise reduction processing is performed on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability.
[0007] In an optional embodiment, the performing noise reduction processing on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability includes: Inputting the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability into a noise prediction model to obtain a predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time; Noise reduction processing is performed on the speech data using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time.
[0008] In an optional embodiment, performing noise reduction processing on the speech data using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time includes: Generate an inverted waveform according to the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time: Calculating the sound pressure of the superimposed signals of the speakers in the smart appliance according to the reverse waveform; The voice data is subjected to noise reduction processing by using the sound pressure of the reverse waveform and the superposition of the loudspeaker signals in the smart appliance.
[0009] In an optional embodiment, determining the corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing includes: Performing speech recognition on the noise-reduced speech data to obtain the initial speech recognition result; The speech data after noise reduction processing is input into a speech feature extraction model to obtain the speech feature data.
[0010] In an optional embodiment, determining the dialect network model corresponding to the user based on the initial speech recognition result and the speech feature data includes: Performing text semantic analysis on the initial speech recognition result to obtain a corresponding revised speech recognition result; Comparing the initial speech recognition result and the revised speech recognition result to determine semantically different text characters and their positions in the initial speech recognition result; The speech feature data determines the dialect network model corresponding to the user through the semantically different text characters and their positions in the initial speech recognition results.
[0011] In an optional embodiment, determining the dialect network model corresponding to the user using the semantically different text characters and their positions in the initial speech recognition results and the speech feature data includes: extracting target feature data from the speech feature data according to the position of the semantically different text character in the initial speech recognition result; Determine a first feature vector based on the semantically different text characters and the corresponding target feature data; determine a second feature vector based on the speech feature data; Inputting the first feature vector and the second feature vector into a dialect recognition model to obtain a dialect region prediction label corresponding to the user; The dialect network model corresponding to the user is determined according to the dialect region prediction label.
[0012] In an optional embodiment, inputting the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result includes: Dividing the noise-reduced speech data into a plurality of speech segments according to pause time; Extracting speech segment feature data corresponding to each speech segment; Combining the speech segment feature data into a speech segment feature vector in chronological order; The speech segment feature vector is input into the dialect network model to obtain the corresponding target speech recognition result.
[0013] In an optional embodiment, the step of forming a speech segment feature vector from the speech segment feature data in chronological order includes: Analyze the semantic association relationship of each speech segment and determine the association weight of each speech segment with other speech segments; The speech segment feature data and their corresponding associated weights are combined into a speech segment feature vector in chronological order.
[0014] An embodiment of the present invention provides a voice recognition control system for intelligent electrical appliances, the system comprising: A noise reduction module, configured to obtain voice data input by the user at the current time through the sound pickup module of the smart appliance and perform noise reduction processing on the voice data; A determination module, configured to determine the corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing; The determination module is further configured to determine the dialect network model corresponding to the user based on the initial speech recognition result and the speech feature data; A prediction module, configured to input the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result; The determination module is further configured to determine whether the semantics of the target speech recognition result conforms to semantic rules; The control module is used to control the execution of the intelligent appliance using the control instruction corresponding to the target text if the semantics of the target speech recognition result conforms to the semantic rules.
[0015] The embodiments of the present invention include the following advantages: The embodiment of the present invention provides a method and system for controlling speech recognition of intelligent electrical appliances. First, a sound pickup module of the intelligent electrical appliance obtains speech data input by a user at the current time and performs noise reduction processing on the speech data; then, an initial speech recognition result and speech feature data corresponding to the speech data after noise reduction processing are determined; then, a dialect network model corresponding to the user is determined based on the initial speech recognition result and the speech feature data; the speech data after noise reduction processing are input into the dialect network model to obtain a corresponding target speech recognition result; it is determined whether the semantics of the target speech recognition result conform to the semantic rules; if the semantics of the target speech recognition result conform to the semantic rules, the control instruction corresponding to the target text is used to control the execution of the intelligent electrical appliance. Because the present application needs to perform noise reduction processing on the speech data input by the user before speech recognition, and determines the corresponding dialect network model based on the noise reduction processing speech data, and then determines the target speech recognition result of the user based on the dialect network model, the present application can solve the problem of inaccurate speech recognition caused by the user's differentiated speech and noise, and can improve the speech control efficiency of the intelligent electrical appliance and enhance the user experience of the intelligent electrical appliance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a method for voice recognition control of an intelligent electrical appliance provided by an embodiment of the present invention; Figure 2 This is a flow chart of a voice data noise reduction process provided by an embodiment of the present invention; Figure 3 The present invention provides a schematic diagram of a voice recognition control system for intelligent electrical appliances. DETAILED DESCRIPTION
[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0018] See also Figure 1, is a method for controlling a voice recognition of an intelligent electrical appliance provided by an embodiment of the present invention, the method specifically comprising S101-S106: S101, obtaining voice data input by a user at a current time through a sound pickup module of the smart appliance, and performing noise reduction processing on the voice data.
[0019] Among them, the smart appliances can be smart home appliances such as sweeping robots, air conditioners, speakers, refrigerators, etc., or they can be smart cars, smart delivery robots, etc. This embodiment does not specifically limit the smart appliances.
[0020] In this embodiment, multiple sound pickup modules can be set inside the smart appliance. For example, the sound pickup modules are set in different directions of the smart appliance. The input voice data can be obtained through the set sound pickup modules, and the voice data obtained by the multiple sound pickup modules are compared. Then, the voice data with the strongest user voice prominence is selected and noise reduction processing is performed on it, so that voice recognition can be performed through the noise-reduced voice data in subsequent steps.
[0021] It's important to note that background noise in a home environment (such as the sound of running appliances, human voices, and echoes) can significantly reduce the signal-to-noise ratio of voice signals, resulting in a 2910% drop in recognition rate. For example, in a multi-person conversation, the device may be unable to distinguish user commands from background noise. Or when a user issues a command from a distance (e.g., 5 meters away), the voice signal is severely attenuated, and room reverberation and echoes further distort the acoustic characteristics, affecting recognition accuracy. To address these issues, this embodiment requires noise reduction processing after obtaining the user's input voice data.
[0022] like Figure 2 As shown, in an optional embodiment provided by the present application, the noise reduction processing on the voice data includes: S1011 , obtaining historical voice data of a historical time before the current time from a database of the smart appliance.
[0023] The database of the smart appliance stores historical user voice data. For example, if a user issues a voice control command to the smart appliance at 15:02, "adjust to 24 degrees Celsius," the historical voice data retrieved from the database at 15:00 is "adjust the air conditioner to 26 degrees Celsius."
[0024] S1012, determining the Mel-frequency cepstral coefficients, wavelet energy distribution data, and noise state transition probability corresponding to the historical speech data.
[0025] Specifically, the historical voice data can first be high-pass filtered (cutoff frequency 50Hz) and normalized to eliminate DC offset and device noise. The normalization can be calculated using the following formula: , in, μ is the sliding window mean, The standard deviation is obtained by dividing the historical speech data into frames (20ms frame length, 10ms frame shift) and performing a short-time Fourier transform (STFT) on the data to extract time-frequency features such as the spectrogram and Mel-frequency cepstral coefficients (13-dimensional coefficients + dynamic difference (39 dimensions in total)). A 5-layer decomposition is performed using the db4 wavelet basis to extract the energy percentage of the frequency band below 1kHz.
[0026] In this embodiment, noise is divided into three states (Markov chain discretization), and the noise transition probabilities are then determined based on these three states. Specifically, this embodiment divides the noise into discrete states (e.g., S1: steady-state low energy, S2: transient medium energy, and S3: impact high energy) to construct the noise state transition probabilities. S1 (steady-state low energy) can be the sound of an air conditioner running continuously; S2 (transient medium energy) can be the short ringing of a doorbell; and S3 (impact high energy) can be the sound of glass breaking.
[0027] For example, the noise state transition probability based on historical data statistics can be shown in Table 1 below: Table 1 , S1013: Perform noise reduction processing on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability.
[0028] Specifically, the denoising process is performed on the speech data based on the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability, including: inputting the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability into a noise prediction model to obtain a predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time; and denoising the speech data using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time.
[0029] In this embodiment, the MFCC (39 dimensions) + wavelet energy distribution (5 dimensions) + Markov state probability (3 dimensions) of the current frame → 47-dimensional vector is input into the noise prediction model. The bidirectional LSTM (256 units) in the noise prediction model captures the temporal dependency, and the temporal convolutional network (TCN) extracts the local pattern to predict the noise waveform parameters within the next 50ms (i.e., the current time predicted by historical speech data): predicted phase, predicted amplitude, and predicted fundamental frequency.
[0030] In this embodiment, the noise prediction model is based on the calculation formula of its loss function during the training process as follows: , in, To control the loss of prediction phase ( ), To control the loss of predicted amplitude ( ), To control the predicted fundamental frequency loss ( ), To control the loss of resistance ( ) contribution weight. Each contribution weight is used to balance the optimization intensity of multiple tasks to avoid a single task dominating the training process. For example, when the input speech signal-to-noise ratio (SNR) is <10dB, increase To enhance the noise resistance of amplitude prediction.
[0031] The formula for calculating the predicted phase loss is: ,in To predict the phase, The present embodiment uses the cosine distance to measure the periodic difference between the predicted phase and the true phase, which can avoid the 2-dimensional error caused by directly using the mean square error. Phase wrapping problem.
[0032] The formula for calculating the predicted amplitude loss is: ,in To predict the amplitude, This embodiment uses the logarithmic domain L1 loss, focusing more on relative energy differences rather than absolute amplitudes, thereby improving the modeling capability of low-energy speech segments (such as voiceless consonants).
[0033] The calculation formula for predicting fundamental frequency loss is: , where K is the number of Gaussian mixture components (i.e. the number of possible distribution modes of the fundamental frequency), For the The mixed weight of the Gaussian components reflects the probability of the distribution contributing to the current fundamental frequency value. is the true value of the fundamental frequency, For the The center frequency value of the Gaussian component, For the The frequency fluctuation range of the Gaussian component.
[0034] More specifically, the noise reduction processing of the voice data using the predicted phase, predicted amplitude and predicted fundamental frequency corresponding to the current time includes: generating an inverted waveform according to the predicted phase, predicted amplitude and predicted fundamental frequency corresponding to the current time; calculating the sound pressure of the superimposed speaker signals in the smart appliance according to the inverted waveform; and performing noise reduction processing on the voice data using the inverted waveform and the sound pressure of the superimposed speaker signals in the smart appliance.
[0035] In this embodiment, the reverse waveform can be generated by the following formula: , in, To predict the amplitude, is the dynamic attenuation coefficient (0.9-1.1), To predict the fundamental frequency, For the predicted phase.
[0036] The sound pressure is calculated using the following formula: , Where N is the number of speakers, For the The amplitude weights of the speakers, It is a reverse waveform. For the The delay time applied to the signal of each speaker. Different delay times are applied to the signal of each speaker so that the sound waves emitted by multiple speakers can achieve coherent superposition at the target location (such as the user's location) and form destructive interference in other areas. Specifically, The calculation formula is as follows: , in, For the The straight-line distance between the loudspeaker and the target position, is the incident angle of the noise or target sound source, c is the speed of sound (343m / s at room temperature), A dynamic correction term used to compensate for environmental reflections and hardware delays (such as speaker response time).
[0037] Amplitude Weight The calculation formula is as follows: , in, is the average distance from all speakers to the target position, is the attenuation coefficient (usually 0.5 to 1.0), which is used to control the contribution weight of the edge speakers. The center speaker has the largest weight, and the edges gradually attenuate, forming a beam with controllable main lobe width.
[0038] In this embodiment, phase inversion and amplitude matching can be used to make the anti-phase sound wave and the original noise superimposed in the time domain to form destructive interference. According to the predicted amplitude A(t), the amplitude of the anti-phase sound wave is dynamically adjusted. ( is the dynamic attenuation coefficient), which compensates for the nonlinear distortion of the loudspeaker.
[0039] Multi-speaker beamforming concentrates the energy of inverted sound waves in the user's area (such as the sofa), avoiding energy waste from whole-house noise reduction. Based on the predicted direction of the noise source (such as the air conditioner on the left), the beamforming parameters are adjusted to focus the inverted sound wave energy on the user's area. The speakers emit inverted sound waves with a phase opposite to the noise, creating a quiet zone in the target area (such as around the user's head) and suppressing low-frequency noise (such as the air conditioner). If the predicted noise is primarily low-frequency (such as the hum of a refrigerator), the active noise cancellation (ANC) is enhanced; if the noise is predominantly high-frequency (such as keyboard tapping), the passive sound-absorbing materials are activated to assist in noise reduction.
[0040] S102: Determine the corresponding initial speech recognition result and speech feature data based on the speech data that has undergone noise reduction processing.
[0041] In an optional implementation provided by the present application, determining the corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing includes: S1021: Perform speech recognition on the noise-reduced speech data to obtain the initial speech recognition result.
[0042] In this embodiment, the speech data that has undergone noise reduction processing may be recognized using speech recognition technology to obtain an initial speech recognition result.
[0043] S1022: Input the noise-reduced speech data into a speech feature extraction model to obtain the speech feature data.
[0044] The speech feature data may include Mel-frequency cepstral coefficients, fundamental frequency extracted by the YIN algorithm, the first three formants (F1-F3) extracted, etc., which are not specifically limited in this embodiment.
[0045] S103: Determine the dialect network model corresponding to the user through the initial speech recognition result and the speech feature data.
[0046] It should be noted that due to a user's accent, speaking speed, dialect (such as Sichuanese or Northeastern dialect), and personalized vocabulary (such as nicknames and professional terms), standard speech recognition technology (initial speech recognition results) or speech recognition models may not accurately identify the user's expressed intent. Therefore, this embodiment also requires determining the user's corresponding dialect network model based on the speech feature data and the initial speech recognition results, so as to further determine the user's true intent based on this dialect network model.
[0047] In an optional implementation provided by the present application, determining the dialect network model corresponding to the user by using the initial speech recognition result and the speech feature data includes: S1031: Perform text semantic analysis on the initial speech recognition result to obtain a corresponding modified speech recognition result.
[0048] In order to solve the problem that the initial speech recognition result may be erroneous due to the presence of dialect or too fast speaking speed in the voice data input by the user, this embodiment can perform semantic analysis based on the context of the input voice data to correct the initial speech recognition result and obtain a corrected speech recognition result.
[0049] S1032 , comparing the initial speech recognition result and the revised speech recognition result to determine semantically different text characters and their positions in the initial speech recognition result.
[0050] Semantically different text characters are those that differ between the initial and revised speech recognition results. For example, the initial speech recognition result is "The weather is too hot, keep the temperature up," while the revised result is "The weather is too hot, lower the temperature." By comparing the two recognition results, we can determine that the different text characters are "maintain," and their corresponding positions are 6 and 7 (excluding punctuation).
[0051] S1033, determining the dialect network model corresponding to the user through the semantically different text characters and their positions in the initial speech recognition result, and the speech feature data.
[0052] Specifically, the method of determining the dialect network model corresponding to the user through the speech feature data using the semantic difference text characters and their positions in the initial speech recognition results includes: extracting target feature data from the speech feature data using the positions of the semantic difference text characters in the initial speech recognition results; determining a first feature vector based on the semantic difference text characters and the corresponding target feature data; determining a second feature vector based on the speech feature data; inputting the first feature vector and the second feature vector into a dialect recognition model to obtain a dialect area prediction label corresponding to the user; and determining the dialect network model corresponding to the user based on the dialect area prediction label.
[0053] The dialect recognition model is trained based on sample data and the corresponding true regional labels for the dialect. This sample data consists of speech samples from major Chinese dialect regions (such as Northern Mandarin, Cantonese, Wu, Minnan, and Hakka), with at least 500 hours of valid speech data for each dialect. The true regional labels for the dialects need to be accurately annotated to the dialect area, such as the Cantonese-Guangfu dialect and the Minnan-Quanzhou-Zhangzhou dialect. The sample data is then used to determine the first and second eigenvectors according to the above method. These first and second eigenvectors are then input into the dialect recognition model to obtain the predicted regional labels for the dialects. The loss value of the dialect regional loss model is then calculated based on the predicted and true regional labels for the dialects. When this loss value is less than a preset value, the training of the dialect recognition model is complete.
[0054] S104: Input the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result.
[0055] In an optional implementation provided in the present application, the inputting the noise-reduced speech data into the dialect network model to obtain the corresponding target speech recognition result includes: dividing the noise-reduced speech data into multiple speech segments according to pause time; extracting speech segment feature data corresponding to each speech segment; organizing the speech segment feature data into a speech segment feature vector in chronological order; and inputting the speech segment feature vector into the dialect network model to obtain the corresponding target speech recognition result.
[0056] Specifically, the method of forming the speech segment feature data into a speech segment feature vector in chronological order includes: analyzing the semantic association relationship of each speech segment to determine the association weight of each speech segment with other speech segments; and forming the speech segment feature data and its corresponding association weight into a speech segment feature vector in chronological order.
[0057] S105: Determine whether the semantics of the target speech recognition result conforms to semantic rules.
[0058] In this embodiment, the voice data provided by the user may not be very accurate or the expression may not be clear enough. If a response is directly made based on the target voice recognition result, it may not be the response the user expects. This embodiment first determines whether the target voice recognition result meets the preset standards after obtaining the target voice recognition result, and then makes different responses based on different judgment results, thereby increasing the probability of giving the response the user expects.
[0059] For example, the target speech recognition result is "Play Half Mountain and Half Sea sung by Su Xing". Analysis determines that the song "Half Mountain and Half Sea" is sung by Wang Zhengliang, so it does not conform to the semantic rules. At this time, this embodiment can revise the target speech recognition result to "Play Half Mountain and Half Sea sung by Wang Zhengliang", so as to realize the execution of playing Half Mountain and Half Sea sung by Su Xing on the smart appliance.
[0060] S106: If the semantics of the target speech recognition result conforms to the semantic rules, the control instruction corresponding to the target text is used to control the execution of the intelligent appliance.
[0061] The embodiment provides a method for controlling speech recognition of smart appliances. First, the method obtains speech data input by a user at the current time through a sound pickup module of the smart appliance and performs noise reduction processing on the speech data; then determines the corresponding initial speech recognition result and speech feature data based on the noise-reduced speech data; then determines the dialect network model corresponding to the user based on the initial speech recognition result and the speech feature data; inputs the noise-reduced speech data into the dialect network model to obtain the corresponding target speech recognition result; determines whether the semantics of the target speech recognition result conforms to the semantic rules; if the semantics of the target speech recognition result conforms to the semantic rules, controls the execution of the smart appliance with the control instruction corresponding to the target text. Because the present application needs to perform noise reduction processing on the speech data input by the user before speech recognition, and determines the corresponding dialect network model based on the noise-reduced speech data, and then determines the target speech recognition result of the user based on the dialect network model, the present application can solve the problem of inaccurate speech recognition caused by the user's differentiated speech and noise, and can improve the speech control efficiency of smart appliances and enhance the user experience.
[0062] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0063] In one embodiment, a voice recognition control system for intelligent electrical appliances is provided. Figure 3As shown, the functional modules of the intelligent appliance voice recognition control system are described in detail as follows: The noise reduction module 31 is used to obtain the voice data input by the user at the current time through the sound pickup module of the smart appliance and perform noise reduction processing on the voice data; A determination module 32 is configured to determine a corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing; The determining module 32 is further configured to determine the dialect network model corresponding to the user based on the initial speech recognition result and the speech feature data; A prediction module 33 is configured to input the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result; The determination module 32 is further configured to determine whether the semantics of the target speech recognition result conforms to semantic rules; The control module 34 is configured to control the execution of the intelligent appliance according to the control instruction corresponding to the target text if the semantics of the target speech recognition result conforms to the semantic rules.
[0064] In an optional embodiment, the noise reduction module 31 is specifically configured to: Acquire historical voice data of a historical time before the current time from a database of the smart appliance; Determining Mel-frequency cepstral coefficients, wavelet energy distribution data, and noise state transition probability corresponding to the historical speech data; Noise reduction processing is performed on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability.
[0065] In an optional embodiment, the noise reduction module 31 is specifically configured to: Inputting the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability into a noise prediction model to obtain a predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time; Noise reduction processing is performed on the speech data using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time.
[0066] In an optional embodiment, the noise reduction module 31 is specifically configured to: Generate an inverted waveform according to the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time: Calculating the sound pressure of the superimposed signals of the speakers in the smart appliance according to the reverse waveform; The voice data is subjected to noise reduction processing by using the sound pressure of the reverse waveform and the superposition of the loudspeaker signals in the smart appliance.
[0067] In an optional embodiment, the determination module 32 is specifically configured to: Performing speech recognition on the noise-reduced speech data to obtain the initial speech recognition result; The speech data after noise reduction processing is input into a speech feature extraction model to obtain the speech feature data.
[0068] In an optional embodiment, the determination module 32 is specifically configured to: Performing text semantic analysis on the initial speech recognition result to obtain a corresponding revised speech recognition result; Comparing the initial speech recognition result and the revised speech recognition result to determine semantically different text characters and their positions in the initial speech recognition result; The speech feature data determines the dialect network model corresponding to the user through the semantically different text characters and their positions in the initial speech recognition results.
[0069] In an optional embodiment, the determination module 32 is specifically configured to: extracting target feature data from the speech feature data according to the position of the semantically different text character in the initial speech recognition result; Determine a first feature vector based on the semantically different text characters and the corresponding target feature data; determine a second feature vector based on the speech feature data; Inputting the first feature vector and the second feature vector into a dialect recognition model to obtain a dialect region prediction label corresponding to the user; The dialect network model corresponding to the user is determined according to the dialect region prediction label.
[0070] In an optional embodiment, the prediction module 33 is specifically configured to: Dividing the noise-reduced speech data into a plurality of speech segments according to pause time; Extracting speech segment feature data corresponding to each speech segment; Combining the speech segment feature data into a speech segment feature vector in chronological order; The speech segment feature vector is input into the dialect network model to obtain the corresponding target speech recognition result.
[0071] In an optional embodiment, the prediction module 33 is specifically configured to: Analyze the semantic association relationship of each speech segment and determine the association weight of each speech segment with other speech segments; The speech segment feature data and their corresponding associated weights are combined into a speech segment feature vector in chronological order.
[0072] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs.
[0073] For the specific definition of the intelligent appliance voice recognition control system, please refer to the definition of the intelligent appliance voice recognition control method above and will not be repeated here. Each module in the above-mentioned device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0074] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for controlling voice recognition of intelligent electrical appliances, characterized in that: The method comprises: Acquire the voice data input by the user at the current time through the sound pickup module of the smart appliance, and perform noise reduction processing on the voice data; Determine the corresponding initial speech recognition result and speech feature data based on the speech data after noise reduction processing; Determining the dialect network model corresponding to the user through the initial speech recognition result and the speech feature data; Inputting the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result; Determining whether the semantics of the target speech recognition result conforms to semantic rules; If the semantics of the target speech recognition result conforms to the semantic rules, the control instruction corresponding to the target text is used to control the execution of the intelligent appliance.
2. The method according to claim 1, characterized in that The performing noise reduction processing on the voice data includes: Acquire historical voice data of a historical time before the current time from a database of the smart appliance; Determining Mel-frequency cepstral coefficients, wavelet energy distribution data, and noise state transition probability corresponding to the historical speech data; Noise reduction processing is performed on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability.
3. The method according to claim 2, characterized in that The performing noise reduction processing on the speech data according to the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability includes: Inputting the Mel-frequency cepstral coefficients, the wavelet energy distribution data, and the noise state transition probability into a noise prediction model to obtain a predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time; Noise reduction processing is performed on the speech data using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time.
4. The method according to claim 3, characterized in that The performing noise reduction processing on the speech data by using the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time includes: Generate an inverted waveform according to the predicted phase, predicted amplitude, and predicted fundamental frequency corresponding to the current time: Calculating the sound pressure of the superimposed signals of the speakers in the smart appliance according to the reverse waveform; The voice data is subjected to noise reduction processing by using the sound pressure of the reverse waveform and the superposition of the loudspeaker signals in the smart appliance.
5. The method according to claim 1, wherein The determining of the corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing includes: Performing speech recognition on the noise-reduced speech data to obtain the initial speech recognition result; The speech data after noise reduction processing is input into a speech feature extraction model to obtain the speech feature data.
6. The method according to claim 5, characterized in that The determining of the dialect network model corresponding to the user by using the initial speech recognition result and the speech feature data includes: Performing text semantic analysis on the initial speech recognition result to obtain a corresponding revised speech recognition result; Comparing the initial speech recognition result and the revised speech recognition result to determine semantically different text characters and their positions in the initial speech recognition result; The speech feature data determines the dialect network model corresponding to the user through the semantically different text characters and their positions in the initial speech recognition results.
7. The method according to claim 6, characterized in that The determining of the dialect network model corresponding to the user by using the semantically different text characters and their positions in the initial speech recognition results and the speech feature data includes: extracting target feature data from the speech feature data according to the position of the semantically different text character in the initial speech recognition result; Determine a first feature vector based on the semantically different text characters and the corresponding target feature data; determine a second feature vector based on the speech feature data; Inputting the first feature vector and the second feature vector into a dialect recognition model to obtain a dialect region prediction label corresponding to the user; The dialect network model corresponding to the user is determined according to the dialect region prediction label.
8. The method according to any one of claims 1 to 7, characterized in that The step of inputting the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result includes: Dividing the noise-reduced speech data into a plurality of speech segments according to pause time; Extracting speech segment feature data corresponding to each speech segment; Combining the speech segment feature data into a speech segment feature vector in chronological order; The speech segment feature vector is input into the dialect network model to obtain the corresponding target speech recognition result.
9. The method according to claim 8, characterized in that The step of forming a speech segment feature vector from the speech segment feature data in chronological order includes: Analyze the semantic association relationship of each speech segment and determine the association weight of each speech segment with other speech segments; The speech segment feature data and their corresponding associated weights are combined into a speech segment feature vector in chronological order.
10. A voice recognition control system for intelligent electrical appliances, characterized in that: The system comprises: A noise reduction module, configured to obtain voice data input by the user at the current time through the sound pickup module of the smart appliance and perform noise reduction processing on the voice data; A determination module, configured to determine the corresponding initial speech recognition result and speech feature data based on the speech data subjected to noise reduction processing; The determination module is further configured to determine the dialect network model corresponding to the user based on the initial speech recognition result and the speech feature data; A prediction module, configured to input the noise-reduced speech data into the dialect network model to obtain a corresponding target speech recognition result; The determination module is further configured to determine whether the semantics of the target speech recognition result conforms to semantic rules; The control module is used to control the execution of the intelligent appliance using the control instruction corresponding to the target text if the semantics of the target speech recognition result conforms to the semantic rules.
Citation Information
Cited By
Voice interaction graphene AI intelligent tea table control method
CN121963725A