Voice instruction recognition system and method based on localization processing
In the voice command recognition system, the home appliance noise and voice segment power collected by the microphone are calculated, combined with the distance between the user and the microphone, the weight coefficient is dynamically updated, and the user's movement direction is judged, which solves the problem of waste of computing power caused by real-time position updates, and realizes efficient localized voice command recognition.
Patent Information
- Application Number
- CN202510771887.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, real-time update of user locations requires too high computing power for the system, which is not conducive to localized processing.
By determining the home appliance noise power and voice segment power collected by the microphone, calculating the weight coefficient, and introducing an attenuation factor based on the distance between the user and the microphone, monitoring the decibel change value in real time, judging the user's movement direction, combining the preset functional area division to generate the area change identification parameter K, and building a dynamic update model, which triggers the distance re-computer system when the model output value exceeds the threshold.
It avoids the waste of computing power caused by real-time calculation of user locations, realizes efficient localized voice command recognition, and reduces the computing burden of system operation.
Smart Images

Figure CN120472929A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice command recognition, and more particularly to a voice command recognition system and method based on localization processing. Background Art
[0002] With the rapid development of smart homes, voice control, as a natural and convenient method of interaction, has become an essential component of modern smart home systems. Users interacting with home devices through voice commands not only improves device usability but also enables more efficient control. However, the understanding and recognition of voice commands vary significantly across different regions and cultural backgrounds. Therefore, the application of voice command recognition systems based on localized processing has become particularly important in smart homes.
[0003] In the existing technology, real-time updating of user locations requires too much computing power for system operation, which is not conducive to localized processing. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a voice command recognition system and method based on localization processing to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: The method for voice command recognition based on localization processing includes the following steps: Determine the power of household appliance noise and speech segments collected by the microphone, and calculate the weight coefficient through weighted summation; introduce an attenuation factor based on the distance between the user and the microphone; and update the weight coefficient through weighted summation. Monitor the decibel change value of the voice segment collected by each microphone in real time, calculate the average decibel change and the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value; judge the user's movement direction according to the change trend of the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value, and generate the area change identification parameter K based on the preset functional area division; build a dynamic update model based on the K value and the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value, and trigger the distance recalculation mechanism when the model output value exceeds the threshold.
[0006] In a preferred embodiment, the attenuation factor is based on the distance between the user and the microphone. , calculate the attenuation factor ; Collocation weighted sum formula ; represents the weight coefficient, Represents the updated weight coefficient.
[0007] In a preferred embodiment, the calculation of the household appliance noise power and voice segment power collected by the microphone includes using an A-weighting algorithm to perform frequency domain weighted processing on the noise signal, and separating the noise segment and the voice segment through short-time energy analysis combined with voice activity detection.
[0008] In a preferred embodiment, the distance between the user and the microphone is determined by using microphone array beamforming technology, and the user's spatial coordinates are calculated through time delay estimation and geometric positioning model.
[0009] In a preferred embodiment, the collected original signal needs to be preprocessed before determining the power of household appliance noise and voice segment power collected by the microphone. The preprocessing includes eliminating DC offset, pre-emphasizing high-frequency components, and frame and window processing.
[0010] In a preferred embodiment, the calculation of the average decibel change and the difference between the decibel value change value of the speech segment collected by the average microphone and the average decibel change value first determines the decibel of the speech segment collected by each microphone, counts the decibel change values of the speech segment collected by each microphone and sums and averages them to calculate the average decibel change value; then calculates the difference between the decibel value change value of the speech segment collected by each microphone and the average decibel change value; desymbolizes the difference between the decibel value change value of the speech segment collected by the microphone and the average decibel change value, and then sums and averages them to obtain the difference between the decibel value change value of the speech segment collected by the average microphone and the average decibel change value.
[0011] In a preferred embodiment, the area change identification parameter K is used to divide the user's area into different areas according to functional areas. When the user changes the area, K=1; otherwise, K=0.
[0012] In a preferred embodiment, the update coefficient is calculated based on the weighted sum of K and the difference between the decibel value change value of the speech segment collected by the average microphone and the average decibel change value. When the update coefficient is greater than the system threshold, the distance between each microphone and the user is recalculated.
[0013] In a preferred embodiment, it is determined to recalculate the distance between each microphone and the user, and the distance between each microphone and the user obtained by recalculation is used to update the original distance between the microphone and the user, and the weight coefficient is recalculated.
[0014] In a preferred embodiment, the present invention further discloses a voice command recognition system based on localization processing, wherein the recognition system is based on the above method and includes the following modules: a signal acquisition and joint preprocessing module, a dynamic weight allocation module, and an update module; The signal acquisition and joint preprocessing module is used to collect the original signal and preprocess the original signal; A weighting is used to further analyze the noise; The dynamic weight allocation module is used to calculate the microphone weight coefficient by weighted summation based on the household appliance noise power and voice segment power collected by each microphone. It also introduces an attenuation factor based on the distance between the user and the microphone. The weight coefficient is updated using the weighted summation formula. The update module is used to calculate the update coefficient based on the weighted sum of the difference between the decibel value change value of the speech segment collected by K and the average microphone and the average decibel change value. When the update coefficient is greater than the system threshold, the distance between each microphone and the user is updated and the weight coefficient is redistributed.
[0015] Technical effects and advantages of the present invention: The present invention is based on a voice command recognition method for localized processing. First, the power of household appliance noise and voice segment power collected by a microphone is determined, and a weight coefficient is calculated by weighted summation. An attenuation factor is introduced according to the distance between the user and the microphone. The weight coefficient is updated by combining the weighted summation. The system monitors the decibel change of each microphone's voice segment in real time, calculates the average decibel change and the difference between the decibel change of the voice segment collected by the average microphone and the average decibel change; determines the user's movement direction based on the trend of the difference between the decibel change of the voice segment collected by the average microphone and the average decibel change, and generates the area change identification parameter K based on the preset functional area division; constructs a dynamic update model based on the K value and the difference between the decibel change of the voice segment collected by the average microphone and the average decibel change; and triggers the distance recalculation mechanism when the model output value exceeds the threshold. This avoids the waste of computing power caused by real-time calculation of user location. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 Schematic diagram of the flow of the voice command recognition method based on localization processing of the present invention; Figure 2 It is a structural diagram of the voice command recognition system based on localization processing of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] Example 1: The present invention provides a method for voice command recognition based on localized processing. First, the power of household appliance noise and voice segment power collected by a microphone is determined, and a weight coefficient is calculated by weighted summation. An attenuation factor is introduced based on the distance between the user and the microphone. The weight coefficient is updated by combining the weighted summation. Monitor the decibel change value of the voice segment collected by each microphone in real time, calculate the average decibel change and the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value; judge the user's movement direction according to the change trend of the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value, and generate the area change identification parameter K based on the preset functional area division; build a dynamic update model based on the K value and the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value, and trigger the distance recalculation mechanism when the model output value exceeds the threshold.
[0019] like Figure 1 As shown, the specific steps include: When analyzing mixed audio signals containing noise and speech, power separation and calculation must be achieved through a systematic data processing process. Data acquisition and preprocessing begin with a microphone capturing the raw signal, which consists of segments containing pure appliance noise (such as the sound of a refrigerator or air conditioner running when no one is speaking) and speech segments (periods where the human voice clearly overrides the noise).
[0020] After the acquisition is completed, the original signal is preprocessed to improve the analysis quality, including the following steps: First, the DC offset is eliminated, the formula is: ; x[n] represents the nth sampling point of the original audio signal (unit: volt or dimensionless value); μ represents the arithmetic mean of all sampling points in the signal segment, which is used to eliminate baseline offset; represents the normalized signal after removal of current.
[0021] Then pre-emphasize the high frequency components, formula: ; Where H(z) represents the transfer function of the filter, which is used to enhance the high-frequency components. Represents the unit delay operator, indicating that the signal is delayed by one sampling point. The coefficient 0.97 represents an empirical value, which balances high-frequency enhancement and stability to avoid excessive amplification of noise.
[0022] Finally, frame segmentation and windowing are used to divide the signal into short time frames of 20-30ms, and Hamming window is used to reduce spectrum leakage.
[0023] Then comes the noise and speech separation stage. The core method is based on short-term energy and voice activity detection (VAD). The energy of each frame signal is calculated using the formula: Energy differences are used to distinguish between noise and speech. E represents the energy of a single frame (unit: squared voltage or normalized energy value), and N represents the frame length (number of samples), which is directly related to the temporal resolution. The dynamic threshold method extracts the energy mean of the pure noise segment in the first 1-2 seconds and sets the threshold: ; The frames that exceed the threshold are marked as speech segments, and the rest are noise segments. represents the average frame energy of the first 1-2 seconds of pure noise segment, which serves as the noise baseline; α represents the threshold coefficient (empirical range 1.5-2.5), which controls the sensitivity of speech detection; to improve robustness, frequency domain features (such as spectral entropy or sub-band energy ratio) can be combined to assist in judgment. For example, the spectral entropy of a speech segment is low due to the concentration of high-frequency energy.
[0024] Finally, perform power calculation and verification: calculate the average power of the frame marked as noise: , where M represents the total number of frames marked as noise; Represents the short-term energy of the kth frame; and can be optionally converted to decibel value ;in Represents the reference power. Voice segment power Then take the mean energy of the speech frame. If you need to estimate the pure speech power, you can do subtraction based on the noise stationarity assumption. ;in This represents the estimated power of pure speech and relies on the noise stationarity assumption. Verification requires confirming the accuracy of noise and speech labeling through time-domain waveform inspection (e.g., plotting in MATLAB / Python). Aural verification is also required by playing the separated audio segments to ensure the absence of crosstalk. This method is suitable for scenarios such as smart home noise monitoring and speech enhancement system design. The key is to balance algorithm complexity with separation accuracy, while also relying on the rationality of the noise stationarity assumption.
[0025] Determine the power of household appliance noise and voice segment collected by each microphone, and calculate the microphone weight coefficient through weighted summation; the specific formula is as follows: ;in Represents the microphone weight coefficient, i represents the microphone serial number; zs represents the household appliance noise power. The greater the household appliance noise power, the smaller the microphone weight coefficient, and vice versa; yy represents the voice segment power. The greater the voice segment power, the greater the microphone weight coefficient, and vice versa; a and b are the weight coefficients of household appliance noise power and voice segment power respectively, both greater than zero.
[0026] Specifically, the microphone installation location and appliance usage result in different appliance noise power levels detected by microphones at different locations and times. Noise analysis often requires simulating the human ear's ability to perceive sounds of different frequencies. The human ear's perception varies across frequencies, meaning some frequencies have a greater impact on the ear, while others are less noticeable. To accurately simulate the human ear's hearing characteristics, we use an A-weighting algorithm. Its purpose is to simulate the human ear's perception in various noise environments.
[0027] A-weighting simulates the human ear's sensitivity to different frequencies by weighting the frequency response. The A-weighting curve resembles the human ear's frequency response to sound intensity. Specifically, the human ear is less sensitive to low frequencies (e.g., 20 Hz to 50 Hz) and high frequencies (e.g., above 10 kHz) and more sensitive to mid-range frequencies (approximately 1 kHz to 5 kHz).
[0028] The A-weighting algorithm reflects the actual perception ability of the human ear by weighting sound signals of different frequencies, and ultimately outputs a weighted noise level value.
[0029] The specific steps include: Collecting noise signals: First, you need to collect noise signals using a microphone or other sensor. This signal can be actual noise captured from the environment or a simulated signal.
[0030] Signal frequency division: The collected noise signal is subjected to spectrum analysis to decompose it into multiple frequency bands. Fast Fourier transform (FFT) is usually used to decompose the signal to obtain the amplitude of the signal in different frequency ranges.
[0031] Applying an A-weighted filter: A-weighting is the process of weighting signals of different frequencies using a specific weighting curve. These weighting coefficients are usually a function of frequency. The specific A-weighting curve is as follows: In the low frequency range (20 Hz to 100 Hz), the A-weighting curve is low, indicating that the human ear is less sensitive to sounds at these frequencies.
[0032] The A-weighting curve has a higher gain in the mid-frequency range (500 Hz to 6 kHz), indicating that the human ear is most sensitive to sounds at these frequencies.
[0033] In the high frequency range (6 kHz to 20 kHz), the A-weighting curve drops again, indicating that the human ear is less sensitive to these high frequency sounds.
[0034] These frequency weighting coefficients can be applied using an A-weighting filter, either by looking up a table or by directly applying a standardized A-weighting filter (for example, using the A-weighting filter built into your audio processing software).
[0035] Calculating the weighted sound level: After applying the A-weighting filter, the weighted total sound level needs to be calculated. This is usually done by synthesizing the weighted spectral amplitudes. The commonly used unit is dB (decibel), and the calculation method is as follows: ;in, is the total sound level after weighting; is the weighted sound level of the ith frequency band; N is the number of frequency bands. This calculation formula combines the weighted sound levels of each frequency band into a comprehensive sound level value, which represents the noise level perceived by the human ear.
[0036] The A-weighting algorithm simulates the human ear's perception of different frequencies and weights the noise signal to produce a noise level that is more consistent with actual perception. This weighted noise level allows for a more accurate assessment of the impact of noise, especially in complex environments with non-stationary components such as sudden percussive sounds.
[0037] Microphone array beamforming is used to obtain the distance between the user and each microphone. Specifically, the microphone array layout (such as linear or circular topologies) must first be designed and calibrated to ensure that the coordinates of each microphone are precisely known and the clock is synchronized. The user's sound source signal is then collected synchronously through multiple channels, and preprocessing such as noise reduction, framing, and frequency domain transformation is combined to improve the signal-to-noise ratio. A time delay estimation algorithm such as generalized cross-correlation is then used to calculate the arrival time difference of the signals between each microphone. This is then combined with delayed summation or adaptive beamforming algorithms to enhance the target direction signal and suppress noise and multipath interference. A set of geometric equations is constructed based on the time difference, and the spatial coordinates of the sound source are solved using the least squares method or numerical optimization to ultimately derive the Euclidean distance from the sound source to each microphone.
[0038] Furthermore, the process requires dynamic calibration of sound velocity (such as temperature and humidity compensation), suppression of reverberation effects, and the use of spherical wave model correction in near-field scenarios. Multi-sensor data can be integrated to improve positioning accuracy in complex environments.
[0039] Based on the distance between the user and the microphone , introducing the attenuation factor ; Collocation weighted sum formula ; The closer the distance, the higher the weight.
[0040] Furthermore, the distance between the user and the microphone , does not require real-time updates, and small-scale user movements will not have much impact on voice recognition; when processing locally, constantly calculating the distance between the microphone and the user wastes computing power.
[0041] Determine the decibel level of the speech segment captured by each microphone, calculate the decibel change values for each microphone, and average them. If the average value changes, it indicates a change in the user's speaking decibel level. Calculate the average decibel change value. Calculate the difference between the decibel change values for each microphone and the average decibel change value. If the difference between the decibel change value for a particular microphone and the average decibel change value increases, it indicates the user is moving closer to that microphone, and vice versa. De-symbolize the difference between the decibel change value for each microphone and the average decibel change value, then average and sum them to obtain the difference between the average decibel change value for the speech segment captured by each microphone and the average decibel change value.
[0042] When the distance from the microphone in a certain area is less than a certain distance, it means that the user may have changed the area.
[0043] The user's location is divided into different areas according to functional areas, such as bedroom A, bedroom B, living room, kitchen, and toilet; when the user changes the area, K=1; otherwise K=0; the update coefficient is calculated by weighted summing the difference between K and the change value of the decibel value of the speech segment collected by the average microphone and the average decibel change value, and the difference between K and the change value of the decibel value of the speech segment collected by the average microphone and the average decibel change value is first normalized; then the update coefficient is calculated using the following formula: G=c*cz+d*k; where G represents the update coefficient; cz represents the difference between the change value of the decibel value of the speech segment collected by the average microphone and the average decibel change value, the greater the difference between the change value of the decibel value of the speech segment collected by the average microphone and the average decibel change value, the greater the update coefficient, and vice versa; c and d are the difference between the change value of the decibel value of the speech segment collected by the average microphone and the average decibel change value and the weight coefficient of k.
[0044] When the update coefficient is greater than the system threshold, the microphone array beamforming method is used to calculate the distance between each microphone and the user.
[0045] Example 2: The design of the voice command recognition system based on localization processing of the present invention is based on the method of Example 1, such as Figure 2 As shown, it specifically includes the following modules: signal acquisition and joint preprocessing module, dynamic weight allocation module, and update module; The signal acquisition and joint preprocessing module is used to collect the original signal and preprocess the original signal; A weighting is used to further analyze the noise; The dynamic weight allocation module is used to calculate the microphone weight coefficient by weighted summation based on the household appliance noise power and voice segment power collected by each microphone. It also introduces an attenuation factor based on the distance between the user and the microphone. The weight coefficient is updated using the weighted summation formula. The update module is used to calculate the update coefficient based on the weighted sum of the difference between the decibel value change value of the speech segment collected by K and the average microphone and the average decibel change value. When the update coefficient is greater than the system threshold, the distance between each microphone and the user is updated and the weight coefficient is redistributed.
[0046] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0047] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0048] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0049] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0050] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A voice command recognition method based on localization processing, characterized in that: The following steps are involved: Determine the power of household appliance noise and speech segments collected by the microphone, and calculate the weight coefficient through weighted summation; introduce an attenuation factor based on the distance between the user and the microphone; Update weight coefficients with weighted summation; Monitor the decibel change value of the speech segment collected by each microphone in real time, calculate the average decibel change value and the difference between the decibel change value of the speech segment collected by the average microphone and the average decibel change value; The user's movement direction is determined based on the trend of the difference between the decibel change value of the voice segment collected by the average microphone and the average decibel change value, and the area change identification parameter K is generated in combination with the preset functional area division; A dynamic update model is constructed based on the K value and the difference between the decibel change value of the speech segment collected by the average microphone and the average decibel change value. When the model output value exceeds the threshold, the distance recalculation mechanism is triggered.
2. The method for voice command recognition based on localization processing according to claim 1, characterized in that: The attenuation factor is based on the distance between the user and the microphone. , calculate the attenuation factor ; Collocation weighted sum formula ; represents the weight coefficient, Represents the updated weight coefficient.
3. The method for voice command recognition based on localization processing according to claim 1, characterized in that: The calculation of the power of household appliance noise and voice segment power collected by the microphone includes using an A-weighting algorithm to perform frequency domain weighted processing on the noise signal, and separating the noise segment and the voice segment through short-time energy analysis combined with voice activity detection.
4. The method for voice command recognition based on localization processing according to claim 1, characterized in that: The distance between the user and the microphone is obtained by adopting microphone array beamforming technology, solving the user's spatial coordinates through time delay estimation and geometric positioning model.
5. The method for voice command recognition based on localization processing according to claim 1, characterized in that: Before determining the power of household appliance noise and the power of the voice segment collected by the microphone, the collected original signal is preprocessed, and the preprocessing includes eliminating DC offset, pre-emphasizing high-frequency components, and framing and windowing processing.
6. The method for voice command recognition based on localization processing according to claim 1, characterized in that: The calculation of the average decibel change and the difference between the decibel change value of the speech segment collected by the average microphone and the average decibel change value first determines the decibel of the speech segment collected by each microphone, counts the decibel change values of the speech segment collected by each microphone and averages them to calculate the average decibel change value; then calculates the difference between the decibel change value of the speech segment collected by each microphone and the average decibel change value; The difference between the decibel value change value of the speech segment collected by the microphone and the average decibel change value is de-signed, and then summed and averaged to obtain the difference between the decibel value change value of the speech segment collected by the average microphone and the average decibel change value.
7. The method for voice command recognition based on localization processing according to claim 1, characterized in that: The area change identification parameter K, the user's range is divided into different areas according to the functional area. When the user changes the area, K=1; otherwise K=0.
8. The method for voice command recognition based on localization processing according to claim 7, characterized in that: The update coefficient is calculated based on the weighted sum of K and the difference between the decibel change value of the speech segment collected by the average microphone and the average decibel change value. When the update coefficient is greater than the system threshold, the distance between each microphone and the user is recalculated.
9. The method for voice command recognition based on localization processing according to claim 8, characterized in that: Determine to recalculate the distance between each microphone and the user, update the original distance between the microphone and the user with the recalculated distance between each microphone and the user, and recalculate the weight coefficient.
10. A voice command recognition system based on localization processing, characterized in that: The identification system is based on the method according to any one of claims 1 to 9, and comprises the following modules: a signal acquisition and joint preprocessing module, a dynamic weight allocation module, and an update module; The signal acquisition and joint preprocessing module is used to collect the original signal and preprocess the original signal; A weighting is used to further analyze the noise; The dynamic weight allocation module is used to calculate the microphone weight coefficient by weighted summation based on the household appliance noise power and voice segment power collected by each microphone, and introduce an attenuation factor based on the distance between the user and the microphone; Update the weight coefficient using the weighted sum formula; The update module is used to calculate the update coefficient based on the weighted sum of the difference between the decibel value change value of the speech segment collected by K and the average microphone and the average decibel change value. When the update coefficient is greater than the system threshold, the distance between each microphone and the user is updated and the weight coefficient is redistributed.