A method and system for switching between multiple intelligent controls for smart doors and windows
By accurately extracting the fundamental frequency and using voiceprint matching technology, the problem of voiceprint matching errors in existing speech recognition technologies has been solved, enabling secure and accurate control of smart doors and windows and enhancing the reliability of user authentication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-04-03
AI Technical Summary
Existing speech recognition technologies rely on short-time Fourier transform and Mel spectrum feature extraction, which can lead to voiceprint matching errors, resulting in false user identity verification and posing security risks.
By acquiring mixed sound signals and template sound signals, decomposing and segmenting signal frames, calculating the energy kurtosis and fundamental frequency of harmonic and subharmonic sets, filtering template sound signals of the same gender, and calculating the degree of voiceprint matching, precise voice control can be achieved.
It improves the security and accuracy of voice control, reduces misidentification issues, enhances user privacy protection, prevents unauthorized access, and provides a safer and more convenient smart home environment.
Smart Images

Figure CN121075329B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice recognition technology. In particular, it relates to a method and system for switching between multiple intelligent controls for smart doors and windows. Background Technology
[0002] Smart home systems are gradually becoming an important part of modern families. By integrating various smart devices, such as smart doors and windows, smart lighting, and smart appliances, smart homes provide users with a convenient, efficient, and comfortable living environment. Among these, smart doors and windows, as a crucial component of smart homes, not only enable remote control but also allow for more natural and convenient operation through voice recognition technology.
[0003] Speech recognition distinguishes different users by analyzing the significant differences in their voice signals at the fundamental frequency and formants. However, existing speech recognition technologies typically rely on Short-Time Fourier Transform (STFT) and Mel-frequency spectra to extract time-frequency features, and then use Mel-frequency cepstral coefficients (MFCC) and x-vectors for voiceprint matching and authentication. While these technologies can extract speech features, their reliance on all frequencies of the Mel-frequency spectrum means that the differences at the fundamental frequency and formants are smoothed out by redundant frequency voiceprint features. This can lead to incorrect voiceprint matching, resulting in false user authentication and posing security risks. Summary of the Invention
[0004] To address the problem that existing speech recognition technologies rely on Short Time Fourier Transform (STFT) and Mel-frequency spectrum feature extraction, which are prone to voiceprint matching errors due to the smoothing of fundamental frequency and formant differences in redundant frequency voiceprint features, leading to false user identity verification and security risks, this invention provides solutions in the following aspects.
[0005] In the first aspect, a method for switching between multiple intelligent controls for smart doors and windows includes: acquiring a mixed sound signal of the scene to be controlled and template sound signals of different commands input by the user; decomposing and segmenting the mixed sound signal to obtain multiple source signals and source signal frames; determining the harmonic set and subharmonic set of each frequency as the target frequency in each source signal frame, and calculating the energy kurtosis and fundamental frequency of the frequencies in each set; selecting the fundamental frequency of each source signal by combining the fundamental frequencies of each frequency in different source signal frames; calculating the voiceprint matching degree based on the fundamental frequencies of the source signals and template sound signals, as well as the energy of the frequencies in the harmonic set and subharmonic set corresponding to the fundamental frequency; determining the template sound signal of the same gender from the template sound signals according to the range of the preset gender fundamental frequency where the fundamental frequency of the source signal is located, calculating the sum of the voiceprint matching degrees of the source signal and the template sound signal of the same gender to determine the credibility of the source signal; determining whether to perform text command recognition on the source signal based on the credibility degree, and transmitting it to the control system, thereby realizing intelligent control of doors and windows.
[0006] By employing precise fundamental frequency extraction and voiceprint matching technology, the security and accuracy of voice control in smart home systems are enhanced. Analyzing the harmonics and subharmonics in mixed sound signals effectively distinguishes voiceprint characteristics between genders, reducing misidentification issues caused by smoothed voiceprint features. Calculating voiceprint matching degree and credibility allows for intelligent filtering of highly reliable voice commands, enabling accurate control of home appliances such as doors and windows. This not only improves system response speed and reliability but also enhances user privacy protection, preventing unauthorized access and providing users with a safer and more convenient smart home environment.
[0007] Preferably, the step of acquiring the plurality of source signals and source signal frames includes:
[0008] Multiple mixed audio signals are combined row by row into a mixed signal matrix. The mixed signal matrix is decomposed to obtain multiple source signals. The source signals are segmented to obtain multiple source signal frames. Each row in the mixed signal matrix is a mixed audio signal, and one source signal corresponds to multiple source signal frames.
[0009] By combining multiple mixed sound signals row-by-row into a mixed signal matrix and then decomposing it, individual source signals can be effectively extracted from complex acoustic environments. These source signals are then segmented to obtain multiple source signal frames, allowing for detailed analysis of each signal. This improves the resolution and accuracy of sound signal processing, providing a clearer and more specific data foundation for subsequent voiceprint recognition and voice command parsing, thereby enhancing the intelligent door and window control system's ability to identify and control different sound sources.
[0010] Preferably, the step of obtaining the harmonic set and subharmonic set of each frequency in the source signal frame includes:
[0011] Using any source signal frame as the target source signal frame, perform a Fourier transform on the target source signal frame to obtain the energy spectrum of the target source signal frame. Using any frequency in the energy spectrum of the target source signal frame as the target frequency, use integer multiples of the target frequency as the harmonic set of the target frequency until it approaches but cannot exceed the maximum frequency of the energy spectrum. Define all frequencies that are less than or equal to the target frequency and are not integer multiples of the target frequency as the subharmonic set of the target frequency.
[0012] By using any source signal frame as the target source signal frame and performing a Fourier transform on it, the energy spectrum of the target source signal frame can be obtained, thereby accurately identifying the target frequency and the sets of harmonics and subharmonics. This provides richer and more detailed frequency information for voiceprint matching, enabling more accurate voiceprint analysis based on the energy characteristics of these frequencies. By distinguishing between harmonics and subharmonics, key features in the sound signal can be captured more effectively, thus improving the accuracy and reliability of voiceprint recognition, which is crucial for the security and response accuracy of intelligent door and window control systems.
[0013] Preferably, the calculation method for the energy kurtosis of each frequency includes:
[0014] Using any source signal frame as the target source signal frame and any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the difference between the energy of the target frequency and the energy of the frequencies adjacent to the target frequency, perform an exponential function mapping, sum the mapping results, and take the average value as the energy kurtosis of each frequency.
[0015] Preferably, the fundamental frequency of each frequency is calculated using the following method:
[0016] Taking any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the average energy kurtosis of all frequencies in the harmonic set of the target frequency, and the average energy kurtosis of all frequencies other than the target frequency in the subharmonic set of the target frequency. Calculate the ratio of the two average values to obtain the fundamental frequency of the target frequency.
[0017] By quantifying the likelihood that a target frequency will become the fundamental frequency of a signal, a more accurate fundamental frequency reference is provided for voiceprint matching. By distinguishing the energy kurtosis of harmonics and subharmonics, the characteristics of different sound sources can be identified and differentiated more effectively, thereby improving the accuracy of voiceprint recognition and the security of the intelligent door and window control system.
[0018] Preferably, the step of selecting the fundamental frequency of each source signal includes:
[0019] When each source signal frame is used as the target source signal frame, the fundamental frequency of the target frequency is averaged to obtain the comprehensive fundamental frequency of the target frequency. The one with the largest comprehensive fundamental frequency is selected as the fundamental frequency of the source signal.
[0020] Preferably, the calculation method for the voiceprint matching degree includes:
[0021] Calculate the frequency difference between the source signal and the template sound signal corresponding to the fundamental frequency; calculate the energy difference between the source signal and the template sound signal at each frequency based on each frequency in the harmonic set and subharmonic set; divide the energy difference at each frequency by the maximum energy value for normalization; sum the normalized energy differences to obtain the total energy difference; multiply the frequency difference by the total energy difference to obtain the comprehensive difference; use a negative exponential function to perform an exponential mapping on the comprehensive difference to obtain the voiceprint matching degree.
[0022] Preferably, the step of determining whether to perform text command recognition on the source signal based on the degree of credibility includes:
[0023] When the confidence level is less than or equal to a preset threshold, the text command for the source signal is not recognized; otherwise, when the confidence level is greater than the preset threshold, the source signal is subjected to Fourier transform to obtain the Mel spectrum and input into the ASR to obtain the text command corresponding to the source signal.
[0024] Preferably, the intelligent control of the doors and windows further includes:
[0025] When the doors and windows do not respond to repeated voice recognition commands in multiple consecutive sound frames, the remote controller is used to control the doors and windows, realizing multi-intelligent control of smart doors and windows.
[0026] Secondly, an intelligent door and window multi-intelligent control switching system includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned intelligent door and window multi-intelligent control switching method is implemented.
[0027] The present invention has the following effects:
[0028] 1. This invention accurately extracts the fundamental frequency of the signal and calculates the voiceprint matching degree based on the harmonics and subharmonics of the fundamental frequency. This avoids the problem in existing technologies where the fundamental frequency and formant differences are smoothed out due to reliance on voiceprint features across all frequencies of the Mel spectrum. This reduces voiceprint matching errors, improves the accuracy of voiceprint recognition, thereby lowering the risk of false user identity verification and enhancing system security.
[0029] 2. This invention determines the credibility of the source signal by filtering template voice signals of the same gender based on the preset gender-specific fundamental frequency range of the signal's fundamental frequency, and calculating the sum of voiceprint matching degrees. This allows for more accurate differentiation of voiceprint characteristics between different genders. The credibility-based judgment mechanism helps prevent unauthorized access, thereby enhancing the security of the intelligent door and window control system. Attached Figure Description
[0030] Figure 1 This is a flowchart of steps S1-S3 in a smart door and window multi-intelligent control switching method according to an embodiment of the present invention.
[0031] Figure 2 This is a structural block diagram of an intelligent door and window multi-intelligent control switching system according to an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0033] Reference Figure 1 A method for switching between multiple intelligent controls for smart doors and windows includes steps S1-S3, as detailed below:
[0034] S1: Acquire the mixed sound signal of the scene to be controlled and the template sound signal of different instructions entered by the user.
[0035] It should be noted that mixed audio signals refer to signals captured by microphones in real-world applications that include the voices of multiple people speaking and background noise; [This is related to] the placement of microphones on doors and windows. Each microphone, for example, is set to 3-4, and can be adjusted according to specific circumstances. Each microphone collects sound signals from the scene to be controlled. These signals may contain the voices of multiple people speaking and background noise. The mixed sound signals collected by each microphone are of the same length, denoted as . It is used for real-time speech recognition and voiceprint matching to determine voice commands and user identity in the current scenario;
[0036] Template audio signals refer to single audio signals pre-recorded by the user, typically free of background noise, and used as standard reference signals. The system acquires template audio signals for different user-recorded commands. An exemplary set of 5-10 template audio signals is used, which can be adjusted according to specific needs. These template audio signals are single audio signals recorded by the user in a quiet environment, free of noise. They serve as a reference for voiceprint matching and command recognition, helping the system more accurately recognize user voice commands and verify user identity.
[0037] The mixed audio signal is acquired in real time, reflecting the sound environment in the actual use scenario. The template audio signal is pre-recorded and used as a standard reference signal to help the system perform voiceprint matching and command recognition.
[0038] S2: Decompose and segment the mixed audio signal to obtain multiple source signals and source signal frames. In each source signal frame, determine the harmonic set and subharmonic set when each frequency is used as the target frequency, and calculate the energy kurtosis and fundamental frequency of the frequencies in each set. Combine the fundamental frequencies of each frequency in different source signal frames to select the fundamental frequency of each source signal. Calculate the voiceprint matching degree based on the fundamental frequencies of the source signals and template audio signals, as well as the energy of the frequencies in the harmonic set and subharmonic set corresponding to the fundamental frequency.
[0039] Multiple mixed audio signals are arranged in rows to form a size of [size]. The mixed signal matrix is obtained, where each row of the mixed signal matrix is a mixed audio signal. The mixed signal matrix is decomposed using ICA (Independent Component Analysis) to obtain multiple source signals. The source signals are then segmented to obtain multiple source signal frames, where one source signal corresponds to multiple source signal frames.
[0040] Using any source signal frame as the target source signal frame, perform a Fourier transform on the target source signal frame to obtain the energy spectrum of the target source signal frame. Using any frequency in the energy spectrum of the target source signal frame as the target frequency, use integer multiples of the target frequency as the harmonic set of the target frequency until it approaches but cannot exceed the maximum frequency of the energy spectrum. Define all frequencies that are less than or equal to the target frequency and are not integer multiples of the target frequency as the subharmonic set of the target frequency.
[0041] For example, if the target frequency is set to Then the target frequency is an exemplary set of harmonics. An exemplary set of subharmonics for the target frequency is as follows: .
[0042] The calculation methods for energy kurtosis at each frequency include:
[0043] Using any source signal frame as the target source signal frame and any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the difference between the energy of the target frequency and the energy of the frequencies adjacent to the target frequency, perform an exponential function mapping, sum the mapping results, and take the average value as the energy kurtosis of each frequency.
[0044] For example, the energy kurtosis satisfies the following relationship:
[0045] ;
[0046] In the formula, Represents frequency energy peak, Represents frequency The number of neighboring frequencies, Represents frequency energy, Represents frequency The Energy of a neighboring frequency Represented by natural numbers An exponential function with base 0.
[0047] It should be noted that adjacent frequencies specifically refer to frequencies... The higher the energy kurtosis of a frequency, the more pronounced its energy kurtosis. This refers to all frequencies contained between the forward and backward nearest neighbor frequencies within the set. The greater the probability of it being an energy peak.
[0048] The methods for calculating the fundamental frequency include:
[0049] Taking any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the average energy kurtosis of all frequencies in the harmonic set of the target frequency, and the average energy kurtosis of all frequencies other than the target frequency in the subharmonic set of the target frequency. Calculate the ratio of the two average values to obtain the fundamental frequency of the target frequency.
[0050] Specifically, the fundamental frequency of the target frequency satisfies the following relationship:
[0051] ;
[0052] In the formula, Indicates target frequency The fundamental frequency is the frequency of the target signal. The higher the fundamental frequency, the greater the probability that the target frequency is the fundamental frequency of the target source signal frame. Indicates target frequency The average energy kurtosis of all frequencies within the harmonic set and the target frequency. Indicates target frequency The average energy kurtosis of all frequencies within the set of subharmonics.
[0053] When each source signal frame is used as the target source signal frame, the fundamental frequency of the target frequency is averaged to obtain the comprehensive fundamental frequency of the target frequency. The one with the largest comprehensive fundamental frequency is selected as the fundamental frequency of the source signal.
[0054] When a frequency is the fundamental frequency, its corresponding harmonic frequencies typically appear as energy peaks in the energy spectrum, while the corresponding subharmonic frequencies tend to appear as energy troughs. Therefore, when assessing whether a frequency is the fundamental frequency, if the energy kurtosis of the subharmonic set of that frequency is low, while the energy kurtosis of its harmonic set is high, then that frequency is more likely to be the fundamental frequency. This characteristic can serve as an important basis for determining whether a frequency is the fundamental frequency.
[0055] For each source signal and each template audio signal, the fundamental frequency was determined, and the corresponding harmonic and subharmonic sets were identified. It must be noted that the number of harmonics and subharmonics contained in the harmonic and subharmonic sets of the source and template audio signals is consistent. This consistency ensures that the comparison between the source and template audio signals is performed on the same basis during subsequent voiceprint matching, thereby improving the accuracy and reliability of the matching.
[0056] The calculation methods for voiceprint matching accuracy include:
[0057] Calculate the frequency difference between the source signal and the template sound signal corresponding to the fundamental frequency; calculate the energy difference between the source signal and the template sound signal at each frequency based on each frequency in the harmonic set and subharmonic set; divide the energy difference at each frequency by the maximum energy value for normalization; sum the normalized energy differences to obtain the total energy difference; multiply the frequency difference by the total energy difference to obtain the comprehensive difference; use a negative exponential function to perform an exponential mapping on the comprehensive difference to obtain the voiceprint matching degree.
[0058] Specifically, the degree of voiceprint matching satisfies the following relationship:
[0059] ;
[0060] In the formula, This indicates the degree of voiceprint matching between a source signal and a template sound signal. The frequency value (dimensionless) represents the fundamental frequency of the source signal. The frequency value (dimensionless) represents the fundamental frequency of the template sound signal. This indicates the degree of matching between the fundamental frequencies of the two signals. This represents the total number of frequencies contained in the harmonic set and the subharmonic set. Represents the first of two sets of source signals Energy at a frequency Represents the first of two sets of template sound signals Energy at a frequency This represents the maximum energy value, used to eliminate dimensions. Represented by natural numbers An exponential function with base 0.
[0061] Considering that the differences in the fundamental frequency and its harmonics and subharmonics of different individuals' sound signals are mainly in the fundamental frequency, the smaller the difference in the fundamental frequency between a source signal and a template sound signal, the higher the probability that the two signals were generated by the same person; at the same time, if the difference in the energy of their harmonics and subharmonics is also smaller, it further indicates that the probability that the two signals were generated by the same person is even greater.
[0062] Existing technologies typically analyze the energy of all frequencies in the Mel spectrum and cepstrum when calculating voiceprint matching accuracy. This method fails to adequately consider the characteristic that voice differences are primarily concentrated in the fundamental frequency and subharmonics. Since the energy of other frequencies may exhibit high similarity in the voice signals of different individuals, this masks the differences in the fundamental frequency and subharmonics, thus reducing the accuracy of voiceprint matching. This embodiment accurately extracts the fundamental frequency of the signal and calculates the matching accuracy based on the subharmonics and harmonics of the fundamental frequency, effectively avoiding interference from redundant frequencies and significantly improving the accuracy and reliability of voiceprint matching.
[0063] S3: Based on the range of the preset gender base frequency where the source signal base frequency is located, determine the template voice signal of the same gender from the template voice signal, calculate the sum of the voiceprint matching degree between the source signal and the template voice signal of the same gender to determine the credibility of the source signal, and determine whether to perform text command recognition on the source signal based on the credibility degree, and transmit it to the control system to realize intelligent control of doors and windows.
[0064] It should be noted that the preset gender base frequencies include male base frequencies and female base frequencies. The male base frequency is in the range of 60-150 Hz, and the female base frequency is in the range of 200-400 Hz. This is well known to those skilled in the art and will not be described in detail here.
[0065] Given that multiple people's voice signals may exist in a scene, and that voice characteristics differ significantly between genders, the accuracy of cross-gender voiceprint matching is typically low. Therefore, this paper analyzes the fundamental frequency range of the source signal to identify and filter template voice signals that match the user's gender. The voiceprint matching degree between the source signal and the same-gender template voice signal is calculated and accumulated to determine the reliability of the source signal. This helps improve the accuracy of voiceprint recognition, ensuring that intelligent control systems can more reliably distinguish between voiceprint features of different genders.
[0066] Specifically, the degree of credibility satisfies the following relationship: ;
[0067] In the formula, Indicates the reliability of a source signal. Template voice signals representing the same gender. Indicates the source signal and the first The degree of voiceprint matching between template voice signals of the same gender.
[0068] When the confidence level is less than or equal to a preset threshold, the text command for the source signal is not recognized. Conversely, when the confidence level is greater than the preset threshold, the source signal is subjected to Fourier transform to obtain the Mel spectrum and input into the ASR to obtain the text command corresponding to the source signal. The command includes two types: open door / open window and close door / close window.
[0069] For example, the preset threshold is 0.7, which can be adjusted according to specific circumstances. The fundamental frequency for men is in the range of 60-150 Hz, and the fundamental frequency for women is in the range of 200-400 Hz.
[0070] When the doors and windows do not respond to repeated voice recognition commands in multiple consecutive sound frames, the remote controller is used to control the doors and windows, realizing multi-intelligent control of smart doors and windows.
[0071] This invention also provides an intelligent door and window multi-intelligent control switching system. For example... Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a multi-intelligent control switching method for intelligent doors and windows according to the first aspect of the present invention. The system also includes other components well-known to those skilled in the art, such as a communication bus and a communication interface. Their configurations and functions are known in the art and will not be described in detail here.
[0072] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for switching between multiple intelligent controls for smart doors and windows, characterized in that, include: Acquire the mixed sound signal of the scene to be controlled and the template sound signal of different commands entered by the user; The mixed audio signal is decomposed and segmented to obtain multiple source signals and source signal frames. In each source signal frame, the harmonic set and subharmonic set when each frequency is used as the target frequency are determined, and the energy kurtosis and fundamental frequency of the frequencies in each set are calculated respectively. The fundamental frequency of each frequency in different source signal frames is combined to select the fundamental frequency of each source signal. Based on the fundamental frequency of the source signal and the template audio signal, as well as the energy of the frequencies in the harmonic set and subharmonic set corresponding to the fundamental frequency, the voiceprint matching degree is calculated. Based on each source signal and each template sound signal, the fundamental frequency of each source signal and each template sound signal was determined, and the harmonic set and subharmonic set corresponding to the fundamental frequency were identified. Based on the range of the preset gender base frequency where the source signal base frequency is located, the template voice signal of the same gender is determined from the template voice signal. The sum of the voiceprint matching degree between the source signal and the template voice signal of the same gender is calculated to determine the credibility of the source signal. Based on the credibility, it is determined whether to perform text command recognition on the source signal and transmit it to the control system, thereby realizing intelligent control of doors and windows.
2. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The steps for acquiring the plurality of source signals and source signal frames include: Multiple mixed audio signals are combined row by row into a mixed signal matrix. The mixed signal matrix is decomposed to obtain multiple source signals. The source signals are segmented to obtain multiple source signal frames. Each row in the mixed signal matrix is a mixed audio signal, and one source signal corresponds to multiple source signal frames.
3. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The steps for obtaining the harmonic set and subharmonic set of each frequency in the source signal frame include: Using any source signal frame as the target source signal frame, perform a Fourier transform on the target source signal frame to obtain the energy spectrum of the target source signal frame. Using any frequency in the energy spectrum of the target source signal frame as the target frequency, use integer multiples of the target frequency as the harmonic set of the target frequency until it approaches but cannot exceed the maximum frequency of the energy spectrum. Define all frequencies that are less than or equal to the target frequency and are not integer multiples of the target frequency as the subharmonic set of the target frequency.
4. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The calculation methods for the energy kurtosis of each frequency include: Using any source signal frame as the target source signal frame and any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the difference between the energy of the target frequency and the energy of the frequencies adjacent to the target frequency, perform an exponential function mapping, sum the mapping results, and take the average value as the energy kurtosis of each frequency.
5. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The calculation methods for the fundamental frequency of each frequency include: Taking any frequency in the energy spectrum of the target source signal frame as the target frequency, calculate the average energy kurtosis of all frequencies in the harmonic set of the target frequency, and the average energy kurtosis of all frequencies other than the target frequency in the subharmonic set of the target frequency. Calculate the ratio of the two average values to obtain the fundamental frequency of the target frequency.
6. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The step of selecting the fundamental frequency of each source signal includes: When each source signal frame is used as the target source signal frame, the fundamental frequency of the target frequency is averaged to obtain the comprehensive fundamental frequency of the target frequency. The one with the largest comprehensive fundamental frequency is selected as the fundamental frequency of the source signal.
7. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The calculation method for the voiceprint matching degree includes: Calculate the frequency difference between the source signal and the template sound signal corresponding to the fundamental frequency; calculate the energy difference between the source signal and the template sound signal at each frequency based on each frequency in the harmonic set and subharmonic set; divide the energy difference at each frequency by the maximum energy value for normalization; sum the normalized energy differences to obtain the total energy difference; multiply the frequency difference by the total energy difference to obtain the comprehensive difference; use a negative exponential function to perform an exponential mapping on the comprehensive difference to obtain the voiceprint matching degree.
8. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The determination of whether to perform text command recognition on the source signal based on the degree of credibility includes: When the confidence level is less than or equal to a preset threshold, the text command for the source signal is not recognized; otherwise, when the confidence level is greater than the preset threshold, the source signal is subjected to Fourier transform to obtain the Mel spectrum and input into the ASR to obtain the text command corresponding to the source signal.
9. The intelligent door and window multi-intelligent control switching method according to claim 1, characterized in that, The intelligent control of the doors and windows also includes: When the doors and windows do not respond to repeated voice recognition commands in multiple consecutive sound frames, the remote controller is used to control the doors and windows, realizing multi-intelligent control of smart doors and windows.
10. A smart door and window multi-intelligent control switching system, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement the intelligent door and window multi-intelligent control switching method according to any one of claims 1-9.
Citation Information
Patent Citations
Determining features of harmonic signals
CN107430850A
Self-adaptive VAD parameter adjusting method and system based on voiceprint recognition
CN120048268A
Intelligent automatic door sound control method and system
CN120089144A