Electric appliance voiceprint recognition management method and electronic equipment thereof
The preset wake-up words enable the electronic voiceprint recognition function, collect and match the voiceprint data to control the electrical appliance, solve the problem that the electrical appliance cannot automatically identify users and operate cumbersomely, realize natural voice interaction and personalized control, and improve intelligence and privacy protection.
Patent Information
- Application Number
- CN202510629425.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-12
AI Technical Summary
The existing electrical voiceprint recognition management methods cannot automatically identify different users and personalize them according to their usage habits. The operation is cumbersome and lacks intelligent interaction and privacy protection mechanisms, especially in high humidity environments, the recognition effect is unstable.
The voiceprint recognition function of the appliance is activated through preset wake-up words, collect the current voiceprint data, mobilize the corresponding execution strategies based on the voiceprint data, and control the appliance to execute the corresponding strategies, combine voiceprint feature extraction and database matching to realize user identity confirmation and personalized configuration.
It realizes natural voice interaction between users and electrical appliances without manual operation, improves user privacy protection capabilities and the intelligence level of equipment, and supports adaptive differentiated control in multi-user scenarios.
Smart Images

Figure CN120472911A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent control, and in particular to a method, device, system, electronic device and storage medium for voiceprint recognition management of electrical appliances. Background Art
[0002] With the development of smart home and voice interaction technology, traditional electrical appliances have gradually evolved towards personalized control and intelligent recognition.
[0003] In the existing technology, some electrical appliances support scheduled appointments, remote control or voice command operations, but there are difficulties in automatically identifying the user's identity and in being able to perform differentiated control based on the usage habits of different users. In addition, facial recognition poses the risk of user privacy leakage and the recognition effect is unstable in high-humidity environments such as bathrooms.
[0004] Therefore, the existing voiceprint recognition and management methods for electrical appliances have the problems of being unable to automatically identify different users and perform personalized settings based on their usage habits, being cumbersome to operate, and lacking intelligent interaction and privacy protection mechanisms. Summary of the Invention
[0005] An embodiment of the present invention provides a method for voiceprint recognition and management of electrical appliances to solve the problems of existing methods for voiceprint recognition and management of electrical appliances, such as the inability to automatically identify different users and perform personalized settings according to their usage habits, cumbersome operation, and lack of intelligent interaction and privacy protection mechanisms.
[0006] In a first aspect, an embodiment of the present invention provides a method for voiceprint recognition and management of an electrical appliance, the method comprising the following steps: Determining a preset wake-up word, where the preset wake-up word is used to activate a voiceprint recognition function of the appliance; When the voiceprint recognition function is activated, the voiceprint data at the current activation moment is collected; Based on the voiceprint data, a corresponding execution strategy is activated and the electrical appliance is controlled to execute the corresponding strategy.
[0007] Optionally, the method for determining a preset wake-up word further includes: When setting a preset wake-up word, collect the current voice audio data of the setting; Extracting voiceprint features from the human voice audio data to obtain corresponding voiceprint feature data; Binding the voiceprint feature data to the preset wake-up word, and when the preset wake-up word is detected, matching it with the corresponding voiceprint feature data to obtain a matching result; Based on the matching result, the activation state of the voiceprint recognition function of the electrical appliance is determined.
[0008] Optionally, when the voiceprint recognition function is activated, collecting voiceprint data at the current moment includes: When the voiceprint recognition function is activated, the surrounding sound data of the appliance at the current moment is collected in real time, and the surrounding sound data includes environmental audio data and human voice audio data; Performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data; The pure human voice audio data is subjected to voiceprint feature extraction processing to obtain corresponding voiceprint data.
[0009] Optionally, the method further comprises: performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data. Decomposing the human voice audio data to obtain multiple human voice audio data; Performing audio feature processing on the plurality of human voice audio data to determine corresponding audio feature data, the audio feature data including volume, pitch, and frequency data; Based on the volume, pitch and frequency data, target human voice audio data is determined from a plurality of human voice audio data and is used as pure human voice audio data.
[0010] Optionally, mobilizing a corresponding execution strategy based on the voiceprint data and controlling the appliance to execute the corresponding strategy includes: Based on the voiceprint data, matching is performed in the database to determine the corresponding target user; According to the target user, the corresponding configuration parameters are called, and the electrical appliance is controlled to execute the parameter configuration.
[0011] Optionally, the method further includes calling corresponding configuration parameters according to the target user and controlling the electrical appliance to perform parameter configuration: If the corresponding voiceprint data cannot be matched in the database, the target user is determined to be a new user, and according to the voice interaction guidance, the new user is guided to enter the corresponding personalized configuration parameters item by item, and the personalized configuration parameters are bound to their voiceprint characteristics and stored. The personalized configuration parameters include at least one of water temperature, water volume, and duration.
[0012] In a second aspect, an embodiment of the present invention further provides a device for voiceprint recognition and management of an electrical appliance, the device comprising: a first determining module, configured to determine a preset wake-up word, wherein the preset wake-up word is used to activate a voiceprint recognition function of the appliance; The first collection module is used to collect the voiceprint data at the current startup moment when the voiceprint recognition function is started; The first calling module is used to mobilize the corresponding execution strategy based on the voiceprint data and control the electrical appliance to execute the corresponding strategy.
[0013] In a third aspect, an embodiment of the present invention provides an electrical appliance voiceprint recognition and management system, which includes: an electrical appliance voiceprint recognition and management device, a server, and a smart appliance.
[0014] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the electrical appliance voiceprint recognition and management method provided in an embodiment of the present invention are implemented.
[0015] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the appliance voiceprint recognition and management method provided in the embodiment of the invention are implemented.
[0016] In an embodiment of the present invention, a preset wake-up word is determined, and the preset wake-up word is used to activate the voiceprint recognition function of the appliance. When the voiceprint recognition function is activated, voiceprint data at the current startup moment is collected. Based on the voiceprint data, a corresponding execution strategy is activated and the appliance is controlled to execute the corresponding strategy. The above method and steps enable natural voice interaction between the user and the appliance without manual operation. In addition, the introduction of voiceprint recognition replaces traditional identity authentication methods, and the user's personalized configuration can be automatically retrieved and executed, achieving adaptive and differentiated control in multi-user scenarios. This avoids the problems of cumbersome setup and frequent misoperation of traditional appliances, and improves user privacy protection capabilities and the intelligence level of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a system architecture diagram of an appliance voiceprint recognition management system provided by an embodiment of the present invention; Figure 2 This is a flow chart of a method for voiceprint recognition and management of electrical appliances provided by an embodiment of the present invention; Figure 3 Schematic diagram of another device for voiceprint recognition and management of electrical appliances provided in an embodiment of the present invention; Figure 4It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] like Figure 1 As shown, Figure 1 This is an architectural diagram of an appliance voiceprint recognition management system 100 provided by an embodiment of the present invention. The appliance voiceprint recognition management system includes: an appliance voiceprint recognition management device 300, a server 101, and a smart appliance 102. The appliance voiceprint recognition management device 300 also includes a first determination module for determining a preset wake-up phrase; a first collection module for collecting voiceprint data at the current startup moment when the voiceprint recognition function is activated; and a first invocation module for activating a corresponding execution strategy based on the voiceprint data and controlling the appliance to execute the corresponding strategy.
[0021] Specifically, the above-mentioned smart appliances may include but are not limited to smart electric water heaters, smart TVs, etc., which can be connected to any smart control device or control system through a preset network protocol, including but not limited to the above-mentioned appliance voiceprint recognition management device and appliance voiceprint recognition management system, and can be remotely controlled through any of the above-mentioned smart control devices, and the above-mentioned smart water heater can recognize the user's voiceprint information and match it in the database to call the corresponding configuration parameters.
[0022] The above-mentioned preset wake-up words may refer to a set of specific voice instructions pre-defined and entered during the initialization of the appliance, the deployment phase of the above-mentioned appliance voiceprint recognition management system, or the first configuration process. It may be a word or a long word, and there is no limit on the number of words. It should be noted that the preset wake-up words can be used to trigger the voiceprint recognition function of the appliance, and the preset wake-up words can be phrases with clear semantic boundaries and high recognizability. For example, when the appliance is turned on, the user wakes up by voice saying "Xiao Wei, Xiao Wei" or "Give me some hot water". The above-mentioned appliance voiceprint recognition management system continuously monitors the ambient sound source and performs keyword matching operations. When a matching preset wake-up word is detected, it enters the subsequent voice recognition and voiceprint analysis stage.
[0023] The voiceprint recognition function may refer to the technical function of the appliance voiceprint recognition management system for identifying the speaker's identity. This function implements the solution function of confirming the speaker's identity by digitally processing the user's voice signal, extracting voiceprint features, and performing database matching. Specifically, this voiceprint recognition function not only identifies and analyzes the speaker's voice characteristics, such as bioacoustic parameters such as "vocal tract length, vocal frequency, and intonation," but also extracts voiceprint feature vectors from the speaker's audio data and calculates similarity matching with voiceprint data in the database. This allows the appliance voiceprint recognition management system to determine the user's identity and then match the corresponding configuration parameters for execution.
[0024] The above-mentioned current startup moment may be the precise time point at which the above-mentioned electrical appliance voiceprint recognition management system enters the voiceprint recognition process after recognizing the above-mentioned preset wake-up word. Specifically, after this moment, the above-mentioned electrical appliance voiceprint recognition management system begins to recognize the user's voice input, including the audio data of the preset wake-up word recognized previously, and inputs it into the subsequent processing process as valid voiceprint data. It should be noted that this time point plays a key role as the starting time of voiceprint recognition, which can ensure that the extracted voice signal truly reflects the user's identity and is not interfered with by environmental noise or non-interactive voice, thereby improving the accuracy and stability of the recognition results.
[0025] The above-mentioned voiceprint data may include but is not limited to any representative digital feature vector extracted from the user's voice signal by the appliance voiceprint recognition management system, and is used to uniquely identify the user's identity.
[0026] Specifically, the above-mentioned user voice signal can be extracted through MFCC (Mel-Frequency Cepstral Coefficients), PLP (Perceptual Linear Prediction), or d-vector / x-vector model generation based on deep neural networks to obtain corresponding voiceprint data. It should be noted that since the above-mentioned voiceprint data has individual uniqueness and time robustness, during the recognition process, this embodiment can perform real-time monitoring and targeted recognition of unique and continuous voiceprint data and corresponding instruction operations, and can perform real-time similarity comparison with templates stored in the system to confirm the identity of the current interactor and serve as the basis for calling control strategies.
[0027] In a possible embodiment, after successfully identifying the user's identity, the above-mentioned appliance voiceprint recognition management system retrieves and extracts the personalized execution strategy bound to the user's identity from the system policy database based on the identification result. Specifically, by querying the database and guiding the appliance to execute the corresponding parameter configuration and setting process based on the query results, the above-mentioned method steps can improve the mobilization response efficiency based on the cache mechanism or local storage, so that the strategy is transferred and driven to be executed in a short time after the voice interaction is completed.
[0028] The above-mentioned execution strategy can be a set of parameter settings bound to different registered users, mainly including dimensions such as water outlet temperature, hot water flow rate, water supply time, energy-saving mode, etc. Specifically, the user can complete the parameter entry through voice interaction during the first use or subsequent use. After detecting the parameter entry or voice interaction, the above-mentioned electrical appliance voiceprint recognition management system will associate it with the voiceprint feature and store it. Each time the above-mentioned electrical appliance voiceprint recognition management system identifies a specific user identity, it will automatically call the execution strategy corresponding to the user, drive the electrical equipment to complete the corresponding control task, and realize a personalized usage experience based on identity.
[0029] For example, when the appliance is awakened by the above-mentioned preset wake-up word, the voiceprint data corresponding to the current preset wake-up word is identified, and then after searching and matching in the database, the corresponding matching strategy or configuration parameters are executed.
[0030] In another possible embodiment, after calling the user's execution strategy, the above-mentioned appliance voiceprint recognition management system sends control instructions to the smart appliance to drive the device to perform corresponding operating behaviors, thereby achieving the purpose of controlling the smart appliance. Specifically, the personalized control needs of the appliance can be realized through control modules including but not limited to heating modules, regulating valve control circuits, start timing modules, etc.
[0031] Through the above method and steps, the natural start of voice interaction is achieved, the user identity is effectively determined, and the personalized execution strategy bound to the user is called based on the recognition result, realizing the "customized and automatic response" intelligent hot water service.
[0032] like Figure 2 As shown, Figure 2 Flowchart of a method for managing voiceprint recognition of an electrical appliance provided by an embodiment of the present invention. The method for managing voiceprint recognition of an electrical appliance includes the following steps: 201. Determine a preset wake-up word.
[0033] In an embodiment of the present invention, the above-mentioned electrical appliance voiceprint recognition management method can be applied to an electrical appliance voiceprint recognition management system. The above-mentioned electrical appliance voiceprint recognition management system has functions such as voiceprint data processing, voiceprint data transmission and reception, and voiceprint data memory storage, and can be constructed based on a server or a server cluster. The above-mentioned server or server cluster can be an electronic device with voiceprint data processing capabilities.
[0034] The above-mentioned preset wake-up words may refer to a set of specific voice instructions pre-defined and entered during the initialization of the appliance, the deployment phase of the above-mentioned appliance voiceprint recognition management system, or the first configuration process. It may be a word or a long word, and there is no limit on the number of words. It should be noted that the preset wake-up words can be used to trigger the voiceprint recognition function of the appliance, and the preset wake-up words can be phrases with clear semantic boundaries and high recognizability. For example, when the appliance is turned on, the user wakes up by voice saying "Xiao Wei, Xiao Wei" or "Give me some hot water". The above-mentioned appliance voiceprint recognition management system continuously monitors the ambient sound source and performs keyword matching operations. When a matching preset wake-up word is detected, it enters the subsequent voice recognition and voiceprint analysis stage.
[0035] 202. When the voiceprint recognition function is activated, the voiceprint data at the current activation moment is collected.
[0036] In embodiments of the present invention, the voiceprint recognition function may refer to the technical function of the appliance voiceprint recognition management system for identifying the speaker's identity. This function implements the solution function of confirming the speaker's identity by digitally processing the user's voice signal, extracting voiceprint features, and performing database matching. Specifically, this voiceprint recognition function not only identifies and analyzes the speaker's voice characteristics, such as bioacoustic parameters such as "vocal tract length, vocal frequency, and intonation," but also extracts voiceprint feature vectors from the speaker's audio data and performs similarity matching calculations with voiceprint data in the database. This allows the appliance voiceprint recognition management system to determine the user's identity, match the corresponding configuration parameters, and execute the function.
[0037] The above-mentioned current startup moment may be the precise time point at which the above-mentioned electrical appliance voiceprint recognition management system enters the voiceprint recognition process after recognizing the above-mentioned preset wake-up word. Specifically, after this moment, the above-mentioned electrical appliance voiceprint recognition management system begins to recognize the user's voice input, including the audio data of the preset wake-up word recognized previously, and inputs it into the subsequent processing process as valid voiceprint data. It should be noted that this time point plays a key role as the starting time of voiceprint recognition, which can ensure that the extracted voice signal truly reflects the user's identity and is not interfered with by environmental noise or non-interactive voice, thereby improving the accuracy and stability of the recognition results.
[0038] The above-mentioned voiceprint data may include but is not limited to any representative digital feature vector extracted from the user's voice signal by the appliance voiceprint recognition management system, and is used to uniquely identify the user's identity.
[0039] Specifically, the above-mentioned user voice signal can be extracted through MFCC (Mel-Frequency Cepstral Coefficients), PLP (Perceptual Linear Prediction), or d-vector / x-vector model generation based on deep neural networks to obtain corresponding voiceprint data. It should be noted that since the above-mentioned voiceprint data has individual uniqueness and time robustness, during the recognition process, this embodiment can perform real-time monitoring and targeted recognition of unique and continuous voiceprint data and corresponding instruction operations, and can perform real-time similarity comparison with templates stored in the system to confirm the identity of the current interactor and serve as the basis for calling control strategies.
[0040] In a possible embodiment, after successfully identifying the user's identity, the above-mentioned appliance voiceprint recognition management system retrieves and extracts the personalized execution strategy bound to the user's identity from the system policy database based on the identification result. Specifically, by querying the database and guiding the appliance to execute the corresponding parameter configuration and setting process based on the query results, the above-mentioned method steps can improve the mobilization response efficiency based on the cache mechanism or local storage, so that the strategy is transferred and driven to be executed in a short time after the voice interaction is completed.
[0041] 203. Based on the voiceprint data, activate the corresponding execution strategy and control the appliance to execute the corresponding strategy.
[0042] In an embodiment of the present invention, the above-mentioned execution strategy can be a set of parameter settings bound to different registered users, mainly including dimensions such as water outlet temperature, hot water flow rate, water supply time, energy-saving mode, etc. Specifically, the user can complete the parameter entry through voice interaction during the first use or subsequent use. After detecting the parameter entry or voice interaction, the above-mentioned electrical appliance voiceprint recognition management system associates it with the voiceprint feature and stores it. Each time the above-mentioned electrical appliance voiceprint recognition management system identifies a specific user identity, it automatically calls the execution strategy corresponding to the user, drives the electrical equipment to complete the corresponding control task, and realizes a personalized usage experience based on identity.
[0043] For example, when the appliance is awakened by the above-mentioned preset wake-up word, the voiceprint data corresponding to the current preset wake-up word is identified, and then after searching and matching in the database, the corresponding matching strategy or configuration parameters are executed.
[0044] In another possible embodiment, after calling the user's execution strategy, the above-mentioned appliance voiceprint recognition management system sends control instructions to the smart appliance to drive the device to perform corresponding operating behaviors, thereby achieving the purpose of controlling the smart appliance. Specifically, the personalized control needs of the appliance can be realized through control modules including but not limited to heating modules, regulating valve control circuits, start timing modules, etc.
[0045] In an embodiment of the present invention, a preset wake-up word is determined and used to activate the voiceprint recognition function of the appliance. When the voiceprint recognition function is activated, voiceprint data at the current activation moment is collected. Based on the voiceprint data, a corresponding execution strategy is activated and the appliance is controlled to execute the corresponding strategy. The above method and steps enable natural voice interaction between the user and the appliance without manual operation. Furthermore, by introducing voiceprint recognition as an alternative to traditional identity verification methods, personalized user configurations can be automatically retrieved and executed, achieving adaptive and differentiated control in multi-user scenarios. This avoids the cumbersome setup and frequent misoperations of traditional appliances, improving user privacy protection capabilities and the intelligence level of the device.
[0046] Optionally, in the step of determining the preset wake-up word, it is also possible to collect the human voice audio data currently being set when setting the preset wake-up word; perform voiceprint feature extraction on the human voice audio data to obtain corresponding voiceprint feature data; bind the voiceprint feature data to the preset wake-up word, and when the preset wake-up word is detected, match it with the corresponding voiceprint feature data to obtain a matching result; based on the matching result, determine the startup status of the voiceprint recognition function of the appliance.
[0047] In an embodiment of the present invention, the surrounding audio can be collected through the built-in microphone matrix module of the smart appliance. Generally speaking, it can be automatically started after entering the setting mode. The goal of the collection is to obtain the continuous audio stream generated by the current user's pronunciation, which is used as the original input for subsequent voiceprint analysis. It should be noted that metadata such as timestamps and volume levels can also be attached during the collection process to assist in signal correction and synchronization judgment in subsequent processing links.
[0048] The above-mentioned human voice audio data may refer to the voice signal emitted by the user when setting the wake-up word, which is converted into a digital audio file in a specific format (such as WAV, PCM, etc.) by the acquisition module. It contains acoustic information such as the speaker's unique voice waveform, frequency components and time structure. It should be noted that this data is different from the text content at the semantic layer, nor is it different from background sound or device noise. Instead, it is a carrier of human vocal tract information used for voiceprint modeling. It is the most basic and critical data source in voiceprint recognition.
[0049] In a possible embodiment, the above-mentioned appliance voiceprint recognition management system performs signal processing and modeling on human voice audio data, and extracts therefrom a numerical vector that can uniquely represent the individual voice characteristics, so as to realize the voiceprint feature extraction process.
[0050] Specifically, during the extraction process, the above-mentioned appliance voiceprint recognition and management system can generate a stable voiceprint feature vector by performing steps such as pre-emphasis, frame windowing, short-time Fourier transform (STFT) on the audio, and using feature extraction algorithms such as MFCC (Mel-frequency cepstral coefficients), PLP or deep neural network encoders (such as x-vector).
[0051] The above-mentioned voiceprint feature data can be extracted from the user's voice signal and is multi-dimensional vector data used to describe the individual physiological and vocal behavior characteristics of the user. It can usually be represented by a set of floating-point arrays with fixed dimensions, reflecting the speaker's vocal tract structure, frequency band energy distribution, pronunciation rhythm and other characteristic parameters within a specific time window. Generally speaking, the above-mentioned electrical voiceprint recognition management system can use it as the core content of the user's voiceprint template, and perform similarity matching with the real-time voiceprint data in the subsequent recognition stage. This process method step can realize the unique identification of individual identity.
[0052] In another possible embodiment, the above-mentioned appliance voiceprint recognition management system establishes a one-to-one correspondence between the wake-up word set by the user and its voiceprint feature data, and stores it in the local or cloud database of the appliance voiceprint recognition management system, so that the binding relationship not only includes the mapping of the voiceprint template and the wake-up word, but also can record the association structure with policy information such as device usage behavior and time preference.
[0053] The above matching process can be the process in which the appliance voiceprint recognition management system detects the wake-up word during actual operation and compares the currently collected voiceprint features with the bound voiceprint template to obtain the matching result. For example, if the similarity score is higher than 90%, the match is successful; otherwise, it is considered a match failure. The above matching result directly affects whether the recognition process is accepted and enters the next step of policy invocation and device control. Generally speaking, a match failure indicates that it is the first time for a new user, and voice guidance can be used for the first-time setup.
[0054] The aforementioned startup state refers to the state judgment result of the appliance voiceprint recognition management system after receiving the wake-up word and completing the voiceprint matching. If the matching result meets the set conditions, the system sets the state to "start" and activates the subsequent user identity recognition and policy retrieval process; if the matching fails, the system remains in a non-interactive state to prevent unauthorized users from accidentally accessing the device.
[0055] In this embodiment, when a user sets a preset wake-up phrase, the appliance voiceprint recognition management system not only saves the user's textual content (e.g., "Xiao Wei") but also simultaneously collects the user's voice signal (i.e., human voice audio data) emitted during the setting process. Voiceprint features are extracted from the collected human voice audio data to generate unique voiceprint feature data. This generated voiceprint feature data is then bound to the corresponding wake-up phrase, forming a "wake-up phrase + voiceprint template" linkage relationship. Thereafter, whenever the system detects a wake-up phrase during operation, it simultaneously extracts the voiceprint features of the current speaker and matches them with the voiceprint template bound to the wake-up phrase. A similarity score is calculated and a match result is generated. If the match result meets a system-defined similarity threshold (e.g., above 90%), the wake-up operation is deemed valid, and the appliance voiceprint recognition function is activated as "activated," officially entering the subsequent user identification and control processes. Conversely, if the match fails, the system does not respond to the wake-up phrase, thereby preventing unauthorized users from accidentally activating the device by imitating the wake-up phrase, enhancing security and exclusivity.
[0056] Optionally, when the voiceprint recognition function is activated, the step of collecting the voiceprint data at the current moment also includes collecting the surrounding sound data of the appliance at the current moment in real time when the voiceprint recognition function is activated, the surrounding sound data including environmental audio data and human voice audio data; performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data; performing voiceprint feature extraction processing on the pure human voice audio data to obtain corresponding voiceprint data.
[0057] In an embodiment of the present invention, the above-mentioned ambient sound data may refer to the set of all audio signals at the current moment collected by the sound pickup device when the voiceprint recognition management system of the electrical appliance starts the voiceprint recognition function, including but not limited to the user's voice (i.e., human voice audio data), various non-human voice components of the environment, such as background noise, electrical appliance operation sound, natural sound sources, etc. Since the ambient sound data reflects the overall state of the actual sound scene in which the device is located, it is necessary to perform noise reduction processing on the ambient sound data to ensure the original input basis for subsequent voice signal cleaning and subsequent voiceprint recognition processing.
[0058] The above-mentioned environmental audio data is a subset of the surrounding sound data, including but not limited to the non-semantic background sound of the user's voice components. It is usually characterized by no language structure, unfixed sound source, and wide spectrum distribution. Common environmental audio includes but is not limited to the sound of fans running, water flow, TV or electrical noise, street noise, etc. It should be noted that the above-mentioned environmental audio data is essentially an interference signal in the recognition process, which will cause model misleading in the voiceprint feature extraction stage and reduce the recognition accuracy. Therefore, the above-mentioned electrical voiceprint recognition management system will model and eliminate the environmental audio data.
[0059] In a possible embodiment, the above-mentioned appliance voiceprint recognition management system separates and suppresses the environmental noise components in the audio data through signal processing or deep learning algorithms, thereby extracting a purer voice signal. The specific noise reduction processing algorithm and process are mentioned in the above-mentioned method steps and will not be explained here. Noise reduction through the above-mentioned method steps can improve the accuracy and robustness of voiceprint recognition.
[0060] The above-mentioned pure human voice audio data refers to the human voice signal obtained by the appliance voiceprint recognition management system after completing the noise reduction processing, which has basically eliminated the interference of environmental noise. This data retains the core acoustic features of the user when speaking, such as pitch, tone, speaking speed and resonance peaks, and has the characteristics of clear signal structure, concentrated spectrum, and can be used for high-quality extraction of voiceprint features.
[0061] In another possible embodiment, when the voiceprint recognition function is activated, the above-mentioned appliance voiceprint recognition management system collects the voice of the current user and simultaneously obtains surrounding audio data, and obtains pure human voice audio data after performing noise reduction processing on the surrounding audio data.
[0062] Through the above method and steps, the robustness and accuracy of voiceprint recognition in noisy environments can be improved, ensuring the reliability and intelligent response capabilities of electrical appliance control in multi-user scenarios.
[0063] Optionally, in the step of performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data, it also includes decomposing the human voice audio data to obtain multiple human voice audio data; performing audio feature processing on the multiple human voice audio data to determine the corresponding audio feature data; based on the volume, pitch and frequency data, determining the target human voice audio data from the multiple human voice audio data, and using it as pure human voice audio data.
[0064] In an embodiment of the present invention, a mixed speech signal having multiple human voice audio data can be decomposed into multiple relatively independent single speaker channels, i.e., multiple human voice audio data, through speaker separation or blind source separation technology (such as DPRNN, Conv-TasNet). Specifically, the specific vocalization audio segment of each human voice audio data can be decomposed and split, and through time-frequency domain modeling and mask learning, distinctive speech channels can be extracted, thereby obtaining multiple relatively independent single speaker channels, i.e., multiple human voice audio data.
[0065] The above-mentioned audio feature processing may refer to the process of extracting structural acoustic parameters of each speech signal segment after the decomposition processing is completed.
[0066] The above-mentioned audio feature data may include but is not limited to structural quantitative data such as volume, pitch and frequency data, and can usually be processed into a multidimensional vector form, including attributes such as {volume, pitch, frequency distribution, speaking speed, signal duration}, for target voice screening.
[0067] The above-mentioned target human voice audio data is identified by comparing audio feature data among multiple separated voice signals to identify the voice channel most likely to be issued by the current wake-up person or user. Generally speaking, this recognition is based on the consistent matching of multiple acoustic parameters, such as the maximum volume, the pitch range that conforms to the personal history model, the stability of the frequency curve and other judgment criteria. The target human voice audio data that is finally selected will be input into the voiceprint recognition module as pure voice to ensure that the identity recognition is accurate and targeted, and to enhance the intelligence of the system interaction and the ability to prevent false triggering.
[0068] In a possible embodiment, after completing the noise reduction processing of the surrounding sound data, the above-mentioned appliance voiceprint recognition management system performs decomposition processing on the extracted human voice audio data, splits it into multiple independent individual audio channels, and performs audio feature processing on these channels respectively, extracts parameters such as volume, pitch and frequency of each voice segment, and forms corresponding audio feature data. Finally, based on these acoustic parameters, the sound characteristics of each sound source in the current scene are compared, and finally the target human voice audio data that is most likely to be the wake-up person is identified, and it is used as pure voice data for subsequent voiceprint recognition.
[0069] The above method and steps can accurately lock the current user even when multiple people are speaking at the same time or there are other voices in the background, effectively improving the recognition robustness, accuracy and practicality of the system's intelligent interaction in a multi-speaker environment.
[0070] Optionally, in the step of mobilizing the corresponding execution strategy and controlling the appliance to execute the corresponding strategy based on the voiceprint data, it also includes matching in the database based on the voiceprint data to determine the corresponding target user; according to the target user, calling the corresponding configuration parameters, and controlling the appliance to execute the parameter configuration.
[0071] In an embodiment of the present invention, the target user may refer to an individual identified as being most consistent with the identity of the current speaker after matching the currently collected voiceprint data with voiceprint templates of multiple registered users in a database through the appliance voiceprint recognition management system. Specifically, a voiceprint vector similarity comparison method (such as cosine similarity and Euclidean distance) may generally be used to calculate the degree of matching between the current voiceprint and the historical template, and to set a recognition threshold to ensure the uniqueness and accuracy of the recognition result. More specifically, the recognition threshold may be dynamically adjusted based on the background noise and the amount of the human voice audio data, i.e., the number of voice channels. Generally speaking, the noisier the background and the more voice channels there are, the higher the recognition threshold, i.e., a higher precision threshold is required to determine whether the user is the target user.
[0072] The above configuration parameters may refer to a set of electrical appliance operation setting values bound to the target user in the above electrical appliance voiceprint recognition management system. Generally speaking, they may be the corresponding parameter settings that the current smart appliance needs to perform operations during operation. For smart water heaters, they usually include but are not limited to target water outlet temperature, water flow rate, usage time, timing mode, energy-saving priority and other setting parameters; or for smart air conditioners, they may include but are not limited to temperature, wind speed and wind direction mode and other setting parameters. This parameter set may be constructed through the user's first setting, system learning or historical behavior recording, and may be automatically called each time the user's identity is confirmed through voiceprint recognition, and sent to the smart appliance control module for execution.
[0073] In another possible embodiment, the above-mentioned electrical appliance voiceprint recognition management system performs a binding operation on the identified target user and the smart water heater that is currently performing the binding operation, and the corresponding operating setting values of the bound smart water heater include but are not limited to parameters such as water temperature and flow. Generally speaking, the operating setting values of the above-mentioned smart water heater can be constructed through the user's first setting, system learning or historical behavior records, and are automatically called each time the user's identity is confirmed through voiceprint recognition, and sent to the control module attached to the smart water heater for execution.
[0074] Through the above method steps, the above appliance voiceprint recognition management system can dynamically adjust the operation strategy based on user preferences, achieve differentiated services, enhance personalized experience and optimize energy efficiency.
[0075] Optionally, in the step of calling the corresponding configuration parameters according to the target user and controlling the electrical appliance to perform parameter configuration, it is also included that if the corresponding voiceprint data cannot be matched in the database, the target user is determined to be a new user, and according to the voice interaction guidance, the new user is guided to enter the corresponding personalized configuration parameters item by item, and the personalized configuration parameters are bound to their voiceprint features and stored.
[0076] In an embodiment of the present invention, the personalized configuration parameters may include but are not limited to at least one or more parameters during the use of the appliance, such as water temperature, water volume, and duration.
[0077] In this embodiment, when the above-mentioned appliance voiceprint recognition management system performs database matching on the currently collected voiceprint data, if no matching record is found in the pre-stored voiceprint templates, the system will automatically determine that the current speaking user is a new user and will initiate a voice interaction guidance process to guide the new user to enter their desired usage parameters item by item, including but not limited to the target water outlet temperature, water flow rate, maximum usage time, and whether to turn on the energy-saving mode. Among them, the user input can be collected step by step through voice question and answer, for example, through natural voice prompts such as "Please tell me what temperature you like the water?" and the validity of the parameters entered by the user is confirmed in real time. After all configuration parameters are collected, the above-mentioned appliance voiceprint recognition management system binds the above-mentioned input information with the user's current voiceprint feature data for storage, and writes it into the database of the appliance voiceprint recognition management system to establish a complete user profile for the new user. Thereafter, when the user again issues the wake-up word and is identified as a "registered user" by the above-mentioned appliance voiceprint recognition management system through voiceprint recognition, its personalized configuration parameters will be automatically called, and the appliance will be driven to execute the corresponding settings, realizing an automated control process that does not require repeated settings.
[0078] Through the above method and steps, not only the problem of lack of personalized parameters in the first-time use scenario is solved, but also the integrity and consistency of subsequent recognition are ensured, and the intelligence level of user experience and system adaptability are improved.
[0079] like Figure 3 As shown, an embodiment of the present invention further provides an appliance voiceprint recognition and management device 300, which includes: A first determining module 301 is configured to determine a preset wake-up word, where the preset wake-up word is used to activate a voiceprint recognition function of the appliance; The first collection module 302 is used to collect voiceprint data at the current startup moment when the voiceprint recognition function is activated; The first calling module 303 is used to call a corresponding execution strategy based on the voiceprint data and control the appliance to execute the corresponding strategy.
[0080] Optionally, the above device further includes: The second collection module is used to collect the current voice audio data when setting the preset wake-up word; A third acquisition module is used to extract voiceprint features from the human voice audio data to obtain corresponding voiceprint feature data; a matching module, configured to bind the voiceprint feature data to the preset wake-up word, and when the preset wake-up word is detected, match it with the corresponding voiceprint feature data to obtain a matching result; The second determining module is used to determine the activation state of the voiceprint recognition function of the electrical appliance based on the matching result.
[0081] Optionally, the first acquisition module 302 includes: The first collection submodule is used to collect the surrounding sound data of the appliance at the current moment in real time when the voiceprint recognition function is activated. The surrounding sound data includes environmental audio data and human voice audio data; A noise reduction submodule, configured to perform noise reduction processing on the surrounding sound data to obtain pure human voice audio data; The first extraction submodule is used to perform voiceprint feature extraction processing on the pure human voice audio data to obtain corresponding voiceprint data.
[0082] Optionally, the above device further includes: a decomposition module, configured to decompose the human voice audio data to obtain a plurality of human voice audio data; a processing module, configured to perform audio feature processing on the plurality of human voice audio data to determine corresponding audio feature data, wherein the audio feature data includes volume, pitch, and frequency data; The third determination module is used to determine target human voice audio data from multiple human voice audio data based on the volume, pitch and frequency data, and use it as pure human voice audio data.
[0083] Optionally, the first calling module 303 includes: A third determination submodule is configured to perform matching in a database based on the voiceprint data to determine a corresponding target user; The control submodule is used to call corresponding configuration parameters according to the target user and control the electrical appliance to execute parameter configuration.
[0084] Optionally, the above device further includes: The guidance module is used to determine that the target user is a new user if no corresponding voiceprint data is matched in the database, and guide the new user to enter the corresponding personalized configuration parameters item by item based on voice interaction guidance, and bind and store the personalized configuration parameters with their voiceprint characteristics. The personalized configuration parameters include at least one of water temperature, water volume, and duration.
[0085] like Figure 4 As shown, an embodiment of the present invention further provides an electronic device 400, including a processor, and the processor can execute any of the above-mentioned electrical appliance voiceprint recognition and management methods.
[0086] Specifically, the system includes a processor 401, a memory 402, and a computer program for executing the method for managing the voiceprint recognition of an electrical appliance, which is stored in the memory 402 and can be run on the processor 401, wherein: The processor 401 runs the computer program of the appliance voiceprint recognition management method stored in the memory 402 and performs the following steps: Determining a preset wake-up word, where the preset wake-up word is used to activate a voiceprint recognition function of the appliance; When the voiceprint recognition function is activated, the voiceprint data at the current activation moment is collected; Based on the voiceprint data, a corresponding execution strategy is activated and the electrical appliance is controlled to execute the corresponding strategy.
[0087] Optionally, the processor 401 executes the determining of the preset wake-up word, and the method further includes: When setting a preset wake-up word, collect the current voice audio data of the setting; Extracting voiceprint features from the human voice audio data to obtain corresponding voiceprint feature data; Binding the voiceprint feature data to the preset wake-up word, and when the preset wake-up word is detected, matching it with the corresponding voiceprint feature data to obtain a matching result; Based on the matching result, the activation state of the voiceprint recognition function of the electrical appliance is determined.
[0088] Optionally, the processor 401 executes the step of collecting the voiceprint data at the current moment when the voiceprint recognition function is activated, including: When the voiceprint recognition function is activated, the surrounding sound data of the appliance at the current moment is collected in real time, and the surrounding sound data includes environmental audio data and human voice audio data; Performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data; The pure human voice audio data is subjected to voiceprint feature extraction processing to obtain corresponding voiceprint data.
[0089] Optionally, the processor 401 performs the noise reduction processing on the surrounding sound data to obtain pure human voice audio data, and the method further includes: Decomposing the human voice audio data to obtain multiple human voice audio data; Performing audio feature processing on the plurality of human voice audio data to determine corresponding audio feature data, the audio feature data including volume, pitch, and frequency data; Based on the volume, pitch and frequency data, target human voice audio data is determined from a plurality of human voice audio data and is used as pure human voice audio data.
[0090] Optionally, the processor 401 further executes the method of mobilizing a corresponding execution strategy based on the voiceprint data and controlling the appliance to execute the corresponding strategy, including: Based on the voiceprint data, matching is performed in the database to determine the corresponding target user; According to the target user, the corresponding configuration parameters are called, and the electrical appliance is controlled to execute the parameter configuration.
[0091] Optionally, the processor 401 further executes the calling of corresponding configuration parameters according to the target user and controls the electrical appliance to perform parameter configuration. The method further includes: If the corresponding voiceprint data cannot be matched in the database, the target user is determined to be a new user, and according to the voice interaction guidance, the new user is guided to enter the corresponding personalized configuration parameters item by item, and the personalized configuration parameters are bound to their voiceprint characteristics and stored. The personalized configuration parameters include at least one of water temperature, water volume, and duration.
[0092] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various processes of the electrical appliance voiceprint recognition management method or the application-end electrical appliance voiceprint recognition management method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0093] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program that instructs related hardware to perform the process, and can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0094] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A method for managing voiceprint recognition of electrical appliances, characterized in that: include: Determining a preset wake-up word, where the preset wake-up word is used to activate a voiceprint recognition function of the appliance; When the voiceprint recognition function is activated, the voiceprint data at the current activation moment is collected; Based on the voiceprint data, a corresponding execution strategy is activated and the electrical appliance is controlled to execute the corresponding strategy.
2. The method for managing the voiceprint recognition of an electrical appliance according to claim 1, wherein: The method of determining a preset wake-up word further includes: When setting a preset wake-up word, collect the current voice audio data of the setting; Extracting voiceprint features from the human voice audio data to obtain corresponding voiceprint feature data; Binding the voiceprint feature data to the preset wake-up word, and when the preset wake-up word is detected, matching it with the corresponding voiceprint feature data to obtain a matching result; Based on the matching result, the activation state of the voiceprint recognition function of the electrical appliance is determined.
3. The method for managing the voiceprint recognition of an electrical appliance according to claim 1, wherein: When the voiceprint recognition function is activated, the voiceprint data at the current moment is collected, including: When the voiceprint recognition function is activated, the surrounding sound data of the appliance at the current moment is collected in real time, and the surrounding sound data includes environmental audio data and human voice audio data; Performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data; The pure human voice audio data is subjected to voiceprint feature extraction processing to obtain corresponding voiceprint data.
4. The method for managing the voiceprint recognition of an electrical appliance according to claim 3, wherein: The method further comprises: performing noise reduction processing on the surrounding sound data to obtain pure human voice audio data; Decomposing the human voice audio data to obtain multiple human voice audio data; Performing audio feature processing on the plurality of human voice audio data to determine corresponding audio feature data, the audio feature data including volume, pitch, and frequency data; Based on the volume, pitch and frequency data, target human voice audio data is determined from a plurality of human voice audio data and is used as pure human voice audio data.
5. The method for managing the voiceprint recognition of an electrical appliance according to claim 1, wherein: The mobilizing a corresponding execution strategy based on the voiceprint data and controlling the electrical appliance to execute the corresponding strategy includes: Based on the voiceprint data, matching is performed in the database to determine the corresponding target user; According to the target user, the corresponding configuration parameters are called, and the electrical appliance is controlled to execute the parameter configuration.
6. The method for managing the voiceprint recognition of an electrical appliance according to claim 5, wherein: The method further includes calling corresponding configuration parameters according to the target user and controlling the electrical appliance to perform parameter configuration. If no corresponding voiceprint data is matched in the database, the target user is determined to be a new user, and according to the voice interaction guidance, the new user is guided to enter the corresponding personalized configuration parameters item by item, and the personalized configuration parameters are bound to their voiceprint features and stored.
7. An electrical appliance voiceprint recognition and management device, characterized in that: include: a first determining module, configured to determine a preset wake-up word, wherein the preset wake-up word is used to activate a voiceprint recognition function of the appliance; The first collection module is used to collect the voiceprint data at the current startup moment when the voiceprint recognition function is started; The first calling module is used to mobilize the corresponding execution strategy based on the voiceprint data and control the electrical appliance to execute the corresponding strategy.
8. An electrical appliance voiceprint recognition and management system, characterized in that: The appliance voiceprint recognition management system includes: an appliance voiceprint recognition management device; The appliance voiceprint recognition and management device implements the appliance voiceprint recognition and management method described in claim 1.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for managing the voiceprint recognition of an electrical appliance as claimed in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for managing voiceprint recognition of electrical appliances according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
AI partner terminal system and method based on cloud edge collaboration
CN121924140A