Electric cooker voice safety control system and method based on voiceprint verification

By using a multimodal sensing array and a hierarchical collaborative authentication mechanism, the security risks of existing voice control systems under recording replay attacks are resolved, achieving effective defense against recording replay attacks and accurate identification of user identities.

CN121506151APending Publication Date: 2026-02-10LINGNAN NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511676426.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing voice control systems lack effective defenses against recording and playback attacks, and single authentication methods have low recognition accuracy in complex environments, making it difficult to ensure security.

Method used

A multimodal sensing array is used to simultaneously acquire acoustic signals and structural vibration signals. By constructing device physical fingerprints and multidimensional feature templates, a hierarchical collaborative authentication mechanism is implemented, including spatial domain verification, physical interaction verification, and user identity verification. The acoustic-vibration coupling characteristics are used to identify non-live attacks.

Benefits of technology

It improves the security and reliability of the voice control system, accurately distinguishes between genuine voice commands from authorized users and recorded playback attacks, and enhances the defense against unauthorized control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506151A_ABST
    Figure CN121506151A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent household electrical appliance control and biological characteristic recognition, and discloses an electric cooker voice safety control system and method based on voiceprint verification, the system comprises a multi-mode sensing array, a storage unit and a core processing unit, the method comprises the following steps: generating and storing equipment physical fingerprints representing unique physical characteristics of equipment, the method comprises the following steps: generating and storing a multi-dimensional feature template comprising an acoustic voiceprint, living body interaction and a spatial vector template for an authorized user, when a voice instruction is received, synchronously acquiring a to-be-detected acoustic and structural vibration signal by a system, executing layered collaborative authentication, and sequentially performing spatial domain verification, physical interaction verification and user identity verification, the physical interaction verification identifies the living body attribute of the sound source by comparing the real-time sound vibration coupling characteristic with the equipment physical fingerprint, so that the non-living body attack is effectively identified and intercepted, and the safety and reliability of voice control are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart home appliance safety control technology, specifically to a voice-based safety control system and method for a rice cooker. Background Technology

[0002] With the development of smart home technology, voice control has been widely applied to various household appliances. Rice cookers, as one of the core kitchen appliances, are also beginning to integrate voice interaction functions to facilitate user operation. Users can control the rice cooker's cooking mode, timer, and other functions through voice commands, simplifying the operation process.

[0003] However, the convenience of voice control also brings security challenges. To prevent unauthorized access, especially when it involves security-related operations such as starting cooking, it is necessary to authenticate the user issuing the command. Currently, the mainstream authentication technology is based on voiceprint recognition, which determines the user's identity by analyzing the acoustic characteristics of the user's voice signal.

[0004] Existing authentication mechanisms that rely solely on acoustic voiceprints have inherent limitations. These mechanisms identify users from a single acoustic dimension, and their discriminative power decreases when faced with high-quality voice imitations or in complex home noise environments, leading to lower authentication accuracy. More seriously, these mechanisms offer virtually no defense against playback attacks; attackers can easily bypass security verification by playing pre-recorded audio from authorized users, thereby gaining unauthorized control of the device.

[0005] To combat playback attacks, some technical solutions have incorporated liveness detection. However, existing liveness detection methods typically rely on analyzing the received signal for subtle distortions introduced by the playback device or environmental reflection artifacts. Their detection performance is limited by the fidelity of the recording and playback devices and is easily affected by varying indoor acoustic environments. These methods lack a stable and objective judgment benchmark, resulting in insufficient adaptability to different application scenarios.

[0006] Furthermore, existing voice control systems typically employ a simplistic security authentication process, lacking a structured, multi-layered defense strategy. These systems often treat user identification and liveness detection as a single unit, failing to layer and progressively verify the spatial physical information of the sound source, the physical interaction characteristics with the device, and the user's biometrics. This unstructured authentication method makes it insufficiently robust against sophisticated deceptive attacks, struggling to prevent the execution of unauthorized commands in all situations. Summary of the Invention

[0007] The purpose of this invention is to provide a voice-based security control system and method for rice cookers, which aims to solve the security risks of existing voice control technologies in terms of identity authentication, especially the problem of insufficient defense against non-liveness attack methods such as recording and playback.

[0008] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of the present invention provides a voice safety control system for rice cookers based on voiceprint verification. The system includes a multimodal sensing array, a storage unit, and a core processing unit.

[0009] The multimodal sensing array is used to simultaneously acquire acoustic signals and structural vibration signals. The array includes an acoustic sensor array for acquiring acoustic signals and a structural vibration sensor for acquiring vibrations of the equipment housing caused by acoustic excitation or physical contact with the user.

[0010] The storage unit is used to store the device physical fingerprint that characterizes the physical characteristics of the rice cooker itself, as well as the multi-dimensional feature templates of one or more authorized users.

[0011] The core processing unit is connected to the multimodal sensing array and the storage unit, and its function is as follows: First, a physical fingerprint of the device is generated and stored. This physical fingerprint is formed based on the device's inherent and unique physical structural characteristics (such as mass, damping, and stiffness distribution). The core processing unit applies a broadband excitation signal to the rice cooker shell through the active excitation unit and simultaneously acquires the structural vibration response signal and the outwardly radiated acoustic response signal caused by this excitation. Using a frequency domain system identification algorithm, the transfer function from structural vibration to acoustic radiation is calculated. This transfer function reflects the unique way the device converts structural vibration energy into acoustic energy. Its calculation can be based on the Welch method, obtained by averaging the power spectral density of multiple signal frames. ; in: For frequency variables, To characterize the structural vibration response signal to acoustic response signal The transfer function of the transformation relationship. Structural vibration response signal Harmony and acoustic response signals Cross-power spectral density between Structural vibration response signal The self-power spectral density.

[0012] A set of stable frequency domain features (such as the frequency of the resonant peak, gain, and quality factor) extracted from the amplitude or phase spectrum of the transfer function constitutes the physical fingerprint of the device.

[0013] Secondly, multi-dimensional feature templates are generated and stored for authorized users. These templates aim to uniquely represent the identity characteristics of authorized users from multiple dimensions, including acoustic voiceprint templates, liveness interaction templates, and spatial vector templates.

[0014] Acoustic voiceprint templates are acoustic feature vectors extracted from a user's speech signal through a deep learning network (such as an x-vector network) that can characterize their vocal tract and vocal habits.

[0015] The live interaction template characterizes the unique physical coupling characteristics between the voice energy of the authorized user and the rice cooker shell when the user actually speaks. This characteristic includes complex acoustic vibration coupling information through multiple paths such as air conduction and bone conduction, which is a key feature that distinguishes it from non-living sound sources such as loudspeakers.

[0016] The template is generated similarly to a device's physical fingerprint. It calculates the transfer function between the first acoustic digital signal and the structural vibration digital signal when the user speaks using a frequency domain system identification algorithm, and extracts key frequency domain features from it. The spatial vector template is used to characterize the authorized user's general spatial orientation when speaking. The time difference of arrival (TDOA) is calculated by simultaneously acquiring the first and second acoustic digital signals from the first and second microphones in the acoustic sensor array. This calculation is implemented using the Generalized Cross-Correlation-Phase Transform (GCC-PHAT) algorithm, which effectively suppresses reverberation and improves the accuracy of time delay estimation. The calculation process is as follows: ; in: For time delay variables, for and The generalized cross-correlation function between them based on phase transformation. For signal Fourier transform, For signal Fourier transform, for The complex conjugate, The normalized cross-power spectral density is used to weight the phase information of the signal. For the complex exponential kernel of the Fourier transform, This is the inverse Fourier transform operation, used to convert the normalized cross-power spectral density in the frequency domain back to the correlation function in the time domain. .

[0017] By searching for the aforementioned generalized cross-correlation function The peak value determines the unique time delay estimate. : ; in: For estimated user latency, To find a generalized cross-correlation function The time delay variable that achieves its maximum value The operation.

[0018] Secondly, real-time collaborative authentication is performed upon receiving a voice command. This authentication is a layered gating process that sequentially verifies the spatial domain, physical interaction, and user identity.

[0019] Spatial domain verification determines the validity of a sound source location by comparing the real-time observed time difference of arrival with the stored spatial vector template.

[0020] Physical interaction verification is one of the core defense mechanisms of this solution. It calculates the liveness interaction features from the signal under test and compares them with the stored device physical fingerprint. Since non-liveness attacks such as audio playback cannot reproduce the physical process of acoustic-vibration coupling of a real user, the liveness interaction features under test will show a significant mismatch with the device physical fingerprint, thus being effectively identified and blocked. This verification uses a cosine similarity algorithm to calculate the similarity between the two.

[0021] User identity verification is performed after the first two levels of verification, and the decision is made by combining information from both sources: The first step is to calculate the acoustic similarity score between the acoustic feature vector to be tested and the acoustic voiceprint template of the authorized user using the cosine similarity algorithm. Second, the cosine similarity algorithm is used to calculate the liveness similarity score between the liveness interaction features of the test subject and the liveness interaction template of the authorized user. Then, the two scores are weighted and fused to obtain the final identity authentication score. ; in: and The similarity calculation function, in this embodiment, specifically refers to the cosine similarity algorithm.

[0022] and The preset weighting coefficients, and .

[0023] Authentication is considered successful only when the final total authentication score is not lower than the preset authentication threshold.

[0024] A second aspect of the present invention provides a voice-based safety control method for a rice cooker, the method comprising the following steps: First, a device physical fingerprint, uniquely representing the physical characteristics of the rice cooker, is generated and stored. This step involves active excitation and synchronization signal acquisition, and the use of a frequency domain system identification algorithm to calculate the device's physical transfer function to extract features.

[0025] Secondly, a unique multi-dimensional feature template is generated and stored for authorized users. The generation process of this template includes: extracting an acoustic voiceprint template that can represent the user's identity; calculating and extracting a liveness interaction template that can represent the physical coupling characteristics between the user and the device; and calculating and storing a spatial vector template that can represent the user's spatial orientation.

[0026] Furthermore, upon receiving a voice command, a multi-stage collaborative authentication decision is executed. This step first simultaneously acquires the acoustic signal and structural vibration signal to be tested, and then sequentially performs spatial domain verification, physical interaction verification, and user identity verification. This multi-stage decision-making process ensures that authentication is only successful when the sound source location, physical interaction characteristics, and user biometrics all conform to the authorized user model.

[0027] Finally, after the multi-stage collaborative authentication decision is passed, the content of the voice command is parsed, and the corresponding device control command is generated and executed.

[0028] This invention constructs a physical fingerprint of the device and a multi-dimensional feature template of the user, and implements a three-layer progressive collaborative authentication mechanism of space, physical and identity, which can accurately distinguish between the real voice commands of authorized users and malicious recording and playback attacks, thereby improving the security and reliability of the voice control system.

[0029] This invention provides a voice-based safety control system and method for rice cookers. It has the following beneficial effects: 1. This invention introduces a multimodal sensing array to simultaneously acquire acoustic and structural vibration signals, and adds a physical interaction verification step to the authentication process. This verification step identifies the liveness attribute of the sound source by judging whether the acoustic-vibration coupling phenomenon generated by the voice command under test conforms to the inherent physical characteristics of the device. Since recording and playback devices can only produce air-conducted sound, their acoustic-vibration coupling characteristics have obvious and measurable physical differences from live voice generation, which includes bone conduction. Therefore, this invention can identify and intercept such non-liveness attacks, solving the problem that existing voice control systems lack effective defense against playback attacks and improving system security.

[0030] 2. This invention establishes a strong hardware-bound authentication mechanism by generating and storing a unique physical fingerprint for each rice cooker during the system initialization phase. This physical fingerprint is formed based on the unique physical structural characteristics of the device due to manufacturing tolerances. During the physical interaction verification phase, the acoustic-vibration coupling characteristics to be tested must match the physical fingerprint of that specific device, rather than a generic model. This mechanism ensures the uniqueness and non-portability of the authentication model, increasing the difficulty of large-scale, generalized attacks targeting a particular model of device.

[0031] 3. This invention improves the accuracy and robustness of authentication decisions by constructing a multi-dimensional feature template that includes spatial vectors, liveness interaction features, and acoustic voiceprints, and implementing a hierarchical gating collaborative authentication process. This process sequentially verifies the spatial location of the sound source, physical liveness attributes, and user biometrics, utilizing the complementarity of information from different dimensions. In the final user identity verification, acoustic similarity scores and liveness interaction similarity scores are integrated, reducing the false positive rate caused by fluctuations in single features under complex environments such as noise, thus achieving higher authentication reliability than a single authentication dimension. Attached Figure Description

[0032] Figure 1 This is a structural block diagram of the voice security control system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system functional modules according to an embodiment of the present invention; Figure 3 This is a flowchart of the device physical fingerprint self-calibration and modeling method according to an embodiment of the present invention; Figure 4 This is a flowchart of the authorized user multi-dimensional feature template registration method according to an embodiment of the present invention; Figure 5 This is a flowchart of the real-time voice command authentication and execution method according to an embodiment of the present invention.

[0033] Among them, 101 is the core processing unit; 102 is the multimodal sensing array; 102a is the acoustic sensor array; 102b is the structural vibration sensor; 103 is the active excitation unit; 104 is the storage unit; 201 is the signal synchronous acquisition module; 202 is the equipment physical fingerprint modeling module; 203 is the user template registration module; 204 is the real-time collaborative authentication engine; and 205 is the instruction parsing and control module. Detailed Implementation

[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] See attached document Figure 1 The present invention provides a voice safety control system for rice cookers based on voiceprint verification. The system may include: a core processing unit 101, a multimodal sensing array 102, an active excitation unit 103, and a storage unit 104.

[0036] The core processing unit 101 is used to execute preset program instructions to control the coordinated operation of various system components, thereby realizing device physical fingerprint modeling, user multi-dimensional feature template registration, and real-time collaborative authentication. In one embodiment, the core processing unit 101 is a microcontroller (MCU) or an embedded microprocessor.

[0037] The core processing unit 101 possesses the computational power required to execute digital signal processing algorithms and integrates a multi-channel synchronous analog-to-digital converter (ADC) module, a timer module, and a general-purpose input / output interface. The core processing unit 101 is electrically connected and communicates with the multimodal sensing array 102, the active excitation unit 103, and the storage unit 104 via a data bus or a dedicated interface.

[0038] In one specific embodiment, the core processing unit 101 communicates with the acoustic sensor array 102a via an I2S (Inter-ICSound) bus to transmit digital audio signals; communicates with the structural vibration sensor 102b via an SPI (Serial Peripheral Interface) or I2C (Inter-Integrated Circuit) bus; drives and controls the active excitation unit 103 via a PWM (Pulse Width Modulation) signal; and reads and writes data with the storage unit 104 via an SPI bus.

[0039] The multimodal sensing array 102, electrically connected to the core processing unit 101, is used to synchronously acquire acoustic signals and structural vibration signals, and convert the acquired analog signals into digital signals before transmitting them to the core processing unit 101. The multimodal sensing array 102 includes an acoustic sensor array 102a and a structural vibration sensor 102b.

[0040] The acoustic sensor array 102a includes at least one first microphone and one second microphone. Both the first and second microphones are microelectromechanical systems (MEMS) microphones, fixedly mounted on the same plane of the rice cooker device, such as the device's control panel, and the geometric center distance between them is a preset distance. The first and second microphones are used to convert the received airborne sound pressure signals into first and second acoustic electrical signals, and output them to the analog-to-digital conversion channels of the core processing unit 101, respectively.

[0041] The structural vibration sensor 102b includes at least one piezoelectric ceramic sensor or microelectromechanical system (MEMS) accelerometer. The structural vibration sensor 102b is rigidly fixed to a predetermined structural node on the rice cooker shell. This predetermined structural node is a location highly sensitive to vibrations generated by external acoustic excitation. The structural vibration sensor 102b converts the mechanical vibration of the rice cooker shell into a structural vibration electrical signal and outputs it to the analog-to-digital conversion channel of the core processing unit 101.

[0042] The active excitation unit 103 is electrically connected to the core processing unit 101 and is used to generate a preset mechanical vibration according to the excitation control signal from the core processing unit 101 and apply it to the rice cooker shell.

[0043] In one embodiment, the active excitation unit 103 is a miniature actuator, such as a piezoelectric buzzer or a linear resonant actuator, which is fixed to the inner wall of the rice cooker shell. In another embodiment, the active excitation unit 103 reuses the same physical device as the structural vibration sensor 102b. When the structural vibration sensor 102b is a piezoelectric ceramic sensor, the core processing unit 101 applies a driving voltage signal to the piezoelectric ceramic sensor during the execution of the device physical fingerprint self-calibration process, causing it to generate mechanical vibration as an actuator.

[0044] Storage unit 104, electrically connected to core processing unit 101, is used to store program instructions, device physical fingerprints, and multidimensional feature templates of one or more authorized users required to execute the method of the present invention. In one embodiment, storage unit 104 is a non-volatile memory, such as flash memory or electrically erasable programmable read-only memory (EEPROM).

[0045] The system in this embodiment of the invention may further include a signal conditioning circuit. The signal conditioning circuit is disposed between the multimodal sensing array 102 and the core processing unit 101, and is used to amplify, filter and perform other processing on the first acoustic electrical signal, the second acoustic electrical signal and the structural vibration electrical signal to improve the signal-to-noise ratio, and to send the processed signal to the analog-to-digital converter inside the core processing unit 101.

[0046] See attached document Figure 2 The functions performed by the system can be implemented by one or more of the following modules, which can be functional units that exist in the form of software, firmware or hardware and are executed by or integrated into the core processing unit 101.

[0047] The signal synchronization acquisition module 201 is used to acquire time-aligned data from multiple signal channels of the multimodal sensing array 102. The signal synchronization acquisition module 201 is configured to synchronously initiate the analog-to-digital conversion process of the first acoustic electrical signal, the second acoustic electrical signal, and the structural vibration electrical signal in response to internal or external triggering events.

[0048] In one specific embodiment, the signal synchronization acquisition module 201 generates a synchronization clock signal using a hardware timer module within the core processing unit 101. This synchronization clock signal is simultaneously routed to the sample-and-hold circuits controlling multiple analog-to-digital converter (ADC) channels, thereby ensuring that signal sampling operations for all channels occur simultaneously at each sampling moment. In this way, the digital signal stream output by the module achieves strict time alignment at the sample points, providing input data with a deterministic time reference for subsequent transfer function calculations and time difference of arrival estimation.

[0049] The device physical fingerprint modeling module 202 is used to generate and store a device physical fingerprint that uniquely represents the physical characteristics of the rice cooker during the system initialization phase. The function of the device physical fingerprint modeling module 202 is achieved by executing a series of preset control and calculation operations.

[0050] First, the device physical fingerprint modeling module 202 outputs a preset excitation control signal to the active excitation unit 103, driving it to generate a broadband excitation signal. The spectral energy of this broadband excitation signal covers the main frequency band of the speech signal.

[0051] Secondly, during the application of the excitation signal, the equipment physical fingerprint modeling module 202 calls the signal synchronization acquisition module 201 to simultaneously capture the structural vibration response signal and acoustic response signal caused by the excitation.

[0052] Next, the equipment physical fingerprint modeling module 202 processes the acquired structural vibration response signal and acoustic response signal to calculate the transfer function between them. This calculation process is implemented using a frequency domain system identification algorithm, which specifically includes framing the signal, windowing it, performing a fast Fourier transform (FFT) to obtain a frequency domain representation, and obtaining the transfer function by calculating the quotient of the cross power spectral density and the self power spectral density.

[0053] Finally, the device physical fingerprint modeling module 202 extracts one or more stable frequency domain features from the calculated transfer function, such as the center frequency, peak gain, or quality factor Q of one or more resonant peaks. The device physical fingerprint modeling module 202 employs a peak detection algorithm, such as a threshold-based or local maximum search algorithm, to automatically identify obvious resonant peaks in the transfer function amplitude spectrum. For each identified resonant peak, its center frequency, peak gain, and quality factor Q are extracted. The feature parameters of all identified resonant peaks are sequentially concatenated to form a variable-dimensional feature vector, which is defined as the device physical fingerprint. These extracted features together constitute a feature vector, which is defined as the device physical fingerprint and stored in the storage unit 104 by the device physical fingerprint modeling module 202.

[0054] See attached document Figure 2 The functions performed by this system may also include: The user template registration module 203 is used to generate and store a unique multi-dimensional feature template for each authorized user when one or more authorized users register. After receiving a user registration instruction, the user template registration module 203 calls the signal synchronization acquisition module 201 to acquire the first acoustic digital signal, the second acoustic digital signal, and the structural vibration digital signal corresponding to the user's sound emission at a preset position.

[0055] Subsequently, the user template registration module 203 processes the acquired signals in parallel to generate a composite template containing three sub-templates: Acoustic voiceprint template: The user template registration module 203 extracts a fixed-dimensional acoustic feature vector from the first acoustic digital signal through a preset voiceprint feature extraction network (e.g., an x-vector network). This acoustic feature vector is defined as the acoustic voiceprint template.

[0056] Live Interaction Template: This module uses a frequency domain system identification algorithm to calculate the transfer function between the first acoustic digital signal and the structural vibration digital signal to characterize the physical interaction between the user and the device. The calculation process of this transfer function is consistent with the method in device physical fingerprint modeling, both based on the analysis of signal power spectral density. The key frequency domain features of this transfer function are extracted and defined as the live interaction template.

[0057] Spatial Vector Template: The user template registration module 203 determines the user's spatial orientation information by calculating the time difference of arrival between the first and second acoustic digital signals. The user template registration module 203 uses the Generalized Cross-Correlation-Phase Transform (GCC-PHAT) algorithm to calculate the cross-correlation function between the two signals and determines the time delay variable τ corresponding to the peak value of this function. This time delay variable τ is defined as the spatial vector template.

[0058] Finally, the user template registration module 203 associates the generated acoustic voiceprint template, live interaction template, and spatial vector template as a whole with the corresponding user identity information and stores them in the storage unit 104.

[0059] The real-time collaborative authentication engine 204 is used to perform real-time, multi-stage collaborative authentication of the legitimacy of the voice command issuer upon receiving a voice command. The engine first calls the signal synchronization acquisition module 201 to acquire the acoustic and vibration signals to be tested, and then uses the same processing method as the user template registration module 203 to extract the acoustic feature vector, the liveness interaction features, and the spatial vector to be tested.

[0060] The engine then executes a layered gating authentication process: Spatial domain verification: The engine first compares the spatial vector to be tested with the authorized user spatial vector template stored in storage unit 104. Only when the difference between the two is within the preset spatial tolerance threshold can the authentication process proceed to the next stage.

[0061] Physical interaction verification: After spatial domain verification, the engine calculates the similarity between the liveness interaction features to be tested and the device physical fingerprint stored in storage unit 104 using a cosine similarity algorithm. Only when the calculated similarity is not lower than the preset physical consistency threshold can the authentication process proceed to the next stage.

[0062] User Identity Verification: After physical interaction verification, the engine first calculates the similarity between the acoustic feature vector to be tested and the acoustic voiceprint template of the authorized user using a cosine similarity algorithm, obtaining an acoustic similarity score. Simultaneously, it calculates the similarity between the liveness interaction feature to be tested and the liveness interaction template of the authorized user using a cosine similarity algorithm, obtaining a liveness interaction similarity score. Subsequently, the engine weighted and fused the two similarity scores to obtain a final identity authentication score. Only when this total score is not lower than a preset identity authentication threshold is the user's identity authentication considered successful.

[0063] If any stage of the verification fails, the real-time collaborative authentication engine 204 immediately interrupts the authentication process and refuses to execute the instruction. Upon successful authentication, the engine outputs an authentication success signal.

[0064] The instruction parsing and control module 205 is used to parse the content of the voice instruction and generate the corresponding device control instruction after receiving the authentication success signal from the real-time collaborative authentication engine 204.

[0065] Upon receiving a successful authentication signal, the command parsing and control module 205 is activated and initiates an automatic speech recognition (ASR) process on the acquired first acoustic digital signal, converting the speech signal into a text string. Subsequently, the command parsing and control module 205 matches the recognized text string with a preset command list to determine the user's intent.

[0066] After determining the user's intent, the instruction parsing and control module 205 converts the intent into a control instruction that conforms to the communication protocol of the rice cooker's main control system, and sends the control instruction to the main control system through an internal interface, thereby realizing safe control of the rice cooker's functions.

[0067] See attached document Figure 3 The initialization and template registration process of the method of the present invention includes a self-calibration and modeling step of the device physical fingerprint, and the detailed implementation of this step is as follows: In step S301, the device physical fingerprint modeling module 202 generates and outputs an excitation control signal to the active excitation unit 103. In one embodiment, the excitation control signal is used to drive the active excitation unit 103 to generate a linear sweep sine wave as mechanical vibration. The frequency range of the linear sweep sine wave covers a preset minimum frequency. Up to the highest frequency For example, from 50Hz to 8000Hz, to cover the main frequency band of human speech. The duration of this signal is a preset length. .

[0068] In step S302, while the active excitation unit 103 applies mechanical vibration, the equipment physical fingerprint modeling module 202 calls the signal synchronization acquisition module 201 to synchronously acquire the structural vibration response signal from the structural vibration sensor 102b caused by the mechanical vibration. and the acoustic response signal from the first microphone in the acoustic sensor array 102a. .in, It is a time variable.

[0069] In step S303, the equipment physical fingerprint modeling module 202 analyzes the collected structural vibration response signals. Harmony and acoustic response signals Preprocessing is performed. This preprocessing includes bandpass filtering of the signal to remove noise outside a preset frequency range; and framing and windowing (e.g., Hamming windowing) of the signal in preparation for subsequent frequency domain analysis.

[0070] In step S304, the device physical fingerprint modeling module 202 calculates the transfer function from structural vibration to acoustic radiation. The transfer function is calculated using the Welch method, which obtains a robust estimate by averaging the power spectral density of multiple signal frames. Its calculation formula is as follows: ; in: For frequency variables, To characterize the structural vibration response signal to acoustic response signal The transfer function of the transformation relationship. Structural vibration response signal Harmony and acoustic response signals Cross-power spectral density between Structural vibration response signal The self-power spectral density.

[0071] In step S305, the device physical fingerprint modeling module 202 obtains the calculated transfer function Extract a set of stable frequency domain features to construct the device's physical fingerprint. In one embodiment, the device physical fingerprint modeling module 202 first analyzes the amplitude spectrum of the transfer function. Perform peak detection to identify A distinct resonance peak. Subsequently, regarding the first... One resonance peak (of which) Extract its center frequency. Peak gain and quality factor These extracted features are organized into a feature vector, which is the device's physical fingerprint. : ; In step S306, the device physical fingerprint modeling module 202 generates the device physical fingerprint. It is stored in storage unit 104 for subsequent use by the real-time collaborative authentication engine 204 when performing physical interaction verification.

[0072] See attached document Figure 4 This registration step can be performed after the device physical fingerprint self-calibration and modeling step, or it can be performed independently.

[0073] In step S401, the user template registration module 203 is activated in response to the user's registration request. The user template registration module 203 prompts the authorized user through the user interface to repeat a preset registration password in a natural volume and speaking speed at a normal operating position on the rice cooker.

[0074] In step S402, during the user's vocalization, the user template registration module 203 calls the signal synchronization acquisition module 201 to synchronously acquire the first acoustic digital signal from the first microphone in the acoustic sensor array 102a. Second acoustic digital signal from the second microphone and digital signals of structural vibration from structural vibration sensor 102b .in, It is a time variable.

[0075] In step S403, the user template registration module 203 processes the three acquired signals in parallel to generate and construct three independent sub-templates: First, the user template registration module 203 generates an acoustic voiceprint template. The user template registration module 203 registers the first acoustic digital signal. Preprocessing is performed, including silence removal, pre-emphasis, and frame segmentation, and Mel-frequency cepstral coefficients (MFCC) or Fbank features are extracted. Subsequently, the extracted acoustic feature sequence is input into a pre-trained speaker feature extraction network (e.g., an x-vector network). Through the network's forward propagation, a fixed-dimensional embedding vector is extracted from the network's pooling or fully connected layers. The embedding vector That is, the acoustic voiceprint template defined for that user. .

[0076] In a preferred embodiment, the user template registration module 203 prompts the user to repeat the registration password multiple times. The user template registration module 203 extracts an embedding vector from the voice signal of each password. Finally, it performs an arithmetic average or weighted average of the multiple extracted embedding vectors, and uses the final average vector as the user's acoustic voiceprint template. This is to improve the stability and noise resistance of the template.

[0077] ; in, For fixed-dimensional embedding vectors extracted from the network, This is for acoustic voiceprint templates. Secondly, the user template registration module 203 generates live interactive templates. The user template registration module 203 uses a frequency domain system identification algorithm to calculate the first acoustic digital signal. Digital signals of structural vibration Transfer function between This is used to characterize the physical coupling characteristics between the user's actual voice output and the device housing. The transfer function is calculated using the same power spectral density-based method as in step S304. ; Subsequently, the user template registration module 203 employs the same feature extraction method as in step S305 for extracting the device's physical fingerprint, such as using a peak detection algorithm to extract the physical fingerprint from the device. Extract key frequency domain features (such as energy gain or phase characteristics of a specific frequency band) and store these features as a live interaction template. .

[0078] Next, the user template registration module 203 generates a spatial vector template. The user template registration module 203 calculates the first acoustic digital signal. Second acoustic digital signal The time difference of arrival (TDOA) between the two points is used to record the generalized spatial orientation of the user's voice. This calculation is performed using the generalized cross-correlation-phase transform (GCC-PHAT) algorithm: ; in: For time delay variables, for and The generalized cross-correlation function between them based on phase transformation. For signal Fourier transform, For signal Fourier transform, for The complex conjugate, The normalized cross-power spectral density is used to weight the phase information of the signal. For the complex exponential kernel of the Fourier transform, This is the inverse Fourier transform operation, used to convert the normalized cross-power spectral density in the frequency domain back to the correlation function in the time domain. .

[0079] The user template registration module 203 searches for a generalized cross-correlation function. The peak value is used to determine the unique time delay estimate. : ; in: For estimated user latency, To find a generalized cross-correlation function The time delay variable that reaches its maximum value The operation.

[0080] The estimated user latency value Defined as a space vector template In step S404, the user template registration module 203 will generate the acoustic texture template. Live interactive template and space vector template The features are combined into a multidimensional feature template, and the multidimensional feature template is associated with the corresponding user identity identifier and stored in the storage unit 104.

[0081] See attached document Figure 5 This process is triggered when the system receives a user's voice command.

[0082] In step S501, the real-time collaborative authentication engine 204 calls the signal synchronization acquisition module 201 to acquire the signal to be tested. This acquisition process is the same as the signal acquisition process during user registration in step S402, synchronously acquiring the first acoustic digital signal to be tested. The second acoustic digital signal to be tested and the digital signal of the structural vibration to be measured. .in, It is a time variable.

[0083] In step S502, the real-time collaborative authentication engine 204 extracts features from the signal under test. This feature extraction process uses the exact same method as the user template generation in step S403 to ensure that the features under test and the template features are comparable within the same feature space. The engine extracts the following features from the signal under test in parallel: An acoustic feature vector to be measured ; A live interaction feature vector to be tested This vector is obtained by first calculating the live interaction transfer function to be tested. Then, the live interaction template generated in step S403 is used. Obtained using the exact same feature extraction methods (e.g., peak detection and parameter extraction).

[0084] A spatial vector to be measured In step S503, the real-time collaborative authentication engine 204 performs a multi-stage collaborative authentication decision. This decision-making process is a serial, gated verification process, comprising the following consecutive stages: First, in step S503a, spatial domain verification is performed. This step aims to confirm whether the physical location of the sound source matches the location where the authorized user registered. The real-time collaborative authentication engine 204 reads the authorized user's spatial vector template from storage unit 104. (i.e., estimated user latency value) ), and perform the following judgment: ; in: The time delay variable observed from the signal under test. For storage in space vector template The estimated user latency value in the data. For the observed time delay variable Compared with the estimated user latency value The absolute difference between them This is a preset spatial tolerance threshold.

[0085] If the above inequality does not hold, the location is determined to be abnormal, and the authentication process terminates immediately. If the inequality holds, the authentication process proceeds to the next stage.

[0086] Secondly, in step S503b, a physical interaction verification is performed. This step aims to confirm whether the current acoustic-vibration coupling phenomenon conforms to the inherent physical characteristics of the rice cooker device, in order to identify non-liveness attacks. The real-time collaborative authentication engine 204 reads the device's physical fingerprint from the storage unit 104. And perform the following judgment: ; in: The live interaction transfer function to be measured is calculated from the signal to be measured. The device physical fingerprint stored in storage unit 104; For calculating the live interaction transfer function to be tested physical fingerprint of the device The function for similarity between them, in one embodiment, is a cosine similarity algorithm, specifically applied to the physical fingerprint of the device. Within the defined key frequency band, calculate the cosine similarity or correlation coefficient of their amplitude spectra. This is a preset physical consistency threshold.

[0087] If the above inequality does not hold, it is determined that there is a physical interaction anomaly, and the authentication process terminates immediately. If the inequality holds, the authentication process proceeds to the final stage.

[0088] Next, in step S503c, user identity verification is performed. This step aims to finally confirm whether the user who issued the instruction is the authorized user. The real-time collaborative authentication engine 204 reads the authorized user's acoustic voiceprint template from the storage unit 104. Interactive templates for live animals And perform fusion authentication: The engine first calculates the acoustic feature vector to be tested. Acoustic voiceprint template Acoustic similarity score between .

[0089] ; The engine simultaneously calculates the live interaction transfer function to be tested. Interact with live objects template Live interaction similarity score between .

[0090] ; Subsequently, the engine calculates the final identity authentication score through weighted fusion. : ; in: and The similarity calculation function, in this embodiment, specifically refers to the cosine similarity algorithm.

[0091] and The preset weighting coefficients, and .

[0092] Finally, the engine performs the following judgment: ; in, A preset authentication threshold is set. If the inequality is true, the user's identity is deemed to have been successfully authenticated, and an authentication success signal is output in step S504; if the inequality is false, the identity is deemed to be mismatched, and the authentication process terminates.

[0093] It should be noted that the spatial tolerance threshold involved in the embodiments of the present invention Physical consistency threshold and identity authentication threshold These thresholds can be preset based on the device's application scenario, security level requirements, and experimental statistical analysis conducted on large-scale test datasets. In some embodiments, these thresholds can also be set as parameters that can be dynamically adjusted by users or administrators in specific modes.

[0094] See attached document Figure 5 After the real-time collaborative authentication engine 204 determines that the user's identity authentication is successful in step S503c and outputs an authentication success signal in step S504, the method of the present invention continues to execute the instruction parsing and execution steps.

[0095] In step S505, the command parsing and control module 205 is activated by the successful authentication signal and begins to execute the command parsing operation. The command parsing and control module 205 uses the first acoustic digital signal to be tested, which was acquired in step S501 and has already been authenticated. As input.

[0096] The instruction parsing and control module 205 integrates or calls an Automatic Speech Recognition (ASR) engine. This ASR engine receives the first acoustic digital signal under test. The process involves converting the contained speech information into one or more candidate text strings. This process includes sub-steps such as acoustic feature extraction, acoustic model matching, and language model decoding.

[0097] In step S506, the instruction parsing and control module 205 performs intent recognition on the candidate text strings output by the ASR engine. The instruction parsing and control module 205 reads a preset command list from the storage unit 104, which contains all legal device control commands and their corresponding text representations. The instruction parsing and control module 205 matches the candidate text strings with the entries in the command list.

[0098] In one embodiment, the matching process employs either exact string matching or keyword-based matching algorithms. When a candidate text string successfully matches a text representation in the command list, the instruction parsing and control module 205 determines the user's control intent.

[0099] In step S507, after successfully determining the user's control intent, the instruction parsing and control module 205 converts the control intent into one or more device control instructions conforming to the communication protocol of the rice cooker's main control system. The device control instruction is a string of binary data in a specific format.

[0100] Finally, in step S508, the instruction parsing and control module 205 sends the generated device control instruction to the rice cooker's main control system via the internal communication interface of the core processing unit 101 (e.g., Serial Peripheral Interface (SPI) or Universal Asynchronous Receiver / Transmitter (UART). Upon receiving the device control instruction, the main control system executes the corresponding physical operation, such as starting the heating program, setting the preset time, or adjusting the operating mode. If no valid instruction is matched in step S506, the process terminates, and no operation is performed.

[0101] Example 1: New Device Initialization and User Registration Scenario See attached document Figure 1 To be continued Figure 4 This embodiment aims to specifically illustrate how a brand-new rice cooker performs its own initialization and completes the registration of multi-dimensional feature templates for authorized users after being powered on for the first time.

[0102] First, when the rice cooker is powered on for the first time, its core processing unit 101 automatically starts the initialization program and calls the device physical fingerprint modeling module 202 to perform the self-calibration and modeling process of the device physical fingerprint.

[0103] The physical fingerprint modeling module 202 of the device controls the active excitation unit 103 to generate a preset linear sweep frequency signal, and simultaneously calls the signal synchronization acquisition module 201 to synchronously acquire the structural vibration response signal generated by the excitation. Harmony and acoustic response signals Subsequently, the device physical fingerprint modeling module 202, according to the formula... Calculate the unique transfer function of the device. Finally, the module receives the data from this transfer function. The amplitude spectrum was used to extract features such as the center frequency, peak gain, and quality factor of multiple resonant peaks. These features were combined into a feature vector, which served as the device's physical fingerprint. The data is then stored in storage unit 104. At this point, the device's initialization is complete.

[0104] After the physical fingerprint model of the device is completed, the system prompts the user to register as an authorized user through its display screen or voice announcement. Following the prompts, the user triggers the registration process, and the core processing unit 101 then calls the user template registration module 203.

[0105] The user template registration module 203 prompts the user to speak the registration command "Start Smart Cooking" at a normal volume from approximately 50 centimeters in front of the rice cooker. During the user's speech, the user template registration module 203 invokes the signal synchronization acquisition module 201 to synchronously acquire the first acoustic digital signal. Second acoustic digital signal and structural vibration digital signals .

[0106] Subsequently, the user template registration module 203 performs parallel processing on the collected signals: The user template registration module 203 will transmit the first acoustic digital signal. The input is fed into a pre-defined x-vector network, which extracts a 512-dimensional embedding vector. and use it as the user's acoustic voiceprint template. .

[0107] User template registration module 203 is based on the formula The first acoustic digital signal was calculated. Digital signals of structural vibration Transfer function between And extract its features as a liveness interaction template for the user. .

[0108] User template registration module 203 is based on the formula and The first acoustic digital signal was calculated. Second acoustic digital signal Estimated user latency values ​​between and use it as the user's spatial vector template. .

[0109] Finally, the user template registration module 203 will generate the acoustic voiceprint template. Live interactive template and space vector template They are integrated into a multi-dimensional feature template bound to the user's identity identifier and stored in storage unit 104.

[0110] After completing the above steps, the rice cooker not only completes its own physical characteristic calibration, but also completes the security template registration for authorized users. The system is ready to receive and authenticate voice commands from the user.

[0111] Example 2: Normal Voice Control Scenario See attached document Figure 1 Appendix Figure 2 and appendix Figure 5 This embodiment aims to specifically illustrate the complete process by which an authorized user successfully uses the method of the present invention to voice control a rice cooker in a typical kitchen environment.

[0112] In this scenario, there is some background noise in the kitchen environment, such as the sound of the range hood running. The authorized user is in front of the rice cooker device that has been initialized and registered, and issues the voice command "Start cooking rice".

[0113] When the voice command is issued, the system's core processing unit 101 detects the voice activity and immediately activates the real-time collaborative authentication engine 204 to begin executing the real-time authentication process.

[0114] First, the real-time collaborative authentication engine 204 calls the signal synchronization acquisition module 201 to synchronously acquire the first acoustic digital signal to be tested through the multimodal sensing array 102. The second acoustic digital signal to be tested and the digital signal of the structural vibration to be measured. .

[0115] Subsequently, the real-time collaborative authentication engine 204 extracts a set of test features in parallel from the acquired test signals, including a test acoustic feature vector. A live interaction transfer function to be tested and an observed time delay variable .

[0116] Next, the real-time collaborative authentication engine 204 initiates a layered, multi-stage collaborative authentication decision-making process: Spatial domain verification: This engine reads the user's spatial vector template from storage unit 104. (i.e., estimated user latency value) ), and calculate Since the user is in their normal operating position, the calculated absolute difference is less than the preset space tolerance threshold. Therefore, the spatial domain verification passed, and the authentication process continues.

[0117] Physical interaction verification: The engine reads the device's physical fingerprint from storage unit 104. and calculate Since the voice command is generated by the user's actual voice, the physical coupling between it and the device casing conforms to the inherent physical laws of the device, and the calculated similarity is higher than the preset physical consistency threshold. Therefore, the physical interaction verification passed, and the authentication process continued.

[0118] User identity verification: The engine reads the user's acoustic voiceprint template from storage unit 104. Interactive templates for live animals The engine calculates the acoustic feature vectors to be tested. Acoustic voiceprint template The similarity is used to obtain an acoustic similarity score. and the live interaction transfer function to be tested Interact with live objects template The similarity is used to obtain the live interaction similarity score. Since the instructions did indeed originate from the user, both scores were high. The engine uses a formula... Calculated final identity authentication score Higher than the preset identity authentication threshold Therefore, the user identity verification passed.

[0119] After successfully passing all three stages of verification, the real-time collaborative authentication engine 204 determines that the voice command is a valid command and outputs an authentication success signal to the command parsing and control module 205.

[0120] After receiving the authentication success signal, the instruction parsing and control module 205 processes the first acoustic digital signal to be tested, which has passed authentication. The Automatic Speech Recognition (ASR) process is executed, parsing the content into the text string "Start cooking". The instruction parsing and control module 205 successfully matches this text string with a preset command list and generates the corresponding device control instruction.

[0121] Finally, the device control command is sent to the rice cooker's main control system, and the rice cooker then executes the physical operation of "start cooking rice." The entire process achieves accurate and secure identification and execution of authorized user commands even in high-noise environments.

[0122] Example 3: Replaying the Attack and Defense Scenario See attached document Figure 1 Appendix Figure 2 and appendix Figure 5 This embodiment aims to illustrate in detail how the method of the present invention can effectively defend against a typical replay attack.

[0123] In this scenario, an unauthorized attacker uses a recording device to play a pre-recorded audio message from an authorized user, stating "Start cooking rice," at a location where the authorized user would normally be operating.

[0124] When the recording playback device emits voice, the system's core processing unit 101 detects the voice activity and activates the real-time collaborative authentication engine 204, starting to execute the same real-time authentication process as in Embodiment 2.

[0125] First, the real-time collaborative authentication engine 204 calls the signal synchronization acquisition module 201 to synchronously acquire the first acoustic digital signal to be tested emitted by the recording and playback device. The second acoustic digital signal to be tested And the resulting digital signal of structural vibration to be measured. .

[0126] Subsequently, the real-time collaborative authentication engine 204 extracts a set of test features in parallel from the acquired test signals, including a test acoustic feature vector. A live interaction transfer function to be tested and an observed time delay variable .

[0127] Next, the real-time collaborative authentication engine 204 initiates a layered, multi-stage collaborative authentication decision-making process: Spatial domain verification: This engine reads the user's spatial vector template from storage unit 104. (i.e., estimated user latency value) ), and calculate Because the attacker placed the recording and playback device in the user's usual operating position, the time difference between the sound waves emitted by the device reaching the two microphones is approximately the same as the time difference when the user speaks. Therefore, the calculated absolute difference is less than the preset spatial tolerance threshold. The spatial domain verification passed, and the authentication process proceeds to the next stage.

[0128] Physical interaction verification: The engine reads the device's physical fingerprint from storage unit 104. And calculate the live interaction transfer function to be tested. physical fingerprint of the device Similarity between Because the sound waves generated by the recording and playback device are purely airborne, the energy transferred to the rice cooker's shell and the resulting structural vibrations are very weak. Furthermore, its acoustic-vibration coupling characteristics are fundamentally different from the physical process of live sound generation (which includes air conduction and bone conduction). This fundamental difference lies in the fact that when an authorized user speaks, sound energy not only travels through the air to the device's microphone and shell, but also some vibrational energy is directly or indirectly transmitted to the device's shell through the user's skull, jawbone, and other body tissues (i.e., the bone conduction path). In contrast, the recording and playback device, as an independent sound source, can only interact with the device through the air medium, lacking the crucial bone conduction path. This results in a significant and distinguishable difference in its acoustic-vibration coupling characteristics compared to live sound generation. Therefore, the calculated live interaction transfer function... Device physical fingerprint, which characterizes the inherent physical properties of a device This discrepancy is significant. The calculated similarity is far below the preset physical consistency threshold. .

[0129] Because the result of physical cross-validation does not meet the requirements Under the given conditions, the real-time collaborative authentication engine 204 determines that the voice event is a physical interaction anomaly event, that is, it is determined to be a non-liveness attack.

[0130] Therefore, the authentication process is immediately terminated at the physical interaction verification stage, and does not proceed to the subsequent user identity verification stage. The real-time collaborative authentication engine 204 does not output an authentication success signal, the instruction parsing and control module 205 is not activated, and the rice cooker's main control system does not receive any control instructions.

[0131] In this way, the method of the present invention successfully intercepted the replay attack, ensuring that even if the attacker obtains the user's voiceprint recording, they cannot illegally control the device, thereby protecting the security of the system.

Claims

1. A voice-based safety control system for rice cookers, characterized in that, include: A multimodal sensing array is used to simultaneously acquire acoustic signals and structural vibration signals. The multimodal sensing array includes an acoustic sensor array for acquiring the acoustic signals and a structural vibration sensor for acquiring the structural vibration signals. Storage unit, used to store the physical fingerprint of the device and the multi-dimensional feature template of the authorized user; as well as The core processing unit is electrically connected to the multimodal sensing array and the storage unit, and the core processing unit is used for: Based on the acoustic signals and structural vibration signals acquired by the multimodal sensing array, the physical fingerprint of the device is generated and stored; Based on the acoustic signals and structural vibration signals acquired by the multimodal sensing array, the multidimensional feature template is generated and stored for the authorized user; Upon receiving a voice command, the device's physical fingerprint and the multi-dimensional feature template are collaboratively invoked to perform real-time collaborative authentication of the voice command based on the test signal collected by the multimodal sensing array.

2. The voice-based safety control system for rice cookers according to claim 1, characterized in that, When generating the physical fingerprint of the device, the core processing unit is specifically used for: The active excitation unit is controlled to generate an excitation signal and apply it to the rice cooker shell; The multimodal sensing array is invoked to simultaneously acquire structural vibration response signals and acoustic response signals caused by the excitation signal; The transfer function between the structural vibration response signal and the acoustic response signal is calculated using a frequency domain system identification algorithm, and frequency domain features are extracted from the transfer function to form the physical fingerprint of the device.

3. The voice-based safety control system for rice cookers according to claim 1, characterized in that, The multidimensional feature template includes: An acoustic voiceprint template, representing the acoustic characteristics of the authorized user; A live interaction template characterizes the physical coupling characteristics between the authorized user and the rice cooker shell when the user speaks. A spatial vector template represents the spatial orientation information of the authorized user when speaking.

4. The voice-based safety control system for rice cookers according to claim 3, characterized in that, When generating the live interaction template, the core processing unit is specifically used for: During the authorized user's vocalization, the first acoustic digital signal and the structural vibration digital signal are simultaneously acquired; The transfer function between the first acoustic digital signal and the structural vibration digital signal is calculated using a frequency domain system identification algorithm, and key frequency domain features are extracted from the transfer function to form the living interaction template.

5. A voice-based safety control system for rice cookers according to claim 3, characterized in that, When generating the spatial vector template, the core processing unit is specifically used for: During the authorized user's vocalization, the first and second acoustic digital signals are simultaneously acquired; The arrival time difference between the first acoustic digital signal and the second acoustic digital signal is calculated using a generalized cross-correlation phase transformation algorithm to form the spatial vector template.

6. The voice-based safety control system for a rice cooker according to claim 1, characterized in that, When performing the real-time collaborative authentication, the core processing unit executes a layered gating authentication process, which includes the following steps: Spatial domain verification is used to confirm that the physical location of the sound source is consistent with the location where the authorized user was registered. Physical interaction verification is used to confirm that the acoustic-vibration coupling phenomenon conforms to the inherent physical characteristics of the rice cooker; and User identity verification is used to ultimately confirm that the user who issued the instruction is the authorized user.

7. A voice-based safety control system for rice cookers according to claim 6, characterized in that, When performing the physical interaction verification, the core processing unit is specifically used for: Calculate the live interaction features to be tested from the signal to be tested; The cosine similarity algorithm is used to calculate the similarity between the live interaction feature to be tested and the physical fingerprint of the device stored in the storage unit, and it is determined whether the similarity is not lower than a preset physical consistency threshold.

8. A voice-based safety control system for a rice cooker according to claim 6, characterized in that, When performing the user identity verification, the core processing unit is specifically used for: The acoustic similarity score between the acoustic feature vector to be tested and the acoustic voiceprint template of the authorized user is calculated using the cosine similarity algorithm. The cosine similarity algorithm is used to calculate the live interaction similarity score between the live interaction feature to be tested and the live interaction template of the authorized user. The acoustic similarity score and the liveness interaction similarity score are weighted and fused to obtain the final identity authentication total score, and it is determined whether the total score is not lower than the preset identity authentication threshold.

9. A voice-based safety control system for a rice cooker according to claim 2, characterized in that, The active excitation unit and the structural vibration sensor reuse the same piezoelectric ceramic sensor device.

10. A voice-based safety control method for a rice cooker, characterized in that, Includes the following steps: S1. Generate and store the device physical fingerprint that uniquely represents the physical characteristics of the rice cooker. S2. Generate and store a unique multi-dimensional feature template for authorized users, the multi-dimensional feature template including an acoustic voiceprint template, a live interaction template and a spatial vector template; S3. In response to the received voice command, simultaneously collect the acoustic signal and structural vibration signal to be tested, and use the stored device physical fingerprint and the multi-dimensional feature template to perform a multi-stage collaborative authentication decision on the voice command. The decision sequentially performs spatial domain verification, physical interaction verification and user identity verification. S4. After the multi-stage collaborative authentication decision is passed, the content of the voice command is parsed to generate and execute the corresponding device control command.

Citation Information

Cited By

  • Voice control method and system based on multi-core heterogeneity, storage medium and chip

    CN121983054A