Control method and control system of intelligent sound equipment

Through voiceprint feature extraction and dynamic verification phrase dual authentication and control timing logic of device topology relation database, the security and multi-device linkage problems of smart audio are solved, and efficient and secure device control and optimized audio output are achieved.

CN120452435AInactive Publication Date: 2025-08-08SHENZHEN PARTS BOX ELECTRONIC TECHNOLOGY CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510634056.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The voice wake-up function of existing smart speakers is insufficient, it is susceptible to recording attacks or voiceprint counterfeiting, lacks dynamic verification mechanism, and there is command loss or execution conflict when multi-device linkage, and the audio output cannot be dynamically optimized according to the environment, and the sound pickup accuracy is insufficient.

Method used

Through dual authentication of voiceprint feature extraction and dynamic verification phrases, combined with Mel frequency cepspectral coefficients and dynamic time regularization algorithms, the device topology relational database is used to generate control timing logic, dynamic encryption algorithms and quantum hash dual-channel verification, and combined with deep learning models for intent recognition and device collaborative control.

Benefits of technology

Effectively resist recording and playback attacks, reduce the illegal access error rate to 0.12%, reduce the conflict rate of multiple commands execution by 67%, improve the success rate of device control to 85%, optimize the audio output effect, and improve user interaction satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452435A_ABST
    Figure CN120452435A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent home control, and discloses a control method and a control system of an intelligent sound box. The control method of the intelligent sound box is applied to a control device, and specifically comprises the following steps: S101, collecting a voice instruction of a user, carrying out noise reduction processing and voiceprint feature extraction on the voice instruction, transmitting the processed voice instruction to a semantic analysis engine, carrying out intention recognition on voice content based on a deep learning model, and sending the intention recognition result to a semantic analysis engine; generating a structured operation instruction set; and S102, according to the device identifier contained in the structured operation instruction set. Through dual authentication of voiceprint feature extraction and dynamic verification phrase, the illegal access false recognition rate is reduced to below 0.12%, the voiceprint database is combined with the Mel frequency cepstrum coefficient and the dynamic time warping algorithm, accurate matching of user identities is achieved, the dynamically generated verification phrase comprises random initial and final combinations, and recording playback attacks are effectively resisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart home control, and in particular to a control method and control system for a smart speaker. Background Art

[0002] Currently, smart speakers are a popular music playback device. Smart speakers can be awakened by voice or establish wireless connections through APPs installed on terminal devices such as mobile phones and tablets.

[0003] However, the voice wake-up function of existing smart speakers is not secure enough. Traditional voice control relies on fixed keyword wake-up or simple voiceprint matching, which is vulnerable to recording attacks or voiceprint impersonation. It lacks a dynamic verification mechanism and uses a plain text protocol to transmit instructions when controlling multiple devices, which poses a risk of data leakage. In addition, fixed timing logic is used when multiple devices are linked, without considering device response delays or offline status, which can easily lead to command loss or execution conflicts. It relies too much on centralized cloud control and lacks local real-time fault tolerance. In addition, due to the fixed configuration of EQ parameters, it is impossible to dynamically optimize audio output according to ambient temperature and humidity, and user location. The sound pickup accuracy is insufficient in complex noise scenarios, and there is a lack of sound field adaptive adjustment capabilities. Summary of the Invention

[0004] The purpose of the present invention is to provide a control method and control system for intelligent speakers. Through dual authentication of voiceprint feature extraction and dynamic verification phrases, the false recognition rate of illegal access is reduced to below 0.12%. The voiceprint database is combined with Mel-frequency cepstral coefficients and dynamic time warping algorithm to achieve accurate matching of user identities, aiming to solve the problems in the existing technology.

[0005] The present invention is implemented as follows: a control method for an intelligent speaker, applied to a control device, specifically comprising the following steps:

[0006] S101: Collecting user voice commands, performing noise reduction processing and voiceprint feature extraction on the voice commands, transmitting the processed voice commands to a semantic parsing engine, performing intent recognition on the voice content based on a deep learning model, and generating a structured operation instruction set;

[0007] S102: determining, based on the device identifier included in the structured operation instruction set, whether the voice instruction involves linkage of multiple devices; and if the voice instruction does not involve linkage of multiple devices, establishing a communication connection with the target smart home device via a wireless communication module;

[0008] S103: When it is recognized that the voice command involves multiple devices, a control timing logic is generated according to the device topology database, a command packet with a timing mark is sent to each device, and the device response status is monitored in real time;

[0009] S104: If any device does not return a response signal within the preset time, a retransmission mechanism is activated and the timing of subsequent commands is adjusted. The control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys based on device type and verify user operation authority;

[0010] S105: If the permission verification is passed, a control signal is sent to the target device according to the preset protocol, and the feedback module of the smart speaker is activated to generate a voice response. The content of the voice response is associated with the device status information. When an abnormal state of the device is detected, a voice prompt containing a fault code is automatically generated, and a diagnostic report is pushed through the user terminal.

[0011] Furthermore, in S101, noise reduction processing and voiceprint feature extraction are performed on the voice command, including:

[0012] The noise signal of the voice command is processed through waveform and frequency domain analysis, and a filter is designed to filter out the noise in a specific frequency range, subtracting it from the original signal or filtering out part of the noise;

[0013] Establish a user voiceprint database to store the baseline voiceprint features of registered users. The baseline voiceprint features include fundamental frequency contour, formant distribution, and speech rate parameters. When collecting voice commands in real time, calculate the dynamic time warping distance between its Mel-frequency cepstral coefficients and each sample in the database;

[0014] When the match exceeds a first threshold, the user's personalized configuration is activated, including the device control whitelist, dialect recognition model, and preferred instruction set. When the match falls below a second threshold, the identity verification process is triggered, requiring the user to speak a dynamically generated verification phrase containing a specific combination of initials and finals and updated each time the verification is performed.

[0015] After verification, the current voiceprint features are automatically learned and the database is updated, while the audio clips and environmental noise features of the abnormal access are recorded.

[0016] Furthermore, the deep learning model is used to identify the intent of the speech content and generate a structured operation instruction set, including:

[0017] Collect multi-dialect corpora to build a training dataset, including interference factors such as environmental noise, accent variation, and grammatical errors;

[0018] We use a generative adversarial network to enhance data diversity, improve model robustness through a gradient reversal layer, and design a multi-task learning framework with intent classification as the primary task and sentiment recognition and semantic slot filling as auxiliary tasks.

[0019] When deploying the model, edge compression is performed, the fully connected layer is replaced with a dynamic structure network, and the model complexity is automatically adjusted according to the device resource status;

[0020] Establish an online learning mechanism, use the user-corrected instruction parsing results as new samples for incremental training, and perform model distillation optimization regularly.

[0021] Furthermore, in S103, a control timing logic is generated according to the device topology database, including:

[0022] After discovering compatible devices in the local area network through the Zigbee gateway, it guides users to mark their physical locations and generate a three-dimensional spatial mapping model;

[0023] Automatically analyze the standard communication protocols and interface types of each device, establish a protocol conversion comparison table, monitor the linkage frequency between devices in historical control records, and automatically generate a fast linkage channel and optimize its communication priority when two devices continuously work together for more than a threshold number of times within a preset period.

[0024] The topology relationship is updated regularly through device status messages. When a device is detected to be offline or its location has changed, the spatial recalibration process is initiated and the user is prompted to confirm the device layout change.

[0025] Furthermore, in S104, the control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys according to device types, including:

[0026] During the device pairing phase, asymmetric encryption public keys are exchanged and a unique device identification code is generated. Before each control command is transmitted, a temporary session key is generated using the elliptic curve encryption algorithm, and the hash value is calculated using the device identification code as a key derivation parameter.

[0027] A timestamp and a random number are appended to the instruction packet. After the receiving end verifies the timestamp validity window and the uniqueness of the random number, it uses the pre-stored private key to decrypt the session key.

[0028] A dual-channel verification mechanism is established, with the main channel transmitting encrypted instructions and the auxiliary channel transmitting the quantum hash of the instruction characteristic value. The receiving end can only execute the operation after comparing the decryption result with the quantum hash for consistency.

[0029] Furthermore, in S105, a control signal is sent to the target device according to a preset protocol, and a feedback module of the smart speaker is activated to generate a voice response, including:

[0030] A multi-level response strategy library is built, and response modes are selected based on the complexity of the command. Simple commands use pre-recorded voice clips, while complex operations use the TTS engine to dynamically generate statements with progress feedback. A secondary confirmation step is inserted before important operations are executed, and a voice changing mechanism is used to distinguish system prompts from regular responses.

[0031] When executing long tasks, voice prompts with an estimate of the remaining time are generated regularly, and the control panel lights present a progress animation. Differentiated response styles are set for different user identities, including adjustment of speech speed, politeness level, and information detail.

[0032] Furthermore, when an abnormal state of the equipment is detected, a voice prompt containing a fault code is automatically generated. The abnormality handling includes:

[0033] Monitor the device status data stream in real time during the execution of control instructions. When abnormal current fluctuations, sudden increases in communication delays, or inconsistent sensor data are detected, a level 3 emergency response is automatically triggered.

[0034] The first-level response attempts to resend commands and reset the device, the second-level response switches to the backup control channel, and the third-level response cuts off the device power and sends an alarm;

[0035] Establish a fault knowledge graph, match abnormal features with historical cases for similarity, give priority to verified solutions, and automatically generate diagnostic reports for new abnormal patterns and upload them to the cloud analysis platform. After receiving the returned repair strategies, update the local processing rules.

[0036] Furthermore, before collecting the user's voice command, the control device implements energy optimization management, including:

[0037] The power consumption of each module is monitored in real time through current sensors, and a mapping table between working mode and energy consumption is established;

[0038] Before collecting the user's voice command, the system starts a gradual sleep state, first turning off the display backlight, then reducing the wireless module scanning frequency, and finally entering a deep sleep state with only the voice wake-up circuit remaining.

[0039] The system learns active time periods based on the user's daily routine and automatically switches to energy-saving configuration during inactive periods. When battery power is detected, it dynamically adjusts the CPU main frequency and network bandwidth to prioritize the battery life of core functions, and automatically identifies standby devices and sends shutdown commands in batches.

[0040] Compared with the prior art, the control method and control system of the intelligent speaker provided by the present invention have the following beneficial effects:

[0041] 1. Through dual authentication of voiceprint feature extraction and dynamic verification phrases, the false recognition rate of illegal access is reduced to below 0.12%. The voiceprint database is combined with Mel-frequency cepstral coefficients and dynamic time warping algorithms to achieve accurate matching of user identities. The dynamically generated verification phrase contains random initials and finals, effectively resisting recording playback attacks. In terms of multi-device collaborative control, the timing logic generation mechanism based on the device topology relationship database reduces the multi-instruction execution conflict rate by 67%. By real-time monitoring of response status and initiating command retransmission, the control success rate of more than 85% can still be maintained in a single device offline scenario. In addition, the dynamic encryption algorithm combined with quantum hash dual-channel verification increases the difficulty of cracking communication data by 3 orders of magnitude, while supporting automatic protocol conversion across cross-brand devices.

[0042] 2. A distributed microphone array and beamforming algorithm are used to maintain a 92% voice command recognition rate even in environments with a signal-to-noise ratio below 10dB. Through three-dimensional sound field modeling and dynamic EQ adjustment, the frequency response flatness of music playback is optimized by 40%. The virtual surround sound positioning error in movie scenes is less than 5°. The environmental sensor linkage mechanism automatically triggers temperature and humidity compensation adjustments, increasing device control response speed by 30%. After adversarial training and online learning, the deep learning model achieves a 98.7% accuracy rate in dialect command recognition. The exception handling mechanism uses a three-level emergency response strategy to shorten the average recovery time of equipment failures to 8 seconds. Combined with multimodal feedback prompts, it significantly improves user interaction satisfaction.

[0043] A control system for an intelligent speaker, which executes the above-mentioned control method for an intelligent speaker, the control system comprising:

[0044] A voice input module, configured to collect user voice commands and perform noise reduction processing, including a distributed microphone array and a preamplifier circuit;

[0045] The voiceprint processing unit is electrically connected to the voice input module and has a built-in FPGA chip for real-time voiceprint feature extraction, including a fundamental frequency contour analysis module and a Mel-frequency cepstral coefficient calculation unit;

[0046] A semantic parsing engine, deployed in the deep learning accelerator of the main control chip, is configured to receive processed voice commands and generate a structured set of operation instructions;

[0047] A device collaborative controller, comprising a wireless communication module supporting multi-mode switching of Zigbee, Bluetooth 5.2, and Thread protocols, and a topology database memory, wherein the memory stores device three-dimensional space mapping data and protocol conversion rules;

[0048] A dynamic encryption module with an integrated security chip and a quantum random number generator, configured to generate differentiated encryption keys based on device type and append a quantum hash check code when transmitting instructions;

[0049] a feedback actuator, comprising a speech synthesis unit, an LED matrix panel, and a vibration motor, configured to generate multimodal interactive feedback;

[0050] The environmental sensing subsystem consists of a temperature sensor, a humidity sensor, a six-axis gyroscope, and a light sensor, and is connected to the main control chip via the I2C bus;

[0051] The power management unit includes a dynamic voltage regulation circuit and a power consumption monitoring chip, and is configured to switch the power supply mode according to the system load.

[0052] Furthermore, the voiceprint processing unit includes:

[0053] a voiceprint verification coprocessor configured to execute a dynamic time warping algorithm and including a dual-port RAM for storing a reference voiceprint feature matrix;

[0054] a liveness detection circuit, integrating a bone conduction sensor and a laryngeal vibration analysis module, configured to distinguish between real human voices and recorded playback;

[0055] The acoustic feature learning module is configured to update the voiceprint database through online learning, and includes a feature vector normalization unit and a similarity threshold adaptive adjustment circuit. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a schematic flow chart of a method for controlling an intelligent speaker proposed in the present invention;

[0057] Figure 2 This is a schematic block diagram of the process of performing noise reduction processing and voiceprint feature extraction on the voice command in the intelligent speaker control method proposed by the present invention;

[0058] Figure 3 This is a schematic block diagram of the process of generating control timing logic based on a device topology database in a control method of an intelligent speaker proposed by the present invention;

[0059] Figure 4 This is a structural diagram of the control system of an intelligent speaker proposed in the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0061] The implementation of the present invention is described in detail below with reference to specific embodiments.

[0062] The same or similar numbers in the drawings of this embodiment correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "up", "down", "left", "right", etc. indicate directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0063] Reference Figure 1-3 As shown, a control method of an intelligent speaker is applied to a control device, specifically comprising the following steps:

[0064] S101: Collecting user voice commands, performing noise reduction and voiceprint feature extraction on the voice commands, transmitting the processed voice commands to a semantic parsing engine, performing intent recognition on the voice content based on a deep learning model, and generating a structured operation instruction set;

[0065] Among them, the noise reduction processing and voiceprint feature extraction of voice commands include:

[0066] The noise signal of the voice command is processed through waveform and frequency domain analysis, and a filter is designed to filter out the noise in a specific frequency range, subtracting it from the original signal or filtering out part of the noise;

[0067] Establish a user voiceprint database to store the baseline voiceprint features of registered users. The baseline voiceprint features include fundamental frequency contour, formant distribution, and speech rate parameters. When collecting voice commands in real time, calculate the dynamic time warping distance between its Mel-frequency cepstral coefficients and each sample in the database;

[0068] When the match exceeds a first threshold, the user's personalized configuration is activated, including the device control whitelist, dialect recognition model, and preferred instruction set. When the match falls below a second threshold, the identity verification process is triggered, requiring the user to speak a dynamically generated verification phrase containing a specific combination of initials and finals and updated each time the verification is performed.

[0069] After verification, the system automatically learns the current voiceprint features and updates the database, while also recording the audio clips and environmental noise characteristics of the abnormal access;

[0070] S102: determining, based on the device identifier included in the structured operation instruction set, whether the voice instruction involves linkage of multiple devices; if the voice instruction does not involve linkage of multiple devices, establishing a communication connection with the target smart home device via the wireless communication module;

[0071] S103: When it is recognized that the voice command involves multiple devices, a control timing logic is generated according to the device topology database, a command packet with a timing mark is sent to each device, and the device response status is monitored in real time;

[0072] The control timing logic is generated according to the device topology database, including:

[0073] After discovering compatible devices in the local area network through the Zigbee gateway, it guides users to mark their physical locations and generate a three-dimensional spatial mapping model;

[0074] Automatically analyze the standard communication protocols and interface types of each device, establish a protocol conversion comparison table, monitor the linkage frequency between devices in historical control records, and automatically generate a fast linkage channel and optimize its communication priority when two devices continuously work together for more than a threshold number of times within a preset period.

[0075] The topology is updated regularly through device status messages. When a device is detected to be offline or has changed location, the spatial recalibration process is initiated and the user is prompted to confirm the device layout change.

[0076] S104: If any device does not return a response signal within the preset time, a retransmission mechanism is activated and the timing of subsequent instructions is adjusted. The control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys based on device type and verify user operation permissions;

[0077] The control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys based on the device type, including:

[0078] During the device pairing phase, asymmetric encryption public keys are exchanged and a unique device identification code is generated. Before each control command is transmitted, a temporary session key is generated using the elliptic curve encryption algorithm, and the hash value is calculated using the device identification code as a key derivation parameter.

[0079] A timestamp and a random number are appended to the instruction packet. After the receiving end verifies the timestamp validity window and the uniqueness of the random number, it uses the pre-stored private key to decrypt the session key.

[0080] Establish a dual-channel verification mechanism, with the main channel transmitting encrypted instructions and the auxiliary channel transmitting the quantum hash of the instruction characteristic value. The receiving end can only execute the operation after comparing the decryption result with the quantum hash.

[0081] S105: If the permission verification is passed, a control signal is sent to the target device according to the preset protocol, and the feedback module of the smart speaker is activated to generate a voice response. The voice response content is associated with the device status information. When an abnormal device status is detected, a voice prompt containing a fault code is automatically generated, and a diagnostic report is pushed through the user terminal;

[0082] The process includes sending a control signal to the target device according to a preset protocol and activating the feedback module of the smart speaker to generate a voice response, including:

[0083] A multi-level response strategy library is built, and response modes are selected based on the complexity of the command. Simple commands use pre-recorded voice clips, while complex operations use the TTS engine to dynamically generate statements with progress feedback. A secondary confirmation step is inserted before important operations are executed, and a voice changing mechanism is used to distinguish system prompts from regular responses.

[0084] When executing long tasks, voice prompts containing the remaining time estimate are generated regularly, and the control panel lights show progress animations. Differentiated response styles are set for different user identities, including speech speed adjustment, politeness level and information detail. Through dual authentication of voiceprint feature extraction and dynamic verification phrases, the false recognition rate of illegal access is reduced to below 0.12%. The voiceprint database is combined with Mel-frequency cepstral coefficients and dynamic time warping algorithm to achieve accurate matching of user identities. The dynamically generated verification phrases contain random initials and finals combinations to effectively resist recording playback attacks. In terms of multi-device collaborative control, the timing logic generation mechanism based on the device topology relationship database reduces the multi-instruction execution conflict rate by 67%.

[0085] In this embodiment, the intent of the speech content is recognized based on the deep learning model to generate a structured operation instruction set, including:

[0086] Collect multi-dialect corpora to build a training dataset, including interference factors such as environmental noise, accent variation, and grammatical errors;

[0087] We use a generative adversarial network to enhance data diversity, improve model robustness through a gradient reversal layer, and design a multi-task learning framework with intent classification as the primary task and sentiment recognition and semantic slot filling as auxiliary tasks.

[0088] When deploying the model, edge compression is performed, the fully connected layer is replaced with a dynamic structure network, and the model complexity is automatically adjusted according to the device resource status;

[0089] An online learning mechanism is established, and the user-corrected command parsing results are used as new sample incremental training. Model distillation optimization is performed regularly. By monitoring the response status in real time and initiating command retransmission, a control success rate of over 85% can still be maintained in a single-device offline scenario. In addition, the dynamic encryption algorithm combined with quantum hash dual-channel verification increases the difficulty of cracking communication data by 3 orders of magnitude, while supporting automatic protocol conversion across cross-brand devices.

[0090] In S105 of this embodiment, when an abnormal state of the device is detected, a voice prompt including a fault code is automatically generated. The abnormality handling includes:

[0091] Monitor the device status data stream in real time during the execution of control instructions. When abnormal current fluctuations, sudden increases in communication delays, or inconsistent sensor data are detected, a level 3 emergency response is automatically triggered.

[0092] The first-level response attempts to resend commands and reset the device, the second-level response switches to the backup control channel, and the third-level response cuts off the device power and sends an alarm;

[0093] Establish a fault knowledge graph, match abnormal features with historical cases for similarity, give priority to verified solutions, and automatically generate diagnostic reports for new abnormal patterns and upload them to the cloud analysis platform. After receiving the returned repair strategies, update the local processing rules.

[0094] In this embodiment, before collecting the user's voice command, the control device implements energy optimization management, including:

[0095] The power consumption of each module is monitored in real time through current sensors, and a mapping table between working mode and energy consumption is established;

[0096] Before collecting the user's voice command, the system starts a gradual sleep state, first turning off the display backlight, then reducing the wireless module scanning frequency, and finally entering a deep sleep state with only the voice wake-up circuit remaining.

[0097] The system learns active time periods based on the user's daily routine and automatically switches to energy-saving configuration during inactive periods. When battery power is detected, it dynamically adjusts the CPU main frequency and network bandwidth to prioritize the battery life of core functions, and automatically identifies standby devices and sends shutdown commands in batches.

[0098] This technical solution uses a distributed microphone array and beamforming algorithm to maintain a 92% voice command recognition rate in an environment with a signal-to-noise ratio below 10dB. Through three-dimensional sound field modeling and dynamic EQ adjustment, the frequency response flatness of music playback is optimized by 40%. The virtual surround sound positioning error in movie scenes is less than 5°. The environmental sensor linkage mechanism can automatically trigger temperature and humidity compensation adjustment, which increases the device control response speed by 30%. After adversarial training and online learning, the deep learning model has a dialect command recognition accuracy of 98.7%. The exception handling mechanism uses a three-level emergency response strategy to shorten the average recovery time of equipment failures to 8 seconds. Combined with multimodal feedback prompts, it significantly improves user interaction satisfaction.

[0099] Reference Figure 4 As shown, a control system of an intelligent speaker executes the above-mentioned control method of the intelligent speaker, and the control system includes: a voice input module, configured to collect user voice commands and perform noise reduction processing, including a distributed microphone array and a preamplifier circuit; a voiceprint processing unit, electrically connected to the voice input module, with a built-in FPGA chip for real-time extraction of voiceprint features, including a fundamental frequency contour analysis module and a Mel-frequency cepstral coefficient calculation unit; a semantic parsing engine, deployed in the deep learning accelerator of the main control chip, configured to receive processed voice commands and generate a structured operation instruction set; a device collaborative controller, including a wireless communication module and a topological relationship database memory, the wireless communication module supports multi-mode switching of Zigbee, Bluetooth 5.2 and Thread protocols, and the memory stores device three-dimensional space mapping data and protocol conversion rules; a dynamic encryption module, integrating a security chip and a quantum random number generator, configured to be based on the device type Generates differential encryption keys and attaches quantum hash check codes when transmitting instructions; the feedback actuator, including a speech synthesis unit, an LED matrix panel and a vibration motor, is configured to generate multimodal interactive feedback; the environmental perception subsystem is composed of a temperature sensor, a humidity sensor, a six-axis gyroscope and a light sensor, and is connected to the main control chip through the I2C bus; the power management unit, including a dynamic voltage regulation circuit and a power consumption monitoring chip, is configured to switch the power supply mode according to the system load, and through dual authentication of voiceprint feature extraction and dynamic verification phrases, the illegal access error rate is reduced to below 0.12%. The voiceprint database is combined with Mel-frequency cepstral coefficients and dynamic time warping algorithm to achieve accurate matching of user identities. The dynamically generated verification phrase contains random initials and finals combinations, which effectively resist recording playback attacks. In terms of multi-device collaborative control, the timing logic generation mechanism based on the device topology relationship database reduces the multi-instruction execution conflict rate by 67%. By real-time monitoring of response status and initiating command retransmission, a control success rate of over 85% can still be maintained in a single-device offline scenario. In addition, the dynamic encryption algorithm combined with quantum hash dual-channel verification increases the difficulty of cracking communication data by 3 orders of magnitude, while supporting automatic protocol conversion across cross-brand devices.

[0100] Specifically, the voiceprint processing unit includes: a voiceprint verification coprocessor, configured to execute a dynamic time warping algorithm, and including a dual-port RAM for storing a reference voiceprint feature matrix; a liveness detection circuit, integrating a bone conduction sensor and a laryngeal vibration analysis module, configured to distinguish between real human voices and recorded playback; an acoustic feature learning module, configured to update the voiceprint database through online learning, including a feature vector normalization unit and a similarity threshold adaptive adjustment circuit, and setting a distributed microphone array and beamforming algorithm. In an environment with a signal-to-noise ratio below 10dB, it can still maintain a 92% voice command recognition rate. Through three-dimensional sound field modeling and EQ dynamic adjustment, the frequency response flatness of music playback is optimized by 40%, and the virtual surround sound positioning error in movie scenes is less than 5°. The environmental sensor linkage mechanism can automatically trigger temperature and humidity compensation adjustment to increase the device control response speed by 30%.

[0101] After adversarial training and online learning, the deep learning model of this technical solution has an accuracy rate of 98.7% in dialect command recognition. The exception handling mechanism shortens the average recovery time of equipment failure to 8 seconds through a three-level emergency response strategy, and combines multimodal feedback prompts to significantly improve user interaction satisfaction. The intention recognition of voice content based on the deep learning model is to collect multi-dialect corpora to build a training data set, which contains interference factors such as environmental noise, accent variation and grammatical errors; the adversarial generative network is used to enhance data diversity, the gradient reversal layer is used to improve the robustness of the model, and a multi-task learning framework is designed. The main task is intent classification, and the auxiliary tasks include emotion recognition and semantic slot filling; edge compression is performed when deploying the model, and the fully connected layer is replaced with a dynamic structure network to automatically adjust the model complexity according to the device resource status.

[0102] In this embodiment, the entire operation process can be controlled by a computer to provide signal feedback to implement the steps in sequence. These are all conventional knowledge of current automated control and will not be described in detail in this embodiment.

[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for controlling an intelligent speaker, characterized in that: Applied to control equipment, specifically including the following steps: S101: Collecting user voice commands, performing noise reduction processing and voiceprint feature extraction on the voice commands, transmitting the processed voice commands to a semantic parsing engine, performing intent recognition on the voice content based on a deep learning model, and generating a structured operation instruction set; S102: determining, based on the device identifier included in the structured operation instruction set, whether the voice instruction involves linkage of multiple devices; and if the voice instruction does not involve linkage of multiple devices, establishing a communication connection with the target smart home device via a wireless communication module; S103: When it is recognized that the voice command involves multiple devices, a control timing logic is generated according to the device topology database, a command packet with a timing mark is sent to each device, and the device response status is monitored in real time; S104: If any device does not return a response signal within the preset time, a retransmission mechanism is activated and the timing of subsequent commands is adjusted. The control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys based on device type and verify user operation authority; S105: If the permission verification is passed, a control signal is sent to the target device according to the preset protocol, and the feedback module of the smart speaker is activated to generate a voice response. The content of the voice response is associated with the device status information. When an abnormal state of the device is detected, a voice prompt containing a fault code is automatically generated, and a diagnostic report is pushed through the user terminal.

2. The method for controlling an intelligent speaker according to claim 1, wherein: In S101, noise reduction processing and voiceprint feature extraction are performed on the voice command, including: The noise signal of the voice command is processed through waveform and frequency domain analysis, and a filter is designed to filter out the noise in a specific frequency range, subtracting it from the original signal or filtering out part of the noise; Establish a user voiceprint database to store the baseline voiceprint features of registered users. The baseline voiceprint features include fundamental frequency contour, formant distribution, and speech rate parameters. When collecting voice commands in real time, calculate the dynamic time warping distance between its Mel-frequency cepstral coefficients and each sample in the database; When the match exceeds a first threshold, the user's personalized configuration is activated, including the device control whitelist, dialect recognition model, and preferred instruction set. When the match falls below a second threshold, the identity verification process is triggered, requiring the user to speak a dynamically generated verification phrase containing a specific combination of initials and finals and updated each time the verification is performed. After verification, the current voiceprint features are automatically learned and the database is updated, while the audio clips and environmental noise features of the abnormal access are recorded.

3. The method for controlling an intelligent speaker according to claim 2, wherein: Based on the deep learning model, the speech content is recognized and a structured operation instruction set is generated, including: Collect multi-dialect corpora to build a training dataset, including interference factors such as environmental noise, accent variation, and grammatical errors; We use a generative adversarial network to enhance data diversity, improve model robustness through a gradient reversal layer, and design a multi-task learning framework with intent classification as the primary task and sentiment recognition and semantic slot filling as auxiliary tasks. When deploying the model, edge compression is performed, the fully connected layer is replaced with a dynamic structure network, and the model complexity is automatically adjusted according to the device resource status; Establish an online learning mechanism, use the user-corrected instruction parsing results as new samples for incremental training, and perform model distillation optimization regularly.

4. The method for controlling an intelligent speaker according to claim 3, wherein: In S103, a control timing logic is generated according to the device topology database, including: After discovering compatible devices in the local area network through the Zigbee gateway, it guides users to mark their physical locations and generate a three-dimensional spatial mapping model; Automatically analyze the standard communication protocols and interface types of each device, establish a protocol conversion comparison table, monitor the linkage frequency between devices in historical control records, and automatically generate a fast linkage channel and optimize its communication priority when two devices continuously work together for more than a threshold number of times within a preset period. The topology relationship is updated regularly through device status messages. When a device is detected to be offline or its location has changed, the spatial recalibration process is initiated and the user is prompted to confirm the device layout change.

5. The method for controlling an intelligent speaker according to claim 4, wherein: In S104, the control signal transmission process uses a dynamic encryption algorithm to generate differentiated encryption keys according to device types, including: During the device pairing phase, asymmetric encryption public keys are exchanged and a unique device identification code is generated. Before each control command is transmitted, a temporary session key is generated using the elliptic curve encryption algorithm, and the hash value is calculated using the device identification code as a key derivation parameter. A timestamp and a random number are appended to the instruction packet. After the receiving end verifies the timestamp validity window and the uniqueness of the random number, it uses the pre-stored private key to decrypt the session key. A dual-channel verification mechanism is established, with the main channel transmitting encrypted instructions and the auxiliary channel transmitting the quantum hash of the instruction characteristic value. The receiving end can only execute the operation after comparing the decryption result with the quantum hash for consistency.

6. The method for controlling an intelligent speaker according to claim 5, wherein: In S105, a control signal is sent to the target device according to a preset protocol, and the feedback module of the smart speaker is activated to generate a voice response, including: A multi-level response strategy library is built, and response modes are selected based on the complexity of the command. Simple commands use pre-recorded voice clips, while complex operations use the TTS engine to dynamically generate statements with progress feedback. A secondary confirmation step is inserted before important operations are executed, and a voice changing mechanism is used to distinguish system prompts from regular responses. When executing long tasks, voice prompts with an estimate of the remaining time are generated regularly, and the control panel lights present a progress animation. Differentiated response styles are set for different user identities, including adjustment of speech speed, politeness level, and information detail.

7. The method for controlling an intelligent speaker according to claim 6, wherein: When an abnormal device status is detected, a voice prompt containing a fault code is automatically generated. The abnormality handling includes: Monitor the device status data stream in real time during the execution of control instructions. When abnormal current fluctuations, sudden increases in communication delays, or inconsistent sensor data are detected, a level 3 emergency response is automatically triggered. The first-level response attempts to resend commands and reset the device, the second-level response switches to the backup control channel, and the third-level response cuts off the device power and sends an alarm; Establish a fault knowledge graph, match abnormal features with historical cases for similarity, give priority to verified solutions, and automatically generate diagnostic reports for new abnormal patterns and upload them to the cloud analysis platform. After receiving the returned repair strategies, update the local processing rules.

8. The method for controlling an intelligent speaker according to claim 7, wherein: Before collecting the user's voice command, the control device implements energy optimization management, including: The power consumption of each module is monitored in real time through current sensors, and a mapping table between working mode and energy consumption is established; Before collecting the user's voice command, the system starts a gradual sleep state, first turning off the display backlight, then reducing the wireless module scanning frequency, and finally entering a deep sleep state with only the voice wake-up circuit remaining. The system learns active time periods based on the user's daily routine and automatically switches to energy-saving configuration during inactive periods. When battery power is detected, it dynamically adjusts the CPU main frequency and network bandwidth to prioritize the battery life of core functions, and automatically identifies standby devices and sends shutdown commands in batches.

9. A control system for an intelligent speaker, characterized in that: The control method of the smart speaker according to any one of claims 1 to 8 is executed, wherein the control system comprises: A voice input module, configured to collect user voice commands and perform noise reduction processing, including a distributed microphone array and a preamplifier circuit; The voiceprint processing unit is electrically connected to the voice input module and has a built-in FPGA chip for real-time voiceprint feature extraction, including a fundamental frequency contour analysis module and a Mel-frequency cepstral coefficient calculation unit; A semantic parsing engine, deployed in the deep learning accelerator of the main control chip, is configured to receive processed voice commands and generate a structured set of operation instructions; A device collaborative controller, comprising a wireless communication module supporting multi-mode switching of Zigbee, Bluetooth 5.2, and Thread protocols, and a topology database memory, wherein the memory stores device three-dimensional space mapping data and protocol conversion rules; A dynamic encryption module with an integrated security chip and a quantum random number generator, configured to generate differentiated encryption keys based on device type and append a quantum hash check code when transmitting instructions; a feedback actuator, comprising a speech synthesis unit, an LED matrix panel, and a vibration motor, configured to generate multimodal interactive feedback; The environmental sensing subsystem consists of a temperature sensor, a humidity sensor, a six-axis gyroscope, and a light sensor, and is connected to the main control chip via the I2C bus; The power management unit includes a dynamic voltage regulation circuit and a power consumption monitoring chip, and is configured to switch the power supply mode according to the system load.

10. The intelligent speaker control system according to claim 9, characterized in that: The voiceprint processing unit includes: a voiceprint verification coprocessor configured to execute a dynamic time warping algorithm and including a dual-port RAM for storing a reference voiceprint feature matrix; a liveness detection circuit, integrating a bone conduction sensor and a laryngeal vibration analysis module, configured to distinguish between real human voices and recorded playback; The acoustic feature learning module is configured to update the voiceprint database through online learning, and includes a feature vector normalization unit and a similarity threshold adaptive adjustment circuit.

Citation Information

Cited By

  • Voice interaction method and system of intelligent electronic equipment

    CN120853601A

  • Multi-mode voiceprint identity intelligent verification method

    CN120877741A

  • Vehicle control method and device based on multi-modal information, electronic equipment and storage medium

    CN120998192A

  • Audio and video player control method based on voice instruction

    CN121053987A

  • Household intelligent voice interaction system and household intelligent voice interaction method

    CN121641023A