Smart home device response method and system based on user positioning and voiceprint recognition

By combining user location and voiceprint recognition, a smart home device response method has been developed, which resolves conflicts and personalization needs in multi-device collaborative response, and achieves accurate, convenient, and personalized smart home control, adapting to dynamic user scenarios and visitor scenarios.

CN122204576APending Publication Date: 2026-06-12ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE
Filing Date
2026-03-04
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing smart home devices suffer from frequent response conflicts, cumbersome operation, lack of personalized services, and shortcomings in adapting to visitor scenarios when responding to multiple devices in a collaborative manner. They are unable to achieve accurate, personalized, and all-scenario adaptability for multi-device smart home responses.

Method used

By combining user positioning with voiceprint recognition, a multi-microphone array is used to collect voice commands, the TDOA algorithm is used to calculate the user's location, and the identity is confirmed by voiceprint feature matching. The central control arbitrator performs a multi-level arbitration strategy to select a unique target device and sends suppression response signals to other devices to achieve precise control.

Benefits of technology

It achieves precise control of "one person, one command, one device", improves resource utilization and ease of operation, meets personalized needs, adapts to visitor scenarios, lowers the threshold for use, and improves the accuracy of positioning and identity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122204576A_ABST
    Figure CN122204576A_ABST
Patent Text Reader

Abstract

The application discloses a smart home device response method and system based on user positioning and voiceprint recognition, and relates to the technical field of smart home devices, comprising the following steps: collecting voice through a microphone array and positioning the user through a TDOA algorithm, confirming whether the user is a family member or a visitor in combination with voiceprint recognition, matching response preferences, selecting a unique target device through a multi-level arbitration strategy according to the physical position, state and instruction type of the device, and sending an execution instruction, while sending a response inhibition signal to other devices. The application solves the response conflict of multiple devices, realizes precise control of "one person, one order and one device", does not require manual intervention, and adapts to dynamic mobile scenarios of users. Customized responses can also be provided according to the identity of the user, and a bottom strategy of triggering public device priority is triggered for visitors, thereby improving the scene adaptability. In addition, the TDOA algorithm fuses multi-array data, has higher positioning accuracy, and the voiceprint recognition adopts multi-feature matching and high threshold screening, thereby effectively avoiding identity misjudgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of smart home devices, and specifically to a smart home device response method and system based on user location and voiceprint recognition. Background Technology

[0002] With the deep integration of IoT and AI technologies, smart home devices have evolved from single-function products to multi-device interconnected systems. Existing technologies have developed preliminary solutions for voice response control of multiple smart devices, mainly including the following three categories: Independent response from a single device: Early smart home devices only supported responding to their own voice commands and lacked cross-device collaboration capabilities. For example, smart speakers could only receive and execute commands such as "play music" or "check the weather" specific to themselves, and could not be linked with other devices such as smart air conditioners and lights. Users had to issue commands to different devices separately, which was cumbersome.

[0003] Signal strength priority response: Some multi-device systems employ a "voice signal strength judgment" strategy. This involves comparing the signal strength of the voice commands received by each device and selecting the device with the strongest signal to execute the command, attempting to avoid multiple devices responding simultaneously. For example, if a user issues a "turn on the lights" command in the living room, and the signal strength received by the living room smart speaker is higher than that of the smart lights in the bedroom, the living room speaker will respond. However, the speaker will be unable to control the lights, ultimately leading to command execution failure or misaligned response.

[0004] Manual Device Grouping Response: Some manufacturers offer device grouping functionality, requiring users to manually divide devices in different areas into multiple groups (such as "Living Room Group" and "Bedroom Group") via a mobile app, and designate a "master responder" for each group. Only the master device can receive and execute voice commands within the group. For example, if a user sets the smart speaker and TV in the living room as the master device in the "Living Room Group," and needs to control the bedroom air conditioner, they must first switch to the "Bedroom Group" in the app, making the process complex.

[0005] However, existing technologies have significant shortcomings and are unable to meet users' core needs for "intelligence, convenience, and personalization" in smart home systems. Specific problems are as follows: Frequent conflicting responses and chaotic user experience: The "signal strength priority" strategy has extremely low reliability in complex indoor environments. It is easily affected by factors such as wall obstruction, sound reflection, and device placement (e.g., devices near the sound source being blocked by furniture), which can lead to situations where "devices not near the user have stronger signals." For example, if a user issues a "turn on the air conditioner" command in the bedroom, the bedroom air conditioner may have a weak signal due to being blocked by a wardrobe, while the living room smart speaker may have a stronger signal due to the open space. Ultimately, the living room speaker may respond incorrectly (unable to execute the air conditioner control command), or the bedroom air conditioner and the living room speaker may respond simultaneously, resulting in wasted resources and operational confusion.

[0006] The operation is cumbersome and lacks flexibility: The "manual device grouping" mode requires users to frequently adjust group settings according to the activity area. For example, when a user moves from the living room to the bedroom, they need to open the app and switch the "response group" to the "bedroom group," otherwise the bedroom devices will not respond to commands. This mode cannot adapt to scenarios where users move dynamically, and its flexibility is severely lacking. It is particularly difficult for users with weaker operational skills, such as the elderly and children, to use.

[0007] Lack of personalized services and insufficient adaptability: Existing technology does not consider the differences in individual user habits. When all users issue the same command, the system adopts a uniform response logic. For example, when a father issues the command "turn on the light," he prefers to turn on the study light, and when a mother issues the same command, she prefers to turn on the kitchen light. However, the existing system cannot distinguish the user's identity and can only respond to the lights in a fixed area, failing to meet the personalized needs of family members.

[0008] Visitor scenario adaptation defects and limited functionality: When a visitor enters the room, the visitor does not have the system's preset identity information and device permissions. The existing system either directly refuses to respond (registration is required to use it) or randomly selects a device to respond (e.g., if the visitor gives a command in the living room, the device in the bedroom will respond), which makes it impossible for the visitor to use the smart home functions normally, resulting in poor scenario adaptability.

[0009] The closest existing technologies to this application, such as the smart home device based on voiceprint recognition in Chinese Patent CN201610025189.9, only implement identity verification and single device control based on voiceprint recognition, without involving arbitration logic for multi-device collaborative response, and thus cannot resolve the problem of multi-device response conflicts; and the smart home intrusion accurate identification and early warning method based on multi-sensor fusion in Chinese Patent CN202510463670.5, focuses on intrusion identification and early warning, without involving collaborative control of user positioning, identity recognition, and device response, and does not consider personalized needs and visitor scenario adaptation. None of the existing technologies have formed a complete technical system of "positioning-identification-arbitration-permission adaptation," and there is an urgent need for a smart home multi-device response solution that can achieve accurate, personalized, and full-scenario adaptation. Summary of the Invention

[0010] The purpose of this invention is to provide a smart home device response method and system based on user location and voiceprint recognition, so as to overcome the above-mentioned defects in the prior art.

[0011] The smart home device response method based on user location and voiceprint recognition includes the following steps: S1: Collect user voice commands and corresponding timestamps through multiple microphone arrays distributed indoors; S2: The central control arbitrator calculates the user's location using the Time Difference of Arrival (TDOA) algorithm based on the timestamps uploaded by each microphone array; S3: The voiceprint recognition module extracts voiceprint features from voice commands and matches them with pre-stored voiceprint models to confirm whether the user is a family member or a visitor. S4: The central control arbitrator obtains the corresponding device response preference based on the user's identity; S5: The central control arbitrator obtains the physical location, current status, and supported command types of each smart device in the room from the device management module; S6: The central control arbitrator selects a unique target device based on the user's location, user identity, device response preferences, and information about smart devices through a multi-level arbitration strategy; S7: The central control arbitrator sends an execution command to the target device and sends a suppression response signal to other smart devices.

[0012] Preferably, step S2 uses the TDOA algorithm to calculate the user's location, specifically including: S2.1: Extract the timestamps of the same voice command collected by each microphone array, and calculate the time difference of the command arriving at different microphone arrays; S2.2: Retrieve the pre-stored physical coordinates of each microphone array; S2.3: Based on the time difference and physical coordinates, the user's location coordinates are calculated using triangulation, with a positioning error of no more than 0.5 meters.

[0013] Preferably, step S3, which confirms the user's identity, specifically includes: S3.1: The voiceprint recognition module extracts the voiceprint frequency, tone, and timbre features of voice commands; S3.2: Perform similarity matching between the extracted voiceprint features and the pre-stored models in the voiceprint model library; S3.3: If the similarity is greater than or equal to 90%, the user is identified as a family member; if the similarity is less than 90% or no model is matched, the user is marked as a visitor.

[0014] Preferably, the multi-level arbitration strategy in step S6 is executed in order of priority, including: First priority: select smart devices that are within 10 meters of the user's location, are currently idle, and support the current voice command as candidate devices; If the number of candidate devices is 1, then the device is selected directly; If the number of candidate devices is greater than or equal to 2, then proceed to the second priority: if the user is a family member, then prioritize the candidate device that matches the device ranked first in the user's preference list; if the user is a visitor or has no preset preferences, then proceed to the third priority: prioritize the public devices preset by the user in the room. If the number of candidate devices is 0, a fallback strategy is implemented: filter the nearest available device that is not offline and send a confirmation prompt to the user, and select a device or terminate the process based on the user's confirmation result.

[0015] Preferably, in step S7, the suppression response signal includes the command ID of the voice command and the suppression duration. After receiving the suppression response signal, other smart devices will not respond to voice commands with the same command ID within the suppression duration.

[0016] A smart home device response system based on user location and voiceprint recognition for implementing the above method includes: A microphone array network, consisting of multiple microphone arrays distributed throughout the indoor area, is used to collect voice commands and timestamps. The voiceprint recognition module is used to store the voiceprint model library and perform voiceprint feature extraction and matching. The device management module is used to register, monitor, and store information about smart devices; The central control arbitrator, connected to the microphone array network, voiceprint recognition module, and device management module, is used to execute user location, arbitration strategy, and generate control commands. The communication module is used to establish communication connections between the central control arbitrator, microphone array network, voiceprint recognition module, device management module, and various smart devices.

[0017] Preferably, the communication module supports multiple communication protocols such as Wi-Fi, Zigbee, and Bluetooth, and can automatically switch to other available communication protocols to maintain data transmission stability when the signal strength of the current protocol is detected to be below -80dBm.

[0018] Preferably, the centrally controlled arbitrator includes: A positioning calculation unit is used to calculate the user's location based on the TDOA algorithm; The strategy execution unit is used to execute multi-level arbitration strategies. The instruction generation unit is used to generate execution instructions and suppression response signals.

[0019] Preferably, the smart device information stored in the device management module includes at least the device type, physical location, supported command types, and real-time status, wherein the real-time status includes idle, working, or offline.

[0020] The beneficial effects achieved by this invention are as follows: 1. This application resolves response conflicts and achieves precise control through a dual mechanism of "user location positioning + device status filtering," combined with a multi-level arbitration strategy to select a unique target device while simultaneously sending a "response suppression signal" to other devices. This fundamentally avoids the problem of multiple devices responding simultaneously or non-target devices responding erroneously, achieving precise control of "one person, one command, one device." Resource utilization is improved, and manual intervention is unnecessary. The system is highly intelligent, completing the entire process automatically without requiring users to manually group, switch devices, or adjust settings. Even if the user moves dynamically indoors, the system can update the responding device through real-time positioning, fully adapting to the user's dynamic scenarios and significantly improving ease of operation, especially reducing the user threshold for elderly and children.

[0021] 2. Personalized Experience Adaptation to Meet Diverse Needs: By linking user preference lists through voiceprint recognition, the system can provide customized responses based on user identity when different users issue the same command. Visitor-Friendly and Highly Adaptable: For visitor scenarios, the system automatically identifies visitors through voiceprint recognition and employs a fallback strategy that prioritizes public facilities in the user's room. This ensures that visitors can use core functions normally through nearby public facilities even without pre-set information, avoiding the problems of existing systems refusing or providing incorrect responses. The system's scenario adaptability is significantly superior to existing technologies.

[0022] 3. Positioning reliability: The TDOA algorithm combined with multi-microphone array data fusion can effectively counteract interference factors such as indoor sound reflection, wall obstruction, and device placement, reducing the user's location positioning accuracy error to a level far higher than that of existing signal strength-based positioning. The voiceprint recognition module uses multi-feature matching and high-threshold filtering to improve the accuracy of identity recognition and avoid misjudgment due to similar voiceprints. Attached Figure Description

[0023] Figure 1 This is a block diagram of the overall system architecture of the present invention.

[0024] Figure 2 This is a flowchart of the method of the present invention.

[0025] Figure 3 This is a schematic diagram illustrating an application scenario of the present invention.

[0026] Figure 4 This is a logic diagram of the arbitration strategy of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.

[0029] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0030] This invention provides a smart home device response method and system based on user location and voiceprint recognition; 1. This system achieves data interaction via Ethernet or wireless communication networks, with a response latency of ≤100ms, ensuring the real-time performance of user operations. Specifically, it includes the five core modules listed in Table 1: Table 1

[0031] 2. Response Method and Steps Based on the above system composition, the response method of this application realizes full-process control from voice command acquisition to device execution through 7 core steps, the specific process is as follows: S1: Voice command acquisition microphone arrays distributed in various indoor areas are in a real-time "standby listening" state (listening sensitivity can be adjusted via APP to avoid accidental triggering). When the user issues a voice command (such as "turn on music", "turn up the temperature", "turn on the lights"), all microphone arrays within the coverage area simultaneously acquire the audio data of the voice command and record their respective acquisition timestamps. Then, the audio data and timestamps are synchronously uploaded to the central control arbitrator through the communication module.

[0032] S2: The central control arbiter for user location positioning receives audio data and timestamps from each microphone array and uses the "Time Difference of Arrival (TDOA) algorithm" to calculate the sound source location. Step 2.1: Extract the timestamps of the same voice command collected by each microphone array, and calculate the time difference between the arrival of the command at different microphone arrays; Step 2.2: Call up the pre-stored physical coordinates of each microphone array (entered by the user through the APP, such as "living room microphone array: X=3.2m, Y=4.5m"); Step 2.3: Based on the time difference and microphone array coordinates, establish a spatial positioning model and calculate the precise location coordinates of the sound source (i.e., the user) using the triangulation method (error ≤ 0.5 meters), for example, "Bedroom area, coordinates X=2.5m, Y=3.8m".

[0033] S3: User Identity Verification The central control arbitrator synchronously sends the collected voice command audio data to the voiceprint recognition module to perform the following operations: Step 3.1: The voiceprint feature processing unit extracts voiceprint features from the audio (including voiceprint frequency range, pitch variation amplitude, timbre uniqueness, etc.). Step 3.2: Perform similarity matching between the extracted voiceprint features and the pre-stored models in the voiceprint model library, setting the matching threshold to 90%; Step 3.3: If the similarity is ≥90%, the user's identity is confirmed (e.g., "family member - Zhang San"), and the identity information is fed back to the central control arbitrator; if the similarity is <90% or no model is matched, the user is marked as "visitor", and the identity information is also fed back.

[0034] S4: User Preference Acquisition The central control arbitrator executes different preference retrieval logic based on the user identity confirmed by S3: If the user is a family member, the central control arbitrator retrieves the user's "device preference list" from the system database (preset by the user through the APP, such as "Zhang San's preferences: play music with the smart speaker in the bedroom, prioritize the smart light in the study to respond to the 'turn on the light' command, and prioritize the smart air conditioner in the bedroom to respond to the 'raise the temperature' command"). If the user is a guest, this step is skipped and the system automatically adopts the default preference policy of "public devices first".

[0035] S5: Device Status and Location Acquisition The central control arbitrator sends a device information query request to the device management module, which then provides real-time information on all responsive smart devices in the room, including: The physical location of the device (e.g., "smart speaker in living room: X=3.0m, Y=4.2m") is compared with the distance to the user's location coordinates calculated by S2. The current status of the device (idle / working / offline) will only include devices in the "idle" state in the candidate scope (devices in the "working" state must complete their current tasks first, while "offline" devices cannot receive instructions). The types of commands supported by the device (such as smart speakers supporting "play music" and "check the weather", smart air conditioners supporting "adjust temperature" and "on / off", and smart lights supporting "on / off" and "dim lighting") will only include devices that support the current voice commands in the candidate scope.

[0036] S6: Intelligent Arbitration Selection Device The central control arbitrator selects a unique target device based on the user location (S2), user identity (S3), preference list (S4), and device information (S5) through a preset multi-level arbitration strategy. The strategy priorities are as follows: First priority: Distance and state filtering Filter candidate devices that meet the criteria of "within 10 meters of the user's location + idle status + support for the current command"; if there is 1 candidate device, select the device directly; if there is 0, execute the fallback strategy; if there are ≥2, proceed to the second priority judgment.

[0037] Second priority: User preference matching If the user is a family member, the device ranked first in the user's preference list will be selected from the candidate devices (for example, if the user prefers to play music through a bedroom speaker, and the candidate devices include a bedroom speaker and a living room speaker, then the bedroom speaker will be selected); if the user has no preset preference (such as if it is the first time using the device), the third priority judgment will be made.

[0038] Third priority: Public facilities take precedence If the user is a visitor or has no preset preferences, the "public equipment in the user's room" will be selected first (public equipment is preset by the user and refers to the equipment in the room that mainly performs general functions, such as the smart speaker in the living room and the smart air conditioner in the bedroom); for example, if the user (visitor) issues the "turn on the light" command in the kitchen, the candidate equipment includes the smart light in the kitchen and the smart light in the living room, then the smart light in the kitchen will be selected.

[0039] Safety net strategy: If there are no available devices within 10 meters, then filter for available devices other than the nearest offline device to the user. Send a prompt to the user via the app (e.g., "The target device is far away (15 meters), do you want to proceed?"). Select the device after the user confirms. If the user cancels, the process ends.

[0040] S7: Instruction Execution and Suppression The central control arbitrator performs the following operations via the communication module: Send an "execution command" to the target device selected by S6. The command contains the specific content of the user's voice command (such as sending the command "play news (finance channel)" to the smart speaker in the bedroom). Send a "suppress response signal" to all other responsive devices (i.e., unselected candidate devices and non-candidate devices). The signal contains the command ID and the suppression duration (default 5 seconds) to ensure that these devices do not respond to the voice command within the suppression duration. After the target device executes the command, it sends the "execution result" (such as "playback successful" or "temperature adjustment to 26℃ completed") back to the central control arbitrator via the communication module. The central control arbitrator then synchronizes the result to the user's APP (optional), completing the entire response process.

[0041] like Figure 1 The overall system architecture diagram shows that the user issues a voice command → microphone array network (living room / bedroom / kitchen microphone array, labeled "audio acquisition + timestamp recording"). The microphone array network is connected to the central control arbitrator via a communication module (labeled "Wi-Fi / Zigbee / Bluetooth, latency ≤100ms"); The central control arbitrator is bidirectionally connected to the voiceprint recognition module (labeled "voiceprint model library + feature matching, accuracy ≥95%) and the device management module (labeled "device registration + status monitoring + information storage"). The central control arbitrator connects to the target smart device (such as a smart speaker in the bedroom, a smart air conditioner in the living room, and a smart light fixture in the kitchen, labeled "execute command / receive suppression signal") via a communication module. Each module is marked with an arrow indicating the data flow direction, with the core modules being the central control arbiter, microphone array network, and voiceprint recognition module.

[0042] like Figure 2 The method flowchart includes the following steps: S1: Voice command acquisition (audio and timestamps are acquired by microphone array and uploaded to the central control arbitrator); S2: User location positioning (TDOA algorithm calculates location, error ≤ 0.5 meters); S3: User identification (voiceprint feature matching to confirm family members / visitors); S4: User Preference Acquisition (Family members retrieve preference list, visitors skip); S5: Device Status and Location Acquisition (Query device location, status, and supported commands); S6: Intelligent Arbitration Device Selection (Multi-level Strategy: Distance Priority → Preference Priority → Public Equipment Priority → Last-Choice Prompt); S7: Instruction execution and suppression (send execution instructions to the target device and suppression signals to other devices); The steps are connected by arrows; S2 and S3 are marked "parallel execution," and S6 is marked "core decision-making step." like Figure 3 A schematic diagram illustrating a specific application scenario, showing a home floor plan with three main areas and equipment deployment marked: Living room: Smart speaker A, microphone array 1 (coordinates X=3.2m, Y=4.5m); Bedroom: Smart air conditioner B, microphone array 2 (coordinates X=2.5m, Y=3.8m); Kitchen: Smart lighting fixture C, microphone array 3 (coordinates X=6.1m, Y=2.3m); The user, located in the bedroom (location coordinates X=2.5m, Y=3.8m), issues the command "Play News"; Data flow direction is indicated by arrows: Microphone arrays 1 and 2 capture audio and timestamps → upload to the central control arbitrator; The central control arbitrator calculates the user's location (bedroom) and identifies the user as "Zhang San (family member)"; Retrieve Zhang San's preferences (bedroom devices preferred), and check device status (Bedroom air conditioner B is available and supports audio playback). Select air conditioner B as the target device, send a "play news" command to B, and send "suppress response" signals to A and C; Equipment status labeling: Air conditioner B is marked with a green checkmark indicating "in progress", while speaker A and light fixture C are marked with a red cross indicating "suppressed".

[0043] like Figure 4 The arbitration strategy logic diagram is shown in the flowchart, which illustrates the multi-level arbitration strategy: Input: User location, user identity (family member / visitor), device information (location + status + supported commands); Step 1: Filter candidate devices that meet the criteria of "within 10 meters + idle + support commands" → Determine the quantity: Quantity = 1 → Select this device; Quantity = 0 → Execute fallback strategy (select the nearest non-offline idle device → prompt user for confirmation → confirmation selects the device, cancellation terminates the process); Quantity ≥ 2 → Determine user identity: Family member → Select the first device in the preference list → Select the device; Visitor / No Preference → Select Public Equipment in User's Room → Select the Equipment; Each decision node is marked with a diamond box, the execution node is marked with a rectangle box, and the fallback strategy is highlighted with a dashed box.

[0044] The embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A smart home device response method based on user location and voiceprint recognition, characterized in that: Includes the following steps: S1: Collect user voice commands and corresponding timestamps through multiple microphone arrays distributed indoors; S2: The central control arbitrator calculates the user's location using the Time Difference of Arrival (TDOA) algorithm based on the timestamps uploaded by each microphone array; S3: The voiceprint recognition module extracts voiceprint features from voice commands and matches them with pre-stored voiceprint models to confirm whether the user is a family member or a visitor. S4: The central control arbitrator obtains the corresponding device response preference based on the user's identity; S5: The central control arbitrator obtains the physical location, current status, and supported command types of each smart device in the room from the device management module; S6: The central control arbitrator selects a unique target device based on the user's location, user identity, device response preferences, and information about smart devices through a multi-level arbitration strategy; S7: The central control arbitrator sends an execution command to the target device and sends a suppression response signal to other smart devices.

2. The smart home device response method based on user location and voiceprint recognition according to claim 1, characterized in that: Step S2 uses the TDOA algorithm to calculate the user's location, specifically including: S2.1: Extract the timestamps of the same voice command collected by each microphone array, and calculate the time difference of the command arriving at different microphone arrays; S2.2: Retrieve the pre-stored physical coordinates of each microphone array; S2.3: Based on the time difference and physical coordinates, the user's location coordinates are calculated using triangulation, with a positioning error of no more than 0.5 meters.

3. The smart home device response method based on user positioning and voiceprint recognition according to claim 1, characterized in that: Step S3 involves verifying the user's identity, specifically including: S3.1: The voiceprint recognition module extracts the voiceprint frequency, tone, and timbre features of voice commands; S3.2: Perform similarity matching between the extracted voiceprint features and the pre-stored models in the voiceprint model library; S3.3: If the similarity is greater than or equal to 90%, the user is identified as a family member; if the similarity is less than 90% or no model is matched, the user is marked as a visitor.

4. The smart home device response method based on user location and voiceprint recognition according to claim 1, characterized in that: The multi-level arbitration strategy in step S6 is executed in order of priority, including: First priority: select smart devices that are within 10 meters of the user's location, are currently idle, and support the current voice command as candidate devices; If the number of candidate devices is 1, then the device is selected directly; If the number of candidate devices is greater than or equal to 2, then proceed to the second priority: if the user is a family member, then prioritize the candidate device that matches the device ranked first in the user's preference list; if the user is a visitor or has no preset preferences, then proceed to the third priority: prioritize the public devices preset by the user in the room. If the number of candidate devices is 0, a fallback strategy is implemented: filter the nearest available device that is not offline and send a confirmation prompt to the user, and select a device or terminate the process based on the user's confirmation result.

5. The smart home device response method based on user location and voiceprint recognition according to claim 1, characterized in that: In step S7, the suppression response signal includes the command ID of the voice command and the suppression duration. After receiving the suppression response signal, other smart devices will no longer respond to voice commands with the same command ID within the suppression duration.

6. A smart home device response system based on user location and voiceprint recognition for implementing the method of any one of claims 1 to 5, characterized in that: include: A microphone array network, consisting of multiple microphone arrays distributed throughout the indoor area, is used to collect voice commands and timestamps. The voiceprint recognition module is used to store the voiceprint model library and perform voiceprint feature extraction and matching. The device management module is used to register, monitor, and store information about smart devices; The central control arbitrator, connected to the microphone array network, voiceprint recognition module, and device management module, is used to execute user location, arbitration strategy, and generate control commands. The communication module is used to establish communication connections between the central control arbitrator, microphone array network, voiceprint recognition module, device management module, and various smart devices.

7. The smart home device response based on user location and voiceprint recognition according to claim 6, characterized in that: The communication module supports multiple communication protocols, including Wi-Fi, Zigbee, and Bluetooth, and can automatically switch to other available communication protocols to maintain data transmission stability when the signal strength of the current protocol is detected to be below -80dBm.

8. The smart home device response based on user location and voiceprint recognition according to claim 6, characterized in that: The centrally controlled arbitrator includes: A positioning calculation unit is used to calculate the user's location based on the TDOA algorithm; The strategy execution unit is used to execute multi-level arbitration strategies. The instruction generation unit is used to generate execution instructions and suppression response signals.

9. The smart home device response based on user location and voiceprint recognition according to claim 6, characterized in that: The device management module stores information about smart devices, including at least the device type, physical location, supported command types, and real-time status, which includes whether the device is idle, in operation, or offline.

Citation Information

Patent Citations

  • Smart home device based on voiceprint recognition

    CN106972990A

  • Smart home intrusion accurate identification and early warning method based on multi-sensor fusion

    CN120199005A