Voice equipment nearby wake-up method and system based on multi-sensor positioning
By using multi-sensor positioning and sound source positioning technologies, the user's coordinates are calculated and the nearest voice device is selected for wake-up, which solves the problem of misjudgment of nearby wake-up under the influence of environmental noise in the existing technology and achieves high-precision and stable nearby wake-up effect.
Patent Information
- Application Number
- CN202511866715.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-02-27
AI Technical Summary
Existing acoustic feature-based proximity wake-up technologies have a high misjudgment rate in noisy environments or multi-person scenarios, making it difficult to achieve stable and high-precision proximity wake-up.
Using multi-sensor positioning technology, a three-dimensional coordinate system is constructed, human presence sensors are deployed, the absolute coordinates of the subject to be woken up are calculated, and the sound source is located by combining the microphone of the voice device, and the nearest device is selected for wake-up.
It achieves stable, high-precision, proximity-based wake-up in complex acoustic environments, avoids the influence of environmental noise, is applicable to any region, requires no image processing, and ensures accurate wake-up in complex environments.
Smart Images

Figure CN121585488A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of nearby wake-up of smart home technology, and relates to a voice device nearby wake-up method based on multi-sensor positioning. BACKGROUND
[0002] With the rapid development of artificial intelligence and Internet of Things technology, whole-house intelligent systems are becoming more and more popular. In a smart home scene, multiple voice devices are usually required to cover the entire scene range. When a user issues a wake-up instruction, the ideal situation is that only the device closest to the user is woken up and responds to avoid interference from a call to a hundred responses, which has given rise to the nearby wake-up technology. At present, the mainstream technology of nearby wake-up is based on acoustic features of nearby wake-up. Its principle is that when multiple devices detect the same wake-up word, the features of the collected audio signals are further analyzed, which are converted into an acoustic index representing the distance. By comparing the index values reported by each device in the local area network, the device closest to the user is selected and woken up. However, this technology has a high misjudgment rate and is easily affected by the environment. In a noisy environment or when multiple people are chatting, the accuracy of the acoustic index will decrease significantly. SUMMARY
[0003] The purpose of the present application is to propose a voice device nearby wake-up method and system based on multi-sensor positioning to solve the problems existing in the prior art.
[0004] To achieve the above purpose, the following technical solutions are adopted: A voice device nearby wake-up method based on multi-sensor positioning, the method comprising: S1. Constructing a three-dimensional coordinate system for a target space; S2. Deploying at least three human presence sensors in the target space, and any coordinate point of a target area in the target space is covered by the detection range of at least three human presence sensors; Recording the sensor coordinates of each human presence sensor in the three-dimensional coordinate system; S3. In response to receiving a voice signal containing a wake-up word, collecting signal measurement values of each human presence sensor; S4. Based on the signal measurement values, calculating the current coordinates of the wake-up subject in the target area; S5. According to the current coordinates and the device coordinates of each voice device receiving the voice signal in the three-dimensional coordinate system, calculating the distance between the wake-up subject and each voice device; S6. Based on the distance, selecting a target voice device and controlling it to enter a wake-up state.
[0005] In the voice device nearby wake-up method based on multi-sensor positioning described above, step S4 specifically comprises: S41. According to the preset signal measurement value and distance relationship, the signal measurement value of each human body presence sensor is converted into the corresponding estimated distance of the wake-up subject to the corresponding human body presence sensor; S42. According to the sensor coordinates of at least three human body presence sensors and the corresponding estimated distances, the current coordinates of the wake-up subject are calculated by using a triangulation algorithm.
[0006] In the voice device nearby wake-up method based on multi-sensor positioning described above, the method further comprises a distance calibration step for the signal measurement value of the human body presence sensor: Collect the signal measurement value of each human body presence sensor on the human body at a plurality of known coordinates; Calculate the calibration distance based on the known coordinates and the sensor coordinates of each human body presence sensor; Based on multiple sets of calibration distances and signal measurement values, a calibration curve of the mapping relationship of each human body presence sensor about signal measurement value and distance is established; In step S41, the preset signal measurement value and distance relationship is determined based on the calibration curve.
[0007] In the voice device nearby wake-up method based on multi-sensor positioning described above, step S6 specifically comprises: S61. From all the voice devices that receive the wake-up voice, the voice devices with a distance less than the effective wake-up distance from the wake-up subject are selected as candidate devices. When there is more than one candidate device, step S62 is executed, otherwise the response is stopped; S62. The device closest to the wake-up subject is selected from the candidate devices as the target voice device, and the target voice device is controlled to enter the wake-up state.
[0008] In the voice device nearby wake-up method based on multi-sensor positioning described above, step S62 further comprises executing wake-up control when the same voice device is determined as the target device in a plurality of consecutive set collection periods.
[0009] In the voice device nearby wake-up method based on multi-sensor positioning described above, the target area is the area space in the target space from the first set height to the second set height from the ground.
[0010] In the voice device nearby wake-up method based on multi-sensor positioning described above, the method further comprises: When multiple wake-up subjects are detected, the multiple wake-up subjects are taken as candidate subjects, and the current coordinates of the multiple candidate subjects in the target area are obtained; determine a source area of the voice signal based on the voice source positioning technology; select a target from the multiple candidate subjects that best fits the source area as the final wake-up subject according to the current coordinates of each candidate subject, and execute step S5 with the final wake-up subject.
[0011] In the above-mentioned voice device wake-up method based on multi-sensor positioning, after the target voice device is woken up, other voice devices in the target space are prohibited from responding to wake-up; After detecting the end of voice interaction of the target voice device, the wake-up state of all voice devices is restored.
[0012] A voice device wake-up system based on multi-sensor positioning, comprising a sensor array comprising at least three human presence sensors deployed in a target space according to preset coordinates, for detecting the presence of a wake-up subject and outputting signal measurement values; at least two voice devices, each voice device having a device coordinate in the target space; a decision center communicatively connected to the sensor array and each voice device, for: in response to receiving a voice signal containing a wake-up word, collecting signal measurement values of each human presence sensor; calculating the current coordinates of the wake-up subject in the target space based on the signal measurement values; calculating the distance between the wake-up subject and each voice device according to the current coordinates and the device coordinates of each voice device that receives the voice signal; selecting a target voice device based on the distance and sending a wake-up instruction to it.
[0013] In the above-mentioned voice device wake-up system based on multi-sensor positioning, the decision center is integrated in a gateway device or the voice device; The human presence sensor is one or more of a millimeter wave radar sensor, an infrared sensor, or a wireless beacon sensor; The decision center is further configured to: When the sensor array detects multiple wake-up subjects, the multiple wake-up subjects are taken as candidate subjects, and the current coordinates of the multiple candidate subjects in the target area are obtained; determine a source area of the voice signal based on the voice source positioning technology; select a target from the multiple candidate subjects that best fits the source area as the final wake-up subject according to the current coordinates of each candidate subject.
[0014] The advantages of this invention are as follows: This solution does not rely on interference-prone acoustic features for distance comparison. It calculates the user's absolute coordinates in the target space through multi-sensor intersection ranging and triangulation algorithms, avoiding the influence of environmental noise on wake-up decisions and achieving stable, high-precision, proximity-based wake-up even in complex acoustic environments. The sensing and positioning process does not rely on visual sensors and does not require the collection or processing of any image or video data involving the user. It can be applied seamlessly to any area, ensuring the applicability of the technology and user acceptance. Through multi-sensor basic positioning and proximity-based wake-up of the wake-up subject, and when facing multiple potential wake-up subjects, it uses the microphone and sound source localization technology of the voice device itself to match the approximate sound source range with the spatial coordinates of each subject, thereby accurately determining the specific target for issuing the voice command and ensuring accurate wake-up in complex crowd environments. Attached Figure Description
[0015] Figure 1 This is a system block diagram of the voice device proximity wake-up system based on multi-sensor positioning according to the present invention; Figure 2 This is a schematic diagram of sensor distance calibration in the voice device proximity wake-up method based on multi-sensor positioning of the present invention; Figure 3 This is an example of a calibration curve showing the received signal strength versus distance for a millimeter-wave radar sensor according to an embodiment of the present invention. Figure 4 This is the general flowchart of the method for the voice device wake-up method based on multi-sensor positioning according to the present invention; Figure 5 This is a schematic diagram illustrating the triangulation calculation principle in the voice device wake-up method based on multi-sensor positioning of the present invention. Figure 6 This is a schematic diagram illustrating the sound source and coordinate matching in a multi-person scenario in the voice device wake-up method based on multi-sensor positioning according to the present invention. Detailed Implementation
[0016] This solution provides a method and system for nearby wake-up of voice devices based on multi-sensor positioning, such as... Figure 1 As shown, this system mainly consists of three parts: The sensor array consists of at least three human presence sensors, referred to simply as sensors. These human presence sensors can be any of millimeter-wave radar sensors, infrared sensors, or wireless beacon sensors; in this embodiment, the former is preferred. The human presence sensors detect the presence of a roused subject and output signal measurements, which, depending on the specific sensor used, can be in the form of signal strength RSSI, time-of-flight, infrared intensity, etc.
[0017] A group of voice devices, typically voice panels, consists of multiple voice devices distributed throughout the target space. These devices receive voice signals and respond according to instructions from the decision center.
[0018] The decision center, integrated into the gateway device or the voice device, is wirelessly connected to each sensor and each voice device. It is used to obtain voice signals and device status information from the voice device, obtain signal measurement values from each human presence sensor, execute positioning algorithms and decision logic, and issue wake-up commands.
[0019] Implementation methods include: First, construct a three-dimensional coordinate system for the target space, with the intersection of a corner (such as the southeast corner) and the ground as the origin (0,0,0), the extension direction of one wall as the X-axis, the extension direction of the other wall as the Y-axis, and the height direction as the Z-axis. Use tools such as a laser rangefinder to measure the room dimensions.
[0020] Based on the installation location of each voice device in space, record the coordinates of its center point in the above three-dimensional coordinate system, denoted as device coordinates, and enter them into the device database of the decision center.
[0021] The sensors are deployed high on the ceiling or wall, ensuring that any point within the target area is covered by the detection range of at least three sensors. The target area is the effective area where wake-up is desired; in this embodiment, it is defined as the space 0.5 meters to 2.0 meters above the ground, covering the sitting and standing height of an adult. The coordinates of each sensor's detection reference point (such as the antenna center) are measured and recorded as sensor coordinates, and entered into the sensor database of the decision center.
[0022] Subsequently, distance calibration was performed on the signal measurements from the sensors present on the human body: The signal measurement values of each human body presence sensor to the human body are collected at multiple known coordinates, and the calibration distance is calculated based on the known coordinates and the sensor coordinates of each human body presence sensor. Based on multiple sets of calibration distances and signal measurements, calibration curves are established to map the relationship between signal measurements and distance for each human body's presence sensor.
[0023] The specific implementation is as follows: like Figure 2As shown, the test human body is placed at N known coordinates (x, y, z) for testing in the target area, and N is preferably greater than 10. At each test coordinate point, all sensors are simultaneously detected, and the signal measurement value V of each sensor in a stable state is collected. In order to reduce random errors, it is preferred to collect multiple data at each test coordinate point and take the average. Record the data pair ((x, y, z), V) of each sensor for the test coordinate point, and calculate the distance of each sensor from the test human body according to the coordinates of each sensor and the known coordinates (x, y, z), respectively, which is called the calibration distance d, and then obtain a new data pair (d, V) about the distance d and the signal measurement value V. After multiple tests at multiple known coordinates, multiple data pairs (d, V) are obtained. A two-dimensional coordinate system with distance d and signal measurement value V as horizontal and vertical coordinates is constructed, and all data pair points are plotted in the two-dimensional coordinate system, and a calibration curve is obtained by curve fitting. Thus, a calibration curve is generated for each sensor, which represents the functional relationship between d and V, and is used to obtain the distance value from the real-time signal measurement value for the corresponding sensor. The more test data points, the stronger the representation ability of the calibration curve, and the more accurate the distance obtained based on the signal measurement value. Figure 3 An example of the calibration curve of the millimeter wave radar sensor outputting the signal strength RSSI in this embodiment is given.
[0024] Specifically, as shown in Figure 4 The near-wakeup method includes: Any voice device detects a voice signal containing a wake-up word, and immediately reports the trigger event to the decision center. The decision center receives the trigger event at the moment and collects a signal measurement value from all sensors. Thereafter, three situations may occur: The first situation is shown in Figure 4 The left branch, the wake-up subject sends a voice signal outside the target area, which may not be detected by enough sensors, i.e. no, only one or only two sensors detect the human body, at this time it is considered that there is no real human body in the target area, the decision center stops responding, and the process ends. This process can avoid the voice signal of external irrelevant personnel being mistakenly recognized as a wake-up voice by the system, leading to false wake-up.
[0025] The second situation is shown in Figure 4 The middle branch, when the wake-up subject initiates a voice wake-up in the target area, the system will obtain the current coordinates of the wake-up subject based on the following methods: According to the calibration curve of each sensor, the signal measurement value of each human body existing sensor is converted into the estimated distance of the wake-up subject to the corresponding human body existing sensor.
[0026] As Figure 5As shown, according to the coordinates of at least three human presence sensors and their corresponding estimated distances, a triangulation calculation is performed, and the coordinates of the three sensors are respectively: Sensor A: , and the estimated distance to the wake-up subject is ; Sensor B: , and the estimated distance to the wake-up subject is ; Sensor C: , and the estimated distance to the wake-up subject is ; The coordinates of the wake-up subject are set as The height of the human body can be assumed to be 1.5m, and solving the above formula can obtain the current coordinates of the wake-up subject .
[0027] The third case is shown in Figure 4 the right branch, when there are multiple possible wake-up subjects in the target area, the system will detect multiple wake-up subjects, which is manifested as positioning the human body in multiple positions, and through the above method, the current coordinates of each possible wake-up subject will be obtained. At this time, the decision center takes multiple wake-up subjects as candidate subjects, and calculates the approximate space of the wake-up word sound source based on the time difference of the voice signal received by each voice device based on the voice sound source positioning technology, which is called the source interval, Figure 6 An example of a source interval is provided, as shown by the dashed circle. According to the current coordinates of each candidate subject, the target that is closest to the center of the source interval, i.e., the candidate subject that is closest to the center of the source interval, is selected as the final wake-up subject, and the subsequent steps are performed with the final wake-up subject. As Figure 6 shown, three candidate subjects are monitored in the target area, and candidate subject 3 is closest to the center of the source interval and is taken as the final wake-up subject.
[0028] According to the current coordinates and the coordinates of all voice devices that have reported trigger events, the distance between the wake-up subject and these voice devices is calculated.
[0029] Based on the calculated distance, the target voice device closest to the wake-up subject is selected to control it to enter the wake-up state, and thus the voice wake-up is completed.
[0030] The decision center sends a wake-up and response instruction to the target device, and at the same time sends a lock instruction to all other voice devices to remain silent during this wake-up period. The target device is woken up and enters a listening state, ready to receive subsequent voice commands. When the decision center detects the end of this voice interaction session, it broadcasts an unlock instruction, and all devices return to the wake-upable listening state, waiting for the next wake-up and positioning.
[0031] The scheme primarily ensures the response efficiency in a single-person scenario, and achieves the nearest wake-up through fast positioning; meanwhile, the wake-up accuracy in a multi-person scenario is further taken into account, and the issuer of the instruction is locked through the spatial matching of the sound source and the coordinates. The design enables the system to perform excellently in different scenarios, and achieves the unity of efficiency and accuracy.
[0032] Further, preferably, after the distances between all the voice devices receiving the wake-up voice and the wake-up subject are obtained, the devices with distances less than the effective wake-up distance are screened out as candidates. If there is only one candidate, it is directly selected; if there are multiple candidates, the one closest to the target device is selected; if there is no candidate, the response is ended. The effective wake-up distance is preferably 3 m. In this way, the system can be effectively prevented from being triggered by irrelevant personnel outside the target area, for example, when the detection range of the sensor extends to the corridor and detects passing personnel, the system can exclude them from the effective wake-up range by comparing the calculated distance with the effective wake-up distance, thereby maintaining the accuracy of the core area interaction.
[0033] Further, the embodiment preferably calculates the results to point to the same target device in the next 2-3 acquisition cycles, such as every 100 ms, and finally issues the wake-up instruction, which can prevent false triggering and improve the stability of wake-up.
[0034] The scheme uses multi-sensor positioning technology in the field of voice device nearest wake-up, and creatively combines it with the basic sound source positioning technology, and solves the problems of nearest wake-up efficiency and accuracy in the nearest wake-up of smart home in steps. Not only is the performance excellent in an ideal quiet single-person scenario; but also in a noisy (such as music, external noise, etc.) single-person scenario, excellent performance can be guaranteed; at the same time, in a complex environment with noise, reverberation, and multiple people, it also exhibits stability and accuracy that cannot be matched by traditional acoustic solutions.
[0035] The specific embodiments described herein are merely illustrative of the spirit of the present application. Those skilled in the art of the present application can make various modifications or supplements to the described specific embodiments or replace them with similar ways, without deviating from the spirit of the present application or exceeding the scope defined by the appended claims.
Claims
1. A method for nearby wake-up of a voice device based on multi-sensor positioning, characterized in that, The method includes: S1. Construct a three-dimensional coordinate system for the target space; S2. Deploy at least three human presence sensors in the target space, and any coordinate point of the target area within the target space is covered by the detection range of at least three human presence sensors; Record the sensor coordinates of each human body sensor in the three-dimensional coordinate system; S3. In response to receiving a voice signal containing a wake word, acquire signal measurement values from each person's presence sensor; S4. Based on the signal measurement values, calculate the current coordinates of the awakened subject within the target area; S5. Calculate the distance between the wake-up subject and each of the voice devices in the three-dimensional coordinate system based on the current coordinates and the device coordinates of each voice device that received the voice signal; S6. Based on the distance, select the target voice device and control it to enter the wake-up state.
2. The method for wake-up of a voice device based on multi-sensor positioning according to claim 1, characterized in that, Step S4 specifically includes: S41. Based on the preset relationship between signal measurement values and distance, convert the signal measurement values of each human presence sensor into the estimated distance from the wake-up subject to the corresponding human presence sensor; S42. Based on the sensor coordinates of at least three human presence sensors and their corresponding estimated distances, the current coordinates of the awakened subject are calculated using a triangulation algorithm.
3. The method for wake-up of a voice device based on multi-sensor positioning according to claim 2, characterized in that, This method also includes a step of distance calibration of the signal measurement values of the human presence sensor: Collect signal measurements of the human body from sensors at multiple known coordinates; The calibration distance is calculated based on the known coordinates and the sensor coordinates of the sensors present on each human body; Based on multiple sets of calibration distances and signal measurement values, calibration curves are established to show the mapping relationship between the sensor values and distance for each human body. In step S41, the relationship between the preset signal measurement value and the distance is determined based on the calibration curve.
4. The method for wake-up of a voice device based on multi-sensor positioning according to claim 1, characterized in that, Step S6 specifically includes: S61. From all voice devices that have received the wake-up voice, select voice devices whose distance from the wake-up subject is less than the effective wake-up distance as candidate devices. If there is more than one candidate device, proceed to step S62; otherwise, stop responding. S62. Select the device closest to the wake-up subject from the candidate devices as the target voice device, and control it to enter the wake-up state.
5. The method for wake-up of a voice device based on multi-sensor positioning according to claim 4, characterized in that, Step S62 further includes performing wake-up control when the same voice device is determined to be the target device in multiple consecutive set acquisition cycles.
6. The method for wake-up of a voice device based on multi-sensor positioning according to claim 1, characterized in that, The target area is the area within the target space, extending from a first predetermined height above the ground to a second predetermined height.
7. The method for wake-up of a voice device based on multi-sensor positioning according to claim 1, characterized in that, This method also includes: When multiple wake-up subjects are detected, the multiple wake-up subjects are used as candidate subjects, and the current coordinates of the multiple candidate subjects in the target area are obtained; The source range of the speech signal is determined based on speech source localization technology; Based on the current coordinates of each candidate subject, the target that best matches the source range is selected from multiple candidate subjects as the final wake-up subject, and step S5 is executed with the final wake-up subject.
8. The method for wake-up of a voice device based on multi-sensor positioning according to claim 1, characterized in that, Once the target voice device is woken up, other voice devices within the target space are prohibited from responding to the wake-up call. After the voice interaction of the target voice device is detected to have ended, the wake-up state of all voice devices is restored.
9. A voice device proximity wake-up system based on multi-sensor positioning, characterized in that, include The sensor array includes at least three human presence sensors deployed at preset coordinates within the target space, used to detect the presence of the awakened subject and output signal measurement values; At least two voice devices, each voice device having device coordinates within the target space; The decision center, which is communicatively connected to the sensor array and each of the voice devices, is used for: In response to receiving a voice signal containing a wake word, the system collects signal measurement values from each person's presence sensor. Calculate the current coordinates of the awakened subject within the target space based on the signal measurement values; Based on the current coordinates and the device coordinates of each voice device that received the voice signal, calculate the distance between the wake-up subject and each of the voice devices; The target voice device is selected based on the distance, and a wake-up command is sent to it.
10. The voice device proximity wake-up system based on multi-sensor positioning according to claim 9, characterized in that, The decision center is integrated into the gateway device or the voice device; The human body presence sensor is one or more of millimeter-wave radar sensors, infrared sensors, or wireless beacon sensors; The decision-making center is also used for: When the sensor array detects multiple wake-up subjects, the multiple wake-up subjects are used as candidate subjects, and the current coordinates of the multiple candidate subjects in the target area are obtained; The source range of the voice signal is determined based on voice source localization technology by utilizing the time difference between the microphone of the voice device and the voice signal received. Based on the current coordinates of each candidate subject, the target that best matches the source range is selected from multiple candidate subjects as the final wake-up subject.