Equipment control method and device, equipment, medium and product

By analyzing voice feature information through a central voice device, the target voice device is identified to respond to voice commands, solving the problem of multiple devices responding simultaneously and improving the user experience.

CN121884797APending Publication Date: 2026-04-17GUANGDONG MURORA INTELLIGENT LIGHTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG MURORA INTELLIGENT LIGHTING CO LTD
Filing Date
2024-10-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing smart home systems, multiple devices respond simultaneously when a user issues a voice command, leading to conflicts and misoperations between devices and affecting the user experience.

Method used

The central voice device acquires voice feature information from multiple voice devices, determines the target voice device from among the multiple voice devices based on the spatial characteristics of the sound source relative to the voice devices, and controls the target voice device to respond to voice commands.

Benefits of technology

This reduces the number of devices responding to voice commands, avoids conflicts and misoperations between devices, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884797A_ABST
    Figure CN121884797A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment control method and device, equipment, a medium and a product, and relates to the field of equipment control. The method comprises the following steps: a central voice device obtains voice feature information obtained by a plurality of voice devices based on a voice command, wherein the plurality of voice devices comprise the central voice device; determining at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice devices; and controlling at least one target voice device to respond to the voice command. According to the method, the number of the devices responding to the voice command is reduced, the problems of conflicts and misoperation caused by the fact that the voice devices receiving the voice command respond to the voice command are avoided, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of equipment control, and in particular to a method, apparatus, equipment, medium and product for controlling equipment. Background Technology

[0002] With the rapid development of IoT technology, voice devices are increasingly being used in smart homes, smart offices, and other fields. Many home devices provide users with a convenient way to control them through voice interaction systems, allowing them to operate the devices with simple voice commands.

[0003] Currently, most smart home systems use voice devices in an instant response mode, meaning that after detecting a user's voice command, all devices in listening mode will attempt to recognize and execute the command.

[0004] However, this response mode has a problem of "one call, a hundred responses" in actual use. That is, when a user issues a voice command, it often triggers multiple devices to respond at the same time. This not only disrupts the user's normal operation, but may also lead to conflicts and misoperations between devices, thus affecting the overall user experience. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and product for controlling a device. The technical solution is as follows:

[0006] On one hand, a method for controlling a device is provided, the method being executed by a central voice device, the method comprising:

[0007] Acquire voice feature information obtained from multiple voice devices based on voice commands, wherein the multiple voice devices include the central voice device;

[0008] Based on the spatial characteristics of the sound source of the voice command relative to the voice device indicated by the voice feature information, at least one target voice device is determined from the plurality of voice devices;

[0009] Control the at least one target voice device to respond to the voice command.

[0010] On the other hand, a control device for an apparatus is provided, the device comprising:

[0011] The acquisition module is used to acquire voice feature information obtained by multiple voice devices based on voice commands, including the central voice device.

[0012] The determining module is used to determine the at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device;

[0013] A control module is used to control the at least one target voice device to respond to the voice command.

[0014] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the device control method as described in any of the embodiments of this application above.

[0015] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the control method of the device as described in any of the embodiments of this application above.

[0016] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the control method of the device described in any of the above embodiments.

[0017] The technical solution provided in this application includes at least the following beneficial effects:

[0018] In a voice control environment comprising multiple voice devices, when a user issues a voice command, and all the voice devices in the environment receive the command, the central voice device analyzes and judges the voice feature information corresponding to each voice device. Based on the spatial characteristics of the sound source indicated by the voice feature information relative to each voice device, the target voice device that can respond to the voice command is determined from among the multiple voice devices, and the target voice device is controlled to respond to the voice command. This reduces the number of responding devices for the voice command, avoids conflicts and misoperations caused by all voice devices receiving the voice command responding to the voice command, and improves the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a structural block diagram of a voice control system provided in an exemplary embodiment of this application;

[0021] Figure 2 This is a flowchart of a device control method provided in an exemplary embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the control process of a device provided in an exemplary embodiment of this application;

[0023] Figure 4 This is a flowchart of a device control method provided in an exemplary embodiment of this application;

[0024] Figure 5 This is a schematic diagram of a voice control scenario provided in an exemplary embodiment of this application;

[0025] Figure 6 This is a schematic diagram illustrating the process of aligning a mobile phone speaker with the microphone array of a voice device to be calibrated, provided in an exemplary embodiment of this application.

[0026] Figure 7 This is a schematic diagram illustrating the calibration of acoustic direction compensation parameters provided in an exemplary embodiment of this application;

[0027] Figure 8 This is a flowchart of a device control method provided in an exemplary embodiment of this application;

[0028] Figure 9 This is a flowchart of a device control method provided in an exemplary embodiment of this application;

[0029] Figure 10 This is a structural block diagram of the control device of a device provided in an exemplary embodiment of this application;

[0030] Figure 11 This is a structural block diagram of the control device of a device provided in an exemplary embodiment of this application;

[0031] Figure 12 This is a structural block diagram of a voice device provided in an exemplary embodiment of this application. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0033] The following is an illustrative description of the voice control system involved in the embodiments of this application. The voice control system provided in the embodiments of this application includes multiple voice devices, including a central voice device and candidate voice devices connected to the central voice device.

[0034] Please refer to Figure 1 This diagram illustrates a structural block diagram of a voice control system provided in an exemplary embodiment of this application. The voice control system 100 includes a central voice device 110 and candidate voice devices 120.

[0035] The central voice device 110 and the candidate voice devices 120 are voice devices deployed in the current voice control environment, capable of responding to received voice commands. Optionally, the aforementioned voice control environment can be a small-scale environment such as a home, office, or vehicle; or a larger-scale environment such as an office building or residential community. Optionally, the aforementioned voice devices can be smart home devices that support voice functionality, such as smart speakers, smart TVs, smart refrigerators, smart curtains, and smart lighting systems.

[0036] In this embodiment, the central voice device 110 and the candidate voice device 120 are connected through a communication network.

[0037] Optionally, there is a wired communication connection between the central voice device 110 and the candidate voice device 120; alternatively, the central voice device 110 and the candidate voice device 120 are connected via short-range wireless communication technology (e.g., Bluetooth, ZigBee, etc.); alternatively, the central voice device 110 and the candidate voice device 120 are connected to the Internet via Wireless Fidelity (WiFi) or cellular mobile communication technology (e.g., 4G, 5G), and connected to the voice control server via the Internet, thereby realizing an indirect communication connection between the central voice device 110 and the candidate voice device 120 through the voice control server.

[0038] Please refer to Figure 2 The diagram illustrates a control method for a device provided in an exemplary embodiment of this application. The method is described using a central voice device as an example, and includes the following steps 210 to 230.

[0039] Step 210: Obtain voice feature information from multiple voice devices based on voice commands.

[0040] In illustrative terms, a voice device refers to an electronic device capable of recognizing and responding to voice commands. In some embodiments, the voice device has built-in voice recognition technology, allowing users to control the device to achieve diverse functions by speaking, without using traditional buttons or touchscreen interfaces. In the embodiments of this application, the voice device is configured in a voice control environment. Optionally, the aforementioned voice control environment can be a small environment such as a home environment, office environment, or vehicle environment; or it can be a larger environment such as an office building or residential community environment.

[0041] Optionally, the aforementioned voice device can be implemented as a smart home device that supports voice functionality, such as a smart speaker, smart TV, smart refrigerator, smart curtains, smart lighting system, etc. In one example, taking a smart speaker as the voice device, a user can control the smart speaker to automatically play music by saying "play music."

[0042] In this embodiment of the application, the multiple voice devices in the voice control environment include a central voice device and at least one candidate voice device connected to the central voice device, wherein the central voice device is a voice device that implements the overall control function of the voice devices in the voice control environment.

[0043] Optionally, the central voice device of the aforementioned multiple voice devices is user-defined; alternatively, the central voice device of the aforementioned multiple voice devices is determined based on the device layout of the multiple voice devices in the current voice control environment; alternatively, the central voice device of the aforementioned multiple voice devices is determined based on the priority of the voice devices.

[0044] In some embodiments, when the central voice device is user-defined, the user can manually select a voice device as the central voice device through a preset application on the mobile terminal. For example, the mobile terminal may run a smart home application that can control voice devices in the environment. Alternatively, when the user adds a voice device in the preset application, the preset application will default to using the first voice device added by the user as the central voice device.

[0045] In some embodiments, when the central voice device is determined based on the device layout of multiple voice devices in the current voice control environment, the central voice device is determined based on the average distance from each voice device to other voice devices. Illustratively, the device distances between the i-th voice device and other voice devices are obtained, the average distance between the i-th voice device and other voice devices is calculated, and the voice device with the smallest average distance among the multiple voice devices is determined as the central voice device, where i is a positive integer.

[0046] In some embodiments, when the central voice device among multiple voice devices is determined based on the priority of the voice devices, the device priorities corresponding to each of the multiple voice devices are obtained, and the voice device with the highest device priority among the multiple voice devices is determined as the central voice device. Optionally, the above-mentioned device priority is determined based on the device functional integration degree, that is, the device functional integration degree is positively correlated with the device priority, and the device functional integration degree is used to measure the number of functions integrated into the device.

[0047] Indicatively, a voice command refers to an instruction issued by a user to a voice device by speaking. The voice command is used to control at least one target voice device among multiple voice devices, wherein the at least one target voice device is the voice device determined by the central voice device in step 220.

[0048] In illustrative terms, speech feature information refers to parameters obtained by a speech device from detecting speech commands, used to describe and analyze speech characteristics. In this embodiment, the speech feature information acquired by the central speech device can characterize the spatial characteristics of the sound source of the speech command relative to the speech device.

[0049] Optionally, the aforementioned speech feature information includes at least one of sound energy value, sound direction angle, and sound arrival time.

[0050] The sound energy value indicates the energy level of the sound when a voice device receives a voice command. Specifically, the sound energy value refers to the energy carried by the sound wave propagating through a preset area per unit time. It can also be understood as the integral of sound intensity over time. The sound energy value reflects the total energy of the sound signal corresponding to the voice command, including factors such as sound intensity, frequency, and duration. For example, the power spectrum corresponding to the voice command is obtained. By performing a Fourier transform on the power spectrum and calculating the square of its amplitude, the power spectral density (PSD) can be obtained, which is used to characterize the sound energy value.

[0051] The sound direction angle is used to indicate the direction in which sound arrives when a voice device receives a voice command. Optionally, the sound direction angle includes at least one of the sound arrival angle or the sound direction deviation angle. The sound arrival angle is the angle between the sound wave ray and a reference direction (such as a horizontal plane or the normal to the horizontal plane) at the receiving point as the sound wave propagates from the sound source to the receiving point. The sound direction deviation angle is the angle between the actual arrival direction of the sound wave and the expected or reference direction. In sound source localization, the sound direction deviation angle describes the degree of deviation of the sound source relative to a reference point (such as the center point of a microphone array) or a reference direction (such as directly in front). For example, sound source localization of voice commands is achieved using microphone array technology, and the sound direction angle is determined based on the sound source localization results.

[0052] The sound arrival time is used to indicate the moment when the voice device receives a voice command. For example, when the voice device receives a voice command, it records the start and end times of the voice command and uses the start time as the sound arrival time.

[0053] Optionally, when receiving a voice command, the voice device detects the voice feature information corresponding to the voice command and sends the voice feature information to the central voice device; or, after receiving a voice command, the voice device sends the voice audio corresponding to the voice command to the central voice device, which then detects the voice audio to obtain the corresponding voice feature information.

[0054] Step 220: Based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device, at least one target voice device is determined from multiple voice devices.

[0055] Optionally, the selection of target voice devices can be achieved in at least one of the following ways:

[0056] The first method involves identifying voice devices whose voice feature information meets specified feature conditions as target voice devices.

[0057] Optionally, if the speech feature information includes sound energy value, the speech device whose sound energy value reaches a preset sound intensity threshold is identified as the target speech device.

[0058] Optionally, if the speech feature information includes the sound direction angle, the speech device with a sound direction angle less than a preset angle threshold is identified as the target speech device.

[0059] Optionally, if the speech feature information includes the time of sound arrival, the N speech devices with the earliest time of sound arrival are identified as the target speech devices, where N is a positive integer, for example, N is 1.

[0060] Optionally, the aforementioned preset sound intensity threshold, preset angle threshold, and value N can be implemented as system preset values ​​or as user-defined values.

[0061] The second method involves sorting multiple voice devices based on their voice feature information to obtain a device queue, and then determining the target voice device based on the device queue.

[0062] Optionally, when the speech feature information includes sound energy values, multiple speech devices are sorted based on their sound energy values ​​to obtain a first device queue. At least one target speech device is then determined from the multiple speech devices based on the first device queue, wherein the order of the speech devices in the first device queue is positively correlated with the magnitude of their sound energy values. In one example, the speech device at the head of the first device queue is determined as the target speech device; that is, the speech device with the highest sound energy value among the multiple speech devices is determined as the target speech device.

[0063] Optionally, when the speech feature information includes acoustic direction angle, multiple speech devices are sorted based on the acoustic direction angle to obtain a second device queue; at least one target speech device is determined from the multiple speech devices based on the second device queue, wherein the order of the speech devices in the second device queue is negatively correlated with the size of the acoustic direction angle. In one example, the speech device at the head of the second device queue is determined as the target speech device, that is, the speech device with the smallest acoustic direction angle among the multiple speech devices is determined as the target speech device.

[0064] Optionally, when the speech feature information includes the sound arrival time, multiple speech devices are sorted based on the sound arrival time to obtain a third device queue; at least one target speech device is determined from the multiple speech devices based on the third device queue, wherein the order of the speech devices in the third device queue is positively correlated with the order of their sound arrival times. In one example, the speech device at the head of the third device queue is determined as the target speech device, that is, the speech device with the earliest sound arrival time among the multiple speech devices is determined as the target speech device.

[0065] In some embodiments, when the speech feature information includes multiple candidate information, the target speech device is determined based on the multiple candidate information. Optionally, the multiple candidate information includes at least two of the following: sound energy value, sound direction angle, and sound arrival time.

[0066] Optionally, when the speech feature information includes multiple candidate information, the determination of the target speech device can be achieved in at least one of the following ways:

[0067] The first approach involves obtaining the information weights corresponding to each candidate information when the voice feature information includes multiple candidate information. These information weights indicate the importance of the candidate information in the judgment of the target voice device. The multiple candidate information corresponding to the i-th voice device are normalized to obtain normalized candidate information, where i is a positive integer. The normalized candidate information is then weighted and summed using the information weights to obtain the voice score corresponding to the i-th voice device. The voice score corresponding to the i-th voice device indicates the probability that the voice command is used to control the i-th voice device. Based on the voice scores corresponding to each voice device, the multiple voice devices are sorted, and the N voice devices with the highest voice scores after sorting are determined as the target voice devices, where N is a positive integer, for example, N = 1.

[0068] The second approach involves inputting a set of candidate information corresponding to multiple voice devices into a pre-trained voice direction prediction model when the voice feature information includes multiple candidate information. The voice direction prediction model then predicts the voice direction probability of the voice command pointing to each voice device based on the multiple candidate information, and obtains the prediction results. Based on the prediction results corresponding to each voice device, the multiple voice devices are sorted, and the N voice devices with the highest voice direction probability indicated by the sorted prediction results are determined as the target voice devices, where N is a positive integer, for example, N is 1.

[0069] The third approach involves determining the target voice device by judging the priority of different candidate information when the voice feature information includes multiple candidate information. Taking the arrangement of multiple candidate information according to information priority as an example, multiple voice devices are sorted based on the highest priority first candidate information to obtain a first queue. The first candidate information of the first voice device at the head of the first queue is compared with the first candidate information of the second voice device. If the difference between the first candidate information of the two is greater than a first threshold, the first voice device is determined as the target voice device. If the difference between the first candidate information of the two is less than or equal to the first threshold, the second candidate information corresponding to the first and second voice devices are obtained respectively, and the difference between the second candidate information of the two is compared. If the difference between the second candidate information of the two is greater than a second threshold, the voice device with higher second candidate information is determined as the target voice device. If the difference between the second candidate information of the two is less than or equal to the second threshold, the third candidate information of the two is further compared. That is, the candidate information of the voice devices is continuously compared according to the priority of the candidate information until the target voice device is determined.

[0070] In some embodiments, before determining at least one target voice device from multiple voice devices based on voice feature information, the central voice device pre-screens multiple voice devices based on the device information of the voice devices.

[0071] Optionally, the aforementioned device information includes at least one of the device function attributes and device status of the voice device. The device function attributes are used to indicate the functions provided by the voice device. For example, when the voice device is a smart speaker, the functions include music playback, weather query, and call functions. The device status is used to indicate the current operating status of the voice device. For example, the current device connection status and load status of the voice device.

[0072] In some embodiments, the central voice device performs intent recognition on the received voice command to obtain the intent recognition result corresponding to the voice command. Based on the intent recognition result and the device functional attributes corresponding to multiple voice devices, it pre-screens multiple voice devices to obtain a filtered set of voice devices. Based on voice feature information, it determines at least one target voice device from the filtered set of voice devices. Illustratively, the intent recognition result indicates the functional requirements corresponding to the voice command, and voice devices that can meet the functional requirements indicated by the intent recognition result are selected from multiple voice devices based on the intent recognition result.

[0073] In some embodiments, the central voice device pre-screens multiple voice devices based on their device status. In one example, when the device status includes load status, the pre-screening is performed based on the functional requirements corresponding to the voice command and the load status of the voice devices. Voice devices whose load status meets the load requirements corresponding to the functional requirements are selected as the pre-screened voice devices. At least one target voice device is then determined from the pre-screened voice devices based on voice feature information. In another example, when the device status includes device connection status, the pre-screening is performed based on the functional requirements corresponding to the voice command and the device connection status of the voice devices. Voice devices whose device connection status meets the connection requirements corresponding to the functional requirements are selected as the pre-screened voice devices. At least one target voice device is then determined from the pre-screened voice devices based on voice feature information. For example, if the voice command is "Call [user's name]", the central voice device will select voice devices with dialing functionality or voice devices connected to dialing devices (e.g., smart speakers connected to the user's mobile phone via Bluetooth) from the multiple voice devices to achieve pre-screening.

[0074] In some embodiments, the number of target voice devices is determined based on the recognition of voice commands. For example, the central voice device identifies the number of response devices in the voice command. If the voice command contains the number of response devices, the number of target voice devices is determined based on the number of response devices. For instance, if the voice command is "turn on three lights," the corresponding number of target voice devices is 3. If the voice command does not contain the number of response devices, the number of target voice devices is determined based on a preset number indication, for example, a unique target voice device is selected.

[0075] Step 230: Control at least one target voice device to respond to voice commands.

[0076] In this embodiment of the application, the multiple voice devices include a central voice device and at least one candidate voice device connected to the central voice device. Optionally, when the selected target voice device is the central voice device, the central voice device responds to the voice command; when the selected target voice device is a candidate voice device, the central voice device sends a voice response instruction to the candidate voice device, which is used to control the candidate voice device to respond to the voice command.

[0077] Optionally, the target voice device's response to the voice command includes at least one of device wake-up response, function execution, and function execution result feedback.

[0078] For example, such as Figure 3 The diagram illustrates the control process of a device provided in an exemplary embodiment of this application. The voice control environment 300 includes multiple voice devices, including a central voice device 310 and candidate voice devices 320. When user 301 issues a voice command, both the central voice device 310 and the candidate voice devices 320 receive the command. The candidate voice device 320 sends the voice feature information corresponding to the voice command to the central voice device 310. After selecting a target voice device 330 from the multiple voice devices based on the voice feature information, the central voice device 310 sends a voice response instruction to the target voice device 330. Upon receiving the voice response instruction, the target voice device 330 responds to the voice command issued by user 301.

[0079] In summary, in a voice control environment comprising multiple voice devices, when a user issues a voice command and all the voice devices in the environment receive the command, the central voice device among the multiple voice devices analyzes and judges based on the voice feature information corresponding to each voice device. Based on the spatial characteristics of the sound source indicated by the voice feature information relative to each voice device, the target voice device that can respond to the voice command is determined from among the multiple voice devices, and the target voice device is controlled to respond to the voice command. This reduces the number of responding devices for the voice command, avoids conflicts and misoperations caused by all voice devices receiving the voice command responding to the voice command, and improves the user experience.

[0080] In some optional embodiments, the speech feature information includes sound energy value and sound direction angle, wherein the sound energy value serves as the primary criterion for determining the target speech device, and the sound direction angle serves as an auxiliary criterion for determining the target speech device. Please refer to... Figure 4 The diagram illustrates a flowchart of a device control method provided in an exemplary embodiment of this application, which includes steps 221 to 224, wherein steps 221 to 224 are subordinate steps to step 220.

[0081] Step 221: If the speech feature information includes sound energy value, sort the multiple speech devices based on the sound energy value to obtain the first device queue.

[0082] In illustrative terms, the sound energy value indicates the magnitude of sound energy when a voice device receives a voice command. Specifically, the sound energy value refers to the energy carried by the sound wave propagating through a preset area per unit time. It can also be understood as the integral of sound intensity over time. The sound energy value reflects the total energy of the sound signal corresponding to the voice command, including factors such as sound intensity, frequency, and duration. For example, the power spectrum corresponding to the voice command is obtained. By performing a Fourier transform on the power spectrum and calculating the square of its amplitude, the power spectral density can be obtained, which is then used to characterize the sound energy value.

[0083] In this embodiment of the application, the order of voice devices in the first device queue obtained by sorting according to the sound energy value is positively correlated with the sound energy value of the voice devices. That is, the voice device with the higher sound energy value is closer to the head of the first device queue, and the voice device with the lower sound energy value is closer to the tail of the first device queue.

[0084] Step 222: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device in the first device queue is greater than the first threshold, the first voice device is identified as the target voice device.

[0085] In this embodiment of the application, the first voice device is the voice device at the head of the first device queue, that is, the first voice device is the voice device with the highest sound energy value among multiple voice devices, and the second voice device is the voice device in the first device queue that is second only to the first voice device, that is, the second voice device is the voice device in the first device queue that is one position after the first voice device.

[0086] In this embodiment of the application, when the first difference in sound energy values ​​between the first voice device and the second voice device is greater than the first threshold, the first voice device is directly identified as the target voice device. That is, if the difference in sound energy values ​​between the first voice device with the highest sound energy value and the second voice device with the second highest sound energy value is large enough, it indicates that the sound source of the voice command is significantly closer to the first voice device than to the other voice devices. Therefore, the first voice device is directly identified as the target voice device.

[0087] Optionally, the first threshold can be a value preset by the system, a value defined by the user, or a value determined based on the sound energy value of the second voice device.

[0088] In one example, the first threshold is 10% of the sound energy value corresponding to the second voice device. That is, when the sound energy value of the first voice device is greater than 10% of the sound energy value of the second voice device, the first voice device is identified as the target voice device.

[0089] Step 223: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold, obtain the first acoustic direction angle corresponding to the first voice device and the second acoustic direction angle corresponding to the second voice device.

[0090] In this embodiment of the application, the voice feature information received by the central voice device also includes a sound direction angle, which is used to indicate the direction in which the sound arrives when the voice device receives a voice command.

[0091] In this embodiment of the application, when the difference between the sound energy value of the first voice device and the sound energy value of the second voice device does not reach the first threshold, the sound direction angle is used to make the judgment. That is, the sound direction angle is introduced to correct the problem that the voice device closer to the sound source may be misjudged due to obstacles if the sound energy value is used to select the voice device closer to the sound source.

[0092] In this embodiment of the application, after determining the first voice device and the second voice device by the sound energy value, the first sound direction angle corresponding to the first voice device and the second sound direction angle corresponding to the second voice device are obtained respectively.

[0093] Step 224: Determine the target voice device from the first voice device and the second voice device based on the first sound direction angle and the second sound direction angle.

[0094] In some embodiments, the voice device with the smallest angle between the first sound direction angle and the second sound direction angle is determined as the target voice device. That is, if the first sound direction angle is less than or equal to the second sound direction angle, the first voice device is determined as the target voice device; if the first sound direction angle is greater than the second sound direction angle, the second voice device is determined as the target voice device.

[0095] In some embodiments, the target voice device is determined based on the difference between the first sound direction angle and the second sound direction angle. Schematic, if the second difference between the first sound direction angle and the second sound direction angle is greater than a second threshold, the voice device with the smallest sound direction angle among the first sound direction angle and the second sound direction angle is determined as the target voice device; if the second difference between the first sound direction angle and the second sound direction angle is less than or equal to the second threshold, the first voice device is determined as the target voice device.

[0096] Optionally, the second threshold can be a system-preset value or a user-defined value.

[0097] In one example, when the second difference between the first sound direction angle and the second sound direction angle is greater than 10°, the speech device with the smallest sound direction angle between the first and second sound direction angles is determined as the target speech device. When the second difference between the first and second sound direction angles is less than or equal to 10°, the speech device with the highest sound energy value (i.e., the first speech device) is still taken as the target speech device.

[0098] Specifically, in different scenarios (e.g., a home setting), sound propagation is affected by reflection and diffraction from obstacles such as walls, floors, and ceilings, leading to a redistribution of sound energy. Furthermore, when speaking from a particular location, the sound energy received at different locations within the environment is influenced by various factors, including but not limited to the directivity of the sound source, the volume of the room, reverberation time, and the shape, materials, and layout of the room. Therefore, due to the complexity of the scenario, simply comparing sound energy values ​​to select the device closest to the user becomes inapplicable. In this embodiment, a sound direction angle is introduced to assist in determining the target voice device when the sound energy values ​​of the corresponding voice devices are relatively similar.

[0099] Indicative, such as Figure 5The diagram illustrates a voice control scenario 500 provided in an exemplary embodiment of this application. The voice control scenario 500 includes voice device A510 and voice device B520. When a user issues a voice command from a voice position 530, due to the influence of an obstacle 511 near voice device A510, the voice energy value of voice device A510 when receiving the voice command is slightly lower than that of voice device B520. Therefore, a sound direction angle is introduced for further judgment. The sound direction angle α corresponding to voice device A510 is 20° (absolute value 20°), and the sound direction angle β corresponding to voice device B520 is -40° (absolute value 40°). Since the absolute value of sound direction angle α is less than the absolute value of sound direction angle β, voice device A510 is determined as the target voice device responding to the voice command.

[0100] Optionally, the sound direction angle includes at least one of the sound arrival angle or the sound direction deviation angle. The sound arrival angle is the angle between the sound wave ray and a reference direction (such as a horizontal plane or the normal to the horizontal plane) at the receiving point as the sound wave propagates from the sound source to the receiving point. The sound direction deviation angle is the angle between the actual arrival direction of the sound wave and the expected or reference direction. In sound source localization, the sound direction deviation angle describes the degree of deviation of the sound source relative to a reference point (such as the center point of a microphone array) or a reference direction (such as directly in front). For example, sound source localization of voice commands is achieved using microphone array technology, and the sound direction angle is determined based on the sound source localization results.

[0101] In some embodiments, the acoustic direction angle includes an acoustic direction offset angle between the acoustic wave arrival direction used to describe the voice command and a preset reference direction. The acoustic direction offset angle is obtained by: obtaining the acoustic wave arrival direction when the voice device receives the voice command; obtaining the preset reference direction corresponding to the voice device; and determining the acoustic direction offset angle based on the directional angle between the acoustic wave arrival direction and the preset reference direction.

[0102] In some embodiments, due to factors such as the layout of other objects in the voice control scenario, environmental noise, hardware differences between voice devices, and positional differences, the sound direction offset angle of the voice device's sound source localization may have errors. Therefore, a sound direction compensation parameter is introduced to determine the sound direction offset angle. Illustratively, the sound direction compensation parameter corresponding to the voice device is obtained, where the sound direction compensation parameter represents the sound offset of the voice device under the influence of environmental factors in its environment. Based on the directional angle between the sound wave arrival direction and a preset reference direction, and the correction of the directional angle by the sound direction compensation parameter, the sound direction offset angle is obtained.

[0103] To address the introduced acoustic direction compensation parameters, this application embodiment also provides a manual calibration method based on the speaker facing the microphone array to determine the acoustic direction compensation parameters, thereby achieving precise calibration of the microphone array in the device, improving the accuracy of sound source localization, and providing a more accurate acoustic direction offset angle. Illustratively, the calibration and verification method for the acoustic direction compensation parameters includes: receiving a first test audio emitted by the auxiliary device at a first position; determining a first acoustic direction offset angle corresponding to the first test audio based on the first test audio; receiving a second test audio emitted by the auxiliary device at a second position; determining a second acoustic direction offset angle corresponding to the second test audio based on the second test audio; correcting the second acoustic direction offset angle using the first acoustic direction offset angle to obtain a third acoustic direction offset angle; acquiring a first angle between the auxiliary device's orientation and a first standard direction when the first test audio is emitted, and a second angle between the auxiliary device's orientation and the first standard direction when the second test audio is emitted; determining the angle difference between the first and second angles; and determining the first acoustic direction offset angle as the acoustic direction compensation parameter if the third acoustic direction offset angle and the angle difference are less than a preset angle threshold.

[0104] In some embodiments, the aforementioned auxiliary device can be implemented as a mobile terminal (e.g., a mobile phone) that establishes a connection with a voice device. For example, taking a mobile phone running an auxiliary calibration program as an example, the calibration process of the acoustic direction compensation parameter is specifically implemented as follows:

[0105] 1. Preparation stage:

[0106] The user launches the mobile auxiliary calibration program and establishes a wireless communication connection (such as Wi-Fi or Bluetooth) with the voice device to be calibrated, which is used for communication between the voice device to be calibrated and the auxiliary calibration program. The user stands at a preset distance directly in front of the microphone array of the voice device to be calibrated (this preset distance can be indicated to the user by the auxiliary calibration program), holds the mobile phone horizontally, and the phone's speaker is facing the microphone array of the voice device to be calibrated.

[0107] Among them, such as Figure 6As shown, the process of ensuring the mobile phone's speaker is aligned with the microphone array of the voice device to be calibrated includes: S1, the mobile phone auxiliary calibration program prompts the user to hold the mobile phone 601 horizontally and place the end of the mobile phone 601 (the main speaker part) against the panel surface of the voice device 602 to be calibrated; S2, the mobile phone auxiliary calibration program displays a level 603, and after maintaining a horizontal position, prompts the user to confirm whether the current orientation is correct. After confirmation, the mobile phone auxiliary calibration program records the current position and starts the inertial navigation algorithm; S3, the mobile phone auxiliary calibration program prompts the user to move. During the movement, the mobile phone auxiliary calibration program displays the horizontal status of the mobile phone 601 and the orientation of the end of the mobile phone 601 (the main speaker part) on the interface in real time through the inertial navigation algorithm, and records the distance between the mobile phone 601 and the voice device 602 to be calibrated (perpendicular to the wall); S4, after reaching the specified distance, and the level indicator on the mobile phone auxiliary calibration program interface remains horizontal and the orientation prompt is accurate, it can be considered that the current position and orientation of the mobile phone 601 are aligned with the microphone array of the voice device 602 to be calibrated.

[0108] 2. Initialize calibration:

[0109] The mobile auxiliary calibration program displays a level and prompts the user to keep the phone level and stable. When the X / Y tilt angles of the level are both less than 1 degree, the auxiliary calibration program calculates and records the offset angle α between the phone's speaker and the north direction when it is facing the voice device to be calibrated. It then actively sends a test request to the voice device to be calibrated and emits a standard test tone for a certain period of time until the voice device to be calibrated finishes the initialization calibration and notifies the mobile auxiliary calibration program to enter the next stage. At this time, the mobile auxiliary calibration program obtains the compass information from the phone system.

[0110] During the initial calibration phase, after receiving a test request from the auxiliary calibration program on the mobile phone, the voice device to be calibrated receives and collects a standard test tone. It calculates and saves the sound direction offset angle θ0 by using the sound source localization method based on time delay difference. This value represents the deviation caused by environmental factors and is also the angle to be corrected and compensated (given that the test sound source is facing the microphone array, in a completely ideal situation the sound direction offset angle should be 0 degrees, i.e. there is no deviation, but in a non-ideal situation it is not 0 degrees, i.e. there is a deviation).

[0111] 3. Correction and Verification:

[0112] After the initial calibration phase ends, the correction and verification phase begins. The mobile auxiliary calibration program prompts the user to move their position, continue holding the phone horizontally, and ensure the phone's speaker is facing the voice device to be calibrated. The user then triggers the correction and verification process.

[0113] After the user triggers the calibration and verification, the mobile auxiliary calibration program calculates and records the offset angle α' of the current mobile phone speaker facing the voice device to be calibrated from due north for the second time, and actively sends a verification request to the voice device to be calibrated, while emitting a standard test tone for a certain period of time until the voice device to be calibrated returns the verification result.

[0114] During the calibration and verification process, the voice device to be calibrated receives a verification request from the mobile auxiliary calibration program, then receives and collects the standard test tone again, and calculates the sound direction offset angle θ1 using the sound source localization method based on time delay difference; θ1 is then compensated with θ0 obtained in the initial calibration stage to eliminate deviations caused by environmental factors, and the corrected sound direction offset angle θ = θ1 - θ0 is obtained, and this value is returned to the mobile auxiliary calibration program as the verification result.

[0115] After receiving the verification result returned by the voice device to be calibrated, the mobile auxiliary calibration program obtains the corrected acoustic offset angle θ from the result. In addition, it calculates the angle θ' = α - α' that the mobile phone speaker turned toward the voice device to be calibrated during the correction verification stage relative to the initialization calibration stage. It compares θ' with θ, and if the difference between the two is less than 5 degrees, the calibration is successful; otherwise, it fails.

[0116] 4. Calibration Completion and Recording:

[0117] After successful calibration, the mobile auxiliary calibration program sends a calibration success request to the voice device to be calibrated; if calibration fails, it sends a calibration failure request to the voice device to be calibrated.

[0118] Upon receiving a successful calibration request, the device under test (DUT) permanently saves the acoustic offset angle θ0 obtained during the initial calibration phase, using it as an acoustic offset parameter for sound source localization. Upon receiving a calibration failure request, the voice device under test deletes the acoustic offset angle θ0 obtained during the initial calibration phase.

[0119] like Figure 7 As shown, this diagram illustrates the calibration of acoustic direction compensation parameters provided in an exemplary embodiment of this application. During the initialization calibration phase 710, the mobile phone speaker located at position 1 faces the voice device 701 to be calibrated, and the offset angle of the mobile phone speaker from the due north direction is α. The acoustic direction offset angle detected by the voice device 701 to be calibrated is θ0. During the correction verification phase 720, the mobile phone speaker located at position 2 faces the voice device 701 to be calibrated, and the offset angle of the mobile phone speaker from the due north direction is α'. The acoustic direction offset angle detected by the voice device 701 to be calibrated is θ1. When the difference between the corrected acoustic direction offset angle θ = θ1 - θ0 and the angle θ' = α - α' rotated by the mobile phone speaker is less than 5 degrees, the calibration is determined to be successful, and the voice device 701 to be calibrated records θ0 as the acoustic direction compensation parameter.

[0120] Optionally, when a user starts the voice device for the first time, the voice device prompts the user to perform initial calibration of the acoustic direction compensation parameters. For example, the user is prompted by voice to perform initial settings on the voice device, which includes initial calibration of the acoustic direction compensation parameters. Alternatively, when the user adds the voice device to a preset application, the preset application prompts the user to perform initial settings. Optionally, the user is prompted to perform initial calibration of the acoustic direction compensation parameters after the voice device is formatted or initialized. Optionally, when the voice device detects that the duration between the first power connection time when the last power connection was made and the second power connection time when the current power connection is made reaches a preset power-off duration, the voice device prompts the user to perform initial calibration of the acoustic direction compensation parameters. Optionally, the device determination accuracy of the target voice device that responds to voice commands is obtained by the central voice device within a historical time period. If the device determination accuracy is lower than a preset accuracy threshold, the central voice device controls multiple voice devices to remind the user to perform initial calibration of the acoustic direction compensation parameters.

[0121] In summary, in a voice control environment comprising multiple voice devices, when a user issues a voice command and all the voice devices in the environment receive the command, the central voice device among the multiple voice devices analyzes and judges based on the voice feature information corresponding to each voice device. Based on the spatial characteristics of the sound source indicated by the voice feature information relative to each voice device, the target voice device that can respond to the voice command is determined from among the multiple voice devices, and the target voice device is controlled to respond to the voice command. This reduces the number of responding devices for the voice command, avoids conflicts and misoperations caused by all voice devices receiving the voice command responding to the voice command, and improves the user experience.

[0122] In some optional embodiments, the speech feature information includes sound energy value and sound arrival time, wherein the sound energy value serves as the primary criterion for determining the target speech device, and the sound arrival time serves as an auxiliary criterion for determining the target speech device. Please refer to [reference needed]. Figure 8 The diagram illustrates a flowchart of a device control method provided in an exemplary embodiment of this application. The method includes steps 221 to 222 and steps 225 to 226, wherein the above steps are subordinate steps to step 220.

[0123] Step 221: If the speech feature information includes sound energy value, sort the multiple speech devices based on the sound energy value to obtain the first device queue.

[0124] In illustrative terms, the sound energy value indicates the magnitude of sound energy when a voice device receives a voice command. Specifically, the sound energy value refers to the energy carried by the sound wave propagating through a preset area per unit time. It can also be understood as the integral of sound intensity over time. The sound energy value reflects the total energy of the sound signal corresponding to the voice command, including factors such as sound intensity, frequency, and duration. For example, the power spectrum corresponding to the voice command is obtained. By performing a Fourier transform on the power spectrum and calculating the square of its amplitude, the power spectral density can be obtained, which is then used to characterize the sound energy value.

[0125] In this embodiment of the application, the order of voice devices in the first device queue obtained by sorting according to the sound energy value is positively correlated with the sound energy value of the voice devices. That is, the voice device with the higher sound energy value is closer to the head of the first device queue, and the voice device with the lower sound energy value is closer to the tail of the first device queue.

[0126] Step 222: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device in the first device queue is greater than the first threshold, the first voice device is identified as the target voice device.

[0127] In this embodiment of the application, the first voice device is the voice device at the head of the first device queue, that is, the first voice device is the voice device with the highest sound energy value among multiple voice devices, and the second voice device is the voice device in the first device queue that is second only to the first voice device, that is, the second voice device is the voice device in the first device queue that is one position after the first voice device.

[0128] In this embodiment of the application, when the first difference in sound energy values ​​between the first voice device and the second voice device is greater than the first threshold, the first voice device is directly identified as the target voice device. That is, if the difference in sound energy values ​​between the first voice device with the highest sound energy value and the second voice device with the second highest sound energy value is large enough, it indicates that the sound source of the voice command is significantly closer to the first voice device than to the other voice devices. Therefore, the first voice device is directly identified as the target voice device.

[0129] Optionally, the first threshold can be a value preset by the system, a value defined by the user, or a value determined based on the sound energy value of the second voice device.

[0130] In one example, the first threshold is 10% of the sound energy value corresponding to the second voice device. That is, when the sound energy value of the first voice device is greater than 10% of the sound energy value of the second voice device, the first voice device is identified as the target voice device.

[0131] Step 225: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold, obtain the first sound arrival time corresponding to the first voice device and the second sound arrival time corresponding to the second voice device.

[0132] In this embodiment of the application, the voice feature information received by the central voice device also includes the sound arrival time. Indicatively, the sound arrival time is used to indicate the moment when the voice device receives the voice command. For example, when the voice device receives a voice command, it records the start and end times of the voice command, using the start time as the sound arrival time.

[0133] In this embodiment of the application, when the difference between the sound energy value of the first voice device and the sound energy value of the second voice device does not reach the first threshold, the time of sound arrival is used to determine the threshold.

[0134] In this embodiment of the application, after determining the first voice device and the second voice device by the sound energy value, the first sound arrival time corresponding to the first voice device and the second sound arrival time corresponding to the second voice device are obtained respectively.

[0135] Step 226: Determine the target voice device from the first voice device and the second voice device based on the first sound arrival time and the second sound arrival time.

[0136] In some embodiments, the voice device with the earliest sound arrival time between the first sound arrival time and the second sound arrival time is determined as the target voice device. That is, if the first sound arrival time is earlier than or equal to the second sound arrival time, the first voice device is determined as the target voice device; if the first sound arrival time is later than the second sound arrival time, the second voice device is determined as the target voice device.

[0137] In some embodiments, the target voice device is determined based on the difference between the first sound arrival time and the second sound arrival time. Schematic, if the third difference between the first sound arrival time and the second sound arrival time is greater than a third threshold, the voice device with the earliest sound arrival time among the first sound arrival time and the second sound arrival time is determined as the target voice device; if the third difference between the first sound arrival time and the second sound arrival time is less than or equal to the third threshold, the first voice device is determined as the target voice device.

[0138] Optionally, the aforementioned third threshold can be a system-preset value or a user-defined value.

[0139] In summary, in a voice control environment comprising multiple voice devices, when a user issues a voice command and all the voice devices in the environment receive the command, the central voice device among the multiple voice devices analyzes and judges based on the voice feature information corresponding to each voice device. Based on the spatial characteristics of the sound source indicated by the voice feature information relative to each voice device, the target voice device that can respond to the voice command is determined from among the multiple voice devices, and the target voice device is controlled to respond to the voice command. This reduces the number of responding devices for the voice command, avoids conflicts and misoperations caused by all voice devices receiving the voice command responding to the voice command, and improves the user experience.

[0140] In some embodiments, before the central voice device makes a decision about the target voice device based on the voice feature information of each voice device in the environment, each voice device uploads its voice feature information based on the matching between the voice command and pre-judgment conditions. Please refer to... Figure 9 The diagram illustrates a flowchart of a control method for a device provided in an exemplary embodiment of this application, the method comprising steps 901 to 909.

[0141] Step 901: The central voice device receives a voice command and detects the first voice feature information corresponding to the voice command when receiving the voice command.

[0142] In this embodiment, the central voice device is a voice device that implements a master control function for the voice devices in the voice control environment.

[0143] Indicatively, a voice command is an instruction issued by a user to a voice device by speaking. Voice commands are used to control at least one target voice device among multiple voice devices.

[0144] Schematic, the first speech feature information is a parameter used to describe and analyze speech characteristics, obtained by the central speech device from detecting the speech command. In this embodiment, the first speech feature information detected by the central speech device can characterize the spatial characteristics of the sound source of the speech command relative to the central speech device.

[0145] Optionally, the aforementioned first speech feature information includes at least one of a first sound energy value, a first sound direction angle, and a first sound arrival time.

[0146] Step 902: The candidate voice device receives a voice command and detects the second voice feature information of the voice command when receiving the voice command.

[0147] In this embodiment of the application, the candidate voice device is a voice device that can be controlled by the central voice device in a voice control environment.

[0148] Schematic, the second speech feature information is a parameter used to describe and analyze speech characteristics, obtained by the candidate speech device from detecting the speech command. In this embodiment, the second speech feature information detected by the candidate speech device can characterize the spatial characteristics of the sound source of the speech command relative to the candidate speech device.

[0149] Optionally, the aforementioned second speech feature information includes at least one of the following: second sound energy value, second sound direction angle, and second sound arrival time.

[0150] Step 903: The candidate speech device detects the matching between the second speech feature information and the pre-judgment conditions. If the second speech feature information fails to match the pre-judgment conditions, it sends the second speech feature information to the central speech device.

[0151] In this embodiment, the candidate voice device performs pre-detection based on the received voice command and determines whether the second voice feature information needs to be uploaded to the central voice device through pre-judgment conditions. That is, the second voice feature information is the information sent by the candidate voice device when it detects that the second voice feature information fails to match the pre-judgment conditions.

[0152] In some embodiments, taking the second voice feature information including the second sound energy value as an example, illustratively, when the candidate voice device detects that the second sound energy value corresponding to the voice command reaches the predicted threshold, the candidate voice device responds to the voice command; when the candidate voice device detects that the second sound energy value corresponding to the voice command is lower than the predicted threshold, the candidate voice device sends the second voice feature information to the central voice device.

[0153] In some embodiments, the pre-judgment conditions further include a judgment based on whether the functional requirements of the voice command are met. For example, the candidate voice device performs intent recognition on the received voice command to obtain the intent recognition result corresponding to the voice command. When the functional requirements indicated by the intent recognition result match the device functional attributes of the candidate voice device itself, the candidate voice device responds to the voice command or sends the second voice feature information to the central voice device based on the second sound energy value. When the functional requirements indicated by the intent recognition result do not match the device functional attributes of the candidate voice device itself, the candidate voice device ignores the voice command.

[0154] In some embodiments, the pre-judgment condition further includes a judgment based on whether the device status of the candidate voice device meets the state requirements for responding to the voice command. In one example, when the device status includes a load status, the candidate voice device obtains the load requirements of the functional requirements corresponding to the voice command. When it is determined that the candidate voice device's own load status meets the load requirements, it determines whether to respond to the voice command or send the second voice feature information to the central voice device based on a second sound energy value. When it is determined that the candidate voice device's own load status does not meet the load requirements, the candidate voice device ignores the voice command. In another example, when the device status includes a connection status, the candidate voice device obtains the connection requirements of the functional requirements corresponding to the voice command. When it is determined that the candidate voice device has connected to the target device indicated by the connection requirements, it determines whether to respond to the voice command or send the second voice feature information to the central voice device based on a second sound energy value. When it is determined that the candidate voice device has not connected to the target device indicated by the connection requirements, the candidate voice device ignores the voice command.

[0155] Indicatively, if the second speech feature information fails to match the pre-judgment condition, the central speech device receives the second speech feature information sent by at least one candidate speech device.

[0156] Step 904: The central voice device detects the matching between the first voice feature information and the pre-judgment conditions.

[0157] Step 905: If the first voice feature information matches the pre-judgment condition, the central voice device responds to the voice command and sends a voice ignore instruction to at least one candidate voice device.

[0158] In this embodiment, the central voice device performs pre-detection based on the received voice command and determines whether to directly respond to the voice command based on pre-judgment conditions.

[0159] In some embodiments, taking the first voice feature information including the first sound energy value as an example, illustratively, when the central voice device detects that the first sound energy value corresponding to the voice command reaches the predicted threshold, the central voice device responds to the voice command; when the central voice device detects that the first sound energy value corresponding to the voice command is lower than the predicted threshold, the central voice device determines the target voice device to respond to the voice command based on the second voice feature information of the received candidate voice devices and its own first voice feature information.

[0160] In some embodiments, the pre-judgment conditions further include a judgment based on whether the functional requirements of the voice command are met. For example, the central voice device performs intent recognition on the received voice command to obtain the intent recognition result corresponding to the voice command. When the functional requirements indicated by the intent recognition result match the device functional attributes of the central voice device itself, the central voice device responds to the voice command based on the first sound energy value or determines the target voice device from the central voice device and candidate voice devices based on the first and second voice feature information. When the functional requirements indicated by the intent recognition result do not match the device functional attributes of the central voice device itself, the central voice device ignores the voice command and determines the target voice device from the candidate voice devices based on the second voice feature information.

[0161] In some embodiments, the pre-judgment condition further includes a judgment based on whether the device status of the central voice device meets the state requirements for responding to the voice command. In one example, when the device status includes a load status, the central voice device obtains the load requirements of the functional requirements corresponding to the voice command. When it is determined that the load status of the central voice device itself meets the load requirements, it determines whether to respond to the voice command based on a first sound energy value or to determine the target voice device from the central voice device and candidate voice devices based on first voice feature information and second voice feature information. When it is determined that the load status of the central voice device itself does not meet the load requirements, the central voice device ignores the voice command and determines the target voice device from the candidate voice devices based on the second voice feature information. In another example, when the device status includes a connection status, the central voice device obtains the connection requirements of the functional requirements corresponding to the voice command. When it is determined that the central voice device has connected to the target device indicated by the connection requirements, it determines whether to respond to the voice command based on a first sound energy value or to determine the target voice device from the central voice device and candidate voice devices based on first voice feature information and second voice feature information. When it is determined that the central voice device has not connected to the target device indicated by the connection requirements, the central voice device ignores the voice command and determines the target voice device from the candidate voice devices based on the second voice feature information.

[0162] In this embodiment, when the first voice feature information matches the pre-judgment condition, the central voice device responds to the voice command. Simultaneously, the central voice device sends a voice ignore instruction to at least one candidate voice device. This voice ignore instruction controls the at least one candidate voice device to ignore the voice command. That is, when the central voice device determines itself to be the device responding to the voice command based on the pre-judgment condition, it no longer makes judgments based on the first and second voice feature information, but directly sends a voice ignore instruction to at least one candidate voice device to prevent the candidate voice device from responding to the voice command.

[0163] Step 906: If the first speech feature information fails to match the pre-judgment condition, the central speech device determines at least one target speech device from the central speech device and at least one candidate speech device based on the first speech feature information and the second speech feature information.

[0164] In this embodiment of the application, when the first voice feature information fails to match the pre-judgment condition, the central voice device determines at least one target voice device that responds to the voice command from the central voice device and at least one candidate voice device based on the first voice feature information and the second voice feature information.

[0165] In some embodiments, if the first voice feature information fails to match the pre-judgment conditions, the central voice device determines the number of candidate voice devices for the second voice feature information; obtains a preset number of candidate voice devices corresponding to the voice command; if the number of devices matches the preset number of devices, at least one target voice device is determined from the central voice device and at least one candidate voice device based on the first and second voice feature information; if the number of devices is less than the preset number of devices, a voice ignore instruction is sent to at least one candidate voice device; and the central voice device is controlled to ignore the voice command. That is, when the central voice device cannot determine whether it needs to respond to a voice command based on the pre-judgment conditions, it first determines the number of candidate voice devices sending the second voice feature information and obtains the preset number of devices corresponding to the voice command. If the number of candidate voice devices sending the second voice feature information matches the preset number of devices, it indicates that no voice device meets the pre-judgment conditions, and the target voice device is confirmed based on the first and second voice feature information. If the number of candidate voice devices sending the second voice feature information does not match the preset number of devices, it indicates that a candidate voice device has already responded to the voice command, and the central voice device directly controls other candidate voice devices and itself to ignore the voice command.

[0166] In other embodiments, if the first voice feature information fails to match the pre-judgment conditions, the central voice device detects whether it has received a voice response message from a candidate voice device. The voice response message is used to indicate that the candidate voice device has responded to the voice command according to the pre-judgment conditions. If the central voice device receives the voice response message, it sends a voice ignore instruction to the candidate voice device that has sent the second voice feature information, and controls itself to ignore the voice command.

[0167] Step 907: When the target voice device is the central voice device, the central voice device responds to the voice command and sends a voice ignore instruction to at least one candidate voice device.

[0168] Step 908: When the target voice device is the i-th candidate voice device, the central voice device ignores the voice command and sends a voice response command to the i-th candidate voice device, and sends a voice ignore command to the candidate voice devices other than the i-th candidate voice device.

[0169] Step 909: When the candidate voice device receives a voice response instruction, the candidate voice device responds to the voice command; or, when the candidate voice device receives a voice ignore instruction, the candidate voice device skips responding to the voice command.

[0170] In summary, in a voice control environment comprising multiple voice devices, when a user issues a voice command and all the voice devices in the environment receive the command, the central voice device among the multiple voice devices analyzes and judges based on the voice feature information corresponding to each voice device. Based on the spatial characteristics of the sound source indicated by the voice feature information relative to each voice device, the target voice device that can respond to the voice command is determined from among the multiple voice devices, and the target voice device is controlled to respond to the voice command. This reduces the number of responding devices for the voice command, avoids conflicts and misoperations caused by all voice devices receiving the voice command responding to the voice command, and improves the user experience.

[0171] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0172] Please refer to Figure 10 The diagram illustrates a control device structure block diagram of an exemplary embodiment of the device provided in this application, which includes the following modules:

[0173] The acquisition module 1010 is used to acquire voice feature information obtained by multiple voice devices based on voice commands, wherein the multiple voice devices include the central voice device;

[0174] The determining module 1020 is used to determine the at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device;

[0175] The control module 1030 is used to control the at least one target voice device to respond to the voice command.

[0176] In some optional embodiments, the voice feature information includes a sound energy value, which is used to indicate the energy level of the sound when the voice device receives the voice command;

[0177] like Figure 11 As shown, the determining module 1020 includes:

[0178] The sorting unit 1021 is used to sort the plurality of voice devices based on the sound energy value when the voice feature information includes the sound energy value, so as to obtain a first device queue;

[0179] The first determining unit 1022 is used to determine the at least one target voice device from the plurality of voice devices based on the first device queue.

[0180] In some optional embodiments, the order of the voice devices in the first device queue is positively correlated with the sound energy value of the voice devices;

[0181] The first determining unit 1022 is further configured to determine the first voice device as the target voice device when the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device in the first device queue is greater than a first threshold. The first voice device is the voice device located at the head of the first device queue, and the second voice device is the voice device in the first device queue that is second only to the first voice device.

[0182] In some optional embodiments, the voice feature information further includes a sound direction angle, which is used to indicate the direction in which the sound arrives when the voice device receives the voice command;

[0183] The determining module 1020 further includes:

[0184] The acquisition unit 1023 is used to acquire a first acoustic direction angle corresponding to the first voice device and a second acoustic direction angle corresponding to the second voice device when the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold.

[0185] The first determining unit 1022 is further configured to determine the target voice device from the first voice device and the second voice device based on the first sound direction angle and the second sound direction angle.

[0186] In some optional embodiments, the first determining unit 1022 is further configured to determine the voice device with the smallest sound direction angle among the first sound direction angle and the second sound direction angle as the target voice device when the second difference between the first sound direction angle and the second sound direction angle is greater than the second threshold.

[0187] The first determining unit 1022 is further configured to determine the first voice device as the target voice device when the second difference between the first sound direction angle and the second sound direction angle is less than or equal to the second threshold.

[0188] In some optional embodiments, the voice feature information further includes a sound arrival time, which is used to indicate the time when the voice device receives the voice command;

[0189] The acquisition unit 1023 is further configured to acquire the first sound arrival time corresponding to the first voice device and the second sound arrival time corresponding to the second voice device when the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold.

[0190] The first determining unit 1022 is further configured to determine the target voice device from the first voice device and the second voice device based on the first sound arrival time and the second sound arrival time.

[0191] In some optional embodiments, the first determining unit 1022 is further configured to determine the voice device with the earliest sound arrival time among the first sound arrival time and the second sound arrival time as the target voice device when the third difference between the first sound arrival time and the second sound arrival time is greater than the third threshold.

[0192] The first determining unit 1022 is further configured to determine the first voice device as the target voice device when the third difference between the first sound arrival time and the second sound arrival time is less than or equal to the third threshold.

[0193] In some optional embodiments, when the voice feature information further includes a sound direction angle, the sound direction angle includes a sound direction offset angle between the direction of arrival of the sound wave used to describe the voice command and a preset reference direction;

[0194] The acquisition module 1010 is further configured to acquire the direction of arrival of the sound wave when the voice device receives the voice command;

[0195] The acquisition module 1010 is further configured to acquire the preset reference direction corresponding to the voice device;

[0196] The acquisition module 1010 is further configured to determine the acoustic offset angle based on the directional angle between the direction of arrival of the sound wave and the preset reference direction.

[0197] In some optional embodiments, the acquisition module 1010 is further configured to acquire the sound direction compensation parameter corresponding to the voice device, wherein the sound direction compensation parameter is the sound deviation of the voice device under the influence of environmental factors in the environment.

[0198] The acquisition module 1010 is further configured to obtain the acoustic offset angle based on the directional angle between the direction of arrival of the sound wave and the preset reference direction, and the correction of the directional angle by the acoustic offset parameter.

[0199] In some alternative embodiments, the apparatus further includes:

[0200] The calibration module 1040 includes:

[0201] The receiving unit 1041 is used to receive the first test audio emitted by the auxiliary device at the first position;

[0202] The second determining unit 1042 is used to determine the first acoustic offset angle corresponding to the first test audio based on the first test audio.

[0203] The receiving unit 1041 is also used to receive a second test audio emitted by the auxiliary device at the second position;

[0204] The second determining unit 1042 is further configured to determine the second acoustic offset angle corresponding to the second test audio based on the second test audio;

[0205] The second determining unit 1042 is further configured to correct the second acoustic offset angle using the first acoustic offset angle to obtain a third acoustic offset angle;

[0206] The receiving unit 1041 is further configured to acquire a first angle between the device orientation of the auxiliary device and the first standard direction when the auxiliary device emits the first test audio, and a second angle between the device orientation of the auxiliary device and the first standard direction when the auxiliary device emits the second test audio;

[0207] The second determining unit 1042 is further configured to determine the angle difference between the first included angle and the second included angle;

[0208] The second determining unit 1042 is further configured to determine the first acoustic offset angle as the acoustic compensation parameter when the difference between the third acoustic offset angle and the included angle is less than a preset included angle threshold.

[0209] In some optional embodiments, the plurality of voice devices further includes at least one candidate voice device connected to the central voice device;

[0210] The acquisition module 1010 is further configured to receive the voice command and, when receiving the voice command, detect the first voice feature information corresponding to the voice command;

[0211] The acquisition module 1010 is further configured to receive second voice feature information sent by the at least one candidate voice device.

[0212] In some optional embodiments, the second voice feature information is information sent by the at least one candidate voice device when it detects that the second voice feature information fails to match the pre-judgment conditions;

[0213] The first determining unit 1022 is further configured to determine the number of candidate voice devices that send the second voice feature information when the first voice feature information fails to match the pre-judgment condition.

[0214] The acquisition unit 1023 is also used to acquire a preset number of candidate voice devices corresponding to the voice command;

[0215] The first determining unit 1022 is further configured to, when the number of devices matches the preset number of devices, determine the at least one target voice device from the central voice device and the at least one candidate voice device based on the first voice feature information and the second voice feature information.

[0216] In some alternative embodiments, the apparatus further includes:

[0217] The sending module 1050 is configured to send a voice ignore instruction to the at least one candidate voice device when the number of devices is less than the preset number of devices, wherein the voice ignore instruction is used to control the at least one candidate voice device to ignore the voice command.

[0218] The control module 1030 is also used to control the central voice device to ignore the voice command.

[0219] In some optional embodiments, the control module 1030 is further configured to control the central voice device to respond to the voice command when the first voice feature information matches the pre-judgment condition;

[0220] The sending module 1050 is further configured to send a voice ignore instruction to the at least one candidate voice device, the voice ignore instruction being used to control the at least one candidate voice device to ignore the voice command.

[0221] It should be noted that the control device of the equipment provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the equipment can be divided into different functional modules to complete all or part of the functions described above. In addition, the control device of the equipment provided in the above embodiments and the control method embodiments of the equipment belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0222] Figure 12 The diagram illustrates a structural block diagram of a voice device 1200 provided in an exemplary embodiment of this application. The voice device 1200 may be a smart speaker, smart TV, smart refrigerator, smart curtains, smart lighting control system, etc.

[0223] Typically, the voice device 1200 includes a processor 1201 and a memory 1202.

[0224] Processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1201 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as a central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1201 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0225] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 are used to store at least one instruction, which is executed by the processor 1201 to implement the device control method provided in the method embodiments of this application.

[0226] Indicatively, the voice device 1200 also includes other components 1203, as will be understood by those skilled in the art. Figure 12 The structure shown does not constitute a limitation on the voice device 1200, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0227] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may be a computer-readable storage medium included in the memory described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into the terminal. The computer-readable storage medium stores at least one instruction, at least one program segment, a code set, or an instruction set. The at least one instruction, the at least one program segment, the code set, or the instruction set is loaded and executed by the processor to implement the control method of the device described in any of the above embodiments.

[0228] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0229] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0230] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for controlling a device, characterized in that, The method is executed by a central voice device, and the method includes: Acquire voice feature information obtained from multiple voice devices based on voice commands, wherein the multiple voice devices include the central voice device; Based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device, at least one target voice device is determined from the plurality of voice devices; Control the at least one target voice device to respond to the voice command.

2. The method according to claim 1, characterized in that, The voice feature information includes a sound energy value, which is used to indicate the energy level of the sound when the voice device receives the voice command. The method of determining at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device includes: When the speech feature information includes the sound energy value, the plurality of speech devices are sorted based on the sound energy value to obtain a first device queue; The at least one target voice device is determined from the plurality of voice devices based on the first device queue.

3. The method according to claim 2, characterized in that, The order of the voice devices in the first device queue is positively correlated with the sound energy value of the voice devices. The step of determining the at least one target voice device from the plurality of voice devices based on the first device queue includes: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device in the first device queue is greater than a first threshold, the first voice device is determined as the target voice device. The first voice device is the voice device located at the head of the first device queue, and the second voice device is the voice device in the first device queue that is second only to the first voice device.

4. The method according to claim 3, characterized in that, The voice feature information also includes a sound direction angle, which is used to indicate the direction in which the sound arrives when the voice device receives the voice command. The method further includes: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold, the first sound direction angle corresponding to the first voice device and the second sound direction angle corresponding to the second voice device are obtained. The target voice device is determined from the first voice device and the second voice device based on the first sound direction angle and the second sound direction angle.

5. The method according to claim 4, characterized in that, Determining the target voice device from the first voice device and the second voice device based on the first sound direction angle and the second sound direction angle includes: If the second difference between the first acoustic direction angle and the second acoustic direction angle is greater than the second threshold, the voice device with the smallest acoustic direction angle among the first acoustic direction angle and the second acoustic direction angle is determined as the target voice device. If the second difference between the first acoustic direction angle and the second acoustic direction angle is less than or equal to the second threshold, the first voice device is determined to be the target voice device.

6. The method according to claim 3, characterized in that, The voice feature information also includes the sound arrival time, which is used to indicate the time when the voice device receives the voice command; The method further includes: If the first difference between the sound energy value of the first voice device and the sound energy value of the second voice device is less than or equal to the first threshold, the first sound arrival time corresponding to the first voice device and the second sound arrival time corresponding to the second voice device are obtained. The target voice device is determined from the first voice device and the second voice device based on the first sound arrival time and the second sound arrival time.

7. The method according to claim 6, characterized in that, Determining the target voice device from the first voice device and the second voice device based on the first sound arrival time and the second sound arrival time includes: If the third difference between the first sound arrival time and the second sound arrival time is greater than the third threshold, the voice device with the earliest sound arrival time among the first sound arrival time and the second sound arrival time is determined as the target voice device. If the third difference between the first sound arrival time and the second sound arrival time is less than or equal to the third threshold, the first voice device is determined to be the target voice device.

8. The method according to any one of claims 1 to 7, characterized in that, When the speech feature information further includes a sound direction angle, the sound direction angle includes a sound direction offset angle between the direction of arrival of the sound wave describing the speech command and a preset reference direction. The sound direction offset angle is obtained in the following ways: Obtain the direction of arrival of the sound wave when the voice device receives the voice command; Obtain the preset reference direction corresponding to the voice device; The acoustic offset angle is determined based on the directional angle between the direction of arrival of the sound wave and the preset reference direction.

9. The method according to claim 8, characterized in that, Determining the acoustic offset angle based on the difference between the direction of arrival of the sound wave and the preset reference direction includes: Obtain the sound direction compensation parameters corresponding to the voice device, wherein the sound direction compensation parameters are the sound deviation of the voice device under the influence of environmental factors in the environment. The acoustic offset angle is obtained based on the directional angle between the direction of arrival of the sound wave and the preset reference direction, and the correction of the directional angle by the acoustic offset parameter.

10. The method according to claim 9, characterized in that, The calibration and verification methods for the acoustic direction compensation parameters include: Receive the first test audio emitted by the auxiliary device at the first position; Determine the first acoustic offset angle corresponding to the first test audio based on the first test audio; Receive the second test audio emitted by the auxiliary device at the second position; Determine the second acoustic offset angle corresponding to the second test audio based on the second test audio; The third acoustic offset angle is obtained by correcting the second acoustic offset angle using the first acoustic offset angle. The first angle between the device orientation of the auxiliary device and the first standard direction is obtained when the auxiliary device emits the first test audio, and the second angle between the device orientation of the auxiliary device and the first standard direction is obtained when the second test audio is emitted; Determine the angle difference between the first included angle and the second included angle; If the difference between the third acoustic offset angle and the included angle is less than a preset included angle threshold, the first acoustic offset angle is determined as the acoustic compensation parameter.

11. The method according to any one of claims 1 to 7, characterized in that, The plurality of voice devices also includes at least one candidate voice device connected to the central voice device; The acquisition of voice feature information obtained from multiple voice devices based on voice commands includes: Receive the voice command, and detect the first voice feature information corresponding to the voice command when receiving the voice command; Receive second voice feature information sent by the at least one candidate voice device.

12. The method according to claim 11, characterized in that, The second voice feature information is the information sent by the at least one candidate voice device when it detects that the second voice feature information fails to match the pre-judgment conditions; The method of determining at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device includes: If the first voice feature information fails to match the pre-judgment condition, the number of candidate voice devices that send the second voice feature information is determined. Obtain the preset number of candidate voice devices corresponding to the voice command; When the number of devices matches the preset number of devices, the at least one target voice device is determined from the central voice device and the at least one candidate voice device based on the first voice feature information and the second voice feature information.

13. The method according to claim 12, characterized in that, The method further includes: If the number of devices is less than the preset number of devices, a voice ignore instruction is sent to the at least one candidate voice device, and the voice ignore instruction is used to control the at least one candidate voice device to ignore the voice command. The central voice device is controlled to ignore the voice command.

14. The method according to claim 12, characterized in that, The method further includes: If the first voice feature information matches the pre-judgment condition, the central voice device is controlled to respond to the voice command; Send a voice ignore instruction to the at least one candidate voice device, the voice ignore instruction being used to control the at least one candidate voice device to ignore the voice command.

15. A control device for an equipment, characterized in that, The device includes: The acquisition module is used to acquire voice feature information obtained by multiple voice devices based on voice commands, including the central voice device. The determining module is used to determine the at least one target voice device from the plurality of voice devices based on the spatial characteristics of the sound source of the voice command indicated by the voice feature information relative to the voice device; A control module is used to control the at least one target voice device to respond to the voice command.

16. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the control method of the device as described in any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the control method of the device as described in any one of claims 1 to 14.

18. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the control method of the device as described in any one of claims 1 to 14.