Device control method and device based on voice interaction
The user's location and target area are determined through voice receiving devices, and home appliances are controlled and provided with feedback in combination with voice commands. This solves the problem that users have difficulty in accurately controlling devices in specific areas in the whole-house smart scenario, and realizes convenient voice interactive control.
Patent Information
- Application Number
- CN202410350616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
In a whole-house smart scenario, it is difficult for users to accurately control home appliances in a specific area through voice interaction. Existing technologies require users to accurately say the device name, resulting in low interaction efficiency.
The user's location and target area are determined through voice receiving devices, and the corresponding home devices are controlled in combination with voice commands. The execution results are fed back through visual or sound effects, or detailed voice feedback is provided in different areas to improve interaction efficiency.
Users can conveniently control home appliances in a specific area through simple voice commands, which improves the efficiency of voice interaction and user experience, and reduces misoperations and additional interference.
Smart Images

Figure CN120704189A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart home, and more specifically, to a device control method and device based on voice interaction. Background Art
[0002] In a whole-home smart scenario, users can control home appliances by interacting with smart devices through voice. For example, users can control home appliances by interacting with smart speakers through voice.
[0003] However, when users interact with smart devices through voice, they need to accurately say the name of the home device to control it. For controlling home devices in a specific area of the house, the current voice interaction method cannot meet this user's needs. Summary of the Invention
[0004] The present application provides a device control method and device based on voice interaction, which can realize convenient control of home appliances in a specific area space.
[0005] In a first aspect, the present application provides a device control method based on voice interaction, which is applied to a voice receiving device and includes: the voice receiving device receiving a user's voice command. The voice receiving device determines a target spatial area based on a target location or information indicating a spatial area in the voice command, wherein the target spatial area is the spatial area where the controlled device is located, and the target spatial area corresponds to the controlled device. The voice receiving device sends the spatial area information indicating the target spatial area and the voice command to a control device. The voice receiving device receives an execution result sent by the control device, which is the result of the controlled device executing the voice command. The voice receiving device provides feedback information based on the execution result. Based on an embodiment of the present application, the voice receiving device can parse the user's voice command to obtain the target spatial area that the user wishes to control, and then send the target spatial area and the user's voice command to the control device to control the corresponding controlled device. In this way, when a user interacts with a home appliance by voice, they can control the home appliance in a specific spatial area that the user wishes to control using a simple voice command, thereby improving the efficiency of voice interaction.
[0006] For example, the user's voice command may directly include the target location, such as here, here, by the wall, etc. The voice command may also include indication information indicating the area space, such as living room, master bedroom, sofa, etc.
[0007] In one implementation, the voice receiving device provides feedback information based on the execution result, including: when the user's position and the controlled device are in the same area space, the voice receiving device provides feedback information in a visual manner and / or in the form of a prompt sound effect based on the execution result. When the user's position and the controlled device are in different area spaces, the voice receiving device provides feedback information based on the execution result by indicating the execution result through voice instructions. Based on the embodiment of the present application, when the user's position and the controlled device are in the same area space, since the user can intuitively feel the execution result of the controlled device, the voice receiving device provides feedback information to the user in the form of visual effects and / or prompt sound effects to avoid causing additional interference to the user. When the user's position and the voice receiving device are not in the same area space, the voice receiving device can output the execution result by voice, which is beneficial for the user to obtain the execution result of the home appliance for his voice instruction.
[0008] In one implementation, the voice receiving device determines the target area space according to the target position in the voice command, including: determining the user's location according to the voice command, and determining the area space corresponding to the user's location as the target area space. Based on the embodiment of the present application, if the voice command is "turn on the light here", the voice receiving device can determine the user's location, and determine the area space corresponding to the user's location as the target area space. In this way, the user does not need to accurately say the area space where the home appliance to be controlled is located. Only simple voice commands are needed to conveniently control home appliances in certain area spaces, thereby improving the efficiency of voice interaction.
[0009] In one implementation, the voice receiving device determines a target area space based on the target location in the voice command, including: the voice receiving device determines that the voice command does not include instruction information, and determines a first area space corresponding to the target location as the target area space, where the first area space includes one or more sub-areas. In this way, the user does not need to accurately state the area space where the home appliance to be controlled is located; a simple voice command can be used to conveniently control home appliances in certain areas, thereby improving the efficiency of voice interaction.
[0010] For example, the voice command is "turn on the light here", which may correspond to the living room area. The living room area may include multiple sub-area spaces such as sofas and TV cabinets. The voice receiving device may use the entire living room area as the target area space.
[0011] In one implementation, the voice receiving device determines the target regional space based on the indication information for indicating the regional space in the voice instruction, including: determining that the voice instruction includes indication information, and determining the regional space indicated by the indication information as the target regional space. Based on the embodiment of the present application, when the user's voice instruction includes indication information for indicating the regional space, the voice receiving device can determine the target regional space as the regional space indicated by the indication information. This helps the voice receiving device determine which devices in the regional space are needed, thereby improving the efficiency of voice interaction.
[0012] For example, if the voice command is "turn on the light in the master bedroom", the instruction information is the master bedroom.
[0013] In one implementation, the voice receiving device determines a target regional space based on indication information indicating a regional space in the voice command, including: determining a target sub-regional space among multiple sub-regional spaces in the regional space based on the indication information, and determining the target sub-regional space as the target regional space. Based on this embodiment of the application, home appliances within the regional space that the user desires to control can be controlled.
[0014] For example, the living room area includes multiple sub-area spaces such as the sofa area, TV cabinet area, and wall area. If the voice command is "turn on the light in the sofa area", the indication information is the sofa area, and the voice receiving device can use the sofa area as the target area space.
[0015] In one implementation, the voice receiving device determines the target area space based on the indication information for indicating the area space in the voice instruction, including: determining that the indication information indicates multiple area spaces. The voice receiving device determines the location of the user. The voice receiving device determines the area space corresponding to the user's location as the target area space, wherein the area space corresponding to the target location is one of the multiple area spaces. Based on the implementation of this application, when the indication information in the voice instruction indicates multiple area spaces, the voice receiving device can also determine the user's location, and determine the area space to which the user's location belongs as the target area space. In this way, although the user's voice instruction is relatively simple, it can still control the home appliances under the area space that the user wants to control.
[0016] Exemplarily, the user's location can be determined by the voice receiving device based on voice instructions, such as performing sound source localization to determine the user's location, or the user's location can be received by the voice receiving device from other devices, such as from a camera or a control device.
[0017] For example, the voice command is "turn on the light next to the wall", then the indication information is the wall, or the direction information is the wall, and the user's entire house includes multiple area spaces, each of which may have a wall, so the indication information can indicate multiple area spaces.
[0018] In one implementation, the voice receiving device determines the target area space according to the target position in the voice instruction, including: the voice receiving device determines the user's movement trajectory according to the voice instruction, and the movement trajectory includes a starting end and an ending end. When the control instruction in the voice instruction is used to turn on the controlled device, the voice receiving device determines the target area space as the area space where the ending end is located. When the control instruction in the voice instruction is used to turn off the controlled device, the voice receiving device determines the target area space as the area space where the starting end is located. Based on the embodiment of the present application, if the user issues a voice instruction while moving, even if there is no indication of the area space in the user's voice instruction, the home appliances in the area space that the user wants to control can be controlled.
[0019] On the second aspect, the present application provides a device control method based on voice interaction, which is applied to a control device, including: a control device receives regional space information and voice instructions sent by a voice receiving device, wherein the voice instruction is a voice instruction received by the voice receiving device, and the regional space information is used to indicate a target regional space, and the target regional space has a corresponding relationship with the controlled device. The control device determines a control instruction based on the voice instruction, and the control instruction is used to control the controlled device in the target regional space. The control device determines the controlled device to be controlled based on the target regional space and the control instruction. The control device sends a control instruction to the controlled device. Based on an embodiment of the present application, the control device can receive the regional space information and voice instructions sent by the voice receiving device, and can determine the controlled device to be controlled based on the regional space information and the voice instruction, and send a control instruction to the controlled device. In this way, when the user interacts with the home appliance by voice, he can control the home appliance in a specific regional space that the user wants to control through simple voice instructions.
[0020] In one implementation, the method further includes: the control device receiving an execution result sent by the controlled device. The control device sending the execution result to the target voice receiving device, or the control device sending instruction information for providing feedback information to the target voice receiving device. Based on this embodiment of the application, the control device may also receive the execution result sent by the controlled device and send the execution result or instruction information to the target voice receiving device, thereby facilitating the target voice receiving device to provide feedback information to the user.
[0021] In one implementation, the control device sends the execution result to the target voice receiving device, including: the control device determines the target voice receiving device and sends the execution result to the target voice receiving device. Based on the embodiment of the present application, the user's home may include multiple voice receiving devices. In this case, the control device may first determine the target voice receiving device and send the execution result to the target voice receiving device, so that the target voice receiving device can output feedback to the user.
[0022] In one implementation, the control device determines the target voice receiving device, including: the control device receives corresponding multiple sound signals sent by multiple voice receiving devices, wherein the positions of the multiple voice receiving devices are different. The control device determines the user's facial orientation based on the intensity distribution of the multiple sound signals at different frequencies. The control device determines the voice receiving devices within the positive and negative first angle range of the direction in which the face is facing as target voice receiving devices. Based on the embodiment of the present application, the control device determines the voice receiving devices within the positive and negative first angle range of the direction in which the user's face is facing as target voice receiving devices, so that the voice receiving devices within the user's line of sight can give the user feedback, which is conducive to the user intuitively obtaining the feedback results.
[0023] For example, a voice receiving device may send a sound signal to the control device, and the control device may determine the user's facial orientation based on multiple sound signals.
[0024] For example, the embodiment of the present application does not limit the specific value of the first angle range. For example, the first angle range is 5 degrees or 8 degrees.
[0025] In one implementation, the method further includes: when there is no voice receiving device within a positive or negative first angle range of the direction in which the user's face is facing, the control device determines the voice receiving device closest to the user as the target voice receiving device. Based on an embodiment of the present application, if there is no voice receiving device within a positive or negative first angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device. In this way, when there is no voice receiving device within the user's line of sight, feedback is given to the user through the voice receiving device closest to the user, which helps the user obtain feedback results.
[0026] In one implementation, the method further includes: when the distance between a voice receiving device within a first positive or negative angle range of the user's facial direction and the user is greater than a first preset distance, the control device determines the voice receiving device closest to the user as the target voice receiving device. Based on this embodiment of the application, if the distance between a voice receiving device within the user's line of sight and the user is greater than the first preset distance, the voice receiving device closest to the user is determined as the target voice receiving device for providing feedback to the user. This technical solution facilitates users to obtain feedback results nearby.
[0027] In one implementation, the control device determines the target voice receiving device by, when the ratio of the distances between the two voice receiving devices closest to the user and the user is less than a first ratio, determining the voice receiving device closest to the user as the target voice receiving device. This technical solution facilitates users obtaining feedback results close to the user.
[0028] In one implementation, the control device sends instruction information for providing feedback information to the target voice receiving device, including: when the user's location and the controlled device are in the same area space, the control device determines the instruction information based on the execution result, and the instruction information is used to instruct the target voice receiving device to provide feedback information in the form of visual effects and / or prompt sound effects. The control device sends instruction information to the target voice receiving device. Alternatively, when the user's location and the controlled device are in different area spaces, the control device determines the instruction information based on the execution result, and the instruction information is used to instruct the target voice receiving device to provide feedback information in the form of indicating the execution result through voice instructions. Based on the embodiment of the present application, the control device can determine different instruction information based on the relationship between the user's location and the location of the controlled device, so that the target voice receiving device can be instructed to output feedback to the user in different ways in different situations, thereby meeting the different needs of the user.
[0029] On the third aspect, the present application provides a device control method based on voice interaction, which is applied to a control device, including: the control device receives a voice instruction sent by a voice receiving device. The control device determines a control instruction based on the voice instruction, wherein the control instruction is used to control the controlled device. The control device determines the target area space based on the target position in the voice instruction or the indication information for indicating the area space in the voice instruction, wherein the target area space is the area space where the controlled device is located, and the target area space has a corresponding relationship with the controlled device. The control device determines the controlled device based on the target area space and the control instruction. The control device sends a control instruction to the controlled device. Based on an embodiment of the present application, the control device can receive the voice instruction sent by the voice receiving device, determine the control instruction and the target area space for controlling the controlled device, and send a control instruction to the controlled device to control the controlled device to execute the user's instruction. In this way, when the user interacts with the home appliance by voice, he can control the home appliance in a specific area space that the user wants to control through simple voice instructions.
[0030] In one implementation, the control device determines the target area based on the target location in the voice command, including: the control device determines the area corresponding to the target location as the target area. This allows users to conveniently control home appliances in certain areas with simple voice commands without having to specify the area where the home appliance they want to control is located, thereby improving the efficiency of voice interaction.
[0031] In one implementation, determining the target regional space based on the indication information for indicating the regional space in the voice instruction includes: the control device determines that the voice instruction includes the indication information, and determines the regional space indicated by the indication information as the target regional space. Based on the embodiment of the present application, when the user's voice instruction includes the indication information for indicating the regional space, the voice receiving device can determine the target regional space as the regional space indicated by the indication information. This helps the voice receiving device determine which devices in the regional space are needed, thereby improving the efficiency of voice interaction.
[0032] In one implementation, the method further includes: the control device receiving an execution result sent by the controlled device; the control device determining a target voice receiving device; and the control device sending the execution result to the target voice receiving device. Based on this embodiment of the present application, the control device may also receive the execution result sent by the controlled device and send the execution result to the target voice receiving device, thereby facilitating the target voice receiving device to provide feedback information to the user.
[0033] In one implementation, the control device determines the target voice receiving device, including: the control device receives corresponding multiple sound signals sent by multiple voice receiving devices, wherein the multiple voice receiving devices are at different positions; and determines the user's facial orientation based on the intensity distribution of the multiple sound signals at different frequencies. The control device determines the voice receiving devices within the positive and negative first angle range of the direction in which the face is facing as target voice receiving devices. Based on the embodiment of the present application, the control device determines the voice receiving devices within the positive and negative first angle range of the direction in which the user's face is facing as target voice receiving devices, so that the voice receiving devices within the user's line of sight can give the user feedback, which is conducive to the user intuitively obtaining the feedback results.
[0034] In one implementation, the method further includes: when there is no voice receiving device within the positive and negative first angle range of the direction in which the face is facing, the control device determines the voice receiving device closest to the user as the target voice receiving device. Based on the embodiment of the present application, if there is no voice receiving device within the positive and negative first angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device. In this way, when there is no voice receiving device within the user's line of sight, feedback is given to the user through the voice receiving device closest to the user, which helps the user obtain feedback results.
[0035] In a fourth aspect, the present application provides a voice receiving device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, wherein when the one or more programs are executed by one or more processors, the device control method based on voice interaction as in the first aspect and any possible implementation thereof is executed.
[0036] In the fifth aspect, the present application provides a controlled device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, wherein when the one or more programs are executed by one or more processors, the device control method based on voice interaction as in the second aspect to the third aspect and any possible implementation thereof is executed.
[0037] In a sixth aspect, the present application provides an apparatus comprising a module for implementing a device control method based on voice interaction as in the first to third aspects and any possible implementation thereof.
[0038] In the seventh aspect, the present application provides a chip including a processor and a communication interface, the communication interface being used to receive signals and transmit the signals to the processor, and the processor processing the signals so that the device control method based on voice interaction as in the first aspect and any possible implementation thereof is executed.
[0039] In the eighth aspect, the present application provides a chip including a processor and a communication interface, the communication interface being used to receive signals and transmit the signals to the processor, and the processor processing the signals so that the device control method based on voice interaction as in the second aspect and any possible implementation thereof is executed.
[0040] In a ninth aspect, the present application provides a readable storage medium, in which instructions are stored. When the instructions are executed on a device, the device control method based on voice interaction as in the first aspect and any possible implementation thereof is executed.
[0041] In the tenth aspect, the present application provides a readable storage medium, which stores instructions. When the instructions are run on the device, the device control method based on voice interaction in the second aspect to the third aspect and any possible implementation thereof is executed.
[0042] In the eleventh aspect, the present application provides a program product, which includes program code. When the program code runs on a device, the device control method based on voice interaction in the first aspect and any possible implementation thereof is executed.
[0043] In the twelfth aspect, the present application provides a program product, which includes program code. When the program code runs on a device, the device control method based on voice interaction in the second to third aspects and any possible implementation thereof is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a schematic diagram of regional spatial information in a whole house provided by an embodiment of the present application.
[0045] Figure 2 This is a schematic diagram of a voice interaction provided in an embodiment of the present application.
[0046] Figure 3 It is a schematic diagram for determining the location of a sound source.
[0047] Figure 4 This is a schematic diagram of determining a user's facial orientation provided by an embodiment of the present application.
[0048] Figure 5 This is a schematic diagram of another voice interaction provided in an embodiment of the present application.
[0049] Figure 6 This is a schematic diagram of another voice interaction provided in an embodiment of the present application.
[0050] Figure 7 This is a schematic diagram of another voice interaction provided in an embodiment of the present application.
[0051] Figure 8 This is a schematic diagram of another voice interaction provided in an embodiment of the present application.
[0052] Figure 9 This is a schematic diagram of another voice interaction provided in an embodiment of the present application.
[0053] Figure 10 This is a schematic interaction diagram of a voice interaction method provided in an embodiment of the present application.
[0054] Figure 11 This is a schematic interaction diagram of a voice interaction method provided in an embodiment of the present application.
[0055] Figure 12 This is a schematic block diagram of a device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The technical solution in this application will be described below with reference to the accompanying drawings.
[0057] The device control method based on voice interaction in the embodiments of this application can be applied to smart home devices. For example, the smart home device can be a whole-house smart host, a whole-house smart central control panel, a smart speaker, a lamp or lamp group, a sensor, a smart curtain, a smart door lock, a robot vacuum, an air conditioner, and the like. The method can also be applied to devices such as smartphones and tablets.
[0058] In a whole-home smart scenario, users can control home appliances by interacting with smart devices through voice. For example, users can control home appliances by interacting with smart speakers through voice.
[0059] However, when users interact with smart devices through voice, they need to accurately say the name of the home device to control it. For controlling home devices in a specific area of the house, the current voice interaction method cannot meet this user's needs.
[0060] In view of this, the present application provides a device control method and device based on voice interaction, which can conveniently realize the control of home appliances in a specific area space.
[0061] The following will be combined Figures 1 to 11 The present invention introduces a device control method based on voice interaction in an embodiment of the present application.
[0062] For example, Figure 1 This is a schematic diagram of regional spatial information in a whole house provided by an embodiment of the present application. Figure 1As shown, the device 1 owned by the user may include regional spatial information of the entire house, and the regional spatial information is used to represent the spatial distribution of the entire house of the user.
[0063] Exemplarily, the device 1 may be a whole-house smart host, a cloud host, or a central control panel, etc. The embodiment of the present application is described by taking the device 1 as a whole-house smart host as an example.
[0064] See also Figure 1 The whole house can be divided into living room area, dining area, kitchen, master bedroom, second bedroom and other areas.
[0065] In some examples, the regional spatial information may be pre-stored in a whole-house smart host or cloud server or device 1.
[0066] It should be understood that the embodiments of the present application do not limit the method for determining or acquiring the spatial information of the region.
[0067] For example, the regional space information can be manually marked by the user, such as the user manually marking the point map or point cloud map of the regional space of the whole house in the display interface in the control application, and storing the point map or point cloud map in the whole house smart host or cloud server or device 1 after marking is completed.
[0068] Alternatively, the regional spatial information may also be configured in the whole-house smart host or cloud server in the form of a configuration file.
[0069] Alternatively, the regional spatial information may also be determined by an intelligent robot or a sweeping robot, and then the intelligent robot or the sweeping robot may send the regional spatial information to the whole-house intelligent host or cloud server or device 1.
[0070] For example, the living room area may also include areas such as a sofa, a TV, and a window area. Accordingly, the living room area is a first-level sub-area space, and the sofa area, TV area, and window area are second-level sub-areas. This can be understood as the living room area including multiple areas such as the sofa area, TV area, and window area.
[0071] The dining area may further include areas such as a dining table and a sideboard. Accordingly, the dining area is a first-level sub-area space, and areas such as a dining table and a sideboard are second-level sub-area spaces.
[0072] The kitchen area may further include areas such as a stove and a wash basin. Accordingly, the kitchen area is a first-level sub-area space, and the stove, wash basin and other areas are second-level sub-area spaces.
[0073] The master bedroom area can further include areas such as the bed, wardrobe, and window area. The bed area can further include bedside A and bedside B areas. Accordingly, the master bedroom area is a first-level sub-area space, the bed area, wardrobe area, and window area are second-level sub-areas, and the area where bedside A and bedside B are located is a third-level sub-area space.
[0074] It should be understood that the aforementioned regional spatial information can be distinguished in the form of spatial coordinates. For example, a spatial coordinate system can be established with the center of the house as the origin, so that the coordinates of the living room, dining room, kitchen, and master bedroom areas are different. The sofa area included in the living room area can also be distinguished by spatial coordinates. The coordinates of each vertex of the sofa can be stored in the living room area, so that the area where the sofa is located can be identified.
[0075] In this way, the entire house is divided into different zones, each of which can correspond to different home appliances. Similarly, the smart host, cloud server, or device 1 can also store information about the zones where each home appliance is located. It should be understood that the above is only an introduction to some of the possible zones in the entire house.
[0076] For example, Table 1 shows the correspondence between home appliances and regional spatial information.
[0077] Table 1
[0078]
[0079]
[0080] Home devices can be grouped together to correspond to spatial information. For example, downlights are typically grouped together, and when controlled by the user, the groups can be turned on or off simultaneously, or partially turned on or off.
[0081] For example, the zones can also have a hierarchical relationship. Referring to Table 1, for lamp group 1, its corresponding zone is the TV area in the living room. For lamp groups 2 and 3, their corresponding zones are the sofa area in the living room. For lamp group 4, its corresponding zone is the dining table area in the dining room area in the living room.
[0082] Home devices can also be individually associated with regional space information. For example, for device 7, its corresponding regional space is the window area of the master bedroom. For device 8, its corresponding regional space is the second bedroom.
[0083] In this way, the user's entire home is divided into multiple zones, and each home appliance corresponds to a zone. Therefore, when the user subsequently controls the home appliance through voice, they can conveniently control the home appliance in a certain zone.
[0084] It should be understood that in other examples, some area spaces may not correspond to devices, some area spaces may correspond to lights or light groups, and some area spaces may correspond to various types of devices such as light groups, smart refrigerators, smart curtains, etc.
[0085] The following will be combined Figure 2-Figure 11 The present invention introduces a device control method based on voice interaction in an embodiment of the present application.
[0086] For example, Figure 2 This is a schematic diagram of a voice interaction provided by an embodiment of the present application. Figure 2 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0087] Figure 2 (a) in FIG. 1 shows a portion of the user's entire house plan, for example, Figure 2 (a) may include the living room area and dining room area in the whole house area spatial information.
[0088] In some examples, the living room area may further include a sofa area and a TV area. The TV area may correspond to lamp group 1, and the sofa area may correspond to lamp group 2 and lamp group 3.
[0089] In some examples, the dining area may include a dining table area and a sideboard area. The dining table area may correspond to light group 4. The sideboard area may correspond to other light groups (not shown).
[0090] Figure 2 (a) may also include multiple voice receiving devices 110.
[0091] For example, the voice receiving device 110 may be a smart speaker, a smart central control screen, a central control panel, a separate microphone module, etc. The embodiment of the present application is described by taking the voice receiving device 110 as a smart speaker as an example.
[0092] In some examples, when a user issues a voice command "Xiao E, light up here" in the sofa area of the living room, the voice receiving device 110 can send the voice command to the whole-house smart host; the whole-house smart host can determine the control instruction "light up" and the regional spatial information "here" in the voice command.
[0093] Since the regional spatial information "here" in the voice command is a vague concept, the whole-house smart host also needs to determine the user's location.
[0094] In one example, the voice receiving device can determine the user's location based on the arrival angle of the received user's voice signal. The whole-house smart host can receive the user's location from the voice receiving device.
[0095] For example, Figure 3 It is a schematic diagram to determine the location of the sound source. Figure 3 As shown in (a) in FIG, the sound source position can be determined by multiple voice receiving devices. Each voice receiving device can include multiple sound signal receiving points. For example, voice receiving device 1 can include receiving point 1 and receiving point 2, and receiving points 1 and 2 can form a microphone array; voice receiving device 2 can include receiving point 3 and receiving point 4, and receiving points 3 and 4 can form a microphone array.
[0096] The voice receiving device 1 can determine the arrival angle θ1 of the sound signal through the sound signals received by receiving point 1 and receiving point 2, and the voice receiving device 2 can determine the arrival angle θ2 of the sound signal through the sound signals received by receiving point 3 and receiving point 4. Since the positions of the voice receiving device 1 and the voice receiving device 2 are known, the position of the sound source can be determined, that is, the position of the user.
[0097] like Figure 3 As shown in (b) of FIG. 5 , the sound source position can also be determined by a voice receiving device 3 , which can include at least three sound signal receiving points. For example, the voice receiving device 3 can include receiving point 5 , receiving point 6 , and receiving point 7 .
[0098] The voice receiving device 3 can determine the arrival angle θ3 of the sound signal through the sound signals received by the receiving points 5 and 6, and can determine the arrival angle θ4 of the sound signal through the sound signals received by the receiving points 6 and 7, thereby determining the position of the sound source.
[0099] Afterwards, the voice receiving device can send the determined location of the sound source to the smart host in the whole house.
[0100] Alternatively, the voice receiving device can also send the arrival angle of the received user's voice signal to the whole-house smart host, and the whole-house smart host will calculate the user's position relative to the voice receiving device. Since the position of the voice receiving device is known, the user's current location can be determined.
[0101] In other examples, the whole-house smart host can also determine the user's current location through data obtained by millimeter wave sensors and voice receiving devices.
[0102] like Figure 3As shown in (c), when there are multiple users, the whole-house smart host can first determine the approximate locations of the multiple users through the millimeter wave sensor, and then determine the arrival angle θ4 of the sound signal through the sound signal received by the voice receiving device 4 through the receiving points 8 and 9, so that the location of the user who issued the voice command can be determined from multiple locations based on the arrival angle θ4.
[0103] After the whole-house smart host determines the user's location, it can determine the area the user is in based on that location. For example, if the whole-house smart host determines that the user is in the sofa area of the living room, the whole-house smart host can send a light-on command to light groups 2 and 3 corresponding to the sofa area; or, if light groups 2 and 3 are already on, the whole-house smart host can control light groups 2 and 3 to increase their brightness.
[0104] See also Figure 2 In (b), the lamp group 2 and lamp group 3 located in the sofa area can be changed from the off state to the on state under the control of the whole-house smart host.
[0105] Based on the embodiment of the present application, the user is located in the sofa area of the living room. When the device receives the voice command "light up here", the light group 2 and light group 3 located in the sofa area light up, thereby realizing the control of the home appliances in the space where the user is located.
[0106] In addition, the whole-house smart host can also determine the voice receiving device that feeds back the results to the user and the way in which the voice receiving device feeds back the results to the user based on the user's orientation.
[0107] For some examples, see Figure 4 , Figure 4 This is a schematic diagram of determining the user's facial orientation provided by an embodiment of the present application. A neural network model can be pre-deployed in the whole-house smart host. The whole-house smart host can use the neural network model to infer and calculate the intensity of the sound signals at various frequencies received by multiple voice receiving devices, thereby determining the user's facial orientation. It should be understood that the multiple voice receiving devices can be located at different locations, such as voice receiving device 1 located at Figure 4 The position of the number 1 is where the voice receiving device 2 is located. Figure 4 The position of the number 2 in the figure, so that the intensity of the sound information received by the voice receiving device 1 and the voice receiving device 2 is different.
[0108] In other examples, the voice receiving devices 1 and 2 can also respectively determine the user's posture and send the determined posture to the whole-house smart host respectively. The whole-house smart host can determine the user's facial orientation based on the multiple postures received.
[0109] Alternatively, the whole-house smart host can also determine the user's facial orientation with the help of images or videos obtained by the camera.
[0110] It should be understood that the whole-house smart host can also be replaced by a cloud server or server or voice receiving device, which is not limited in the embodiments of this application.
[0111] The whole-house smart host can determine the voice receiving device within the positive and negative angle α of the user's facial orientation as the voice receiving device that feeds back the result to the user. The voice receiving device that feeds back the result to the user can be called a voice response device.
[0112] For example, the angle α may be 15 degrees or 10 degrees, etc. The embodiment of the present application does not limit the specific value of the angle α.
[0113] When the whole-house smart host determines how the voice response device feeds back the results to the user, the final feedback result can be determined by whether the user's location and the area space where the home device to be controlled is located are in the same area space.
[0114] For some examples, see Figure 2 In (b), when the user's location and the area space where the home device to be controlled are located are in the same area space, the user can immediately obtain the execution result of the home device on the voice command, so the whole-house smart host can instruct the voice response device to simplify the feedback or not feedback.
[0115] For example, the whole-house smart host can instruct the voice response device to only provide ringtone feedback, such as playing a short prompt tone; or, the whole-house smart host can instruct the voice response device to provide flashing indicator light feedback; or, the whole-house smart host can also instruct the voice response device not to provide feedback to reduce interference to the user.
[0116] In other examples, if the user's location and the area space where the home device to be controlled are located are in different areas, the user cannot immediately obtain the execution results of the home device on the voice command, so the whole-house smart host can instruct the voice response device to provide detailed feedback on the execution results.
[0117] For example, if a user issues a voice command from the sofa area, "Turn off the air conditioner in other rooms," and the whole-home smart host executes the corresponding action based on the voice command, the whole-home smart host can instruct the voice response device to output, "The air conditioner in other rooms has been turned off for you." This allows users to provide detailed feedback on the execution results of home devices when they cannot immediately obtain the execution results of the home devices, helping them understand the execution results of the home devices in response to the user's voice command.
[0118] Figure 5This is a schematic diagram of another voice interaction provided by an embodiment of the present application. Figure 5 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0119] Figure 5 (a) in FIG. 1 shows a portion of the user's entire house plan. It should be understood that Figure 5 The regional spatial information included in (a) can be found in Figure 2 For the sake of brevity, the description of (a) will not be repeated here.
[0120] See also Figure 5 (a) in Figure 2 The difference between the above example and (a) is that the user is located at the dining table in the restaurant area. When the user issues a voice command "Xiao E, light up here" at the dining table, the voice receiving device 110 can send the voice command to the whole-house smart host. The whole-house smart host can then determine the control instruction "light up" and the spatial area information "here" in the voice command.
[0121] It should be understood that the whole-house smart host can determine the user's location by the method described above. When it is determined that the user is located in the dining table area, a light-on command can be sent to the light group 4 corresponding to the dining table area; or, when the light group 4 is already in the light-on state, the whole-house smart host can control the light group 4 to increase the brightness.
[0122] See also Figure 5 In (b), the light group 4 located in the dining table area can be changed from off state to on state under the control of the whole-house smart host.
[0123] Based on the embodiment of the present application, the user is located in the dining table area of the restaurant area. When the user issues a voice command of "light up here", the light group 4 located in the dining table area lights up, thereby realizing the control of the home appliances in the space where the user is located.
[0124] In addition, the whole-house smart host can also determine the user's facial orientation using the method mentioned above.
[0125] Continue to see Figure 5 In (b), there is no voice receiving device within the range of positive and negative angles α of the user's facial orientation. In this case, the whole-house smart host can determine the voice receiving device closest to the user's location as the voice response device to provide feedback to the user.
[0126] In other embodiments, there is a voice receiving device within the range of positive and negative angles α of the user's facial orientation, but when the distance between the user's location and the voice receiving device is greater than a preset distance, the whole-house smart host can determine the voice receiving device closest to the user's location as the voice response device for providing feedback to the user.
[0127] For example, the preset distance is 10 meters or 8 meters, etc. The embodiment of the present application does not limit the specific value of the preset distance.
[0128] In other embodiments, if the ratio of the distance between the user's location and the two nearest voice receiving devices is less than a preset ratio, the whole-house smart host can determine the voice receiving device closest to the user's location as the voice response device for providing feedback to the user.
[0129] For example, the preset ratio may be 1:2 or 1:3, etc. The embodiment of the present application does not limit the specific value of the preset ratio.
[0130] In some examples, since the user's location and the area where the home device to be controlled are located are in the same area, the user can immediately obtain the execution results of the home device on the voice command, so the whole-house smart host can instruct the voice response device to simplify feedback or not provide feedback to reduce interference to the user.
[0131] In other examples, if the user's location is in a different area than the area where the home appliance to be controlled is located, the user cannot immediately obtain the execution result of the home appliance on the voice command, so the whole house smart host can instruct the voice response device to provide feedback on the execution result. Figure 6 Introduce the technical solution.
[0132] For example, Figure 6 This is a schematic diagram of another voice interaction provided by an embodiment of the present application. Figure 6 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0133] It should be understood that Figure 6 The regional spatial information included in (a) can be found in Figure 2 For the sake of brevity, the description of (a) will not be repeated here.
[0134] See also Figure 6 (a) in Figure 2The difference in (a) is that the user is located in the living room's sofa area. When the user issues a voice command from the sofa area, "Xiao E, turn off the master bedroom light," the voice receiving device can send this voice command to the whole-house smart host. The whole-house smart host can then determine the control instruction "turn off" and the spatial area information, "master bedroom," in the voice command. Therefore, the whole-house smart host can control the master bedroom light to turn off.
[0135] The whole-home smart host can determine the user's facial orientation. Figure 6 In (b), if there is a voice receiving device within the range of positive and negative angles α of the user's facial orientation, the whole-house smart host can determine the voice receiving device as a voice response device that provides feedback to the user.
[0136] After determining that the user is in the sofa area, the whole-house smart host can determine that the user's location is in a different area space from the area space where the home appliances to be controlled are located. The whole-house smart host can then instruct the voice response device to provide feedback on the execution results.
[0137] In some examples, after the whole-house smart host controls the light in the master bedroom to turn off, the whole-house smart host can instruct the voice response device to output the voice "OK, the light in the master bedroom has been turned off for you."
[0138] In this way, since the user's location is not in the same area as the area where the home appliance to be controlled is located, the voice response device can output the execution result by voice, which helps the user understand the execution result of the home appliance to the user's voice command.
[0139] In some cases, the user's voice command may not contain an indication of a regional space. In this case, the whole-house smart host cannot determine which regional spaces the user wants to control the home devices in. In this case, the whole-house smart host can determine the regional space where the user is located and use the first-level regional space to which the regional space belongs as the regional space for the home devices to be controlled. Figure 7 Introduce the technical solution.
[0140] For example, Figure 7 This is a schematic diagram of another voice interaction provided by an embodiment of the present application. Figure 7 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0141] It should be understood that Figure 7 The regional spatial information included in (a) can be found in Figure 2 For the sake of brevity, the description of (a) will not be repeated here.
[0142] See also Figure 7 In (a), when the user issues a voice command "Xiao E, turn on the light" in the sofa area, the voice receiving device can send the voice command to the whole-house smart host; the whole-house smart host can determine the control command "turn on the light" in the voice command.
[0143] The user is located in the sofa area of the second-level sub-region. Since the user's voice command does not specify any specific sub-region, the whole-home smart host can determine the sub-region corresponding to the home device to be controlled as the "living room area," the first-level sub-region to which the sofa area belongs. The whole-home smart host can then control all lights in the living room or activate the default lighting mode for the living room. For example, this lighting mode can be normal lighting mode or guest reception mode.
[0144] See also Figure 7 In (b), light group 1, light group 2 and light group 3 located in the living room area can be changed from off state to on state under the control of the whole-house smart host.
[0145] In other examples, some of the light groups in the living room area may be changed from the off state to the on state, for example, light group 2 and light group 3 may be changed from the off state to the on state under the control of the whole-house smart host, or light group 2 may be changed from the off state to the on state under the control of the whole-house smart host.
[0146] Based on the embodiment of the present application, when there is no indication of a regional space in the user's voice command, the whole-house smart host can determine the regional space where the user is located, and use the regional space where the user is located or the regional space where the user is located and the adjacent regional space as the regional space where the home appliances need to be controlled, and control the corresponding home appliances according to the user's voice command. In this way, the user does not need to specify the space where the home appliances need to be controlled, but only needs simple voice commands to conveniently control home appliances in certain regional spaces, thereby improving the efficiency of voice interaction.
[0147] Combined with the above Figure 2-7 This section describes situations where a user's location may remain fixed or change slightly when interacting with home devices via voice. However, in some cases, the user may be moving while issuing a voice command, or issuing a voice command while moving. In this case, the user's location may change. For example, a user may move from the living room to the dining room while issuing a voice command. In this case, the area the user wishes to control may be the living room or the dining room.
[0148] The following will be combined Figure 8-Figure 9 Introduce the technical solution.
[0149] For example, Figure 8This is a schematic diagram of another voice interaction provided by an embodiment of the present application. Figure 8 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0150] It should be understood that Figure 8 The regional spatial information included in (a) can be found in Figure 2 For the sake of brevity, the description of (a) will not be repeated here.
[0151] See also Figure 8 In (a), the user's current location 1 is near the sofa area in the living room, and light groups 2 and 3 in the sofa area are on. The user issues the voice command "Xiao E, turn off the lights" while moving from the sofa area in the living room to the dining area.
[0152] After receiving the voice command, the voice receiving device can send the voice command to the whole-house smart host. The whole-house smart host can determine that the control command in the voice command is "off", and the whole-house smart host also needs to determine that the user issued the voice command while moving.
[0153] In some examples, the home device can determine that the user issued a voice command while moving and send this result to the whole-house smart host. Alternatively, the whole-house smart host can determine that the user issued a voice command while moving based on information obtained by the home device.
[0154] For example, the camera can determine that the user issued a voice command while moving through the captured audio and video images. Alternatively, the camera can send the captured audio and video images to the whole-house smart host, which can determine that the user issued a voice command while moving through the audio and video images.
[0155] For example, multiple millimeter wave sensors can be distributed throughout the user's house. These multiple millimeter wave sensors can determine the user's movement through the acquired data. The voice receiving device can receive the sound signals emitted by the user, thereby determining that the user issues voice commands while moving.
[0156] In addition, the home appliance can also predict the user's movement trajectory when determining the user's movement. It should be understood that the movement trajectory can also be called a walking trajectory, a sports trajectory, etc.
[0157] Exemplarily, the user's movement trajectory can be predicted by the above-mentioned camera or millimeter wave sensor.
[0158] Alternatively, the user's movement trajectory can be predicted using the sound source localization method described above. For example, the voice receiving device can obtain the initial and final segments of the user's voice command and perform sound source localization on each segment. This can determine the user's movement trajectory during the voice command, and further predict the user's subsequent movement trajectory.
[0159] continue Figure 8 In (a), the movement trajectory may be a trajectory from the sofa area in the living room to the dining area.
[0160] At this point, the whole-house smart host can determine that the user is moving from the sofa area in the living room to the dining area, and combined with the control command "turn off the lights" in the user's voice command, it can be determined that the user wants to turn off light group 2 and light group 3 in the sofa area. After that, the whole-house smart host can control the turning off of light group 2 and light group 3.
[0161] See also Figure 8 In (b), the lamp group 2 and lamp group 3 located in the sofa area can be changed from the on state to the off state under the control of the whole house smart host.
[0162] Based on the embodiment of the present application, when the user issues the voice command "turn off the lights" while moving, the whole-house smart host can control to turn off the lights or light groups in the area space that the user has left, so that even if there is no indication of the area space in the user's voice command, the home appliances in the area space that the user wants to control can be controlled.
[0163] In another example, when a user issues the voice command "Xiao E, turn off the lights" while moving from the sofa area in the living room to the dining room, if the whole-house smart host determines that only light groups 2 and 3 are on in the living room and dining room, and all other lights are off, and it is night time or the data obtained by the ambient light sensor is less than a preset value, the whole-house smart host can control the turning off of light groups 2 and 3 while turning on light group 4 in the dining table area in the dining room, where the user is about to move. This prevents users from walking in the dark and improves the intelligence of home appliances.
[0164] For example, Figure 9 This is a schematic diagram of another voice interaction provided by an embodiment of the present application. Figure 9 (a) to (b) show the process of a user controlling home appliances through voice commands.
[0165] It should be understood that Figure 9 The regional spatial information included in (a) can be found in Figure 2 For the sake of brevity, the description of (a) will not be repeated here.
[0166] See also Figure 9 In (a), the user's current position 1 is near the sofa area in the living room area, and the user moves toward the dining area while issuing a voice command "Xiao E, turn on the light".
[0167] After receiving the voice command, the voice receiving device can send the voice command to the whole-house smart host. The whole-house smart host can determine that the control command in the voice command is "turn on the light", and the whole-house smart host also needs to determine that the user issued the voice command while moving.
[0168] In addition, the home device can also predict the user's movement trajectory when determining the user's movement. For example, the whole-house smart host can predict the user's movement trajectory through the above-mentioned camera or millimeter wave sensor.
[0169] It should be understood that the way in which the whole-house smart host determines the user's voice command during movement and the way in which the user's movement trajectory is predicted can be found in Figure 8 Related description in .
[0170] continue Figure 9 In (a), the movement trajectory may be from the sofa area in the living room area to the dining table area in the dining room area.
[0171] At this point, the whole-house smart host can determine that the user is about to move from the sofa area in the living room to the dining table area in the dining room, and combined with the control command "turn on the lights" in the user's voice command, it can be determined that the user wants to turn on the light group 4 in the dining table area that is about to be reached. After that, the whole-house smart host can control the light group 4 to turn on.
[0172] See also Figure 9 In (b), the lamp group 4 located in the sofa area can be changed from the off state to the on state under the control of the whole-house smart host.
[0173] Based on the embodiment of the present application, when the user issues a voice command "turn on the lights" while moving, the whole-house smart host can control the lights or groups of lights in the area space that the user is about to arrive at to turn on, so that even if there is no indication of the area space in the user's voice command, the home appliances in the area space that the user wants to control can be controlled.
[0174] In some embodiments, when the result of the regional space indicated by the indication information contained in the user's voice command is not unique, the whole-house smart host or voice receiving device can determine the user's location and use the regional space of the user's location as the regional space where the device to be controlled is located. Afterwards, the whole-house smart host or voice receiving device can control the devices in the regional space according to the control instructions in the user's voice command.
[0175] The way in which the whole-house smart host or voice receiving device can determine the user's location can be found in the relevant description in the previous article. For the sake of brevity, it will not be repeated here.
[0176] It should be understood that the result of the regional space indicated by the indication information is not unique, which can also be understood as the result indicated by the direction information contained in the user's voice instruction is not unique.
[0177] For example, the user is located in the sofa area of the living room and issues a voice command "turn on the wall light". The direction information in the voice command is "wall", which indicates multiple area spaces (such as the wall of the living room, the wall of the master bedroom, etc.). At this time, the whole-house smart host or voice receiving device can determine that the user is located in the sofa area of the living room, and then determine that the light to be turned on is in the living room. The whole-house smart host can then control the light to turn on the wall of the living room.
[0178] In this way, when the area space indicated by the user's voice command is not unique, the technical solution can use the area space where the user is located as the area space where the device to be controlled is located, thereby controlling the home devices in the area space that the user wants to control.
[0179] Figure 10 This is a schematic flow chart of a device control method based on voice interaction provided by an embodiment of the present application. Figure 10 As shown, the method 200 may include steps 210 to 295 .
[0180] It should be understood that the embodiments of the present application are described using voice receiving devices such as smart speakers and control panels, controlling devices such as whole-house smart hosts and smart gateways, and controlled devices such as lamps or lamp groups as examples.
[0181] In an embodiment of the present application, the voice receiving device, the control device, and the controlled device can be connected to each other via wireless communication, such as via WiFi, Bluetooth, ZigBee, etc., or can be connected to each other via wired communication, such as power line connection; or, the voice receiving device, the control device, and the controlled device can also log in to the same account, or be in the same local area network.
[0182] It should be understood that the voice receiving device and the control device can also be the same device, and the embodiments of the present application are not limited to this.
[0183] 210. The voice receiving device receives a user's voice instruction.
[0184] For example, the voice command may be a command issued by the user to turn on the lights, turn off the lights, turn off the air conditioner, turn on the air conditioner, etc., which is used to control home appliances.
[0185] 220. The voice receiving device determines the target area space according to the target position in the voice instruction or the indication information for indicating the area space in the voice instruction.
[0186] For example, if the voice command is “turn off the lights in the master bedroom”, the area space indicated by the instruction information is “master bedroom area”, and the voice receiving device can use the master bedroom area as the target area space.
[0187] For example, if the voice command is “turn on the light here”, the target location is “here”, and the voice receiving device can use the area space corresponding to the user's location as the target area space.
[0188] 230. The voice receiving device sends the target area space and the voice instruction to the control device. Correspondingly, the control device receives the target area space and the voice instruction.
[0189] In some examples, the voice receiving device may also perform preliminary processing on the voice instructions, such as filtering, noise reduction, etc., and send the preliminarily processed voice instructions to the control device.
[0190] It should be understood that step 220 may also be performed by the control device. In this case, the voice receiving device may first send the voice instruction to the control device, and the voice receiving device does not need to perform step 230.
[0191] 240. The control device determines a control instruction through a voice instruction.
[0192] For example, if the voice command is “turn on the light here”, the control device may determine that the control command is “turn on the light”.
[0193] 250. The control device determines the controlled device according to the target area space and the control instruction.
[0194] Since the control device can pre-store the area spaces corresponding to the home appliances in the whole house, when the control device obtains the target area space, it can determine which home appliances the target area space corresponds to.
[0195] For example, when the voice command is "turn on the light in the master bedroom", the control command is "turn on the light", and the control device can determine the light or light group in the master bedroom area that needs to be controlled.
[0196] 260. The control device sends a control instruction to the controlled device. Correspondingly, the controlled device receives the control instruction.
[0197] For example, when the voice command is "turn on the light here", the control command is "turn on the light", then the control device can send a "turn on the light" command to the corresponding controlled device (such as a light or a light group).
[0198] 270. The controlled device performs corresponding operations according to the control instruction.
[0199] When the control instruction is "turn on the light", the controlled device can change from the light-off state to the light-on state according to the "turn on the light" instruction.
[0200] When the control instruction is "turn off the lights", the controlled device can change from the light-on state to the light-off state according to the "turn off the lights" instruction.
[0201] 280. The controlled device sends the execution result to the control device. Correspondingly, the control device receives the execution result.
[0202] After executing the corresponding operation according to the control instruction, the controlled device can send the execution result to the control device.
[0203] For example, an execution result of 1 indicates successful execution, and an execution result of 0 indicates failed execution.
[0204] For example, after the controlled device changes from the off state to the on state according to the "turn on light" instruction, the execution result 1 can be fed back to the control device, thereby facilitating the control device to determine whether the controlled device executes the control instruction successfully.
[0205] It is understandable that step 280 may not be performed, and this embodiment of the present application is not limited thereto.
[0206] 290. The control device sends the execution result to the voice receiving device. Correspondingly, the voice receiving device receives the execution result.
[0207] The control device can send the execution result of the controlled device to the voice receiving device, thereby facilitating the voice receiving device to provide feedback information to the user.
[0208] 295. The voice receiving device provides feedback information based on the execution result.
[0209] In some examples, the voice receiving device may directly output feedback based on the execution result. For example, when the execution result is successful, the voice receiving device may directly output the successful execution.
[0210] In some examples, when the execution result is successful, the voice receiving device can also determine a method for providing feedback to the user based on the user's location and regional space.
[0211] For example, when the user's location and the area space where the home device to be controlled are located are in the same area space, the user can immediately obtain the execution result of the home device on the voice command, so the voice receiving device can simplify the feedback or not provide feedback.
[0212] In other examples, if the user's location and the area space where the home appliance to be controlled is located are in different areas, the user cannot immediately obtain the execution results of the home appliance on the voice command, so the voice receiving device needs to provide feedback on the execution results in the form of voice.
[0213] Based on the embodiments of the present application, the voice receiving device can determine the target area space where the controlled device is located based on the target position and indication information in the user's voice command, and send area space information and voice commands used to indicate the target area space to the control device; the controlled device can determine the control command based on the voice command, and determine the controlled device based on the control command and the target area space, and send the control command to the controlled device; the voice receiving device receives the result of the controlled device executing the voice command sent by the control device, and provides corresponding feedback information to the user based on the result.
[0214] In this way, when users interact with home appliances by voice, they only need simple voice commands to control the home appliances in a specific area space that the user wants to control.
[0215] It is understood that the embodiment of the present application does not limit the specific execution order of the above steps 210-290. In some examples, some steps in the steps 210-290 may not be executed or may be replaced by other steps.
[0216] In other examples, the voice receiving device may also output feedback according to the instruction information of the control device. In this case, step 290 may be replaced by the following steps:
[0217] The control device sends instruction information to the voice receiving device according to the execution result.
[0218] In some examples, the control device may directly send instruction information to the voice receiving device based on the execution result. For example, when the execution result is successful, the control device may directly instruct the voice receiving device to output the successful execution.
[0219] In some examples, when the execution result is successful, the control device may also determine indication information based on the user's location and regional space.
[0220] For example, when the user's location and the area space where the home device to be controlled are located are in the same area space, the user can immediately obtain the execution result of the home device on the voice command, so the control device can instruct the voice receiving device to simplify the feedback or not to feedback, such as visual methods such as flashing or prompt sound effects.
[0221] In other examples, if the user's location and the area space where the home device to be controlled are located are in different areas, the user cannot immediately obtain the execution results of the home device on the voice command, so the control device can instruct the voice receiving device to provide feedback on the execution results.
[0222] The step 295 can be replaced by the following steps:
[0223] The voice receiving device provides feedback information according to the instruction information.
[0224] In this way, the voice receiving device can provide feedback to the user in different feedback modes according to the instruction information output of the control device, thereby providing feedback to the user in an appropriate manner, thereby improving the intelligence level of home appliances.
[0225] In some embodiments, the voice receiving device determines the target area space according to the target position in the voice instruction, including: determining the user's position according to the voice instruction; and determining the area space corresponding to the user's position as the target area space.
[0226] Based on the embodiment of the present application, if the voice command is "Turn on the light here", the voice receiving device can determine the user's location and determine the area space corresponding to the user's location as the target area space. In this way, the user does not need to accurately say the area space where the home appliance to be controlled is located. Only simple voice commands can be used to conveniently control home appliances in certain areas and spaces, thereby improving the efficiency of voice interaction.
[0227] In some embodiments, the voice receiving device determines the target area space according to the target position in the voice command, including:
[0228] Determining that the voice command does not include instruction information;
[0229] A first region space corresponding to the target position is determined as the target region space, wherein the first region space includes one or more sub-region spaces.
[0230] For example, the voice command is "turn on the light here", which may correspond to the living room area. The living room area may include multiple sub-area spaces such as sofas and TV cabinets. The voice receiving device may use the entire living room area as the target area space.
[0231] In this way, users do not need to accurately say the area space where the home appliances that need to be controlled are located. They only need simple voice commands to conveniently control the home appliances in certain areas, thereby improving the efficiency of voice interaction.
[0232] In some embodiments, the voice receiving device determines the target area space according to the indication information for indicating the area space in the voice instruction, including:
[0233] Determining that the voice command includes instruction information;
[0234] The region space indicated by the instruction information is determined as the target region space.
[0235] For example, if the voice command is "turn on the light in the master bedroom", the instruction information is the master bedroom, and the master bedroom area can be used as the target area space.
[0236] According to the embodiments of the present application, when a user's voice command includes information indicating a regional space, the voice receiving device can determine that the target regional space is the regional space indicated by the information. This helps the voice receiving device determine which devices in the regional space are needed, thereby improving the efficiency of voice interaction.
[0237] In some embodiments, the voice receiving device determines the target area space based on the indication information used to indicate the area space in the voice instruction, including: determining the target sub-area space among multiple sub-area spaces in the area space based on the indication information; and determining the target sub-area space as the target area space.
[0238] For example, the living room area includes multiple sub-area spaces such as the sofa area, TV cabinet area, and wall area. The voice command is "turn on the light in the sofa area", then the indication information is the sofa area, and the target sub-area space is the sofa area. The voice receiving device can use the sofa area as the target area space.
[0239] Based on the embodiments of the present application, home appliances in the area space that the user wants to control can be controlled.
[0240] In some embodiments, the voice receiving device determines the target area space according to the indication information for indicating the area space in the voice instruction, including:
[0241] Determining that the indication information indicates a plurality of regional spaces;
[0242] Determine the user's location;
[0243] The regional space corresponding to the location of the user is determined as the target regional space, wherein the regional space corresponding to the target location is one of the multiple regional spaces.
[0244] Exemplarily, the user's location can be determined by the voice receiving device based on voice instructions, such as performing sound source localization to determine the user's location, or the user's location can be received by the voice receiving device from other devices, such as from a camera or a control device.
[0245] For example, the voice command is "turn on the light next to the wall", then the indication information is the wall, or the direction information is the wall, and the user's entire house includes multiple area spaces, each of which may have a wall, so the indication information can indicate multiple area spaces.
[0246] Based on the implementation of this application, when the instruction information in the voice command indicates multiple regional spaces, the voice receiving device can also determine the user's location and determine the regional space to which the user's location belongs as the target regional space. In this way, even if the user's voice command is relatively simple, it can still control the home devices in the regional space that the user wants to control.
[0247] In some embodiments, the voice receiving device determines the target area space according to the target position in the voice command, including:
[0248] Determine the user's movement trajectory based on the voice command, the movement trajectory including the starting end and the ending end;
[0249] When the control command in the voice command is used to turn on the controlled device, the target area space is determined to be the area space where the end end is located;
[0250] When the control command in the voice command is used to shut down the controlled device, the target area space is determined to be the area space where the starting end is located.
[0251] For example, see Figure 8-9 Related description in .
[0252] Based on the embodiment of the present application, if the user issues a voice command while moving, even if there is no indication of a regional space in the user's voice command, the home appliances in the regional space that the user wants to control can be controlled.
[0253] In some embodiments, the control device determines the target voice receiving device, including:
[0254] receiving corresponding multiple sound signals sent by multiple voice receiving devices, wherein the multiple voice receiving devices are located at different locations;
[0255] Determining the user's facial orientation based on the intensity distribution of multiple sound signals at different frequencies;
[0256] A voice receiving device within a positive and negative first angle range of the direction in which the face is facing is determined as a target voice receiving device.
[0257] For example, a voice receiving device may send a sound signal to the control device, and the control device may determine the user's facial orientation based on multiple sound signals.
[0258] For example, the embodiment of the present application does not limit the specific value of the first angle range. For example, the first angle range is 5 degrees or 8 degrees.
[0259] Based on the embodiment of the present application, the control device determines the voice receiving device within the positive and negative first angle range of the direction in which the user's face is facing as the target voice receiving device, so that the voice receiving device within the user's line of sight can give the user feedback, which is conducive to the user intuitively obtaining the feedback result.
[0260] In some embodiments, the method 200 may further include:
[0261] When there is no voice receiving device within the positive and negative first angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device.
[0262] According to an embodiment of the present application, if there is no voice receiving device within the first positive and negative angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device. In this way, when there is no voice receiving device within the user's line of sight, feedback is provided to the user through the voice receiving device closest to the user, which helps the user obtain feedback results.
[0263] In some embodiments, the method 200 may further include: when the distance between the voice receiving device within the positive and negative first angle range of the user's facial direction and the user is greater than a first preset distance, the voice receiving device closest to the user is determined as the target voice receiving device.
[0264] According to the embodiment of the present application, if the distance between the voice receiving device within the user's line of sight and the user is greater than a first preset distance, the voice receiving device closest to the user is determined as the target voice receiving device for providing feedback to the user. This technical solution helps users obtain feedback results nearby.
[0265] In some embodiments, the control device determines the target voice receiving device, including:
[0266] When the ratio of the distances between the two voice receiving devices closest to the user and the user is less than a first ratio, the voice receiving device closest to the user is determined as the target voice receiving device.
[0267] This technical solution helps users obtain feedback results nearby.
[0268] Figure 11 This is a device control method based on voice interaction provided in an embodiment of the present application. The method 300 is applied to control a device and includes steps 310-350.
[0269] 310, receiving a voice instruction sent by a voice receiving device.
[0270] For this step, please refer to the relevant description of step 210 above.
[0271] 320. Determine a control instruction according to the voice instruction, where the control instruction is used to control the controlled device.
[0272] It should be understood that the step 320 can refer to the relevant description of the step 240 in the above text.
[0273] 330. Determine a target area space according to a target position in the voice command or indication information for indicating an area space in the voice command, wherein the target area space is an area space where the controlled device is located, and the target area space has a corresponding relationship with the controlled device.
[0274] It should be understood that the technical solution for the control device to determine the target area space based on the target position in the voice instruction or the indication information for indicating the area space in the voice instruction can be referred to the description of the various technical solutions for the voice receiving device to determine the target area space in method 200. For the sake of brevity, it will not be repeated here.
[0275] 340, determining the controlled device according to the target area space and the control instruction.
[0276] It should be understood that the step 340 can refer to the relevant description of the step 250 in the above text.
[0277] 350, sending a control instruction to the controlled device.
[0278] Based on the embodiments of the present application, based on the embodiments of the present application, the control device can receive voice instructions sent by the voice receiving device, determine the control instructions and target area space for controlling the controlled device, determine the controlled device according to the target area space and the control instructions, and send control instructions to the controlled device to control the controlled device to execute the user's instructions.
[0279] In this way, when users interact with home appliances by voice, they can control the home appliances in a specific area that they want to control through simple voice commands.
[0280] In some embodiments, the control device determines the target area space according to the target position in the voice command, including:
[0281] The region space corresponding to the target position is determined as the target region space.
[0282] In this way, users do not need to accurately say the area space where the home appliances that need to be controlled are located. They only need simple voice commands to conveniently control the home appliances in certain areas, thereby improving the efficiency of voice interaction.
[0283] In some embodiments, the control device determines the target area space according to the indication information for indicating the area space in the voice instruction, including:
[0284] Determining that the voice command includes instruction information;
[0285] The region space indicated by the instruction information is determined as the target region space.
[0286] According to the embodiments of the present application, when a user's voice command includes information indicating a regional space, the voice receiving device can determine that the target regional space is the regional space indicated by the information. This helps the voice receiving device determine which devices in the regional space are needed, thereby improving the efficiency of voice interaction.
[0287] In some embodiments, the method 300 further includes:
[0288] Receive the execution result sent by the controlled device;
[0289] Determine the target voice receiving device;
[0290] Send the execution result to the target voice receiving device.
[0291] Based on the embodiment of the present application, the control device can also receive the execution result sent by the controlled device and send the execution result to the target voice receiving device, thereby facilitating the target voice receiving device to provide feedback information to the user.
[0292] In some embodiments, the control device determines the target voice receiving device, including:
[0293] receiving corresponding multiple sound signals sent by multiple voice receiving devices, wherein the multiple voice receiving devices are located at different locations;
[0294] Determining the user's facial orientation based on the intensity distribution of multiple sound signals at different frequencies;
[0295] A voice receiving device within a positive and negative first angle range of the direction in which the face is facing is determined as a target voice receiving device.
[0296] Based on the embodiment of the present application, the control device determines the voice receiving device within the positive and negative first angle range of the direction in which the user's face is facing as the target voice receiving device, so that the voice receiving device within the user's line of sight can give the user feedback, which is conducive to the user intuitively obtaining the feedback result.
[0297] In some embodiments, the method 300 further includes:
[0298] When there is no voice receiving device within the positive and negative first angle range of the direction in which the face is facing, the voice receiving device closest to the user is determined as the target voice receiving device.
[0299] According to an embodiment of the present application, if there is no voice receiving device within the first positive and negative angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device. In this way, when there is no voice receiving device within the user's line of sight, feedback is provided to the user through the voice receiving device closest to the user, which helps the user obtain feedback results.
[0300] The present application also provides a device control method based on voice interaction, which can be applied to a voice receiving device. The method includes:
[0301] The voice receiving device receives the user's voice command;
[0302] The voice receiving device controls the command according to the voice command, and the control command is used to control the controlled device;
[0303] The voice receiving device determines a target area space according to a target position in the voice command or indication information for indicating an area space in the voice command, wherein the target area space is an area space where the controlled device is located, and the target area space has a corresponding relationship with the controlled device;
[0304] The voice receiving device determines the controlled device based on the target area space and control instructions;
[0305] The voice receiving device sends control instructions to the controlled device;
[0306] The voice receiving device receives the result of the controlled device executing the control instruction;
[0307] The voice receiving device outputs feedback based on the execution result.
[0308] Based on the embodiments of the present application, the voice receiving device can receive the user's voice instructions, process the voice instructions, control the controlled device to perform corresponding operations, and output feedback to the user based on the results of the controlled device executing the control instructions.
[0309] In this way, when users interact with home devices by voice, they can control the home devices in a specific area space that the user wants to control through simple voice commands.
[0310] In other examples, the voice receiving device can also convert the user's voice commands into text information and send the text information to other devices (such as smart hosts, servers, etc.), which analyze the text information and control the controlled devices to perform corresponding operations. This embodiment of the present application is not limited to this.
[0311] Figure 12 This is a schematic block diagram of a device provided in an embodiment of the present application. Figure 11 As shown, the device 900 may include one or more processors 910; one or more memories 920; the one or more memories 920 store one or more instructions, and when the instructions are executed by the one or more processors 910, the device control method based on voice interaction as described in any possible implementation method described above is executed.
[0312] Exemplarily, the device 900 can be the whole-house smart host, voice receiving device, server or cloud server mentioned above.
[0313] In some examples, device 900 may also be a mobile phone or tablet. It should be understood that the mobile phone or tablet may also be used to receive user voice commands and execute the device control method based on voice interaction described in any of the possible implementations above. Alternatively, the mobile phone or tablet may also be used to receive voice commands sent by a voice receiving device and process the voice commands to control corresponding home appliances. Alternatively, the mobile phone or tablet may also provide feedback information to the user.
[0314] An embodiment of the present application also provides a device including a processor and a communication interface, the communication interface being used to receive signals and transmit the signals to the processor, and the processor processing the signals so that the device control method based on voice interaction as described in any possible implementation method described above is executed.
[0315] The device may be a chip. For example, the chip may be a chip system or an independent chip.
[0316] An embodiment of the present application also provides a readable storage medium, which stores instructions. When the instructions are executed on a device, the device executes the above-mentioned related method steps to implement the device control method based on voice interaction in the above-mentioned embodiment.
[0317] An embodiment of the present application also provides a program product, which, when executed on a device, enables the device to execute the aforementioned related steps to implement the device control method based on voice interaction in the aforementioned embodiment.
[0318] An embodiment of the present application also provides a voice interaction device, including a module for implementing the device control method based on voice interaction as described in any of the embodiments above.
[0319] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store instructions, and when the device is running, the processor can execute the instructions stored in the memory to enable the device to execute the device control method based on voice interaction in the above-mentioned method embodiments.
[0320] Among them, the equipment, readable storage medium, program product or device provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.
[0321] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0322] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0323] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0324] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0325] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0326] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0327] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A device control method based on voice interaction, characterized in that: The method is applied to a voice receiving device, and the method includes: Receive user's voice commands; determining a target area space according to a target position in the voice command or indication information for indicating an area space in the voice command, wherein the target area space is an area space where the controlled device is located, and the target area space has a corresponding relationship with the controlled device; sending the regional space information indicating the target regional space and the voice instruction to a control device; receiving an execution result sent by the control device, where the execution result is a result of the controlled device executing the voice command; Feedback information is provided according to the execution result.
2. The method according to claim 1, characterized in that Providing feedback information according to the execution result includes: When the user is located in the same area as the controlled device, feedback information is provided in a visual manner and / or in a prompt sound effect manner according to the execution result; When the user is located in a different area from the controlled device, feedback information is provided according to the execution result by indicating the execution result through voice instructions.
3. The method according to claim 1 or 2, characterized in that The determining of the target area space according to the target position in the voice instruction includes: determining that the voice instruction does not include the instruction information; A first region space corresponding to the target position is determined as the target region space, wherein the first region space includes one or more sub-region spaces.
4. The method according to claim 1 or 2, characterized in that The determining the target area space according to the indication information for indicating the area space in the voice instruction includes: Determining that the voice instruction includes the instruction information; The area space indicated by the indication information is determined as the target area space.
5. The method according to claim 4, characterized in that The determining of the target area space according to the indication information for indicating the area space in the voice instruction includes: determining, according to the indication information, a target sub-region space among the plurality of sub-region spaces in the region space; The target sub-region space is determined as the target region space.
6. The method according to claim 1 or 2, characterized in that The determining the target area space according to the indication information for indicating the area space in the voice instruction includes: Determining that the indication information indicates multiple area spaces; Determine the user's location; The area space corresponding to the location of the user is determined as the target area space, wherein the area space corresponding to the target location is one of the multiple area spaces.
7. The method according to claim 1 or 2, characterized in that The determining of the target area space according to the target position in the voice instruction includes: Determine the user's movement trajectory according to the voice command, wherein the movement trajectory includes a starting end and an ending end; When the control instruction in the voice instruction is used to turn on the controlled device, determining that the target area space is the area space where the end terminal is located; When the control instruction in the voice instruction is used to shut down the controlled device, the target area space is determined to be the area space where the starting end is located.
8. A device control method based on voice interaction, characterized in that: The method is applied to a control device, and the method includes: receiving regional space information and a voice command sent by a voice receiving device, wherein the voice command is a voice command received by the voice receiving device, the regional space information is used to indicate a target regional space, and the target regional space has a corresponding relationship with the controlled device; Determining a control instruction according to the voice instruction, wherein the control instruction is used to control a controlled device in the target area space; Determining a controlled device to be controlled according to the target area space and the control instruction; Sending the control instruction to the controlled device.
9. The method according to claim 8, characterized in that The method further comprises: receiving an execution result sent by the controlled device; The execution result is sent to the target voice receiving device, or instruction information for providing feedback information is sent to the target voice receiving device.
10. The method according to claim 9, characterized in that The sending the execution result to the target voice receiving device includes: determining the target voice receiving device; The execution result is sent to the target voice receiving device.
11. The method according to claim 10, characterized in that The determining the target voice receiving device includes: receiving corresponding multiple sound signals sent by multiple voice receiving devices, wherein the multiple voice receiving devices are located at different positions; determining a facial orientation of the user based on intensity distribution of the multiple sound signals at different frequencies; A voice receiving device within a positive and negative first angle range of the direction in which the face is facing is determined as the target voice receiving device.
12. The method according to claim 11, characterized in that The method further comprises: When the voice receiving device does not exist within the positive and negative first angle range of the direction in which the user's face is facing, the voice receiving device closest to the user is determined as the target voice receiving device.
13. The method according to claim 11, characterized in that The method further comprises: When the distance between the voice receiving device and the user within the positive and negative first angle range of the user's facial direction is greater than a first preset distance, the voice receiving device closest to the user is determined as the target voice receiving device.
14. The method according to claim 10, characterized in that The determining the target voice receiving device includes: When the ratio of the distances between the two voice receiving devices closest to the user and the user is less than a first ratio, the voice receiving device closest to the user is determined as the target voice receiving device.
15. The method according to any one of claims 9 to 14, characterized in that The sending instruction information for providing feedback information to the target voice receiving device includes: When the user is located in the same area as the controlled device, determining the instruction information according to the execution result, the instruction information is used to instruct the target voice receiving device to provide feedback information in the form of visual effects and / or prompt sound effects; Sending the instruction information to the target voice receiving device; or, When the user is located in a different area space from the controlled device, the indication information is determined according to the execution result, and the indication information is used to instruct the target voice receiving device to provide feedback information by indicating the execution result through voice instructions.
16. A device control method based on voice interaction, characterized in that: The method is applied to a control device, and the method includes: Receive voice commands sent by a voice receiving device; determining a control instruction according to the voice instruction, wherein the control instruction is used to control a controlled device; determining a target area space according to a target position in the voice command or indication information for indicating an area space in the voice command, wherein the target area space is an area space where the controlled device is located, and the target area space has a corresponding relationship with the controlled device; Determining the controlled device according to the target area space and the control instruction; Sending the control instruction to the controlled device.
17. The method according to claim 16, characterized in that The determining the target area space according to the indication information for indicating the area space in the voice instruction includes: Determining that the voice instruction includes the instruction information; The area space indicated by the indication information is determined as the target area space.
18. The method according to claim 16 or 17, characterized in that The method further comprises: receiving an execution result sent by the controlled device; Determine the target voice receiving device; The execution result is sent to the target voice receiving device.
19. The method according to claim 18, characterized in that The step of determining a target voice receiving device includes: receiving corresponding multiple sound signals sent by multiple voice receiving devices, wherein the multiple voice receiving devices are located at different positions; determining a facial orientation of the user based on intensity distribution of the multiple sound signals at different frequencies; A voice receiving device within a positive and negative first angle range of the direction in which the face is facing is determined as the target voice receiving device.
20. The method according to claim 19, characterized in that The method further comprises: When the voice receiving device does not exist within the positive and negative first angle range of the direction in which the face is facing, the voice receiving device closest to the user is determined as the target voice receiving device.
21. A voice receiving device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by one or more processors, the device control method based on voice interaction as described in any one of claims 1 to 7 is executed.
22. A control device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by one or more processors, the device control method based on voice interaction as described in any one of claims 8 to 20 is executed.
23. A readable storage medium, characterized in that The readable storage medium stores instructions, and when the instructions are executed on the device, the device control method based on voice interaction according to any one of claims 1 to 20 is executed.
24. A program product, characterized in that The program product includes program code, and when the program code is run on a device, the device control method based on voice interaction according to any one of claims 1 to 20 is executed.
Citation Information
Cited By
Method for performing device control on basis of speech interaction, and device
EP4726484A1
Method for performing device control on basis of speech interaction, and device
WO2025201141A1