Voice processing method and electronic device
By combining voice commands and scene information, the positional relationship of target devices is determined, solving the problem that users have difficulty accurately controlling multiple devices and achieving more efficient and accurate voice control.
Patent Information
- Application Number
- PCT/CN2025/077944
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-27
- Filing Date
- 2025-02-19
- Publication Date
- 2025-12-04
AI Technical Summary
Users find it difficult to accurately control multiple electronic devices via voice commands. Existing technologies have obvious limitations, making it difficult to improve the efficiency and accuracy of identifying target devices.
By acquiring the user's voice commands and scene information, and combining the positional relationship between the controllable device and the reference object, the target device is determined, and the voice commands are processed using a language reconstruction model to improve processing efficiency and accuracy.
It improves the efficiency of electronic devices in processing user voice commands, accurately identifies target devices, and enhances user experience and the reliability of device control.
Smart Images

Figure CN2025077944_04122025_PF_FP_ABST
Abstract
Description
Methods and electronic devices for speech processing
[0001] This application claims priority to Chinese Patent Application No. 202410668858.9, filed on May 27, 2024, entitled "Method and Electronic Device for Speech Processing", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal device software, and more specifically, to a voice processing method and an electronic device. Background Technology
[0003] The types of electronic devices are increasing year by year, and users need a certain learning curve to become familiar with and use them. However, with the help of natural language processing technology, users can control electronic devices by interacting with voice assistants. Through voice control, users can easily turn on the TV, adjust the room temperature, play music, etc., without needing to operate a remote control or mobile phone, which greatly improves the user experience. Generally speaking, users can control different electronic devices through specified keywords or voice commands. However, as the number and types of controllable devices in the usage scenario increase, it becomes difficult for users to accurately describe the control instructions for all controllable devices. Therefore, the limitations of this method are becoming increasingly apparent.
[0004] Enhancing the ability of electronic devices to process natural language commands and improving the efficiency of identifying controllable devices that can fulfill user intentions will help solve the above problems. Summary of the Invention
[0005] This application provides a voice processing method in which an electronic device can determine the target device for realizing the user's voice command by combining the user's voice command and the positional relationship between the controllable device and the reference object indicated by scene information. The electronic device has high processing efficiency for the user's voice command, more accurate determination of the target device, and is more conducive to realizing the user's intention.
[0006] In a first aspect, a voice processing method is provided, the method comprising: acquiring a user's voice command, the voice command including a reference object; determining a target device based on scene information and the voice command, the scene information being used to indicate the positional relationship between a controllable device and a reference object in the target scene, the controllable device including the target device, and the target device being used to execute control commands corresponding to the voice command.
[0007] In some scenarios, this technical solution can be applied to terminal devices such as voice assistant devices or central control devices. The target device can be understood as a controllable device indicated by the user's voice command, which the target device can execute after receiving the control command.
[0008] In some scenarios, this technical solution can also be applied to devices such as cloud servers. For example, after receiving a user's voice command, the terminal device can send the voice command to the cloud server, which will then process the voice command, determine the target device, and return the information of the determined target device to the terminal device. Thus, the terminal device can perform operations such as sending control commands to the target device.
[0009] In one possible implementation, the control instructions can be determined based on the user's voice commands. The control instructions may include a task field, an operation field, and a device description field. The device description field may include the device identifier of the target device.
[0010] In this technical solution, the electronic device can determine the target device based on the reference objects contained in the user's voice command and the positional relationship between the controllable device and the reference objects indicated by the scene information of the target scene. Compared with the method of controlling the device by specifying keywords, this technical solution can be widely applied to the user's commonly used voice commands. The electronic device has higher efficiency in processing the user's voice commands, higher efficiency in determining the target device, and a better user experience.
[0011] In conjunction with the first aspect, in some implementations of the first aspect, the scene information includes the position of the reference object and the position of the controllable device, or the scene information includes the positional relationship between the reference object and the controllable device.
[0012] Scene information can directly include the positions of the reference object and the controllable device, or the relative positions of the reference object and the controllable device. Electronic devices can directly determine the target device based on this information, which simplifies the processing flow of scene information and improves the efficiency of electronic devices in determining the target device.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the reference object includes a fixed reference object and / or a movable reference object.
[0014] In the process of determining the target device, the user can select either a fixed reference object or a movable reference object in the target scene. This allows the user to use a wider variety and number of voice commands, which is beneficial to improving the applicability of the voice processing method provided in this application to different voice commands.
[0015] In one possible implementation, the fixed reference object includes a wall in the target scene.
[0016] Walls are commonly used as reference objects in home settings. Using walls as reference objects is helpful in determining the distance between controllable devices and reference objects, in determining the positional relationship between controllable devices and reference objects, and in improving the efficiency of electronic devices in identifying target devices.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the movable reference includes movable objects and / or users in the target scene.
[0018] In this technical solution, the electronic device can use the user as a reference to determine the target device. Voice commands such as "turn off the lights here" can also be recognized by the electronic device and the target device can be accurately determined, which helps to improve the applicability of the voice processing method provided in this application.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, determining the target device based on scene information and voice commands includes: determining the target device based on scene information, voice commands, and the type and capability information of the controllable device, wherein the capability information is used to indicate the performance of the controllable device in executing control commands.
[0020] In one possible implementation, the type and capabilities of the controllable device can be included in the scene information. Alternatively, the type and capabilities of the controllable device can be obtained by the electronic device from the controllable device via network communication.
[0021] In this technical solution, the electronic device can combine more information to determine the target device, which is conducive to more accurate determination of the target device and more reliable realization of the user's intention.
[0022] In conjunction with the first aspect, in some implementations of the first aspect, the capability information includes the range of action of the controllable device and / or the direction of action of the controllable device.
[0023] In some examples, the controllable devices here may include: spotlights, fans, air conditioners, cameras, etc.
[0024] Due to limitations in the direction and / or range of action, different controllable devices may not be able to operate within the area indicated by the user's voice command. When selecting controllable devices, considering factors such as the direction and / or range of action can improve the efficiency of device selection and increase the likelihood that the controllable device will fulfill the user's intent.
[0025] In conjunction with the first aspect, in some implementations of the first aspect, the positional relationship includes any of the following: the controllable device is located inside the reference object, the controllable device is located on one side of the reference object, or the controllable device is located near the reference object.
[0026] The controllable device being located on one side of the reference object includes the controllable device being located above or below the reference object. The controllable device being located on one side or near the reference object can be understood as the controllable device being located outside the reference object.
[0027] In one possible implementation, the positional relationship between the controllable device and the reference object can be roughly determined by the distance between the center of the controllable device and the center of the reference object.
[0028] This technical solution can be applied to different positional relationships between controllable equipment and reference objects, which helps to improve the applicability of this solution in different scenarios.
[0029] In conjunction with the first aspect, in certain implementations of the first aspect, scene information is determined based on one or more of the following: point cloud information of objects in the target scene, bounding box information of objects in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, wherein the point cloud information of the objects is used to indicate the position and / or orientation of the objects in the target scene.
[0030] By determining scene information through multiple methods, and by allowing these methods to complement each other, the implementation of this technical solution is beneficial to improving the feasibility of the speech processing method provided in this application in practical applications.
[0031] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: sending control commands to the target device.
[0032] In conjunction with the first aspect, in some implementations of the first aspect, the voice command may also include one or more of the following: the user's operational intent, information about the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
[0033] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: processing speech commands through a language reconstruction model, wherein the output of the language reconstruction model includes one or more of the following: the user's operational intent, a reference object, information about the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
[0034] The use of language reconstruction models can improve the processing efficiency of voice assistant devices or central control devices for user-inputted voice commands, and can also improve the efficiency of controllable devices.
[0035] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: in response to adding a new reference object in the target scene, obtaining updated scene information, the updated scene information including the position information of the new reference object; or, in response to removing an existing reference object in the target scene, obtaining updated scene information, the updated scene information not including information of the removed existing reference object; or, in response to moving an existing reference object in the target scene, obtaining updated scene information, the updated scene information including the position information of the moved existing reference object; determining the target device based on the scene information and the voice command, including: determining the target device based on the updated scene information and the voice command.
[0036] When the reference objects in the target scene change, the electronic device can obtain the updated scene information in a timely manner. The electronic device is less likely to make mistakes in determining the target device, the efficiency of determining the target device is higher, and the user experience is better.
[0037] The relevant explanations and descriptions of the beneficial effects in the following technical solutions can be found in the relevant descriptions in the first aspect.
[0038] Secondly, a method for controlling a device based on voice commands is provided, comprising: acquiring voice commands; processing the voice commands through a language reconstruction model, and determining one or more of the following: a reference object, the user's operational intent, the device type of the target device, or the positional relationship between the target device and the reference object; determining the controllable device as the target device when the number of controllable devices corresponding to the device type is unique; or, when the number of controllable devices corresponding to the device type is multiple, determining the controllable device that satisfies the positional relationship among the multiple controllable devices as the target device based on scene information, wherein the scene information is used to indicate the positional relationship between the multiple controllable devices and the reference object; and sending a control command corresponding to the operational intent to the target device.
[0039] The operations described above can be performed by one or more devices.
[0040] In one possible implementation, the voice assistant device can receive and process user voice commands. The central control device can receive the processing results of the voice assistant device, determine the target device based on the structure, and send control commands to the target device.
[0041] In one possible implementation, the voice assistant device can perform all the operations in this scheme, or the voice assistant device can be integrated with the central control device into one device.
[0042] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: when there are multiple controllable devices that satisfy the positional relationship, acquiring capability information of the multiple controllable devices, the capability information being used to indicate the performance of the controllable devices in executing control commands;
[0043] Based on the equipment capability information, the controllable equipment that meets the performance requirements among multiple controllable devices is identified as the target equipment.
[0044] In conjunction with the second aspect, in some implementations of the second aspect, the scene information is determined based on one or more of the following: point cloud information of objects in the target scene, bounding box information of objects in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, wherein the objects include movable objects and fixed objects, and movable objects include users.
[0045] Thirdly, a voice processing apparatus is provided, comprising an acquisition module and a processing module. The acquisition module is used to acquire a user's voice command, the voice command including a reference object. The processing module is used to determine a target device based on scene information and the voice command. The scene information is used to indicate the positional relationship between a controllable device and a reference object in the target scene. The controllable device includes the target device, and the target device is used to execute control commands corresponding to the voice command.
[0046] In conjunction with the third aspect, in some implementations of the third aspect, the scene information includes the position of the reference object and the position of the controllable device, or the scene information includes the positional relationship between the reference object and the controllable device.
[0047] In conjunction with the third aspect, in some implementations of the third aspect, the reference object includes a fixed reference object and / or a movable reference object.
[0048] In conjunction with the third aspect, in some implementations of the third aspect, the movable reference includes movable objects and / or users in the target scene.
[0049] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is specifically used to: determine the target device based on scene information, voice commands, and the type and capability information of the controllable device, wherein the capability information is used to indicate the performance of the controllable device in executing control commands.
[0050] In conjunction with the third aspect, in some implementations of the third aspect, the capability information includes the range of action of the controllable device and / or the direction of action of the controllable device.
[0051] In conjunction with the third aspect, in some implementations of the third aspect, the positional relationship includes any of the following: the controllable device is located inside the reference object, the controllable device is located on one side of the reference object, or the controllable device is located near the reference object.
[0052] In conjunction with the third aspect, in some implementations of the third aspect, the scene information is determined based on one or more of the following: point cloud information of objects in the target scene, bounding box information of objects in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, wherein the point cloud information of the objects is used to indicate the position and / or orientation of the objects in the target scene.
[0053] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is also used to: send control commands to the target device.
[0054] In conjunction with the third aspect, in some implementations of the third aspect, the voice command may also include one or more of the following: the user's operational intent, information about the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
[0055] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is also used to: process speech commands through a language reconstruction model, the output of which includes one or more of the following: the user's operational intent, reference objects, information about the target scene, the positional relationship between the controllable device and the reference objects, or the type of the controllable device.
[0056] In conjunction with the third aspect, in some implementations of the third aspect, the acquisition module is further configured to: acquire updated scene information in response to adding a new reference object in the target scene, the updated scene information including the position information of the new reference object; or, acquire updated scene information in response to removing an existing reference object in the target scene, the updated scene information not including information of the removed existing reference object; or, acquire updated scene information in response to moving an existing reference object in the target scene, the updated scene information including the position information of the moved existing reference object; the processing module is specifically configured to: determine the target device based on the updated scene information and the voice command.
[0057] Fourthly, an electronic device is provided, which further includes a processor and a memory. The memory is used to store program instructions, and the processor is used to: acquire a user's voice command, the voice command including a reference object; and determine a target device based on scene information and the voice command, the scene information indicating the positional relationship between a controllable device and a reference object in the target scene, the controllable device including the target device, and the target device executing control commands corresponding to the voice command.
[0058] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the scene information includes the position of the reference object and the position of the controllable device, or the scene information includes the positional relationship between the reference object and the controllable device.
[0059] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the reference object includes a fixed reference object and / or a movable reference object.
[0060] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the movable reference includes movable objects and / or users in the target scene.
[0061] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processor is specifically used to: determine the target device based on voice commands, positional relationships, and the type and / or capability information of the controllable device, wherein the capability information is used to indicate the performance of the controllable device in executing control commands.
[0062] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the capability information includes the range of action of the controllable device and / or the direction of action of the controllable device.
[0063] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the positional relationship includes any of the following: the controllable device is located inside the reference object, the controllable device is located above or below the reference object, or the controllable device is located near the reference object.
[0064] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the scene information is determined based on one or more of the following: point cloud information of objects in the target scene, bounding box information of objects in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, wherein the point cloud information of the objects is used to indicate the position and / or orientation of the objects in the target scene.
[0065] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processor is also used to: send control instructions to the target device.
[0066] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the voice command may also include one or more of the following: the user's operational intent, information about the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
[0067] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processor is also used to: process speech commands through a language reconstruction model, the output of which includes one or more of the following: the user's operational intent, a reference object, information about the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
[0068] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processor is further configured to: in response to adding a new reference object in the target scene, acquire updated scene information, the updated scene information including the position information of the new reference object; or, in response to removing an existing reference object in the target scene, acquire updated scene information, the updated scene information not including information of the removed existing reference object; or, in response to moving an existing reference object in the target scene, acquire updated scene information, the updated scene information including the position information of the moved existing reference object; and determine the target device based on the updated scene information and the voice command.
[0069] Fifthly, an apparatus for controlling a device is provided, the apparatus including modules for implementing the second aspect and any possible implementation thereof.
[0070] In a sixth aspect, an electronic device is provided, comprising a processor and a memory for storing program instructions, and the processor for executing the methods of the second aspect and any possible implementation thereof.
[0071] In a seventh aspect, a computer program product is provided, the computer program product including computer program code, which, when run on a computer, causes the method in the first aspect and any possible implementation thereof or the method in the second aspect and any possible implementation thereof to be executed.
[0072] Eighthly, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the methods in the first aspect and any possible implementation thereof, or the methods in the second aspect and any possible implementation thereof, to be executed.
[0073] A ninth aspect provides a chip including a processor for reading instructions stored in a memory, wherein when the processor executes the instructions, the chip performs a method of the first aspect and any possible implementation thereof, or a method of the second aspect and any possible implementation thereof. Attached Figure Description
[0074] Figure 1 is a schematic diagram of a system architecture applicable to an embodiment of this application.
[0075] Figure 2 is a flowchart of a method for controlling a device provided in an embodiment of this application.
[0076] Figure 3 is a schematic diagram of a method for controlling a device provided in an embodiment of this application.
[0077] Figure 4 is a flowchart of another method for controlling a device provided in an embodiment of this application.
[0078] Figures 5 to 7 are flowcharts of the method for updating scenario information provided in the embodiments of this application.
[0079] Figure 8 is a schematic diagram of another method for controlling a device provided in an embodiment of this application.
[0080] Figure 9 is a flowchart of another speech processing method provided in an embodiment of this application.
[0081] Figure 10 is a schematic diagram of another method for controlling a device provided in an embodiment of this application.
[0082] Figure 11 is a flowchart of another speech processing method provided in an embodiment of this application.
[0083] Figure 12 is a flowchart of a model training and inference process provided in an embodiment of this application.
[0084] Figure 13 is a schematic diagram of a speech processing device provided in an embodiment of this application.
[0085] Figure 14 is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0086] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0087] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. In the description of this application, it should be understood that the terms “center,” “longitudinal,” “lateral,” “upper,” “lower,” “front,” “rear,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” and “outer,” etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0088] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0089] Before formally introducing the embodiments of this application, some terms that may be used in the following examples will be explained and described.
[0090] Automatic speech recognition (ASR), also known as speech-to-text (STT), aims to use computers to automatically convert human speech into corresponding text.
[0091] Natural language processing (NLP) is a branch of artificial intelligence and linguistics. This field explores how to process and utilize natural language. NLP includes multiple aspects and steps, fundamentally comprising cognition, understanding, and generation.
[0092] Point cloud is a collection of point data on the surface of a product obtained through measuring instruments.
[0093] Point cloud segmentation sequence number: In the point cloud segmentation process, a point cloud dataset containing a large number of points is usually divided into multiple sub-regions or objects. Each sub-region represents an independent entity in the point cloud, and each sub-region can correspond to a point cloud segmentation sequence number.
[0094] A bounding box (BBOX) is a method for approximating complex geometric objects. It simplifies computation by enclosing the object with a slightly larger, simpler geometric shape. Bounding boxes are widely used in computer graphics, physics simulations, and collision detection.
[0095] Axis-aligned bounding box (AABB) collision detection is a collision detection method used in computer graphics and physics simulations. In this method, the minimum hexahedron that encloses the 3D object is used to approximate the 3D object.
[0096] The voice processing method provided in this application can be applied to multi-device scenarios such as smart homes and smart cockpits. Figure 1 shows a schematic diagram of a system architecture provided in an embodiment of this application. This system architecture may involve one or more controllable devices, one or more voice assistant devices, and a central control device. The controllable devices and the central control device can be connected via wired and / or wireless means, and the voice assistant devices and the central control device can be connected via wired and / or wireless means.
[0097] A voice assistant device receives the user's voice and generates voice control commands, which are then transmitted to a central control device. The central control device can then control the corresponding controllable devices based on these voice commands, such as turning them on or off, adjusting volume, brightness, etc. In some scenarios, controllable devices may also be referred to as controlled devices or devices subject to control.
[0098] In smart home scenarios, voice assistant devices refer to devices with microphones that can receive user voice commands, such as speakers, televisions with microphones, and alarm clocks with microphones. Controllable devices can include one or more of the following: audio-visual equipment, lighting equipment, and security equipment, such as lights, air conditioners, televisions, speakers, alarm clocks, and curtains.
[0099] In the context of smart cockpits, voice assistant devices refer to devices equipped with microphones that can receive user voice commands, such as speakers and audio-visual entertainment devices with microphones. Controllable devices can include lights, air conditioning, audio-visual entertainment devices, various sensors such as cameras, adjustable seats, and so on.
[0100] In some examples, the voice assistant device, the central control device, and the controllable device can be three separate devices. For instance, a user issues voice commands to a smart speaker (equivalent to a voice assistant device) and adjusts the temperature of a bedroom air conditioner (equivalent to a controllable device) through a gateway (equivalent to a central control device).
[0101] In some examples, the voice assistant device, central control device, and controllable device can also be integrated into a single device. For instance, a user issues a voice command to a smart speaker to adjust its volume.
[0102] In some examples, the voice assistant device, the central control device, and the controllable device can be any two devices integrated into one device. For example, a user issues a voice command to a smart speaker, which recognizes the user's voice command and sends it directly to the bedroom air conditioner to adjust the temperature. In this scenario, the smart speaker is equivalent to both the voice assistant device and the central control device, while the bedroom air conditioner is equivalent to the controllable device.
[0103] Figure 2 is a schematic flowchart of a method for controlling a device provided in an embodiment of this application.
[0104] After the voice assistant device receives the user's voice command, the voice assistant device or central control device generates text from the user's voice command through automatic speech recognition. The voice assistant device or central control device processes the above text through natural language processing to generate control instructions, such as text containing <task, operation, device description>. In this control instruction, the device description can refer to the identifier of the controllable device, the operation can be used to indicate the user's operation intention, and the task can be used to indicate the device type to which the controllable device belongs. After identifying the controllable device, the central control device can control the controllable device according to the control instructions, or send the control instructions to the controllable device, and the controllable device executes the control instructions.
[0105] In the above-mentioned equipment control process, the difficulty lies in determining the controllable equipment that can realize the user's intentions.
[0106] In order to more efficiently identify target devices from controllable devices that can be used to realize user intentions, embodiments of this application provide a voice processing method, which is described below.
[0107] In the following examples, the controllable device that can be used to realize the user's intent is referred to as the target device. In some scenarios, the target device can also be understood as the controllable device indicated by the user's voice command.
[0108] In some home scenarios, referring to Figures 3 and 4, the user's voice command can be "Xiao E, turn on the light next to the dining table against the wall", hereinafter referred to as voice command V1.
[0109] According to the control device method shown in Figure 2, the processing flow of the voice command by the voice assistant device, the central control device, and the controllable device can be as follows: the voice assistant device obtains the voice command V1; the voice assistant device or the central control device can use ASR to generate text from the user's voice command; the voice assistant device or the central control device can use NLP to process and parse the aforementioned text, and obtain one or more of the following information: operation intention, room information, reference object information, the positional relationship between the reference object and the controllable device, or the type of controllable device, etc.
[0110] For example, by parsing the above voice command V1, we can obtain the user's operation intention as "open", the reference object information as "dining table", the positional relationship between the reference object and the controllable device as "dining table and controllable devices inside and near the dining table", and the type of controllable device as "lamp".
[0111] In some examples, the voice assistant device or central control device can acquire scene information that can be used to indicate the positional relationships of different objects (e.g., furniture, equipment, objects, etc.) in a home scene or cabin scene.
[0112] For example, the scene information may include positional information of different objects in the scene, thereby indicating the positional relationship between different objects in the scene.
[0113] For example, the scene information can also directly include the positional relationships of different objects in the scene.
[0114] One possibility is that the aforementioned scene information can be stored on the local disk of the voice assistant device or the central control device. In this case, the voice assistant device or the central control device obtaining scene information can also be understood as the voice assistant device or the central control device reading scene information from the local disk.
[0115] One possibility is that the aforementioned scene information can be stored on servers or other electronic devices. In this case, the voice assistant device or central control device obtaining scene information can also be understood as the voice assistant device or central control device receiving scene information from other electronic devices via the network.
[0116] The aforementioned scene information can be represented in different ways, or in other words, the aforementioned scene information can originate from different data. For example, the aforementioned scene information can be determined based on one or more of the following: point cloud information of different objects in the scene, bounding box information of different objects, image information from different perspectives in the scene, video information, text information used to describe the positions of different objects in the scene or the positional relationships between different objects, etc.
[0117] As an example and not a limitation, the image information from different perspectives in the above-mentioned home or cabin scenarios can be a floor plan of the house or vehicle space, an image taken by the user using an electronic device (such as a mobile phone), etc.
[0118] Methods for determining the positional relationships between different objects in a scene using one or more of the above information will be described in detail below and will not be elaborated here.
[0119] In some examples, based on the parsing results of the voice command V1, the voice assistant device or central control device can determine the controllable device, i.e., the target device, for implementing the voice command V1.
[0120] For example, a voice assistant device or a central control device can determine the target device based on information about the type of controllable device.
[0121] For example, in a user's home or cabin scenario, there might be only one controllable device, D1, whose controllable device type is "light." Based on this, the voice assistant device or central control device can identify controllable device D1 as the target device.
[0122] It should be noted that the home or cockpit scene here exemplifies the inclusion of only one controllable device of type "light," and is not limited to a single light in the home or cockpit scene. One possibility is that the home or cockpit scene contains multiple lights, but only controllable device D1 can be controlled via a voice assistant device or a central control device.
[0123] In some examples, by combining the parsing results of voice command V1 with the above-mentioned scenario information, the voice assistant device or central control device can identify the target device.
[0124] One possibility is that the voice assistant device or central control device determines, based on the type information of the controllable devices, that there are multiple controllable devices of the same type; that is, the controllable device type determined by the voice command is not unique. Based on this, the voice assistant device or central control device can select the controllable device that satisfies the positional relationship with the reference object as the target device, according to the scene information.
[0125] In some examples, the voice assistant device or central control device can determine the controllable device based on the reference object information contained in the parsing result of the voice command V1, the positional relationship between the reference object and the controllable device, and the scene information.
[0126] For example, the voice assistant device or central control device can first determine the reference object indicated in the voice command V1 and the position of the reference object in the scene information, and then determine the controllable device that satisfies the positional relationship with the reference object.
[0127] One possibility is that the scene information contains the positional relationship between the controllable device and the reference object indicated in the voice command V1. In this case, the voice assistant device or the central control device can match controllable devices that satisfy the positional relationship with the reference object from the scene information.
[0128] For example, in the home scene shown in Figure 3, the controllable devices of type "light" include: light group 1, light group 2, light group 3, and light group 4. Based on the scene information, the voice assistant device or the central control device can determine that the dining table is located below the schematic diagram shown in Figure 3, and that light group 4 is located near the dining table. Based on this, the voice assistant device or the central control device can identify light group 4 as the target device.
[0129] Another possibility is that the scene information does not contain the positional relationship between the controllable device and the reference object indicated in the voice command V1. In this case, the voice assistant device or the central control device can determine the controllable device in the scene information that satisfies the positional relationship with the reference object based on the positional relationship between the reference object and the controllable device contained in the parsing result of the voice command V1, as well as the positional information of multiple controllable devices contained in the scene information. Then, it can combine other information parsed from the voice command V1 to determine the target device.
[0130] For example, in the parsing result of the aforementioned voice command V1, the reference object is "dining table," the positional relationship between the controllable device and the reference object is "located next to the dining table," and the type of the controllable device is "lamp." Referring to Figure 3, the controllable devices next to the dining table include speaker 2, router, and lamp group 4. Based on the type of controllable device, the voice assistant device or central control device can determine that lamp group 4 is the target device.
[0131] After the controllable device is identified, the voice assistant device or the central control device can generate an instruction M1 that acts on the controllable device based on the parsing result of the voice instruction V1.
[0132] For example, the control command M1 can be: <Turn on light, turn on, D1>.
[0133] Among them, "turn on the light" is used to indicate the task information contained in the control command M1, "open" is used to indicate the operation information applied to the controllable device, and "D1" is used to indicate the specific controllable device.
[0134] In the above example, the voice assistant device or central control device can obtain the positional relationship between the already constructed controllable device and the reference object, and use the positional relationship between the controllable device and the reference object to determine the controllable device.
[0135] In some examples, the positional relationship between the controllable device and the reference object can be roughly divided into two categories: the controllable device is located inside the reference object and the controllable device is located outside the reference object.
[0136] Among them, the controllable device located outside the reference object can be further divided into: the controllable device located above the reference object, the controllable device located below the reference object, the controllable device located to the left of the reference object, the controllable device located to the right of the reference object, the controllable device located in front of the reference object, and the controllable device located behind the reference object.
[0137] For example, in a furniture scenario, the reference object can be a wardrobe, a TV cabinet, etc., and the controllable device can be a speaker, a light, etc. The speaker can be located near the wardrobe, and the light can be located inside the TV cabinet. In a cockpit scenario, the reference object can be a car door, a glove box, etc., and the controllable device can be an air conditioner, a speaker, a light, etc. The air conditioner can be located near the car door, and the speaker can be located inside the glove box.
[0138] In some examples, the positional relationship between the controllable device and the reference object can be determined based on the minimum distance between the controllable device and the reference object.
[0139] For example, the distance between the center point of the reference object and the center point of the controllable device can be used to represent the minimum distance between the controllable device and the reference object. The distance d1 between the center point of the reference object and the center point of the controllable device in three-dimensional space can be calculated using the following formula.
[0140] The distance d2 between the center point of the reference object and the center point of the controllable device in the plane can be calculated using the following formula.
[0141] Where (x0, y0, z0) represents the three-dimensional coordinates of the center point of the controllable device, (x0, y0, z0) b y b , z b (s) represents the three-dimensional coordinates of the center point of the reference object. x s y s z ) represent the dimensions of the reference object in the X-axis, Y-axis and Z-axis directions in three-dimensional space, respectively.
[0142] According to the above formula, when the distance between the center point of the reference object and the center point of the controllable device is less than or equal to half the dimension of the reference object in the X, Y, or Z axis direction, the distance between the center point of the reference object and the center point of the controllable device in that axis direction is recorded as zero. In this case, the controllable device can be considered to be located inside the reference object in that axis direction. When the minimum distance d1 between the center point of the controllable device and the center point of the reference object is zero, the controllable device can be considered to be located inside the reference object.
[0143] In some examples, when the controllable device is located outside the reference object, the vertical relationship between the controllable device and the reference object can be determined based on the positional relationship between the center point of the reference object and the center point of the controllable device in the height direction.
[0144] For example, taking the Z-axis direction as the height direction, if z0 is less than z when d1 and / or d2 are not zero. b Then it can be determined that the controllable device is located below the reference object; if z0 is greater than z b This confirms that the controllable device is located above the reference object.
[0145] One possible scenario is that the controllable device is located above or below the reference object, but the controllable device and the reference object are not on the same floor. To avoid errors in selecting the controllable device in this situation, the aforementioned z0 and z... bThe difference between them should also be less than or equal to a preset distance, for example, |z0-z b |≤A1, where A1 is greater than 0.
[0146] In some examples, when the controllable device is located outside the reference object, the positional relationship (front-back relationship, left-right relationship) between the controllable device and the reference object in the same plane can be determined based on the distance between the center point of the reference object and the center point of the controllable device in the same plane.
[0147] One possibility is that the positional relationship between the controllable device and the reference object in the same plane can be represented by whether the controllable device is near the reference object. For example, for the distance d2 between the center point of the reference object and the center point of the controllable device in the plane, if d2 is less than or equal to a preset threshold, it can be determined that the controllable device is near the reference object; otherwise, it can be determined that the controllable device is not near the reference object.
[0148] For example, in Figure 3, the distances in the plane between the center point of speaker 2, the center point of router, or the center point of light group 4 and the center point of dining table are all less than a preset threshold. Therefore, speaker 2, router, and light group 4 can all be considered to be located near the dining table. The distances in the plane between the center point of speaker 1, the center point of light group 1, and the center point of light group 2 and the center point of dining table are all greater than the preset threshold. Therefore, speaker 1, light group 1, and light group 2 are not considered to be located near the dining table.
[0149] The above evaluation method only considers the positional relationship between the controllable device and the reference object in the same plane. For controllable devices located near the reference object, considering that these controllable devices may also have vertical positional relationships with the reference object, in order to reduce the situation where the same controllable device has multiple positional relationships with the reference object and improve the efficiency of identifying controllable devices, when d2 is less than or equal to a preset threshold, the distance difference in the height direction between the center point of the reference object and the center point of the controllable device can be further determined, i.e., |z0-z b |。 In |z0-z b If |≤A2 (A2 is greater than 0), it can be determined that the controllable device is located near the reference object.
[0150] For example, in Figure 3, the distance between the center point of the speaker 2 or the center point of the light group 3 and the center point of the robot vacuum cleaner in the plane is less than a preset threshold. However, the difference in the distance between the center point of the speaker 2 and the center point of the robot vacuum cleaner in the height direction meets the above requirements, while the difference in the distance between the center point of the light group 3 and the center point of the robot vacuum cleaner in the height direction does not meet the above requirements. In this case, the speaker 2 can be regarded as being located near the robot vacuum cleaner, but the light group 3 can not be regarded as being located near the robot vacuum cleaner.
[0151] In some examples, the distance d2 between the center point of the reference object and the center point of the controllable device in the plane satisfies: d2 ≤ A3 (A3 > 0), then the controllable device is determined to be near the reference object. In some examples, the distance d2 between the center point of the reference object and the center point of the controllable device in the plane satisfies: d2 ≤ A3 (A3 > 0), and the difference in distance between the center point of the reference object and the center point of the controllable device in the height direction satisfies: |z0 - z b If |≤A2 (A2 is greater than 0), then the controllable device is located near the reference object.
[0152] One possible approach is to determine the front-back and left-right relationships between the controllable device and the reference object in the same plane based on the size relationship between the center point of the reference object and the center point of the controllable device in the X-axis and / or Y-axis directions.
[0153] For example, the coordinates of the reference object and the controllable device in the X-axis direction are x b x0, if x b If x > 0, then it can be determined that the controllable device is located behind the reference object in the X-axis direction; if x b If x < x0, then it can be determined that the controllable device is in front of the reference object in the X-axis direction.
[0154] For example, in the coordinate system shown in Figure 3, speaker 2 is located in front of the robot vacuum cleaner.
[0155] For example, the coordinates of the reference object and the controllable device in the Y-axis direction are y and y, respectively. b y0, if y b If y > 0, then the controllable device is located to the left of the reference object in the Y-axis direction; if y b If y < 0, then it can be determined that the controllable device is located to the right of the reference object in the Y-axis direction.
[0156] For example, in the coordinate system shown in Figure 3, the dining table is located to the left of the router.
[0157] By combining the coordinates of the reference object and the controllable device in the X and Y axes, it is also possible to determine whether the controllable device is to the right front, right rear, left front, or left rear of the reference object. For example, if x b >x0 and y b If x > y0, then the controllable device is located to the left rear of the reference object; if x b <x0 and y b If y < 0, then the controllable device is located to the right front of the reference object; if x b >x0 and y b If x < y0, then the controllable device is located to the right rear of the reference object; if x b <x0 and y bIf y < 0, then it can be determined that the controllable device is located to the left front of the reference object.
[0158] For example, in the coordinate system shown in Figure 3, speaker 2 is located to the right front of the robot vacuum cleaner, and light group 2 is located to the left rear of the router.
[0159] The reference objects used to determine the target device can be either fixed reference objects in a home scene or a cockpit scene, or movable reference objects in a home scene or a cockpit scene.
[0160] For example, in a home setting, walls, floors, or ceilings can serve as fixed reference points. Voice assistant devices or central control devices can distinguish between different controllable devices based on their positional relationship with these fixed reference points.
[0161] For example, based on whether the distance between the controllable device and the wall is less than or equal to a preset threshold, multiple controllable devices can be classified as "close to the wall" and "not close to the wall"; based on whether the distance between the controllable device and the ground is less than or equal to a preset threshold, multiple controllable devices can be classified as "close to the ground" and "not close to the ground"; based on whether the distance between the controllable device and the ceiling is less than or equal to a preset threshold, multiple controllable devices can be classified as "close to the ceiling" and "not close to the ceiling".
[0162] Regarding the aforementioned voice command V1, the voice assistant device or central control device can combine the positional relationship between the controllable device and the dining table, as well as whether the controllable device is against a wall, to select the wall-adjacent light group 4 from light groups 1, 2, 3, and 4 as the controllable device indicated by voice command V1. The voice assistant device or central control device is more efficient in determining the controllable device.
[0163] Similarly, in a cockpit scenario, the vehicle's roof, chassis, or doors can serve as fixed reference points. The voice assistant or central control device can distinguish between different controllable devices based on their positional relationship to these fixed reference points. For example, based on whether the distance between a controllable device and the vehicle's roof is less than or equal to a preset threshold, multiple controllable devices can be categorized as "close to the roof" and "not close to the roof"; based on whether the distance between a controllable device and the vehicle's chassis is less than or equal to a preset threshold, multiple controllable devices can be categorized as "close to the chassis" and "not close to the chassis"; and based on whether the distance between a controllable device and the door is less than or equal to a preset threshold, multiple controllable devices can be categorized as "close to the door" and "not close to the door".
[0164] Voice assistant devices or central control devices can combine the positional relationship between the controllable device and reference objects, as well as whether the controllable device is close to the roof (or chassis, doors, etc.) of the vehicle, to determine the controllable device that can realize the user's intention, thereby improving the efficiency of determining the controllable device.
[0165] Besides the aforementioned reference objects such as walls, ceilings, floors, car roofs, and doors, most reference objects in home or cabin scenarios are movable. For example, in a home scenario, a user might buy a new bed. In a cabin scenario, a user might take away a toy figurine. Unlike the fixed reference objects mentioned above, the "bed," "toy figurine," and "dining table" in the V1 voice command can all be considered movable reference objects. For these movable reference objects, timely updating their location information and the relative positions of controllable devices and movable reference objects can improve the efficiency of voice assistant devices or central control devices in identifying target devices.
[0166] Here, "movable" can include changes in the reference object from non-existent to existing (i.e., adding), changes in the reference object from existing to non-existent (i.e., removing), or changes in the position of the reference object (i.e., moving).
[0167] In some examples, when a new reference object (e.g., referred to as reference object R1) is added to a home scene or cockpit scene, the existing scene information can be updated by adding relevant information about reference object R1 to the existing scene information.
[0168] For example, the relevant information of the reference object R1 may include one or more of the following: point cloud information of the reference object R1, position of the reference object R1, orientation of the reference object R1, or relative position of the reference object R1 to one or more controllable devices.
[0169] Figure 5 is an exemplary schematic diagram of a method for updating existing scene information when adding new furniture to a home scene.
[0170] In one possible implementation, the point cloud information of the reference object R1 can be retrieved from an existing object library by searching for information such as the name and identifier of the reference object R1. Alternatively, the point cloud information of the reference object R1 can also be entered into the computer based on the physical object of the reference object R1, for example, generated after calculation based on multi-view photos of the reference object R1.
[0171] In the process of adding the point cloud of a new object to the point clouds of multiple objects in the existing scene information, the point cloud of the new object can be appropriately modified based on information such as the placement and orientation of the new object. The modified point cloud of the new object can then be merged with the point clouds of multiple objects in the existing scene information to obtain an updated scene point cloud. Based on the modified point cloud of the new object, information such as the bounding box of the new object, the category of the new object, and the point cloud segmentation sequence number of the new object can also be determined. Furthermore, the positional relationship between the new object and other objects in the existing scene information can be determined. The updated scene point cloud, the bounding box of the new object, its category, and the positional relationship between the new object and existing objects can all together constitute the updated (after adding the object) scene information.
[0172] In some examples, when an existing reference object (e.g., referred to as reference object R2) is removed from a home or cockpit scene, the existing scene information can be updated by removing the relevant information of reference object R2 from the existing scene information.
[0173] For example, the relevant information of the reference object R2 may include one or more of the following: point cloud information of the reference object R2, position of the reference object R2, orientation of the reference object R2, or relative position of the reference object R2 with one or more controllable devices.
[0174] Figure 6 is an exemplary schematic diagram of a method for updating existing scene information when an existing object is removed in a cockpit scene.
[0175] In the process of removing information about existing objects from existing scene information, the identifier of the object to be removed can be obtained first. Based on the identifier, the point cloud segmentation sequence number corresponding to the identifier can be found in the existing scene information. Then, the point cloud of the object to be removed, the bounding box of the object to be removed, and the positional relationship between the object to be removed and other objects can be removed, thereby obtaining the updated (after removing the object) scene information.
[0176] In some examples, when an existing reference object (e.g., referred to as reference object R3) is moved in a home or cockpit scene, the existing scene information can be updated by adjusting the relevant information of reference object R3 before the move to the relevant information of reference object R3 after the move in the scene information.
[0177] For example, the relevant information of the reference object R3 may include one or more of the following: point cloud information of the reference object R3, position of the reference object R3, orientation of the reference object R3, or relative position of the reference object R3 with one or more controllable devices.
[0178] Figure 7 is an exemplary schematic diagram of a method for updating existing scene information when moving existing objects in a home scene.
[0179] The process of updating scene information when an object moves can be performed by referring to the process of updating scene information when removing an existing object and the process of updating scene information when adding a new object.
[0180] For example, during the process of updating scene information, information such as the identifier of the moving object, its new position, and its new orientation can be obtained. Based on the identifier of the moving object, the point cloud segmentation number of the moving object can be obtained from the existing scene information. Afterward, the point cloud of the moving object can be transformed according to its new position and orientation, and the original point cloud of the moving object can be deleted from the existing scene information. A new scene point cloud can be obtained by merging the transformed point cloud of the moving object with the point clouds of other objects. Similarly, using the transformed point cloud of the moving object, information such as the bounding box, segmentation number, and positional relationship with other objects of the moving object can be recalculated, and the original bounding box, category, and positional relationship with other objects of the moving object can be deleted from the existing scene information, thereby obtaining the updated (after moving the object) scene information.
[0181] In some examples, the voice assistant device or central control device can determine the target device by combining the positional relationship between the controllable device and a movable reference object as well as the positional relationship between the controllable device and a fixed reference object.
[0182] For example, in the relevant examples based on the above voice command V1, the voice assistant device or central control device can determine the target device by combining the positional relationship between the controllable device and the dining table and the positional relationship between the controllable device and the wall.
[0183] In some examples, a voice assistant device or central control device can determine the target device by combining the positional relationship between the controllable device and multiple reference objects.
[0184] For example, in the relevant examples based on the above voice command V1, the voice assistant device or central control device can determine the target device by combining the positional relationship between the controllable device and the dining table and the positional relationship between the controllable device and the wall.
[0185] In some examples, the user in a home or cockpit setting can also be considered a movable reference object. In some scenarios, referring to Figures 8 and 9, the user's voice command could be "Xiao E, turn off the speaker here," hereinafter referred to as Voice Command V2.
[0186] In some examples, after the voice assistant device receives voice command V2, the voice assistant device or central control device can use ASR to generate text from the user's voice command, and convert various pronouns in the generated text into "person," "user," or "person's" and "user's," resulting in the converted text. For example, converting "I" in voice command V2 into "person" transforms voice command V2 into "Little E, turn off the speaker here." Another example is converting "Tom," "Jane," and "Johnson" into "person," and "brother's," "my," and "sister's" into "user's."
[0187] After the voice command V2 is converted, the voice assistant device or central control device can use NLP to process and parse the user's voice command to obtain one or more of the following information: operation intention, room information, reference object information, positional relationship between the controllable device and the reference object, or type of controllable device, etc.
[0188] For example, by parsing the above voice command V2, we can obtain that the user's operation intention is "turn off", the reference object information can be "person (user)", the positional relationship between the controllable device and the reference object can be "controllable device near the person", and the type of controllable device can be "speaker".
[0189] In some examples, the voice assistant device or central control device can acquire scene information that can be used to indicate the positional relationships of different objects (e.g., furniture, equipment, objects, etc.) in a home scene or cabin scene.
[0190] For example, the scene information may include positional information of different objects in the scene, thereby indicating the positional relationship between different objects in the scene.
[0191] For example, the scene information can also directly include the positional relationships of different objects in the scene.
[0192] For details on how scene information is acquired and what it contains, please refer to the description of the relevant section on voice command V1 above. We will not repeat it here.
[0193] In some examples, the scene information may include not only the positions of objects and their relative positions in a home or cockpit scene, but also the position of a reference character and their relative positions to other objects. The reference character's information can have default values; for example, the default position of the reference character could be the center of the living room, and the default orientation could be due south.
[0194] Based on the parsing results of the aforementioned voice command V2, the voice assistant device or central control device can obtain the location information of the user (target person) who issued the voice command V2, the target person's orientation information, etc., and update the scene information accordingly. The process of updating scene information can be referenced in the previous description of the process of updating scene information when the reference object R3 moves.
[0195] For example, the point cloud of a reference person in the existing scene information is transformed using the target person's position and orientation information to obtain the target person's point cloud. The reference person's point cloud in the existing scene information is then deleted. The target person's point cloud is then merged with the point clouds of other objects in the existing scene information to obtain a new scene point cloud. Using the target person's point cloud, information such as the target person's bounding box, segmentation number, and positional relationship with other objects can be calculated. By deleting the reference person's bounding box, category, and positional relationship with other objects in the existing scene information, scene information containing the target person's point cloud, the target person's position, and the target person's positional relationship with other objects can be obtained; this is the updated scene information (including the target person).
[0196] In some examples, based on the parsing results of the voice command V2, the voice assistant device or central control device can determine the controllable device indicated by the voice command V2, i.e., the target device.
[0197] For example, a voice assistant device or a central control device can determine the target device based on information about the type of controllable device.
[0198] For example, in a user's home or cabin scenario, there might be only one controllable device Y1, and its controllable device type is "speaker". Based on this, the voice assistant device or central control device can identify controllable device Y1 as the target device.
[0199] It should be noted that the home or cockpit scenario here exemplifies the inclusion of only one controllable device of type "speaker," and is not limited to a single speaker in the home or cockpit scenario. One possibility is that the home or cockpit scenario includes multiple speakers, but only controllable device Y1 can be controlled via a voice assistant device or a central control device.
[0200] In some examples, by combining the parsing results of voice command V2 with the updated scenario information mentioned above, the voice assistant device or central control device can identify the target device.
[0201] One possibility is that the voice assistant device or central control device determines, based on the information about the type of controllable device, that there are multiple controllable devices of the same type. Based on this, the voice assistant device or central control device can determine the controllable device that satisfies the positional relationship between the controllable device and the target person based on the updated scene information.
[0202] In some examples, the voice assistant device or central control device can determine the target device based on the information of the target person contained in the parsing result of the voice command V2 and the positional relationship between the controllable device and the target person, combined with the updated scene information.
[0203] For example, a voice assistant device or a central control device can first determine the location of the target person in the scene information, and then determine a controllable device that satisfies the positional relationship with the target person.
[0204] One possibility is that the updated scene information includes the positional relationship between the controllable device and the target person. In this case, the voice assistant device or the central control device can match controllable devices that meet the positional relationship with the target person from the scene information.
[0205] For example, in the home scene shown in Figure 8, the controllable devices of type "speaker" include speaker 1 and speaker 2. Based on the scene information, the voice assistant device or central control device can determine that the target person is located above the schematic diagram shown in Figure 8, and that speaker 1 is located near the target person. Based on this, the voice assistant device or central control device can identify speaker 1 as the target device.
[0206] Another possibility is that the scene information does not contain the positional relationship between the controllable device and the target person. In this case, the voice assistant device or the central control device can determine the controllable device that satisfies the positional relationship with the target person based on the positional relationship between the target person and the controllable device contained in the parsing result of the voice command V2, and determine the target device based on the positional information of multiple controllable devices contained in the updated scene information, and then combine it with other information parsed from the voice command V2.
[0207] For example, in the parsing result of the aforementioned voice command V2, the positional relationship between the controllable device and the target person is "near the person," and the type of the controllable device is "speaker." Referring to Figure 8, the controllable devices near the target person include speaker 1 and light group 1. Based on the type of controllable device, the voice assistant device or central control device can determine that speaker 1 is the target device.
[0208] In some examples, the distance between the center coordinates of the controllable device and the center coordinates of the target person can be used to determine whether the controllable device is near the target person. For details, please refer to the description above.
[0209] After the controllable device is identified, the voice assistant device or the central control device can generate a command M2 that acts on the controllable device based on the parsing result of the voice command V2.
[0210] For example, the control command M2 can be: <speaker, off, Y1>.
[0211] Among them, "speaker" is used to indicate the task information contained in the control command M2, that is, the equipment type to which the controllable device belongs; "off" is used to indicate the operation information applied to the controllable device; and "Y1" is used to indicate the specific controllable device.
[0212] For voice commands such as "turn on the lamp next to your younger brother" and "turn up the volume of the speaker next to your mother's feet", the processing method can refer to the processing method of the aforementioned voice command V2, and will not be repeated here.
[0213] In some examples, the scenario information provided in this application embodiment may also include relevant information about the reference animal, such as the default position and default orientation of the reference animal. Based on this, this application embodiment may also use the animal as a reference to determine the controllable device. The specific implementation method can be referred to the processing method of voice command V2 mentioned above, which will not be elaborated here.
[0214] In home or cabin settings, various controllable devices with specific directions and / or ranges of action may be involved, such as spotlights, fans, air conditioners, and cameras. Due to limitations in their direction and / or range, different controllable devices may not be able to operate within the area indicated by the user's voice commands. When selecting controllable devices, considering their capabilities (such as direction and / or range of action) can improve efficiency and increase the likelihood of the devices fulfilling the user's intentions.
[0215] In some scenarios, referring to Figures 10 and 11, the user's voice command can be "Xiao E, turn the TV screen up a little brighter", hereinafter referred to as Voice Command V3.
[0216] According to the control device method shown in Figure 2, the processing flow of the voice assistant device, central control device, and controllable device for the above-mentioned voice command can be as follows: the voice assistant device obtains the above-mentioned voice command V3; the voice assistant device or central control device can use ASR to generate text from the user's voice command; the voice assistant device or central control device can use NLP to process and parse the aforementioned text, and obtain one or more of the following information: operation intention, room information, reference object information, positional relationship between the controllable device and the reference object, or type of controllable device, etc.
[0217] For example, by parsing the above voice command V3, we can obtain that the user's intention is "to turn off", the reference object information can be "television", the positional relationship between the controllable device and the reference object can be "controllable device near the television", and the type of controllable device can be "lamp".
[0218] In some examples, the voice assistant device or central control device can acquire scene information that can be used to indicate the positional relationships of different objects (e.g., furniture, equipment, objects, etc.) in a home scene or cabin scene.
[0219] For example, the scene information may include positional information of different objects in the scene, thereby indicating the positional relationship between different objects in the scene.
[0220] For example, the scene information can also directly include the positional relationships of different objects in the scene.
[0221] For details on how scene information is acquired and what it contains, please refer to the description of the relevant section on voice command V1 above. We will not repeat it here.
[0222] In some examples, based on the parsing results of voice command V3, the voice assistant device or central control device can determine the controllable device indicated by voice command V3, i.e., the target device.
[0223] For example, a voice assistant device or a central control device can determine the target device based on information about the type of controllable device.
[0224] For example, in a user's home or cabin scenario, there might be only one controllable device, D2, whose controllable device type is "light." Based on this, the voice assistant device or central control device can identify controllable device D2 as the target device.
[0225] It should be noted that the home scene or cockpit scene here exemplifies only one controllable device of type "lamp", and is not limited to only one lamp in the home scene or cockpit scene.
[0226] In some examples, by combining the parsing results of voice command V3 with the above-mentioned scenario information, the voice assistant device or central control device can identify the target device.
[0227] For example, the voice assistant device or central control device determines that there are multiple controllable devices of the same type based on the information of the controllable device type, and then determines the controllable device that satisfies the positional relationship with the reference object by combining the scene information.
[0228] Taking the aforementioned voice command V3 as an example, the voice assistant device or central control device can determine the location of the television, and then determine controllable devices that satisfy the positional relationship with the television. Referring to Figure 10, the four spotlights located above the television are spotlight 1, spotlight 2, spotlight 3, and spotlight 4, and no light is located near the television. After determining the location of the television, the voice assistant device or central control device can select the aforementioned spotlights 1, 2, 3, and 4. In this case, spotlights 1, 2, 3, and 4 can be used as target devices.
[0229] In some examples, by combining the parsing results of the voice command V3, the aforementioned scenario information, and information on the device capabilities of multiple controllable devices, the voice assistant device or central control device can identify the target device.
[0230] For example, a voice assistant device or a central control device can determine multiple controllable devices based on information about the type of controllable device and the positional relationship between the controllable device and the reference object in the scene information. Among these controllable devices, some can realize the intent of voice command V3, while some controllable devices cannot realize the intent of voice command V3. In this case, the voice assistant device or the central control device can obtain information about the device capabilities of multiple controllable devices and determine the target device based on the capability information of the controllable devices.
[0231] Here, the capability information of the controllable device can be used to instruct the controllable device on its performance in executing control commands to implement voice command V3.
[0232] In one possible implementation, the type and capabilities of the controllable device can be included in the scene information. Alternatively, the type and capabilities of the controllable device can be obtained by the voice assistant device or the central control device through network communication.
[0233] For example, the capacity information of an air conditioner may include: cooling, heating, blowing, dehumidifying, ventilation, etc. The capacity information of an air conditioner may also include the temperature range for cooling and the temperature range for heating. The capacity information of an air conditioner may also include the vertical oscillation, horizontal oscillation, and blowing direction of the airflow.
[0234] For example, fan capability information may include: air blowing, and may also include: fan oscillation angle, fan oscillation angle, and fan noise level.
[0235] For example, camera capability information may include: taking photos and recording videos. Camera capability information may also include: camera angle, camera pixels, and whether it can shoot at night.
[0236] For example, the capability information of a spotlight can include: illumination. The capability information of a spotlight can also include beam angle (range of action), adjustable angle (direction of action), luminous flux, and light color temperature.
[0237] Referring to Figure 10, the directions and ranges of action of spotlights 1, 2, 3 and 4 can be represented as (dr1, ra1), (dr2, ra2), (dr3, ra3) and (dr4, ra4), respectively.
[0238] The voice assistant device or central control device can determine, by combining the effective range and direction of the aforementioned four spotlights with the position of the television, that the light emitted by spotlights 2 and 3 can reach the television, while the light emitted by spotlights 1 and 4 cannot reach the television. Based on this, the voice assistant device or central control device can select spotlights 1 and 2 as the target device.
[0239] In one possible implementation, the television set can be approximated as a cuboid T. The process of determining whether a spotlight can illuminate the vicinity of the television set can be equivalent to determining whether the light emitted by the spotlight intersects with the six faces of the aforementioned cuboid T.
[0240] In some examples, the AABB collision detection method can be used to determine whether the light emitted by the spotlight intersects with the cuboid T. The following is a brief explanation of this method.
[0241] Any point on the light ray emitted by the spotlight can be represented by a vector. express, in, The light source of the spotlight is represented by t, which is greater than or equal to zero. Different values of t can represent any point within the illumination range of the light emitted by the spotlight. The aforementioned cuboid T can be approximated by three sets of mutually parallel planes, namely X-slab, Y-slab, and Z-slab. Assuming vectors... If it intersects with the cuboid, then the vector It can intersect with X-slab, Y-slab and Z-slab to obtain 6 intersection points.
[0242] X-slab can be expressed as x = X min and x = X max This means that Y-slab can be represented by y = Y min and y=Y max This means that Z-slab can be represented by z = Z min and x = Z max Indicate. Utilize Combining the above six equations, we can obtain the t values corresponding to the six intersection points, i.e.: t x,1 , t x,2 , ty,1 , t y,2 , t z,1 and t z,2 If the following inequality holds for these six t values, it can be determined that the light emitted by the spotlight intersects the cuboid T. max(t) x1 ,t y1 ,t z1 )≤min(t x2 ,t y2 ,t z2 )
[0243] Imagine the illumination range of a spotlight as a cone. Generally, the distance from the light source to any point on the outer periphery of the cone's base is the greatest. In other words, using... This represents any point on the outer periphery of the base of the vertebral body. With light source The distance between them is denoted as d. L Q, any point in space, and the light source distance d Q Under the condition that the following inequality is satisfied, it can be determined that point Q is located within the illumination range of the light emitted by the spotlight. Q ≤d L
[0244] If both of the above inequalities are satisfied, it can be determined that the light emitted by the spotlight can illuminate the vicinity of the television set.
[0245] It should be noted that the above-described method of using AABB collision detection to determine whether the light emitted by the spotlight can illuminate the vicinity of the television is merely exemplary, and this application does not impose any limitations on it.
[0246] To improve the efficiency of voice assistant devices or central control devices in parsing user-inputted voice commands, and to enable devices to process a wider range of natural voice commands, such as abbreviated commands (e.g., "I'm bright here") and complex commands (e.g., "The night light on the wall next to the TV cabinet"), in some examples, a language reconstruction model can be introduced to process user-inputted voice commands.
[0247] Figure 12 provides an exemplary schematic diagram of the training and inference processes of the above-mentioned language reconstruction model.
[0248] In the training process of the language reconstruction model, commonly used user voice commands are organized and broken down according to the controlled object, operation method, etc., to form different breakdown results. Different parts of the breakdown results can be divided into different fields according to attributes, such as prefix fields, intent fields, room fields, person fields, reference object fields, etc. A set of multiple different fields can form an output template, and the output template with filled data can serve as the test dataset for training the language reconstruction model. The user's input voice commands are used as input to the language reconstruction model during training, and the filled template is used as the model's output to train the language reconstruction model.
[0249] During the inference process of the language reconstruction model, the voice assistant device or central control device can input the text of the user's voice command into the trained language reconstruction model. The voice reconstruction model is used to reconstruct the language of the task, intent, reference object, controllable device, etc. in the voice command, and obtain prefix fields, intent fields, room fields, person fields, reference object fields, etc. Some or all of these information can be used to determine the target device.
[0250] In some examples, the input information during the training process of the language reconstruction model can be a language description of the user's commonly used voice commands, such as: "Turn on the floor lamp", "Turn off the light for me", "Can you turn up the living room lights for me", "Turn down the night light on the wall next to the TV in the living room", "It's too dark, turn on the bedside lamp", "It's too dark, turn on the table lamp behind the bookshelf", "Please turn down the light on my desk", or "Make the sink area a little brighter", etc.
[0251] In some examples, the output information during the training process of the language reconstruction model can be the result of splitting the input information. The split result of the input information can consist of multiple different fields, and these multiple different fields can form the output template.
[0252] For example, the splitting results of the input information or the output template may include one or more of the following: prefix field, intent field, room field, person field, reference object field, relationship field, quantity field, or controllable device type field, etc.
[0253] Taking the commonly used voice commands mentioned above as an example, the prefix fields in the split results of these commands can include one or more of the following: "I want", "help me", "please help me", "give me", "please give me", "please" and "can I?", etc.
[0254] The intent field in the breakdown results of these commands can include one or more of the following: "on", "off", "brighten", and "dimen". This intent field can correspond to the operation intent in the parsing results of the voice command.
[0255] The room field in the split results of these instructions can include one or more of the following: "living room", "kitchen", "bedroom", "study", "entryway", "toilet", and "bathroom". This room field can be used to indicate the room where the target device is located.
[0256] The person field in the splitting results of these instructions can include one or more of the following: "this is me," "this is me," "here I am," and "here I am." This person field can serve as a basis for determining reference information in the voice instruction parsing results.
[0257] The reference object field in the splitting results of these instructions can include one or more of the following: "door", "window", "bed", "sofa", "television", "refrigerator", "chair", "stool", and "desk". This reference object field can serve as another basis for determining reference object information in the voice instruction parsing results.
[0258] The relational field in the decomposition results of these instructions may include one or more of the following: "of", "inside", "above", "below", "nearby", and "against the wall". This relational field can be used as the basis for determining the positional relationship between the controllable device and the reference object in the voice instruction parsing results.
[0259] The quantity field in the splitting results of these instructions may include one or more of the following: “one”, “one side”, and “all”.
[0260] The controllable device type field in the splitting results of these commands can include one or more of the following: "lamp", "chandelier", "table lamp", "spotlight", "speaker", "air conditioner", and "fan", etc. This controllable device type field can correspond to the type of controllable device in the voice command parsing results.
[0261] In some examples, the user's voice command "turn off all the nightlights in the TV cabinet" is input into the language reconstruction model described above. The output of this model may include the following:
[0262] Task field: Lamp; Intent field: Off; Reference object field: TV cabinet; Relationship field: Inside; Quantity field: All; Controllable device type field: Night light.
[0263] The voice assistant device or central control device can input the text obtained by converting the user's voice command into the above-mentioned language reconstruction model, and determine one or more of the following information based on the information of different fields output by the language reconstruction model: the user's voice command operation intention, room information, reference object information, positional relationship between the controllable device and the reference object, or the type of controllable device.
[0264] The use of language reconstruction models can improve the processing efficiency of voice assistant devices or central control devices for user-inputted voice commands, and can also improve the efficiency of controllable devices.
[0265] It should be noted that, unless otherwise specified, this application does not limit the execution order of different steps in the multiple flowcharts provided in the above examples.
[0266] Based on the same inventive concept, this application also provides a voice processing apparatus 1300, as shown in FIG13. This apparatus 1300 may possess the functions of a voice assistant device or central control device as described in the above method embodiments, and can be used to execute the steps performed by the functions of the voice assistant device or central control device in the above method embodiments. This function can be implemented by hardware, or by software or hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0267] In one possible implementation, the speech processing apparatus 1300 may include an acquisition module 1310 and a processing module 1320, which are coupled to each other.
[0268] In some examples, the acquisition module 1310 can be used to support the voice assistant device in the foregoing embodiments in acquiring user-inputted voice commands, etc.
[0269] The processing module 1320 is used to support electronic devices in performing the processing actions in the above method embodiments, such as determining controllable devices based on voice commands.
[0270] Optionally, the speech processing apparatus 1300 may further include a storage unit 1330 for storing the program code and data of the speech processing apparatus 1300.
[0271] Figure 14 illustrates an electronic device 1400 provided in an embodiment of this application. As shown, the electronic device 1400 includes at least one processor 1410 and a transceiver 1420. The processor 1410 is coupled to a memory and is used to execute instructions stored in the memory to control the transceiver 1420 to transmit and / or receive signals.
[0272] Optionally, the electronic device 1400 also includes a memory 1430 for storing instructions.
[0273] In some embodiments, the processor 1410 and the memory 1430 can be combined into a single processing device, with the processor 1410 executing program code stored in the memory 1430 to achieve the aforementioned functions. Specifically, the memory 1430 can be integrated into the processor 1410 or independent of it.
[0274] In some embodiments, transceiver 1420 may include a receiver (or receiver unit) and a transmitter (or transmitter unit).
[0275] The transceiver 1420 may further include an antenna, and the number of antennas may be one or more. The transceiver 1420 may be a communication interface or an interface circuit.
[0276] When the electronic device 1400 is a chip, the chip includes a transceiver module and a processing module. The transceiver module can be an input / output circuit or a communication interface; the processing module can be a processor, microprocessor, or integrated circuit integrated on the chip.
[0277] This embodiment also provides a device for controlling a device. This device may possess the functions described in the method embodiments for a voice assistant device or a central control device, and may be used to execute the steps performed by the functions of the voice assistant device or central control device in the method embodiments. This function may be implemented in hardware, or in software, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the described functions. In one possible implementation, the device for controlling the device may include an acquisition module and a processing module, which are coupled to each other.
[0278] In some examples, the acquisition module can be used to support the voice assistant device in acquiring user-inputted voice commands and the central control device in issuing control commands to controllable devices, as described in the foregoing embodiments.
[0279] The processing module is used to support electronic devices in performing the processing actions in the above method embodiments, such as determining controllable devices based on voice commands.
[0280] Optionally, the control device may further include a storage unit for storing the program code and data of the control device.
[0281] This embodiment also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the voice processing or control device method in the above embodiment.
[0282] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the voice processing or control device method described in the above embodiment.
[0283] Furthermore, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory. The memory stores computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to perform the voice processing or device control methods described in the above-described method embodiments.
[0284] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0285] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0286] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0287] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0288] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0289] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0290] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method of speech processing, characterized by, The method comprises: obtaining a voice instruction of a user, the voice instruction comprising a reference object; determining a target device according to scene information and the voice instruction, the scene information being used to indicate a positional relationship between a controllable device and the reference object in a target scene, the controllable device comprising the target device, and the target device being used to execute a control instruction corresponding to the voice instruction.
2. The method of claim 1, wherein, The scene information comprises a position of the reference object and a position of the controllable device, or the scene information comprises the positional relationship between the reference object and the controllable device.
3. The method according to claim 1 or 2, characterized in that, The reference object comprises a fixed reference object and / or a movable reference object.
4. The method of claim 3, wherein, The movable reference object comprises a movable object in the target scene and / or the user.
5. The method according to any one of claims 1 to 4, characterized in that, The determination of the target device according to the scene information and the voice instruction comprises: determining the target device according to the scene information, the voice instruction, and type information of the controllable device and capability information of the controllable device, the capability information being used to indicate performance of the controllable device in executing the control instruction.
6. The method of claim 5, wherein, The capability information comprises an action range of the controllable device and / or an action direction of the controllable device.
7. The method according to any one of claims 1 to 6, characterized in that, The positional relationship comprises any one of the following: the controllable device is located inside the reference object, the controllable device is located on one side of the reference object, or the controllable device is located near the reference object.
8. The method according to any one of claims 1 to 7, characterized in that, The scene information is determined according to one or more of the following: point cloud information of an object in the target scene, bounding box information of the object in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, the point cloud information of the object being used to indicate a position and / or an orientation of the object in the target scene.
9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: sending the control instruction to the target device.
10. The method according to any one of claims 1 to 9, characterized in that, The voice instruction further comprises one or more of the following: an operation intention of the user, information of the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
11. The method of claim 10, wherein, The method further comprises: processing the voice instruction through a language reconstruction model, an output of the language reconstruction model comprising one or more of the following: the operation intention of the user, the reference object, information of the target scene, the positional relationship between the controllable device and the reference object, or the type of the controllable device.
12. The method according to any one of claims 1 to 11, characterized in that, The method further comprises: in response to adding a new reference object in the target scene, obtaining updated scene information, the updated scene information comprising position information of the new reference object; or in response to removing an existing reference object in the target scene, obtaining updated scene information, the updated scene information not comprising information of the removed existing reference object; or in response to moving an existing reference object in the target scene, obtaining updated scene information, the updated scene information comprising position information of the moved existing reference object; The determination of the target device according to the scene information and the voice instruction comprises: determining the target device according to the updated scene information and the voice instruction.
13. A method of controlling a device based on a voice instruction, the method comprising: comprises: obtaining a voice instruction; processing the voice instruction through a language reconstruction model, and determining one or more of the following: a reference object, an operation intention of a user, a device type of a target device, or a positional relationship between the target device and the reference object; in a case where the number of controllable devices corresponding to the device type is unique, determining the controllable device as the target device; or, in a case where the number of controllable devices corresponding to the device type is multiple, determining, according to scene information, a controllable device that satisfies the positional relationship from the multiple controllable devices as the target device, the scene information being used to indicate the positional relationship between the multiple controllable devices and the reference object; sending a control instruction corresponding to the operation intention to the target device.
14. The method of claim 13, wherein, The method further includes: in a case where the number of controllable devices that satisfy the positional relationship is multiple, obtaining capability information of the multiple controllable devices, the capability information being used to indicate performance of the controllable devices in executing the control instruction; determining, according to the capability information, a controllable device that satisfies a performance requirement from the multiple controllable devices as the target device.
15. The method according to claim 13 or 14, characterized in that, The scene information is determined according to one or more of the following: point cloud information of an object in the target scene, bounding box information of the object in the target scene, image information of the target scene, video information of the target scene, or text information of the target scene, wherein the object includes a movable object and a fixed object, and the movable object includes the user.
16. An electronic device, comprising: A processor and a memory, the memory being used to store program instructions, and the processor being used to invoke the program instructions to execute the method of any one of claims 1 to 12, or the method of any one of claims 13 to 15.
17. An apparatus for speech processing, the apparatus comprising: A module for implementing the method of any one of claims 1 to 12, or the method of any one of claims 13 to 15.
18. A computer-readable storage medium, characterized in that, A computer program stored thereon, the computer program being executed by a computer to implement the method of any one of claims 1 to 12, or the method of any one of claims 13 to 15.
Citation Information
Patent Citations
Intelligent household voice control method, intelligent device and apparatus with storage function
CN107528753A
Equipment control method, device and system based on position information
CN110989372A
Intelligent device control method, electronic device and system
CN113823280A
Voice interaction method, server and computer readable storage medium
CN115457959A
Voice instruction processing method and device and terminal equipment
CN115981168A