A control method, device, and storage medium

By collecting and analyzing the fusion process of the user's directional posture action information and voice control instructions, the problem of device control inaccuracy caused by voice control instructions is solved, and a simpler and more accurate multi-device control is achieved.

CN114171019BActive Publication Date: 2025-08-29GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111340879.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-08-29
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

In multi-device intelligent interaction scenarios, when there is ambiguity in voice control instructions, it leads to the problem of cumbersome control steps and low accuracy.

Method used

The attitude action analysis module collects directional attitude action information, combines the voice command recognition module to recognize voice control instructions, and the decision module performs time alignment and recognition feature fusion to determine the target device and control parameter values.

Benefits of technology

Reduced equipment control steps, improving the accuracy and user experience of equipment control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114171019B_ABST
    Figure CN114171019B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a control method, device, and storage medium. The device includes: a posture motion analysis module, which is used to collect directional posture motion information and determine posture control instruction information based on the directional posture motion information; a voice instruction recognition module, which is used to recognize voice control instruction information; a decision module, which is used to time-align the posture control instruction information and the voice control instruction information to obtain a correspondence between the posture control instruction information and the voice control instruction information; based on the type of features to be recognized, respectively recognize the first posture control instruction information and the first voice control instruction information with a corresponding relationship to obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information; determine the target device and the device control parameter value for the target device based on the first recognition result data and the second recognition result data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet of Things, and in particular to a control method and device, and a storage medium. Background Art

[0002] With the continuous development, iteration and enrichment of smart Internet of Things (IOT) devices, the integration of all things has gradually become a smart interactive scenario for smart homes, smart offices, etc. Multi-device, multi-modality, interconnection and intelligence are the new features of current smart interactive scenarios. In existing smart interactive scenarios, the control of multiple devices is achieved through voice interaction. Specifically, control devices such as smart speakers receive the voice control instructions spoken by the user, and parse the voice control instructions to obtain the device that the user intends to control and the control parameter values ​​of the device. When the user's expression is unclear and there is ambiguity, the control device and the user need to eliminate the ambiguity through multiple rounds of dialogue or probabilistic selection methods, which will lead to cumbersome steps for device control and low accuracy of device control. Summary of the Invention

[0003] The embodiments of the present application provide a control method, apparatus, and storage medium, which can reduce the steps of device control and improve the accuracy of device control.

[0004] The technical solution of this application is achieved as follows:

[0005] In a first aspect, an embodiment of the present application provides a control device, comprising: a posture and motion analysis module, a voice command recognition module, and a decision module; wherein,

[0006] The posture motion analysis module is used to collect directional posture motion information and determine posture control instruction information based on the directional posture motion information;

[0007] The voice command recognition module is used to recognize voice control command information;

[0008] The decision module is used to time-align the posture control instruction information and the voice control instruction information to obtain the correspondence between the posture control instruction information and the voice control instruction information; based on the feature type to be identified, the first posture control instruction information and the first voice control instruction information with the correspondence are respectively identified to obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information; determine the target device and the device control parameter value for the target device based on the first recognition result data and the second recognition result data; and control the target device by using the device control parameter value.

[0009] In a second aspect, an embodiment of the present application provides a control method, which is applied to the above-mentioned control device, and the method includes:

[0010] Collecting directional posture motion information, and determining posture control instruction information based on the directional posture motion information;

[0011] Recognize voice control command information;

[0012] Performing time alignment on the posture control instruction information and the voice control instruction information to obtain a corresponding relationship between the posture control instruction information and the voice control instruction information;

[0013] Based on the type of the feature to be identified, the first posture control instruction information and the first voice control instruction information having a corresponding relationship are respectively identified to obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information;

[0014] A target device and a device control parameter value for the target device are determined according to the first recognition result data and the second recognition result data; and the target device is controlled by using the device control parameter value.

[0015] In a third aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, and when the computer program is executed by a processor, the control method as described above is implemented.

[0016] An embodiment of the present application provides a control method, device, and storage medium, wherein the device includes: a posture motion analysis module, a voice command recognition module, and a decision module; the posture motion analysis module is used to collect directional posture motion information and determine posture control command information based on the directional posture motion information; the voice command recognition module is used to recognize voice control command information; the decision module is used to time-align the posture control command information and the voice control command information to obtain a corresponding relationship between the posture control command information and the voice control command information; based on the feature type to be recognized, the first posture control command information and the first voice control command information with a corresponding relationship are respectively recognized to obtain first recognition result data corresponding to the first posture control command information and second recognition result data corresponding to the first voice control command information; the target device and the device control parameter value for the target device are determined based on the first recognition result data and the second recognition result data; the target device is controlled by using the device control parameter value. Using the above-mentioned device implementation scheme, the control device also includes a posture action analysis module, which can collect posture action information; when there is ambiguity in the voice control command information, the decision module can directly determine the target device and the control parameter value of the target device in combination with the posture control command information corresponding to the posture action information during the determination process, which can reduce the steps of device control and improve the accuracy of device control. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic structural diagram of a control device provided in an embodiment of the present application;

[0018] Figure 2 An exemplary deployment method of a voice command recognition module provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of the structure of an exemplary voice command recognition module provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of a posture and motion analysis module provided in an embodiment of the present application;

[0021] Figure 5 A schematic diagram of an exemplary arrangement queue of decision-making devices provided in an embodiment of the present application;

[0022] Figure 6 A flow chart of a control method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0025] In the following description, reference is made to "some embodiments," which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. It should also be noted that the terms "first, second, and third" in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequential order where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0026] Currently, the following deficiencies exist in the process of controlling multiple executable devices based on voice signals:

[0027] 1. When there are multiple executable devices in a scene at the same time, for example, when a user asks to "turn on the air conditioner" but the scene includes multiple air conditioners, it is impossible to determine which air conditioner to turn on. At this time, there is a possibility of decision-making errors, resulting in the executable device decided to be different from the executable device the user intended, thus affecting the user experience.

[0028] 2. If the user does not clearly express the executable device to be controlled, such as "adjust to 25 degrees", the voice signal expressed is insufficient to determine the device the user wants to control. In this case, the decision fails, which in turn affects the performance of the controlled device.

[0029] In order to solve the above problems, the present application provides a control device, such as Figure 1 As shown, the device 1 may include: a gesture analysis module 10, a voice command recognition module 11 and a decision module 12; wherein,

[0030] The posture motion analysis module 10 is used to collect directional posture motion information and determine posture control instruction information based on the directional posture motion information;

[0031] The voice command recognition module 11 is used to recognize voice control command information;

[0032] The decision module 12 is used to time-align the posture control instruction information and the voice control instruction information to obtain the corresponding relationship between the posture control instruction information and the voice control instruction information; based on the feature type to be identified, the first posture control instruction information and the first voice control instruction information with the corresponding relationship are respectively identified to obtain the first recognition result data corresponding to the first posture control instruction information and the second recognition result data corresponding to the first voice control instruction information; determine the target device and the device control parameter value for the target device according to the first recognition result data and the second recognition result data; and control the target device by using the device control parameter value.

[0033] A control device proposed in an embodiment of the present application is suitable for scenarios where posture information and voice information are used to jointly control an execution device to perform a corresponding function.

[0034] In the embodiments of the present application, the execution device can be a smart home device such as a smart desk lamp, a smart speaker, a smart home appliance such as a smart air conditioner, a smart refrigerator, a smart TV, etc., and can also be an executable smart terminal such as a smart phone, a tablet computer, a PDA, a mobile station (Mobile Station, MS), a mobile terminal (Mobile Terminal), etc. The specific one can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations.

[0035] In an embodiment of the present application, the posture action analysis module in the control device is responsible for collecting the user's directional posture action information and determining the posture control instruction information based on the directional posture action information, and then sending the posture control instruction information to the decision module; the voice instruction recognition module in the control device is responsible for collecting the user's voice signal and identifying it. When the information type corresponding to the voice signal is identified as voice control, the information corresponding to the voice signal is determined as voice control instruction information, and then the voice control instruction information is sent to the decision module; the decision module jointly determines the target device to be executed and the device control parameter value for the target device based on the voice control instruction information and the posture control instruction information, and the decision module performs corresponding control operations on the target device based on the device control parameter value.

[0036] It should be noted that in the embodiment of the present application, the posture action analysis module and the voice command recognition module operate independently, and can be synchronized or asynchronous in timing. The decision module is responsible for aligning, mapping and fusing the voice control command information and the posture control command information to obtain the user's true intention and then decide on the execution device that needs to be controlled.

[0037] In an embodiment of the present application, a voice command recognition module is located on the audio processing unit of each smart terminal device with a microphone, such as a smart speaker, a smart TV, a smart air conditioner, etc. The number of voice command recognition modules can be one or more, wherein multiple voice command recognition modules are distributed and deployed in the same space, and each voice command recognition module operates independently of each other. Since each voice command recognition module operates independently of each other, multiple voice command recognition modules may be awake at the same time and output the same voice control command information, which will facilitate users to control devices remotely.

[0038] For example, Figure 2 As shown, a smart TV, smart speakers, and smart air conditioner are deployed in the living room, and a smart speaker is deployed in the bedroom. At this time, the user sends a voice signal in the bedroom to control the smart terminal devices in the living room.

[0039] In the embodiments of this application, Figure 3 As shown, each voice command recognition module can include a wake-up detection submodule, a voice recognition submodule, and a semantic recognition submodule. The wake-up detection device matches the received voice information with a pre-set wake-up word upon receiving it. If a match is successful, it outputs a wake-up activation signal and wakes up the voice command recognition module. The wake-up detection device may or may not include voiceprint verification. If voiceprint verification is included, the wake-up detection device's execution logic is modified to output a wake-up activation signal upon detecting the correct wake-up word and a valid voiceprint. The voice recognition device then continues operating after receiving the wake-up activation signal, converting the received voice signal into text data. After receiving the last segment of voice data, it waits for a threshold period of time. If no new voice signal is received, the voice recognition device ceases operation and returns to sleep. The semantic recognition device classifies the text data and determines whether the information type corresponding to the text data is voice control. If the text data corresponds to voice control, it outputs the text data to the decision module.

[0040] Optional, such as Figure 4 As shown, the posture and motion analysis module 10 includes: a posture and motion acquisition submodule 100 and a posture and motion detection submodule 101; wherein,

[0041] The posture and motion collection submodule 100 is used to collect the directional posture and motion information of the target object;

[0042] The gesture action detection submodule 101 is configured to filter out first directional gesture action information that satisfies a preset device control gesture from the directional gesture action information, and use the first directional gesture action information as the gesture control instruction information.

[0043] In the embodiment of the present application, the posture and motion acquisition submodule can be an ordinary camera, a depth camera, an inertial measurement unit (IMU) on a wearable device, or other device capable of collecting posture and motion information. The specific one can be selected according to actual conditions, and the embodiment of the present application does not make any specific limitations.

[0044] In the embodiment of the present application, the directional gesture action information of the target object may include the facing gesture action information of the target object and / or the pointing gesture action information of the target object. The specific one can be selected according to actual conditions, and the embodiment of the present application does not make any specific limitation.

[0045] It should be noted that ordinary cameras and depth cameras can effectively detect facing posture motion information and pointing posture motion information, while IMU can detect pointing posture motion information but cannot detect facing posture motion information.

[0046] It should be noted that when a user is active in a scene with distributed gesture and motion collection devices, not all directional gesture and motion information generated at all times is used to control the device. This information may include a large amount of invalid gesture and motion information for a long period of time. Therefore, after the gesture and motion collection device collects directional gesture and motion information corresponding to the target object, the gesture and motion detection device needs to filter out the first directional gesture and motion information from the directional gesture and motion information that satisfies the preset device control gesture. The preset device control gesture is gesture information that represents the control of the executing device.

[0047] It should be noted that the posture motion detection device can use a lightweight posture detection deep learning network in the process of screening the first directional posture motion information that meets the preset device control posture. The embodiments of this application do not limit the specific network model type. For example, the screening process of the first directional posture motion information based on the IMU can be performed using a correlation filter (CF) algorithm, or other fast signal detection algorithms.

[0048] It should be noted that the ordinary camera collects RGB images, and the corresponding first directional posture motion information is obtained by inputting the RGB image into the detection network; the depth camera collects RGDB images, and the corresponding first directional posture motion information is obtained by inputting the RGDB image into the detection network; the IMU sensor collects angular rate and acceleration parameters, and the corresponding first directional posture motion information is obtained by inputting the angular rate and acceleration parameters into the digital signal detection algorithm.

[0049] Optionally, the posture and motion acquisition submodule includes at least one group of posture and motion acquisition devices, each group of posture and motion acquisition devices in the at least one group of posture and motion acquisition devices has the same device type, and each group of posture and motion acquisition devices corresponds to a posture and motion detection submodule.

[0050] In an embodiment of the present application, directional gesture and motion information of a target object can be collected through at least one group of gesture and motion acquisition devices, wherein each group of gesture and motion acquisition devices may include one gesture and motion acquisition device, or multiple gesture and motion acquisition devices. The specific number of gesture and motion acquisition devices in each group is not specifically limited, and it is only necessary to ensure that the device type of each group of gesture and motion acquisition devices is consistent.

[0051] For example, if there is a depth camera, two ordinary cameras and a smart watch in total, the two ordinary cameras together constitute the posture and motion acquisition device 1, the depth camera constitutes the posture and motion acquisition device 2, and the smart watch constitutes the posture and motion acquisition device 3.

[0052] Optionally, if a group of posture motion acquisition devices corresponding to the first posture motion detection submodule is multiple posture motion acquisition devices, then the first posture motion detection submodule is also used to respectively determine the angle difference between each two first directional posture motion information in the multiple first directional posture motion information to obtain multiple posture angle differences; based on the multiple posture angle differences, determine abnormal directional posture motion information from the multiple first directional posture motion information; and delete the abnormal directional posture motion information from the multiple first directional posture motion information to obtain second directional posture motion information; and use the second directional posture motion information as the posture control instruction information.

[0053] It should be noted that if the two directional gesture action acquisition devices in the same group capture the user facing two different directions, then abnormal directional gesture action information will be present in the two first directional gesture action information determined by the corresponding first gesture action detection submodule. Based on this, for a group of gesture action detection devices consisting of multiple gesture action acquisition devices, a first gesture action detection submodule associated with the multiple gesture action acquisition devices is not only used to filter out a first directional gesture action information from the directional gesture action information collected by each gesture action acquisition device, but is also used to delete the abnormal directional gesture action information from the multiple first directional gesture action information to obtain second directional gesture action information, and then use the second directional gesture action information as gesture control instruction information.

[0054] Exemplarily, a group of gesture detection devices includes i gesture acquisition devices, wherein the sum of the directional gesture information collected by the i gesture acquisition devices at time t is as shown in formula (1),

[0055]

[0056] in, is the set of directional gesture information collected by the i gesture collection devices at time t, X i (t) is the first directional gesture information collected by the i-th gesture collection device at time t.

[0057] Afterwards, from Find any two first directional gesture action information as Subset of As shown in formula (2),

[0058]

[0059] Afterwards, The angle difference between the two first directional gesture action information is calculated, and the first directional gesture action information with an angle difference less than or equal to 15 degrees is determined as the second directional gesture action information and added to In the set, the first directional gesture action information with an angle difference greater than 15 degrees is abnormal directional gesture action information, and then The second directional gesture action information in the set is sent to the subsequent module for processing. At this point, the process of determining abnormal directional gesture action information from multiple first directional gesture action information is completed. Specifically, as shown in formula (3),

[0060]

[0061] Among them, x1, x2 are , θ(x1, x2) is the angle difference between the two first directional gesture action information. It should be noted that the angle difference between the two first directional gesture action information can be the difference between the two pointing gesture action information or the difference between the two facing gesture action information.

[0062] It should be noted that 15 degrees is only an exemplary angle difference threshold, and the specific one can be selected according to actual conditions. The embodiments of this application do not make any specific limitations.

[0063] Optionally, if the gesture action detection submodule 101 includes multiple gesture action detection submodules, refer to Figure 4 , the posture and action analysis module 10 further includes: a posture and action alignment submodule 102;

[0064] The posture action alignment submodule 102 is also used to obtain multiple directional posture action information from the multiple posture action detection submodules; and based on a preset time threshold, time align the multiple directional posture action information to obtain the posture control instruction information; the multiple directional posture action information includes at least one of the first directional posture action information and the second directional posture action information.

[0065] It should be noted that the posture motion alignment submodule is responsible for aligning the posture motion information collected by different types of posture motion acquisition devices in terms of time and magnitude. Generally speaking, the processing speed of the posture motion detection submodule based on the IMU sensor of the wearable device is much faster than the processing speed of the posture motion detection submodule based on the camera and / or depth camera. Therefore, the posture motion alignment submodule will wait for receiving the directional posture motion information from the posture motion detection submodule based on the camera and / or depth camera within a preset time threshold after receiving the directional posture motion information from the posture motion detection submodule based on the IMU sensor. If the directional posture motion information from the posture motion detection submodule based on the camera and / or depth camera is received within the preset time threshold, the directional posture motion data within the entire time period will be recorded as a segment of posture motion information, and the segment of directional posture motion information will be determined as posture control instruction information and transmitted to the decision module.

[0066] It should be noted that the specific value of the preset time threshold can be adjusted based on actual conditions to ensure that a section of directional gesture action information only contains directional gesture action information used to control the device once. The specific value can be selected according to actual conditions, and the embodiments of this application do not make specific limitations.

[0067] In the embodiment of the present application, correspondingly, the posture action alignment submodule defines a section of directional posture action information record as follows: the time node when the posture action alignment submodule receives directional posture action information from any posture action detection submodule is the starting node. Within the preset time threshold, if the directional posture action information transmitted by other posture action detection submodules is received, a preset time threshold will be monitored again until no directional posture action information transmitted by other posture action detection submodules is received within the preset time threshold. Starting from the starting node, all the directional posture action information received will be determined as a section of directional posture action information.

[0068] Optionally, the control device includes at least one decision-making device, and the decision module is a decision-making device whose device status is idle and has the highest device performance among the at least one decision-making device.

[0069] In the embodiments of the present application, common decision-making devices include distributed intelligent terminal devices with decision-making processing capabilities, such as smart phones, tablet computers, smart TVs, and smart speakers. The specific ones can be selected according to actual conditions, and the embodiments of the present application do not make specific limitations.

[0070] In the embodiment of the present application, during the initialization process of the decision node, multiple decision devices will participate in building a device priority queue, such as Figure 5 As shown, multiple decision devices are constructed into a binary tree queue according to priority, wherein the priority can be determined according to the device performance and device status, such as advancing the priority of the decision device whose device status is idle, and sorting them in order from high to low according to the device performance, and then sorting the decision devices whose device status is occupied, closed or down in order. It should be noted that the sorting of the decision devices in the binary tree queue is updated in real time, ensuring that the decision module determined from the binary tree queue is always available. When the decision module is needed to perform the fusion between the posture control instruction information and the voice control instruction information and determine the target device and the device control parameter value for the target device, the decision device with the highest current priority is found from the binary tree queue.

[0071] Optionally, the decision module includes: a speech analysis submodule, a posture analysis submodule, an information alignment submodule and an execution decision submodule;

[0072] The information alignment submodule is used to perform time alignment on the posture control instruction information and the voice control instruction information to obtain a corresponding relationship between the posture control instruction information and the voice control instruction information;

[0073] The posture analysis submodule is configured to perform posture recognition on the first posture control instruction information based on the feature type to be recognized, and obtain the first recognition result data including at least the target device;

[0074] The voice analysis submodule is configured to perform semantic recognition on the first voice control instruction information corresponding to the first posture control instruction information based on the feature type to be recognized, to obtain the second recognition result data including at least the device control type;

[0075] The execution decision submodule is used to fuse the first recognition result data and the second recognition result data to obtain the target device and the device control parameter value for the target device; and use the control parameter value to control the target device.

[0076] In an embodiment of the present application, after the decision module receives the posture control instruction information and the voice control instruction information, the information alignment submodule is mainly responsible for the time alignment operation of the posture control instruction information and the voice control instruction information. For the case where only one posture control instruction information and one voice control instruction information are received in a short period of time, it is determined to be a one-to-one correspondence. For the case where multiple consecutive posture control instruction information and multiple voice control instruction information are received, a preset time alignment algorithm, such as the dynamic time warping (DTW) algorithm, can be used for time alignment.

[0077] It should be noted that the corresponding relationship is any one of: one gesture control instruction information corresponds to multiple voice control instruction information, one gesture control instruction information corresponds to one voice control instruction information, and multiple gesture control instruction information corresponds to one voice control instruction information.

[0078] In actual application, a scenario where one gesture control command message corresponds to multiple voice control command messages can be: pointing to the TV and outputting "start and tune to channel 32" by voice; a scenario where multiple gesture control command messages correspond to one voice control command message can be: pointing to the refrigerator and air conditioner respectively and outputting "adjust the temperature to 8 degrees" by voice.

[0079] In the embodiment of the present application, the speech analysis submodule first performs a text hash on the text information corresponding to the first voice control instruction information and removes the executed instructions. Then, the text information from multiple voice instruction recognition modules is aligned and merged. After alignment, only one voice control instruction information is retained for the same text information. Then, a long short-term memory network (LSTM) model is used to perform natural language understanding on the text information and identify the second recognition result data corresponding to the feature type to be identified, where the feature type to be identified includes three types: execution device, execution action, and execution parameter. The second recognition result data of the first voice control instruction information must include the result data corresponding to the execution action.

[0080] For example, for the command "Turn up the volume on the speaker," the corresponding second recognition result data is: "Execution device: speaker; Execution action: adjust volume; Execution parameter: turn up." Another example is the recognition slot result for the command "Turn this off" is "Execution device: none; Execution action: turn off; Execution parameter: none." Since the execution device is unknown, further fusion decision-making is required based on the gesture control command information.

[0081] In an embodiment of the present application, the decision module performs posture recognition on the first posture control instruction information based on the feature type to be identified, and obtains first recognition result data including at least the target device. At this time, the first recognition result data and the second recognition result data can be combined to determine the target device and the device control parameter value for the target device.

[0082] For example, if there's a smart refrigerator in the kitchen and a smart refrigerator in the living room, and the user points to the kitchen smart refrigerator and says, "Raise the temperature by one degree," the decision module determines, based on the gesture control command information, that the executing device is the kitchen smart refrigerator, and, based on the voice control command information, that the executing action is to adjust the temperature, with the execution parameter being to raise it by one degree. For another example, if there's a smart refrigerator in the kitchen and a smart refrigerator in the living room, and the user points to the kitchen smart refrigerator, raises their hand, and says, "Adjust the temperature," the decision module determines, based on the gesture control command information, that the executing device is the kitchen smart refrigerator, that the execution parameter is to raise it, and, based on the voice control command information, that the executing action is to adjust the temperature.

[0083] It can be understood that the control device also includes a posture action analysis module that can collect posture action information; when there is ambiguity in the voice control command information, the decision module can directly determine the target device and the control parameter value of the target device in combination with the posture control command information corresponding to the posture action information during the determination process, which can reduce the steps of device control and improve the accuracy of device control.

[0084] Based on the above embodiments, the present application also provides a device control method. Figure 6 As shown, the method includes:

[0085] S101: Collect directional posture motion information, and determine posture control instruction information according to the directional posture motion information.

[0086] In the embodiment of the present application, the specific process of collecting directional posture and action information can be found in the description of the posture and action collection submodule, which will not be repeated here.

[0087] In the embodiment of the present application, first directional gesture action information that satisfies a preset device control gesture is selected from the directional gesture action information; and the first directional gesture action information is used as gesture control instruction information. For details, please refer to the description of the gesture action detection submodule, which will not be repeated here.

[0088] In an embodiment of the present application, if the first directional gesture action information is a plurality of first directional gesture action information corresponding to a plurality of gesture action acquisition devices belonging to a group of gesture action acquisition devices in the gesture action analysis module, then after screening out the first directional gesture action information that satisfies the preset device control gesture from the directional gesture action information, the angle difference between each two first directional gesture action information in the plurality of first directional gesture action information is determined respectively to obtain a plurality of gesture angle differences; based on the plurality of gesture angle differences, abnormal directional gesture action information is determined from the plurality of first directional gesture action information; and the abnormal directional gesture action information is deleted from the plurality of first directional gesture action information to obtain the second directional gesture action information; and the second directional gesture action information is used as the gesture control instruction information. For details, refer to the description of the first gesture action detection submodule, which will not be repeated here.

[0089] In an embodiment of the present application, if the directional gesture action information includes multiple directional gesture action information including first directional gesture action information and / or second directional gesture action information, the multiple directional gesture action information are time-aligned based on a preset time threshold to obtain gesture control instruction information. For details, see the description of the gesture action alignment submodule, which will not be repeated here.

[0090] S102: Recognize voice control instruction information.

[0091] It should be noted that the specific recognition process of the voice control command information can be found in the description of the voice command recognition module, which will not be repeated here.

[0092] S103: Time-align the posture control instruction information and the voice control instruction information to obtain a corresponding relationship between the posture control instruction information and the voice control instruction information.

[0093] It should be noted that the specific process of determining the correspondence between the posture control instruction information and the voice control instruction information can be found in the description of the information alignment submodule, which will not be repeated here.

[0094] S104: Based on the feature type to be identified, identify the first posture control instruction information and the first voice control instruction information that have a corresponding relationship, and obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information.

[0095] It should be noted that, specifically, the process of obtaining the first recognition result data corresponding to the first posture control instruction information and the second recognition result data corresponding to the first voice control instruction information can be found in the description of the posture analysis submodule and the voice analysis submodule, which will not be repeated here.

[0096] S105 , determining a target device and a device control parameter value for the target device according to the first recognition result data and the second recognition result data; and controlling the target device using the device control parameter value.

[0097] It should be noted that the specific process of determining the target device and the device control parameter value for the target device can be found in the description of the execution decision submodule, which will not be repeated here.

[0098] It is understandable that this application provides a natural, seamless, and smooth solution for controlling multiple devices by integrating gesture action information with voice control command information. It uses cameras, depth cameras, and IMU sensors of smart wearable devices to identify gesture control command information such as pointing and facing when the user issues voice control command information. Combined with the natural language understanding of the voice control command information, it accurately perceives the true intention that the user wants to express. The gesture action acquisition submodule used in this application can arrange different gesture action recognition schemes according to scene restrictions, cost, space and other factors. The distributed voice command recognition module adopted meets the current multi-device, multi-scene smart home and smart office needs. The interaction method of this technical solution is a natural and seamless interaction method that conforms to human intuition. It uses pointing gesture action information and facing gesture action information to enable voice control of multiple devices, with almost zero learning cost for the user end. This solution has high accuracy, and facing gesture action information and pointing gesture action information are strong selection intentions with almost no possibility of error. It is used to eliminate the recognition ambiguity of pure voice signals and can significantly improve the user experience.

[0099] An embodiment of the present application provides a storage medium on which a computer program is stored. The computer-readable storage medium stores one or more programs. The one or more programs can be executed by one or more processors and applied to a control device. The computer program implements the control method as described above.

[0100] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0101] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling an image display device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0102] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.

Claims

1. A control device, characterized in that: The device includes: a posture and action analysis module, a voice command recognition module and a decision module; wherein, The posture motion analysis module is used to collect directional posture motion information and determine posture control instruction information based on the directional posture motion information; The voice command recognition module is used to recognize voice control command information; The decision module is configured to perform time alignment on the posture control instruction information and the voice control instruction information to obtain a correspondence between the posture control instruction information and the voice control instruction information; based on the type of features to be identified, respectively identify first posture control instruction information and first voice control instruction information having a correspondence to obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information; determine a target device and a device control parameter value for the target device based on the first recognition result data and the second recognition result data; and control the target device using the device control parameter value; The gesture action analysis module includes: a gesture action collection submodule and a gesture action detection submodule; wherein the gesture action collection submodule is used to collect directional gesture action information of the target object; the gesture action detection submodule is used to filter out first directional gesture action information that meets the preset device control gesture from the directional gesture action information, and use the first directional gesture action information as the gesture control instruction information; The gesture action acquisition submodule includes at least one group of gesture action acquisition devices, wherein each group of gesture action acquisition devices in the at least one group of gesture action acquisition devices has the same device type, and each group of gesture action acquisition devices corresponds to one gesture action detection submodule; If a group of posture action acquisition devices corresponding to the first posture action detection submodule is multiple posture action acquisition devices, then the first posture action detection submodule is also used to respectively determine the angle difference between each two first directional posture action information in the multiple first directional posture action information to obtain multiple posture angle differences; based on the multiple posture angle differences, determine abnormal directional posture action information from the multiple first directional posture action information; and delete the abnormal directional posture action information from the multiple first directional posture action information to obtain second directional posture action information; and use the second directional posture action information as the posture control instruction information.

2. The control device according to claim 1, characterized in that If the gesture action detection submodule includes a plurality of gesture action detection submodules, the gesture action analysis module further includes: a gesture action alignment submodule; The posture action alignment submodule is also used to obtain multiple directional posture action information from the multiple posture action detection submodules; and based on a preset time threshold, time align the multiple directional posture action information to obtain the posture control instruction information; the multiple directional posture action information includes at least one of the first directional posture action information and the second directional posture action information.

3. The control device according to claim 1, characterized in that The control device includes at least one decision-making device, and the decision module is a decision-making device with an idle state and the highest device performance among the at least one decision-making device.

4. The control device according to claim 3, characterized in that The decision module includes: a speech analysis submodule, a posture analysis submodule, an information alignment submodule and an execution decision submodule; The information alignment submodule is used to perform time alignment on the posture control instruction information and the voice control instruction information to obtain a corresponding relationship between the posture control instruction information and the voice control instruction information; The posture analysis submodule is configured to perform posture recognition on the first posture control instruction information based on the feature type to be recognized, and obtain the first recognition result data including at least the target device; The voice analysis submodule is configured to perform semantic recognition on the first voice control instruction information corresponding to the first posture control instruction information based on the feature type to be recognized, to obtain the second recognition result data including at least the device control type; The execution decision submodule is used to fuse the first recognition result data and the second recognition result data to obtain the target device and the device control parameter value for the target device; and use the control parameter value to control the target device.

5. The control device according to claim 1 or 4, characterized in that: The corresponding relationship is any one of: one gesture control instruction information corresponds to multiple voice control instruction information, one gesture control instruction information corresponds to one voice control instruction information, and multiple gesture control instruction information corresponds to one voice control instruction information.

6. A control method, characterized in that: Applied to the control device according to any one of claims 1 to 5, the method comprises: Collecting directional posture motion information, and determining posture control instruction information based on the directional posture motion information; Recognize voice control command information; Performing time alignment on the posture control instruction information and the voice control instruction information to obtain a corresponding relationship between the posture control instruction information and the voice control instruction information; Based on the type of the feature to be identified, the first posture control instruction information and the first voice control instruction information having a corresponding relationship are respectively identified to obtain first recognition result data corresponding to the first posture control instruction information and second recognition result data corresponding to the first voice control instruction information; determining a target device and a device control parameter value for the target device according to the first recognition result data and the second recognition result data; and controlling the target device using the device control parameter value; Wherein, determining the posture control instruction information according to the directional posture action information includes: Filtering out first directional gesture action information that satisfies a preset device control gesture from the directional gesture action information; and using the first directional gesture action information as the gesture control instruction information; The first directional gesture action information is a plurality of first directional gesture action information corresponding to a plurality of gesture action acquisition devices belonging to a group of gesture action acquisition devices in the gesture action analysis module. After screening out the first directional gesture action information that satisfies the preset device control gesture from the directional gesture action information, the method further includes: Determine the angle difference between each two first directional posture action information in the multiple first directional posture action information respectively to obtain multiple posture angle differences; determine abnormal directional posture action information from the multiple first directional posture action information based on the multiple posture angle differences; delete the abnormal directional posture action information from the multiple first directional posture action information to obtain second directional posture action information; and use the second directional posture action information as the posture control instruction information.

7. The control method according to claim 6, characterized in that: The directional gesture action information is a plurality of directional gesture action information including the first directional gesture action information and / or the second directional gesture action information, and determining the gesture control instruction information according to the directional gesture action information includes: Based on a preset time threshold, the plurality of directional gesture action information are time-aligned to obtain the gesture control instruction information.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the control method according to claim 6 or 7 is implemented.

Citation Information

Patent Citations

  • Voice instruction processing method, device and control system

    CN112053683A