Robot Task Planning Method and Device Based on Body Movements and Gaze Tracking

Through a robot task planning method based on body movement and line of sight tracking, combined with user preference knowledge base and environmental perception information, the problem of existing smart home robots being unable to serve users who cannot clearly express language needs is achieved, and personalized services to more users are achieved.

CN118650635BActive Publication Date: 2025-06-24WESTLAKE INTERACTIVE ROBOT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411140079.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2025-06-24
Estimated Expiration
2044-08-20

AI Technical Summary

Technical Problem

Existing smart home robots are unable to effectively serve users who cannot clearly express language needs, such as the elderly, children and those with language disorders, resulting in their application limitations.

Method used

A robot task planning method based on limb movement and line of sight tracking is adopted. By building a user preference knowledge base, combining limb movement information, line of sight information and environment perception information, user needs are decomposed into sub-task sequences and executed.

Benefits of technology

It enables the analysis of user needs and perform tasks without explicit language expression. It is suitable for users who are difficult to express their needs clearly, enhancing the universality of robots in home applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118650635B_ABST
    Figure CN118650635B_ABST
Patent Text Reader

Abstract

This application proposes a robot task planning method and device based on limb movements and gaze tracking, including the following steps: Obtain user identification information, and obtain the user preference information of the current user from the user preference knowledge base based on the user identification information. Obtain the limb movement information and gaze information when the user issues a demand instruction, and perform intention understanding based on the limb movement information and gaze information to obtain intention understanding information; Obtain environmental perception information in real time, and obtain user demand tasks based on the intention understanding information, environmental perception information, and user preference information; Decompose and execute the user demand tasks. This solution analyzes the user's needs through body language and gaze tracking technology, and then combines the personalized content of the user to plan and execute tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a robot task planning method and device based on body movements and gaze tracking. Background Art

[0002] With the rapid development of artificial intelligence and robotics technologies, robots are increasingly widely used in various fields such as home, medical, and industrial. To enable robots to better adapt to complex and changing environments, intelligent task planning has become a research hotspot. Robot intelligent task planning refers to designing algorithms and processes for robots so that they can autonomously or semi-autonomously determine the steps required to complete tasks. Existing technologies generally give the robot clear task instructions input by users, and then the robot uses large language models or knowledge graphs to understand the task instructions and perform task planning before executing the corresponding tasks. Task planning based on large language models uses natural language processing (NLP) technology to decompose complex task instructions into a series of executable subtasks, and optimizes the execution process of these tasks through the understanding ability of the large language model. Correspondingly, this requires users to input a clear language description of the task requirements, which means that users need to describe their needs in a relatively rigorous language. This is not friendly to users who have difficulty expressing clearly or have lost their language ability. However, the target users of household intelligent robots are not entirely users who can clearly and rigorously express their needs in language. At this time, the user group may be: the elderly who cannot communicate in Mandarin, children who are temporarily unable to organize language logic, language-disabled people with speech difficulties, etc. Such people cannot clearly convey their needs to the robot, resulting in the current household robots being unable to serve such people; and often this group of people has more personalized needs than other service groups. Therefore, their personal preferences need to be considered during task planning. So this group of people is the most applicable scenario for household intelligent robots.

[0003] In summary, the current intelligent household robots that execute tasks only based on clear language task expressions cannot meet the actual household needs, restricting the application of household robots. Summary of the Invention

[0004] The embodiments of this application provide a robot task planning method and device based on body movements and gaze tracking, which analyze the user's needs through body language and gaze tracking technologies, and at the same time combine the user's preference information to perform task planning and execute the corresponding tasks, without the need for the user to clearly express the task requirements in language.

[0005] In a first aspect, the embodiments of this application provide a robot task planning method based on body movements and gaze tracking, the method includes:

[0006] Construct a user preference knowledge base based on the personal preference information of each user, where the personal preference information records the different preferences of different users for the same event and the different instruction information of different users when issuing requirements for the same event;

[0007] Obtain the body movement information and line-of-sight information when the user issues a requirement instruction, obtain the user preference information of the user who issues the requirement instruction in the user preference knowledge base, and judge the user requirements expressed by the body movement information and the line-of-sight information according to the user preference information;

[0008] Obtain the house layout information and house item information in real time, and decompose the user requirements into a sub-task sequence based on the house layout information, house item information, and user preference information, and sequentially execute each sub-task in the sub-task sequence.

[0009] In a second aspect, an embodiment of the present application provides a robot task planning device based on body movement and eye tracking, including:

[0010] A construction module that constructs a user preference knowledge base based on the personal preference information of each user, where the personal preference information records the different preferences of different users for the same event and the different instruction information of different users when issuing requirements for the same event;

[0011] A requirement judgment module, configured to obtain the body movement information and line-of-sight information when the user issues a requirement instruction, obtain the user preference information of the user who issues the requirement instruction in the user preference knowledge base, and judge the user requirements expressed by the body movement information and the line-of-sight information according to the user preference information;

[0012] An execution module, configured to obtain the house layout information and house item information in real time, and decompose the user requirements into a sub-task sequence based on the house layout information, house item information, and user preference information, and sequentially execute each sub-task in the sub-task sequence.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute a robot task planning method based on body movement and eye tracking.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, where a computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process, and the process includes a robot task planning method based on body movement and eye tracking.

[0015] The main contributions and innovations of the present invention are as follows:

[0016] The robot task planning method based on limb movement and gaze tracking proposed in the solution of this application accurately analyzes the user's needs by analyzing the user's preference information and combining the user's limb movement information and gaze information. Combining the information of these three modalities replaces the traditional rigorous language expression information, so as to provide various services for elderly and child users and those users with inconvenient mobility or difficulty in clear expression; this solution decomposes the user's needs into multiple subtasks to be executed sequentially, so as to better understand various information in the user's needs and achieve better execution results; this solution monitors the execution results of each subtask in real time, avoids the repeated execution of the same subtask, and updates the user preference knowledge base in real time based on the execution results of each subtask, which can provide a reference for the next task planning.

[0017] The details of one or more embodiments of this application are set forth in the following drawings and description, so that the other features, objects, and advantages of this application will become more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0019] Figure 1 is a flowchart of a robot task planning method based on limb movement and gaze tracking according to an embodiment of this application;

[0020] Figure 2 is a logic diagram of a robot task planning method based on limb movement and gaze tracking according to an embodiment of this application;

[0021] Figure 3 is a schematic diagram of obtaining limb movement information according to an embodiment of this application;

[0022] Figure 4 is a schematic diagram of the gaze information output by a human eye tracking module according to an embodiment of this application;

[0023] Figure 5 is a logic flowchart of obtaining user needs through limb movement information and gaze information according to an embodiment of this application;

[0024] Figure 6 is a structural block diagram of a robot task planning device based on limb movement and gaze tracking according to an embodiment of this application;

[0025] Figure 7 is a schematic hardware structure diagram of an electronic device according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0027] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or fewer than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0028] Embodiment 1

[0029] An embodiment of the present application provides a robot task planning method based on limb movements and gaze tracking, enabling the robot to analyze the user's needs through body language and gaze tracking technology, and then plan and execute tasks in combination with the user's personalized content. Specifically, referring to Figure 1 and Figure 2 , the method includes:

[0030] Obtain user identification information, and based on the user identification information, obtain the user preference information of the current user from the user preference knowledge base, where the user preference information at least records the historical behavior preferences of the current user for different demand instructions;

[0031] Obtain the limb movement information and gaze information when the user issues a demand instruction, and perform intention understanding based on the limb movement information and gaze information to obtain intention understanding information, where the intention understanding information records the demand instruction that the current user intends to complete;

[0032] Obtain environmental perception information in real time, where the environmental perception information records the scene state in the environmental scene where the robot is currently located;

[0033] Obtain the user demand task based on the intention understanding information, environmental perception information, and user preference information, decompose the user demand task, and execute it.

[0034] It should be noted that different from the traditional intelligent robots that need to execute tasks through explicit language instructions, the present solution aims to control the robot to execute the user's demand instructions based on the user's body movement information, sight information, and the user's personal preference information, so that the household robot can serve more users.

[0035] In the step of "obtaining user identification information", the user identification information is the information that uniquely identifies the user. Specifically, the user identification information can be one or any combination of the user's face information, user account information, user ID card information, and user number information.

[0036] In the embodiment of the present solution, the user identification information is set as the user's face information. At this time, a camera is installed on the robot of the present solution. The camera captures the user's face image, and the face recognition degree is used to identify the face image to obtain the user identification information that uniquely marks each user. Specifically, the user's face information is obtained by identifying the face image through the face recognition program. The face recognition degree can be DeepFace, which is a lightweight face recognition and facial attribute analysis framework.

[0037] Of course, in some other embodiments, the robot can obtain the user identification information in other ways. For example, the user can select his own user identification information by himself, or tell the robot his own user identification information, etc. The present solution does not make special restrictions on the way of obtaining the user identification information.

[0038] In the step of "obtaining the user's preference information of the current user from the user preference knowledge base based on the user identification information", the user preference knowledge base records the user preference information of different users. The user identification information is used to locate the current user, and the user preference information of the current user is retrieved from the user preference knowledge base.

[0039] For household robots, the user preference information of different family members is obtained and a user preference knowledge base is constructed, and the user preference knowledge base can be continuously updated during the use of the robot.

[0040] It should be noted that the user preference information proposed in the present solution records at least the historical behavior preferences of the current user for different demand instructions. The historical behavior preference is the historical behavior and preference of the current user for the current demand instruction, and the historical behavior preference is updated by the demand task that the current user actually requires the robot to execute when the same demand instruction appears at a historical moment; it can also be updated by the user's self-defined behavior that is expected to occur for each demand instruction.

[0041] For example, for a demand instruction of "pick up a water cup", the historical behavior preference can be "pick up the yellow water cup exclusive to the current user"; for a demand instruction of "turn on the air conditioner", the historical behavior preference can be "set to 26°C".

[0042] In some other embodiments, the user preference information records the personality difference information of the current user, where the personality difference information records the personality information of the current user for different demand instructions. At this time, the personality difference information is updated by the user's self-defined behavior required for each demand instruction.

[0043] Exemplarily, for an elderly person at home, when the demand instruction is "go to the toilet", the corresponding personality difference information of the current user is "fetch a walking stick and help the elderly person go to the toilet".

[0044] The method of obtaining user preference information in this solution can combine the subsequent intention understanding information to obtain a demand task that better meets the actual needs of the current user, and comprehensively considers the personalities of different users, aiming to provide personalized home robot service functions for users.

[0045] In the step of "obtaining the limb movement information when the user issues a demand instruction", a series of consecutive frames of images when the user issues a demand instruction are obtained, feature extraction is performed on the series of consecutive frames of images and position encoding is performed to obtain a position encoding result, and then the position encoding result is input into a pre-trained limb movement recognition encoder to obtain limb movement information.

[0046] In some embodiments, the limb movement recognition encoder is obtained by connecting multiple encoding modules in series, and each encoding module is sequentially composed of a multi-head attention module, a fully connected layer, and a normalization layer.

[0047] Specifically, the schematic diagram of obtaining limb movement information is as Figure 3 shown. As can be Figure 3 seen, the limb movement recognition encoder can accurately recognize the limb movements of the user's hand, so as to analyze the user's intention understanding information in combination with the line-of-sight information later.

[0048] In the step of "obtaining the line-of-sight information when the user issues a demand instruction", a series of consecutive frames of images when the user issues a demand instruction are obtained, feature extraction is performed on the series of consecutive frames of images and human eye detection is performed to obtain a human eye detection result, and then the human eye detection result is input into a pre-trained human eye tracking module to obtain line-of-sight information.

[0049] In some embodiments, the human eye tracking module is sequentially composed of a convolutional layer, a max pooling layer, and a fully connected layer. Specifically, the line-of-sight information output by the human eye tracking module is as Figure 4As shown, the line-of-sight information includes the pitch angle P and yaw angle Y of the human eye. Through the line-of-sight information, the direction of the human eye can be accurately obtained.

[0050] Preferably, the resolution of each frame of the continuously acquired multi-frame images in this solution is 1920×1080, and the refresh rate is 30 frames of images.

[0051] Preferably, this solution uses the convolutional neural network ResNet34 to extract features from the continuously acquired multi-frame images.

[0052] In the step of "performing intention understanding based on limb movement information and line-of-sight information to obtain intention understanding information", the pre-trained semantic analysis model is used to perform intention understanding on the limb movement information and line-of-sight information to obtain intention understanding information.

[0053] It should be noted that the input information of the semantic analysis model in this solution is limb movement information and line-of-sight information, and the output is the semantic understanding text representing the intention understanding information. Preferably, the pre-trained ChatGLM3 model is used as the semantic analysis model, and this semantic analysis model is trained to obtain intention understanding information.

[0054] The reason why this solution emphasizes using limb movement information and line-of-sight information for intention understanding is that often the user's limb movements and line of sight can represent the user's intention. In particular, for users who cannot use the language system smoothly or rigorously, limb movements and line of sight are the most crucial information for users to express their intentions.

[0055] Specifically, the logic flow chart for obtaining the user's needs through limb movement information and line-of-sight information is as Figure 5 As shown, the intention understanding information in this solution records the demand instructions for the current user's intention to complete. The demand instructions include the target objects that the current user may be interested in and the need for the robot's help. For example, if the user's line-of-sight information looks at a water cup and the limb movement information also points to the water cup, then at this time, it is understood that the user's demand instruction is "want to drink water and need the robot to help get the water cup".

[0056] In the step of "real-time obtaining environmental perception information", the color image and depth image in the environmental scene where the robot is currently located are obtained, and the pre-trained environmental perception model is used to analyze the color image and depth image in the environmental scene to obtain environmental perception information.

[0057] In this solution, the environmental perception information includes the house layout information and the house item information. The house layout information includes the structural information of each room in the house, and the house item information includes the placement positions of the items in each room and the state of each item.

[0058] Exemplarily, a camera equipped on the robot is used to read the color image and depth image of the house.

[0059] Specifically, the house layout information can be directly obtained from the color image and depth image. The CogVLM multimodal large model is used as the environmental perception model to analyze the color image and depth image to obtain environmental perception information. The environmental perception information is saved in the form of an item list. The item list records the placement position and status of each item. For example, it can be obtained from the item list that there is an apple on the table in the living room and its status is stationary, and there is a pot in the kitchen and its status is open.

[0060] In the step of "obtaining the user's needs task based on the intention understanding information, environmental perception information, and user preference information", the intention understanding information, environmental perception information, and user preference information are combined to generate a specific task. For example, when the user looks at and points to a water cup, the intention understanding information inferred by the intention understanding model is that the user wants to drink water. The user preference is to drink hot water and is used to holding the water cup with the left hand. The environmental perception information is the position of the water cup and hot water:

[0061] Generation of the user's needs task: According to the user's intention and preference, as well as the current environmental situation, a specific execution task is generated. For example, when the user looks at the water cup and wants to drink water, the robot needs to hand the water cup to the user.

[0062] Task refinement and optimization: Refine the task execution steps. For example, search for the water cup based on the environmental perception information. If not found, get the kettle from the kitchen, pour water into the water cup, and then hand it to the user's left hand position. The above content uses the ChatGLM model to output the user's needs task.

[0063] In the step of "decomposing and executing the user's needs task", the user's needs task is decomposed into a sequence of subtasks. The sequence of subtasks contains the operations to be performed to complete the user's needs task, and each subtask in the sequence of subtasks is executed in order.

[0064] Specifically, if the user's needs task is "clean the living room and kitchen", it is decomposed into subtasks such as "sweep the living room floor" and "clean the kitchen countertop", and the specific operation steps and requirements of each subtask are clarified.

[0065] Specifically, the pre-trained ChatGLM model is used for semantic analysis, extracting task information, and decomposing the user's needs into executable subtasks, aiming to better understand the language expression in the user's needs and the implicit requirements of the context.

[0066] That is to say, in this solution, the user identification information is first obtained through a camera to determine which user is using the robot, and then the user preference information of the user who issues the demand instruction is obtained from the user preference knowledge base. Finally, the user demand task is decomposed into a subtask sequence in combination with the real-time obtained environmental perception information, and each subtask in the subtask sequence is sequentially executed to complete the user demand.

[0067] In some specific embodiments, a disabled user sits in a wheelchair, gazes at a certain area in the living room, and makes instruction information that conforms to the user's preferences. The robot first obtains the user preference information from the user preference knowledge base through the user identification information, and then analyzes the disabled user's instruction information to confirm that the user points to the garbage and sundries on the scene, and the user's line of sight moves back and forth among different sundries, so as to confirm that the user's demand is to clean the garbage. Then, by obtaining the environmental perception information in real time, it is obtained from the house layout information and the house item information that there is a lot of garbage and sundries on the living room floor, and there is a trash can in the corner of the living room. The garbage needs to be thrown into the trash can and the sundries need to be arranged neatly; based on the above-obtained information, the user demand task is decomposed into subtask sequences such as "go to the garbage location", "pick up the garbage", "discard the garbage", "identify the sundry location", "sort out the sundries", etc., and each subtask is sequentially executed.

[0068] In this solution, a prior rule base is constructed, and the usage logic and basic common sense of each item are recorded in the prior rule base. Each subtask in the subtask sequence is sequentially executed based on the prior rule base.

[0069] For example, it is recorded in the prior rule base that for container-type household appliances such as pots, microwave ovens, refrigerators, etc., they should be turned on first and then turned off during operation. If the current state is already on, it cannot be set to on again. In addition, the basic common sense is that all foods are default in the refrigerator and the kitchen. When giving location information, navigate to that location first and then search for whether there is such an object. The robot has only two hands, and at most two grasping operations can be performed in a single subtask, etc.

[0070] In this solution, during the execution of each subtask, the execution status of each subtask is monitored in real time to record the execution result of each subtask.

[0071] Specifically, there is a subtask of "grab an apple from the table". If it is judged that the subtask is successfully executed, "YES" is output, and if it fails, "NO" is output. By judging the execution status of each subtask in real time, the repeated execution of subtasks can be avoided.

[0072] In this solution, the user preference knowledge base is updated in real time based on the execution result of each subtask for use as a reference during the next task planning.

[0073] In some specific embodiments, there is a disabled user who wants to drink water. The user looks at the water cup with their eyes and makes command information that conforms to the user's preferences. First, the robot obtains the user's preference information from the user preference knowledge base through the user identification information, and then analyzes the command information of the disabled person to confirm that the user wants to drink hot water with the water cup. Then, by obtaining the user's house layout information and house item information in real time, it is obtained from the house layout information and house item information that there is a water cup and a kettle on the living room table, and the hot water is in the kettle. Based on the above information, the obtained user needs are decomposed into subtasks such as "go near the living room table", "grab the kettle", "pour water into the water cup", and "hand the water cup to the user", and then these subtasks are executed based on the prior rule base.

[0074] Embodiment 2

[0075] Based on the same concept, referring to Figure 6 , this application also proposes a robot task planning device based on limb movements and gaze tracking, including:

[0076] A user identification module, configured to obtain user identification information, and obtain the user preference information of the current user from the user preference knowledge base based on the user identification information, where the user preference information at least records the historical behavior preferences of the current user for different demand commands;

[0077] An intention understanding module, configured to obtain the limb movement information and gaze information when the user issues a demand command, and perform intention understanding based on the limb movement information and gaze information to obtain intention understanding information, where the intention understanding information records the demand command that the current user intends to complete;

[0078] An environment perception module, configured to obtain environment perception information in real time, where the environment perception information records the scene state in the environment scene where the robot is currently located;

[0079] An execution module, which obtains the user demand task based on the intention understanding information, environment perception information, and user preference information, decomposes the user demand task, and executes it.

[0080] Embodiment 3

[0081] This embodiment also provides an electronic device, referring to Figure 7 , including a memory 404 and a processor 402. The memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0082] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0083] Among them, the memory 404 may include a mass memory 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0084] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0085] By reading and executing the computer program instructions stored in the memory 404, the processor 402 implements any one of the robot task planning methods based on limb movements and gaze tracking in the above embodiments.

[0086] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.

[0087] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0088] The input / output device 408 is used to input or output information. In this embodiment, the input information can be personal preference information, instruction information, etc., and the output information can be a subtask sequence, etc.

[0089] Optionally, in this embodiment, the above processor 402 can be set to execute the following steps through a computer program:

[0090] Obtain user identification information, and based on the user identification information, obtain the user preference information of the current user from the user preference knowledge base, where the user preference information at least records the historical behavior preferences of the current user for different demand instructions;

[0091] Obtain the limb movement information and gaze information when the user issues a demand instruction, and perform intention understanding based on the limb movement information and gaze information to obtain intention understanding information, where the intention understanding information records the demand instruction that the current user intends to complete;

[0092] Obtain environmental perception information in real time, where the environmental perception information records the scene state in the environmental scene where the robot is currently located;

[0093] Obtain the user demand task based on the intention understanding information, environmental perception information, and user preference information, decompose the user demand task and execute it.

[0094] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated here.

[0095] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, the blocks, devices, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers, or other computing devices, or some combination thereof.

[0096] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components that are configured to perform the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any block in the logical flow, as Figure 7 described, can represent a program step, or an interconnected logical circuit, block, and function, or a combination of a program step and a logical circuit, block, and function. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.

[0097] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0098] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A robot task planning method based on body movement and eye tracking, characterized in that: The following steps are involved: Acquire user identification information, and acquire user preference information of the current user from a user preference knowledge base based on the user identification information, wherein the user preference information at least records the historical behavior preference of the current user for different demand instructions, wherein the historical behavior preference is the historical behavior and preference of the current user for the current demand instruction; Obtain the body movement information and sight line information when the user issues a demand instruction, and perform intention understanding based on the body movement information and sight line information to obtain intention understanding information, wherein the intention understanding information records the demand instruction that the current user intends to complete, and the demand instruction includes the target object that the current user may be interested in and the help of the robot required. The body movement information and sight line information are used to understand the intention to obtain the intention understanding information through the semantic analysis model obtained by pre-training; Acquire environmental perception information in real time, wherein the environmental perception information records the scene state in the environment scene where the robot is currently located, the environmental perception information includes house layout information and house item information, the house layout information includes the structural information of each room in the house, and the house item information includes the location of items in each room and the state of each item; Obtain user demand tasks based on intention understanding information, environmental perception information and user preference information; Decompose the user demand task and execute it, and update the user preference knowledge base in real time based on the execution result of each subtask, decompose the user demand task into a subtask sequence, wherein the subtask sequence includes the operations to be performed to complete the user demand task, and execute each subtask in the subtask sequence sequentially, wherein a priori rule base is constructed, wherein the priori rule base records the usage logic and basic common sense of each item, and executes each subtask in the subtask sequence sequentially based on the priori rule base.

2. A robot task planning method based on body movement and line of sight tracking according to claim 1, characterized in that: In the step of "obtaining user preference information of the current user from the user preference knowledge base based on user identification information", the user preference knowledge base records user preference information of different users, the current user is located using the user identification information, and the user preference information of the current user is retrieved from the user preference knowledge base.

3. A robot task planning method based on body movement and eye tracking according to claim 1, characterized in that: In the step of "obtaining body movement information when the user issues a demand instruction", multiple consecutive frames of images are obtained when the user issues a demand instruction, features are extracted from the multiple consecutive frames of images, and position encoding is performed to obtain position encoding results, and then the position encoding results are input into the pre-trained body movement recognition encoder to obtain body movement information.

4. A robot task planning method based on body movement and eye tracking according to claim 1, characterized in that: In the step of "obtaining line of sight information when the user issues a demand instruction", a plurality of continuous frames of images are obtained when the user issues a demand instruction, and features are extracted from the plurality of continuous frames of images and human eye detection is performed to obtain human eye detection results, and then the human eye detection results are input into a pre-trained human eye tracking module to obtain line of sight information.

5. A robot task planning device based on body movement and eye tracking, characterized in that: include: A user identification module, used to obtain user identification information, and obtain user preference information of the current user from a user preference knowledge base based on the user identification information, wherein the user preference information at least records the historical behavior preference of the current user for different demand instructions, wherein the historical behavior preference is the historical behavior and preference of the current user for the current demand instruction; An intention understanding module is used to obtain the body movement information and sight line information when the user issues a demand instruction, and to understand the intention based on the body movement information and sight line information to obtain the intention understanding information, wherein the intention understanding information records the demand instruction that the current user intends to complete, and the demand instruction includes the target object that the current user may be interested in and the help required by the robot, and the intention understanding information is obtained by understanding the body movement information and sight line information through the semantic analysis model obtained by pre-training; The environment perception module is used to obtain environment perception information in real time, where the environment perception information records the scene state of the robot's current environment scene; An execution module obtains user demand tasks based on intention understanding information, environmental perception information and user preference information, decomposes user demand tasks and executes them, decomposes the user demand tasks into subtask sequences, wherein the subtask sequences include operations to be performed to complete the user demand tasks, and sequentially executes each subtask in the subtask sequence, wherein a priori rule library is constructed, wherein the priori rule library records the usage logic and basic common sense of each item, and sequentially executes each subtask in the subtask sequence based on the priori rule library.

6. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute a robot task planning method based on limb motion and line of sight tracking as described in any one of claims 1-4.

7. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a robot task planning method based on limb movement and line of sight tracking according to any one of claims 1-4.

Citation Information

Patent Citations

  • Determination method and apparatus of target object, storage medium and robot

    CN108491790A

  • Information processing method and device, electronic equipment and readable storage medium

    CN111444982A

  • Robot man-machine action interaction system and method based on human behavior recognition perception

    CN114131610A