Robot control method and system

By acquiring multimodal environmental information in real time and combining it with historical data analysis, the robot can efficiently identify and plan tasks in complex environments, solving the problem of traditional robots failing to execute tasks in changing environments and improving user experience and task completion quality.

CN120985633APending Publication Date: 2025-11-21RUERMAN INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510980965.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional robot control methods lack the ability to fuse and process multimodal information when facing complex and ever-changing environments, making it difficult to acquire and analyze environmental data in real time, leading to task execution failures and a degraded user experience.

Method used

By acquiring multimodal environmental information in real time, including images, sounds, touch, and environmental conditions, and combining it with historical data for in-depth analysis and associative memory, the system can identify the current task and plan execution strategies, and select the best strategy using breadth-first or depth-first search.

Benefits of technology

It enables robots to efficiently and accurately identify and execute tasks in complex environments, improving user experience and task completion quality, and enhancing the flexibility of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120985633A_ABST
    Figure CN120985633A_ABST
Patent Text Reader

Abstract

The invention discloses a robot control method and system. The method comprises the steps that multi-mode environment information in a detection range is obtained in real time; analyzing the multi-modal environment information, and retrieving associated memory data according to the multi-modal environment information; identifying the current task according to the analysis result of the multi-modal environment information and the retrieved associated memory data; planning an execution strategy according to the current task; wherein the multi-mode environment information comprises environment scene information and / or environment condition information, the information type of the environment scene information comprises part or all of images, sounds and touch, and the information type of the environment condition information comprises part or all of environment temperature, environment humidity, environment brightness and environment air. The robot actively acquires multi-modal data including scene and condition information in real time, performs fusion analysis and retrieval of associated memories, accurately recognizes tasks and plans strategies, and overcomes the defects that a traditional robot depends on presetting, is one-sided in perception, is difficult to strain and understand user intentions and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot control, and particularly relates to a robot control method and system. BACKGROUND

[0002] In today's era of rapid technological development, robot technology is increasingly becoming the focus of attention in various fields. With the continuous expansion of robot application scenarios, from industrial production lines to daily life services, to complex and dangerous rescue, exploration and other special tasks, people have put forward higher and higher requirements for the intelligent level of robots.

[0003] Traditional robot control methods often rely on preset programs, and are not capable of dealing with complex and variable real-world environments. For example, early robots could only perform single tasks in fixed structured environments according to pre-written precise instructions. Once the work scene changed slightly, such as the position of the placed items deviated slightly or the light conditions changed, the task execution might fail. These robots lack comprehensive and in-depth perception of real-time environmental information, can only process simple and limited sensor data, and are difficult to adapt to dynamic environments.

[0004] In the field of service robots, similar difficulties are also encountered. Taking a home service robot as an example, when it needs to deliver items to the user in the living room, it may not be able to accurately identify the target items and the path in the case of dark light or partial obstruction, and it is even more difficult to flexibly adjust the service strategy according to the immediate needs of family members, if it only relies on single visual image information. Moreover, different environmental conditions, such as changes in temperature and humidity, can also affect the accuracy and reliability of the robot during task execution.

[0005] In addition, when robots face complex scenes, they cannot obtain real-time environmental data containing multiple information types, including images, sounds, tactile feedback, and environmental temperature, humidity, brightness, air quality, etc., and quickly correlate and analyze them. Therefore, the robot will have difficulty in effectively identifying the key target of the task.

[0006] Further, current robot control methods have significant shortcomings in understanding and responding to user instructions and intentions, and lack the ability to autonomously discover and solve problems. In daily life scenarios, a user may issue an instruction such as "Help me get a cup of warm water on the kitchen table". Due to the lack of multi-modal information fusion processing capability, a traditional robot may misjudge the user's location or the characteristics of the target object based on simple speech recognition. If there are multiple similar cups in the kitchen and the light causes a visual misjudgment, the robot cannot accurately locate the cup of warm water required by the user and cannot provide correct feedback in a timely manner. For example, in an office scenario, a user urgently needs a robot to help find a specific file and deliver it to the conference room. However, the robot cannot accurately understand the user's urgency at the moment due to its inability to perceive real-time changes in the surrounding personnel flow and temporary adjustments in the office area layout, resulting in delays or incorrect delivery, greatly reducing the user experience and failing to meet the timeliness and accuracy requirements in actual use. SUMMARY

[0007] An embodiment of the present application provides a robot control method and system that actively and real-time acquires multi-modal data including scene and condition information, fuses and analyzes the data, retrieves associated memories, accurately identifies tasks, and plans strategies based on the tasks, overcoming the shortcomings of traditional robots that rely on pre-sets, one-sided perception, difficulty in adapting to changes, and difficulty in understanding user intentions.

[0008] To solve the above technical problems, a first aspect of an embodiment of the present application provides a robot control method, comprising the following steps:

[0009] real-time acquisition of multi-modal environmental information within a detection range;

[0010] analyzing the multi-modal environmental information and retrieving associated memory data based on the multi-modal environmental information;

[0011] identifying a current task based on the analysis results of the multi-modal environmental information and the retrieved associated memory data;

[0012] planning an execution strategy based on the current task;

[0013] wherein the multi-modal environmental information includes environmental scene information and / or environmental condition information, the information types of the environmental scene information include some or all of images, sounds, and tactile sensations, and the information types of the environmental condition information include some or all of environmental temperature, environmental humidity, environmental brightness, and environmental air.

[0014] Further, the analyzing the multi-modal environmental information comprises:

[0015] Based on the multi-modal information analysis model and historical environmental data, determine part or all of the data of semantic, timbre, sound source and decibel of sound information, analyze voice interaction demand scene and / or environmental sound processing trend;

[0016] And / or,

[0017] Based on the multi-modal information analysis model and historical environmental data, determine the scene change result of image information and / or tactile information;

[0018] And / or,

[0019] Based on the multi-modal information analysis model, analyze the environmental change result of the environmental condition information.

[0020] Further, the associative memory data includes an environmental data set and a task execution data set, and the environmental data set and the task execution data set are stored in association;

[0021] According to the analysis result of the multi-modal environmental information and the retrieved associative memory data, the current task is identified, which includes:

[0022] Retrieving the environmental data set matching the analysis result of the multi-modal environmental information in the environmental data set;

[0023] According to the corresponding associated task execution data set of the retrieved environmental data set, the target task to be executed is identified.

[0024] Further, according to the corresponding associated task execution data set of the retrieved environmental data set, the target task to be executed is identified, which includes:

[0025] In the case of retrieving multiple environmental data sets, according to the historical task execution time and / or historical task execution frequency, the task execution data set currently adapted is filtered out from the corresponding multiple task execution data sets;

[0026] According to the filtered task execution data set, the target task to be executed is identified.

[0027] Further, the associative memory data includes environmental data and task execution data associated and stored based on pre-training, and environmental data and task execution data stored in historical operation.

[0028] Further, according to the current task, the execution strategy is planned, which includes:

[0029] According to the task type and task object, the current task is decomposed into multiple unit tasks;

[0030] determining initial parameters and target parameters of each of the unit tasks;

[0031] generating multiple sets of task planning strategies according to the initial parameters and the target parameters of all the unit tasks;

[0032] traversing the task planning strategies by a breadth-first search method or a depth-first search method to select the task planning strategy currently executed.

[0033] Further, the generating multiple sets of task planning strategies according to the initial parameters and the target parameters of all the unit tasks comprises:

[0034] obtaining running parameters of each execution unit of the robot and initial position data of the robot;

[0035] generating multiple sets of task planning strategies according to the running parameters and the initial position data, in combination with the initial parameters and the target parameters of all the unit tasks.

[0036] Further, the analyzing the multi-modal environment information and retrieving associated memory data according to the multi-modal environment information comprises:

[0037] determining whether the multi-modal environment information matches a preset scene triggering condition;

[0038] if the multi-modal environment information does not match the preset scene triggering condition, then analyzing the multi-modal environment information and retrieving the associated memory data;

[0039] if the multi-modal environment information matches the preset scene triggering condition, then executing a task setting command corresponding to a preset scene information combination;

[0040] the preset scene triggering condition comprises at least one environment information setting parameter associated with a task setting command corresponding thereto.

[0041] Further, after obtaining the multi-modal environment information within the detection range in real time and before implementing the execution strategy, the control method further comprises:

[0042] obtaining a preset voice instruction in real time;

[0043] if the preset voice instruction is not obtained, then maintaining an execution state of a previous execution step and continuing to obtain the preset voice instruction in real time;

[0044] if the preset voice instruction is obtained, then stopping the previous execution step and executing a task command corresponding to the preset voice instruction.

[0045] Correspondingly, a second aspect of the embodiment of the present application provides a robot control system, comprising:

[0046] An information acquisition module is configured to acquire multi-modal environment information in a detection range in real time, wherein the multi-modal environment information comprises environment scene information and / or environment condition information, the information types of the environment scene information comprise part or all of images, sounds and tactile sensations, and the information types of the environment condition information comprise part or all of environment temperature, environment humidity, environment brightness and environment air;

[0047] A task decision module is configured to analyze the multi-modal environment information, retrieve associated memory data according to the multi-modal environment information, and identify a current task according to the analysis result of the multi-modal environment information and the retrieved associated memory data.

[0048] A task planning module is configured to plan an execution strategy according to the current task.

[0049] Correspondingly, a third aspect of the embodiment of the present application provides an electronic device, comprising at least one processor and a memory connected with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the robot control method.

[0050] Correspondingly, a fourth aspect of the embodiment of the present application provides a computer readable storage medium having computer instructions stored thereon, wherein the instructions are executed by a processor to implement the robot control method.

[0051] The above technical solutions of the embodiment of the present application have the following beneficial technical effects:

[0052] 1. By acquiring multi-modal environment information covering images, sounds, tactile sensations and various environment conditions in real time, and combining historical data with multi-modal information analysis models for deep analysis, the environment changes can be accurately understood, the current task can be quickly and accurately identified according to the analysis result and associated memory data, the complex and variable scene can be adapted, and the scientificity and accuracy of robot decision-making can be effectively improved.

[0053] 2. The identified current task is disassembled into multiple unit tasks according to the task type and object, the running parameters of each execution unit of the robot, the initial position and the initial and target parameters of the unit task are comprehensively considered, multiple sets of strategies are generated and intelligently screened, the task execution is ensured to be efficient and smooth, the performance advantages of the robot are maximized, and the task completion quality and efficiency are improved.

[0054] 3. Real-time acquisition of preset voice instructions and immediate adjustment of task execution process accordingly, ensuring original task continuity when no instructions are received, and prioritizing response when instructions are received, cooperating with multi-modal information processing mechanism, realizing high flexibility of human-machine interaction, allowing the robot to adapt to user needs at any time, and enhancing user experience. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 is a robot control method flowchart provided by an embodiment of the present application;

[0056] Figure 2 is a robot control method module block diagram provided by an embodiment of the present application.

[0057] REFERENCE NUMERALS:

[0058] 1, information acquisition module, 2, task decision module, 3, task planning module. DETAILED DESCRIPTION

[0059] To make the purpose, technical solutions and advantages of the present application clearer and more intelligible, the present application is further described in detail below with specific embodiments and with reference to the drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. In addition, in the following description, the description of known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present application.

[0060] Please refer to Figure 1 The first aspect of the embodiment of the present application provides a robot control method, comprising the following steps:

[0061] Step S100, real-time acquisition of multi-modal environmental information in the detection range. Wherein, the multi-modal environmental information includes environmental scene information and / or environmental condition information, the information type of the environmental scene information includes part or all of image, sound and touch, and the information type of the environmental condition information includes part or all of environmental temperature, environmental humidity, environmental brightness and environmental air.

[0062] During the operation of the robot, the multi-modal environmental information provides a basis for its comprehensive understanding of the surrounding situation. First, the environmental scene information can be rich images captured by a high-definition camera, whether it is the placement of furniture at home, the action posture of a person, or the topographic features of complex terrain outdoors, etc., which can be intuitively presented; sound information is the surrounding environmental information monitored by the robot, such as user's conversation instructions, machine operation noise, or natural wind and rain sounds; touch information is the information perceived by the robot through touch, such as feeling the texture and hardness of an object when grabbing it, and understanding the roughness of the surface when touching an obstacle, etc. The above image, sound and touch information are not all acquired every time, but according to the specific task demand and scene, the robot can flexibly select part of them to construct accurate scene cognition.

[0063] Take a home robot as an example. Its high-definition camera can accurately capture images of every corner of the home, whether it is a toy, a book placed randomly in the living room, or the state of the tableware on the kitchen stove. It can clearly image and identify the category, location, and posture of the object. Its microphone listens to the surrounding sound at any time. The user calls the name of the robot, the daily conversation, and even the abnormal sound emitted by the operation of the electrical appliance. The tactile sensor distributed on the surface of the shell allows the robot to feel the hardness and smoothness of the surface when touching furniture and grabbing objects, and to perceive the home space in all directions.

[0064] For example, in the evening, the camera of the home robot captures the gradually darkening light in the living room and identifies the location of the lamp. The microphone hears the child in the bedroom say "I am thirsty." At the same time, the tactile sensor feels the joint gap between the carpet and the floor during the movement. These different modal information is real-time converged to provide a basis for subsequent service decision-making, allowing the robot to know the current environmental conditions.

[0065] Step S200, analyze the multi-modal environmental information, and retrieve the associated memory data according to the multi-modal environmental information.

[0066] The collected information needs to be quickly sorted out in the complex and variable home environment. Through intelligent image algorithms, the robot can accurately classify the people and objects in the camera picture and extract the key targets from the cluttered background. The voice recognition and semantic analysis of the sound understand the user's instruction intention. Based on these analysis results, similar scene memories such as the user's habit demand at this time period in the past, the regularity of the placement of objects in different rooms, etc. are immediately searched in the memory library of the robot to mine the experience that can be learned.

[0067] For example, the robot hears the sound of water flowing in the kitchen for a long time, and combined with the image analysis that there is water overflowing at the edge of the sink, it retrieves the memory to find that there has been a similar situation before, that is, the faucet is not closed tightly, and then associates it with the relevant processing flow memory module of closing the faucet and cleaning the accumulated water to find the right direction for the next action.

[0068] Step S300, according to the analysis results of the multi-modal environmental information and the retrieved associated memory data, identify the current task.

[0069] By integrating various information, the task that best fits the current situation is determined. By weighing the user's demand, the safety and comfort of the home environment, etc., at least one task with the highest priority is locked from numerous possible tasks; the immediate instruction is given priority, and the potential demand is also considered.

[0070] Step S400, according to the current task, plan the execution strategy.

[0071] After the task is clear, the home robot plans an efficient execution plan according to the layout of the home, its own movement ability, and the remaining energy. Consider how to avoid obstacles such as tables, chairs, and pets, choose the most convenient path, and reasonably allocate the actions of components such as mechanical arms and wheels to ensure that the task is completed smoothly and quickly.

[0072] Specifically, the analysis of the multi-modal environment information in step S200 includes:

[0073] Step S210, based on the multi-modal information analysis model and historical environment data, determine part or all of the data of semantic, timbre, sound source, and decibel of the sound information, analyze the voice interaction demand scene and / or environmental sound processing trend.

[0074] In a home environment, sound information is diverse and complex. Based on the multi-modal information analysis model, the robot combines historical environment data accumulated in the past, such as the voice characteristics of family communication at different times, and the sound spectrum of common appliances, to deeply analyze the real-time collected sound. Determining the semantic can make the robot understand the content of the owner's instructions, such as "help turn on the TV"; identifying the timbre can distinguish which family member is speaking in order to provide personalized service; locating the sound source helps to determine the direction of the sound, quickly respond to the demand; and paying attention to the decibel size can detect environmental noise abnormalities, determine whether there is an emergency, and then analyze whether it is a daily voice interaction assistance or an environmental sound processing scene such as a baby crying or an appliance malfunctioning.

[0075] Step S220, based on the multi-modal information analysis model and historical environment data, determine the scene change result of image information and / or tactile information.

[0076] The home scene is constantly changing, and the focus is on visual and tactile perception results analysis. Using the analysis model and historical memory, the images captured by the camera are dynamically compared to determine whether the objects have been moved, added or reduced, identify the activity state and expression of the personnel, etc.; tactile information reflects the details of the robot's contact with the environment, such as touching the door handle from smooth to rough, which may mean there is a stain or wear. By integrating these, the overall scene change is accurately grasped, providing intuitive visual and tactile basis for subsequent decision-making.

[0077] Step S230, based on the multi-modal information analysis model, analyze the environmental change result of the environmental condition information.

[0078] The environmental conditions in the home, although easily overlooked, have a significant impact on the comfort of life. Through a specialized analysis module for environmental condition information such as temperature, humidity, brightness, and air quality, the robot can acutely perceive changes in the environment. For example, a sudden drop in temperature may indicate that a window is not properly closed; an increase in humidity may indicate that there is water accumulation in the bathroom that has not yet dried; a decrease in light may require adjusting the brightness of the lamps. Real-time mastery of the changing state of these environmental conditions can enable the robot to better identify and handle tasks.

[0079] The association memory data includes an environmental data set and a task execution data set. The environmental data set and the task execution data set are stored in association.

[0080] In the robot control system, the association memory data is very important. The environmental data set stores various environmental data collected by the robot in various scenarios in the past, including image details, sound characteristics, tactile perception, and environmental condition information such as temperature, humidity, and brightness at different times and in different areas. These data are stored in an orderly manner in groups. The task execution data set stores details of each task execution, including identified task targets, planned and implemented execution strategies, and final task completion feedback. When the robot real-time acquires and analyzes the current multi-modal environmental information, it can quickly find a similar past scenario data group in the environmental data set, and then seamlessly link to the corresponding task execution data group, learn from experience, quickly and accurately determine the task and action plan to be executed, greatly improve the robot's autonomous decision-making ability in complex and variable home environments, make its service more in line with user needs, avoid repeated groping, and efficiently complete various tasks.

[0081] Correspondingly, the step S300 of identifying the current task according to the analysis result of the multi-modal environmental information and the retrieved association memory data further includes:

[0082] Step S310: searching for an environmental data group matching the analysis result of the multi-modal environmental information in the environmental data set.

[0083] When the robot completes the collection and analysis of the multi-modal environmental information, it performs a precise search in the environmental data set. The environmental data set stores a large amount of environmental information experienced by the robot in the past, and each environmental data group is a comprehensive record of multi-modal information in a specific scenario. At this moment, the robot compares the multi-modal environmental information features obtained through real-time analysis, such as the current environmental image key elements, sound spectrum characteristics, tactile perception details, and environmental condition value ranges, in the vast environmental data set, finds the most suitable past scenario data group, provides a key environmental reference for subsequent accurate decision-making, and ensures that the robot can learn from past experience to cope with the current situation.

[0084] Step S320, according to the retrieved environment data set corresponding to the associated task execution data set, identify the target task to be executed at present.

[0085] Since the environment data set and the task execution data set are closely associated and stored, the matching environment data set is successfully retrieved in step S310, and the robot can be directly positioned to the task execution data set corresponding thereto. This task execution data set records in detail the task details executed in the similar environment in the past, including the task target, the execution step, the final result and other key information. The robot analyzes these associated task execution data in depth, and screens and extracts the most suitable and most urgent target task to be executed at present. For example, if the retrieved past environment data set corresponds to the scene of kitchen smoke diffusing, temperature rising and accompanied by alarm sound, and the associated task execution data set shows that the task executed at that time is to close the stove, open the ventilation device and call the host, when similar environmental characteristics are identified again at present, the robot can quickly lock the same target task, and efficiently and accurately respond, greatly improving the accuracy and timeliness of task identification.

[0086] Further, in step S320, according to the retrieved environment data set corresponding to the associated task execution data set, identifying the target task to be executed at present, includes the following steps:

[0087] Step S321, in the case of retrieving multiple environment data sets, according to the historical task execution time and / or the historical task execution frequency, screening the current adaptive task execution data set from the corresponding multiple task execution data sets.

[0088] In a complex and variable actual scene, when the robot searches according to the multi-modal environment information analysis result, it is extremely likely to find multiple environment data groups with certain similarity to the current environment. At this time, the historical task execution time and the historical task execution frequency become important basis for accurate screening. The historical task execution time contains rich information, for example, at different times of the day, some tasks have higher priority or necessity. Taking the family scene as an example, in the morning, if multiple environment data groups related to the bedroom environment are retrieved, and the task execution data group related to opening the curtains and adjusting the indoor light is retrieved, because it is often executed at this time in the past, its execution time characteristics are consistent with the present, so it is more likely to be screened out first; looking at the historical task execution frequency, it reflects the frequency of different tasks in similar environments in the past experience. If the kitchen often appears smoke alarm due to cooking, when the kitchen is detected again under similar environmental characteristics such as smoke, and multiple related environment data groups are retrieved, the task execution data group corresponding to the task of turning off the stove and opening the ventilation equipment, which is frequently executed, will be given priority and considered as a candidate solution, laying a foundation for the next step of accurately identifying the target task.

[0089] Step S322, according to the screened task execution data group, the target task to be executed at present is identified.

[0090] When the most suitable one for the current situation is screened out from the many possible task execution data groups by using time and frequency and other key factors through step S321, the process of target task locking can be carried out. This screened task execution data group records the whole process information of the past successful task execution in detail, and the robot only needs to analyze the core content such as the setting of the task target and the planning of the task step in it, so that the target task that needs to be immediately started at present can be clearly and definitely determined. For example, the screened task execution data group shows that in the environment of a dark and loud music in the living room, the task that has been executed many times before is to adjust the light and the volume to meet the needs of people's communication, so when the similar living room gathering scene characteristics are identified at present, the robot can quickly determine that the target task to be executed at present is to adjust the light and the volume, so as to efficiently and accurately integrate into the scene.

[0091] In addition, in one specific embodiment of the present application, the associative memory data includes environment data and task execution data associated and stored based on pre-training, and environment data and task execution data stored in historical operation.

[0092] In one aspect, the pre-training manner associated storage of environment data and task execution data plays a key role in the basic guidance. In the pre-training stage, a large amount of simulation data covering various typical scenes and environmental conditions is input to the robot, and the robot is repeatedly trained in a virtual but highly simulated environment. For example, simulate different seasons and different time periods of home environment, from the low temperature and dim of the bedroom in the cold winter morning to the high temperature and strong light of the living room in the summer afternoon; simulate various daily activity scenes, such as kitchen cooking with smoke, and lively and noisy living room gathering. For each simulation scene, a corresponding task execution scheme is carefully set, such as turning on the heating and opening the curtains in the cold winter morning, adjusting the air conditioning temperature and pulling up the sunshade in the summer afternoon, closing the stove and opening the ventilation when the kitchen smoke alarm goes off, and other tasks. These environment data and task execution data are closely associated and stored. Through such large-scale pre-training, the robot has pre-constructed a huge and clear knowledge system, so that it can quickly find similar references from the pre-training accumulated experience library when facing real scenes, and provide strong support for task decision-making.

[0093] On the other hand, the environment data and task execution data stored in the history running are an important supplement and dynamic optimization of the pre-training system. When the robot is actually put into actual use environment, it will record the real environment details and the task details executed in real time. Taking a home robot as an example, it will store the specific environment conditions when the owner assigns a task each time, such as the change of the layout of the items in the room, the activity state of the personnel, the actual temperature and humidity values, and the task execution steps taken to complete the owner's instructions, such as taking items from which position, delivering by which path, and operating which household appliances. With the passage of time, the above data will become more and more rich, and the robot can adapt to the unique living habits and environmental characteristics of each family more accurately according to these constantly updated historical running data. For example, if it is known that a family often holds a small party in the evening on weekends, the robot can make preparations in advance according to the environment data and corresponding service task execution data in the past party, such as adjusting the light atmosphere in advance, preparing common drinks and tableware, and further improving the personalization and intelligence level of service, realizing the transformation from general intelligence to exclusive intelligence.

[0094] Specifically, the execution strategy planning in step S400 according to the current task includes:

[0095] In step S410, the current task is decomposed into multiple unit tasks according to the task type and the task object.

[0096] When the robot accurately identifies the current task, it first makes fine planning of the task, and decomposes the current task into multiple unit tasks; the task type determines the general direction and category of the task, such as cleaning task, delivery task or environment adjustment task, and the execution logic and focus of different types of tasks are also different. Secondly, the task object further determines the specific direction of the task implementation, such as the cleaning task is aimed at the living room floor, the kitchen table, or the bedroom clothes, etc. Taking the family scene as an example, if the current task is to prepare the dining table for family dinner, it belongs to the arrangement task type, and the task object is the dining table, then multiple unit tasks such as wiping the table top, placing tableware, and arranging tablecloth can be decomposed, through this subdivision, the complex task is organized, which is convenient for subsequent targeted arrangement of execution details, reduces the difficulty of task execution, and improves the efficiency.

[0097] Step S420, determine the initial parameters and target parameters of each unit task.

[0098] After the unit tasks are determined, the initial parameters and target parameters of each unit task should be further determined. The initial parameters reflect the state condition at the beginning of the unit task, for example, in the unit task of wiping the table top, the initial parameters include the initial dirtiness of the table top, the material of the table top (wood, glass, etc.), the current placement of the objects, etc., which determine the starting point and required resources; the target parameters are the final state that the unit task is expected to achieve, such as the table top reaching a bright and clean state without stains, the tableware being placed neatly and in accordance with the etiquette norms for dining, etc. Accurate definition of the above parameters sets a clear starting point and ending point for each unit task, ensuring that the task execution does not deviate from the target.

[0099] Step S430, generate multiple sets of task planning strategies according to the initial parameters and target parameters of all unit tasks.

[0100] With the unit tasks and their parameters, multiple sets of task planning strategies are generated to realize the flexibility of robot intelligent decision-making. Due to the complexity and variability of the actual environment, a single strategy may not be able to cope with various unexpected situations or optimally achieve the target. For example, in the case of preparing the dining table for family dinner, one strategy may be to wipe the table top first, then place the tableware, and finally arrange the tablecloth; another strategy may be to clean the table top objects to one side, simultaneously arrange the tablecloth, then wipe the table top, and finally place the tableware. Each strategy considers the sequence of unit tasks, resource allocation, time arrangement and other factors, and by generating multiple sets, more possibilities are provided for subsequent selection of the best solution to adapt to different scene requirements.

[0101] Step S440, traverse the task planning strategies by breadth-first search or depth-first search to select the current execution task planning strategy.

[0102] Finally, the breadth-first search method or the depth-first search method is used to traverse the task planning strategy, and the currently executed task planning strategy is selected from the task planning strategy.

[0103] The breadth-first search starts from the top layer of the tree structure and explores all branches horizontally layer by layer to comprehensively evaluate the advantages and disadvantages of different strategies in the initial stage, and then gradually deepens. The depth-first search selects a branch and deepens it to the bottom, first focuses on exploring the complete execution process and final effect of a strategy, and then explores other branches. Taking the preparation of a dining table for a family dinner as an example, if the breadth-first search is used, the robot will first compare the differences in the first few steps (such as the way and speed of cleaning the table surface) of several strategies, preliminarily select a set of strategies with better initial steps, and then further compare the subsequent steps. If the depth-first search is used, the robot will completely simulate the execution of a strategy, evaluate its final presentation effect, and then try other strategies one by one.

[0104] Through this search traversal, combined with real-time factors such as the current environment, resource limitations, etc., the robot selects the task planning strategy that best fits the current needs, ensures the smooth and efficient execution of the task, and achieves the best service effect.

[0105] Further, in step S430, generating a plurality of task planning strategies according to the initial parameters and target parameters of all unit tasks, comprising:

[0106] In step S431, the running parameters of each execution unit of the robot and the initial position data of the robot are obtained.

[0107] The execution units of the robot include different components such as mechanical arms, mobile chassis, sensor assemblies, etc., and each execution unit has its unique running parameters. Taking the mechanical arm as an example, the running parameters include joint range of motion, upper limit of force output, current extension state, motion accuracy, etc., which limit the type of motion that the mechanical arm can complete, the degree of precision of operation, and the range of force control. The running parameters of the mobile chassis include the range of movement speed, turning flexibility, terrain adaptability, current power or energy reserve, etc., which determine the movement characteristics of the robot. At the same time, it is crucial to determine the initial position data of the robot, which is equivalent to a coordinate origin, allowing the robot to know its starting position in the environment space, such as knowing its position in the corner of the living room, how far it is from the dining table, and the relative position relationship with the surrounding obstacles, etc. Obtaining these information provides basic data support for subsequent rational allocation of robot resources and planning of action routes based on unit task requirements, ensuring the feasibility of the generated strategy.

[0108] In step S432, a plurality of task planning strategies are generated according to the running parameters and initial position data, combined with the initial parameters and target parameters of all unit tasks.

[0109] Based on the above collected data, combined with the initial parameters and target parameters of all unit tasks, multiple sets of task planning strategies can be generated. On the one hand, the initial parameters of the unit task depict the scene details at the start of the task, such as the initial state of the table in the table wiping task mentioned earlier, which requires the planning strategy to fully consider how the robot starts from the current state and reasonably intervenes using its own execution units. For example, if the table has heavy stains and is made of wood, combined with the upper limit of the mechanical arm's force output and the adaptability of the cleaning tool, the appropriate wiping force and cleaning supplies are selected. On the other hand, the target parameters indicate the end point of the task, such as the placement of tableware in an orderly and standardized manner, considering the precise trajectory and sequence of the mechanical arm placing the tableware. In addition to the initial position of the robot and the running parameters of each execution unit, one strategy may be for the robot to first move to one side of the table, use the mechanical arm to grab the cleaning cloth at a specific angle, adjust the wiping force according to the degree of dirt on the table, clean the table, then move to the storage area to pick up the tableware, and place them in order according to the dining etiquette requirements. Another strategy may be for the robot to first adjust the mechanical arm posture in place to prepare to clean the items around the table, while moving the chassis slowly towards the table, then directly lay the tablecloth after cleaning the items, and then wipe the table and place the tableware. Based on the above multiple factors, multiple sets of task planning strategies are generated to adapt to different scene changes and fully utilize the performance advantages of the robot to meet the complex and variable actual task requirements.

[0110] In addition, in step S200, the multi-modal environment information is analyzed, and the associated memory data is retrieved according to the multi-modal environment information.

[0111] In step S241, it is judged whether the multi-modal environment information matches the preset scene trigger condition. The preset scene trigger condition includes at least one environment information setting parameter associated with the corresponding task setting command.

[0112] Each preset scene trigger condition is closely bound to at least one task setting command, ensuring that the robot accurately knows what to do after identifying the scene. This one-to-one correspondence design guarantees the accuracy and reliability of the entire control system.

[0113] At the initial stage of the robot processing the multi-modal environment information flow, it can be determined whether the multi-modal environment information matches the preset scene trigger condition. The preset scene trigger condition is a key judgment basis that is set in advance based on the robot's past learning, training, and induction of common task scenarios, covering various specific and detailed environment information setting parameters that are closely related to specific task requirements. For example, in a home scenario, the preset "evening movie watching scene trigger condition" may be set as a combination of parameters such as the environment brightness being lower than a certain value, a person staying in the living room sofa area for a long time, and nearby electronic devices being turned on to emit specific audio signals. By comparing the multi-modal environment information collected in real time, including the brightness of the living room light captured by the camera, the sound characteristics listened to by the microphone, and the sofa area pressure changes sensed by the sensor, the robot can quickly determine whether the current environment preliminarily matches this preset movie watching scene feature, thereby determining the subsequent processing flow and greatly improving the relevance and efficiency of information processing.

[0114] Step S242, if the multi-modal environment information does not match the preset scene trigger condition, then the multi-modal environment information is analyzed and the associated memory data is retrieved.

[0115] When it is determined that the multi-modal environment information does not match the preset scene trigger condition, it means that the current environment is not one of those with clear, specific preset task scenarios. At this time, the regular process of analyzing multi-modal environment information and retrieving associated memory data can be entered to avoid the robot doing useless work in non-typical scenarios. In a complex daily environment, most of the time the environment is in a relatively random, non-specific task-oriented state, such as the regular static images of each room when no one is at home, and the slight environmental background sound. At this time, the robot needs to analyze the object layout in the image, the semantic content of the sound, and the tactile perception details using the multi-modal information analysis model as usual, while deeply retrieving the vast amount of past environment data and task execution experience stored in the memory library to mine potential task clues and prepare for general tasks such as daily cleaning and item organization, ensuring that the robot can adapt to diverse daily environmental needs.

[0116] Step S243, if the multi-modal environment information matches the preset scene trigger condition, the task setting command corresponding to the preset scene information combination is executed.

[0117] Once the multi-modal environment information matches the preset scene trigger condition, the robot directly executes the task setting command corresponding to the preset scene information combination, skips the complex regular analysis and retrieval process, realizes fast response, and improves the efficient processing capability of the robot for high-frequency and typical scenes. Taking the evening viewing scene as an example, if the environment information matches the corresponding trigger condition, the robot will directly execute the task setting command associated therewith, such as automatically dimming other irrelevant lights, pulling up the curtains to enhance the viewing atmosphere, and adjusting the sound to the best viewing sound mode. The robot seamlessly connects the potential needs of the user in a specific scene, greatly improves the user experience, and makes the robot show high intelligence and humanized service characteristics at the critical moment.

[0118] Further, after the multi-modal environment information in the detection range is acquired in real time in step S100 and before the execution strategy is implemented, the control method further comprises:

[0119] S110, acquiring a preset voice instruction in real time.

[0120] After the robot has successfully acquired the multi-modal environment information in the detection range in real time, the preset voice instruction is acquired in real time, which enhances the flexibility and immediacy of human-computer interaction. The preset voice instruction is a voice information that the user sets in advance and can directly convey an explicit task requirement. For example, in a home scene, the user may have set in advance simple and clear voice instructions such as “pause cleaning”, “help me get a book”, and “turn on the air purifier”. The robot continuously monitors the environmental sound through the built-in high-sensitivity microphone and uses advanced voice recognition technology to filter and extract the voice instruction conforming to the preset format from the environmental sound, ensuring that any immediate needs of the user are not missed, and laying a foundation for subsequent flexible adjustment of the task execution strategy according to the voice instruction.

[0121] S120, if the preset voice instruction is not acquired, maintaining the execution state of the original execution step and continuing to acquire the preset voice instruction in real time.

[0122] When the robot does not acquire the preset voice instruction, the strategy of maintaining the execution state of the original execution step and continuing to acquire the preset voice instruction in real time is adopted, which realizes the stability of the robot task execution process and the instruction intelligent waiting mechanism. In the daily environment, most of the time, the robot executes the task according to the established process, such as cleaning the living room floor according to the planned route in the home cleaning task. At this time, if no new voice instruction of the user is received, the robot continues to focus on the current cleaning action, while maintaining the monitoring of the voice instruction, avoiding the blind action or task interruption of the robot when there is no additional instruction, ensuring the continuity of the task, and ensuring that the regular task can proceed smoothly, adapting to the relatively stable operation demand of the daily environment.

[0123] S130, if the preset voice instruction is acquired, the original execution step is stopped, and a task command corresponding to the preset voice instruction is executed.

[0124] If the robot acquires the preset voice instruction, the original execution step is stopped, and a task command corresponding to the preset voice instruction is executed, which improves the rapid response ability to user demand, ensures the highest priority of user instruction, and pauses the task being executed at present as long as the preset voice instruction is recognized, regardless of what complex task process the robot is currently executing. For example, the robot is originally arranging tableware in the kitchen, and when the preset voice instruction "help me get a book" is heard, it will quickly stop the arrangement action, determine the location of the target book according to the instruction, plan the optimal path to the study or bookshelf, use the mechanical arm to accurately grasp the book and deliver it to the user's hand, greatly improve the user experience, and realize efficient interaction of human-computer collaboration.

[0125] Correspondingly, please refer to Figure 2 The second aspect of the embodiment of the application provides a robot control system, which comprises:

[0126] The information acquisition module 1 is used for acquiring multi-modal environment information in a detection range in real time, the multi-modal environment information comprises environment scene information and / or environment condition information, the information types of the environment scene information comprise part or all of images, sounds and tactile sensations, and the information types of the environment condition information comprise part or all of environment temperature, environment humidity, environment brightness and environment air;

[0127] The task decision module 2 is used for analyzing the multi-modal environment information, retrieving associated memory data according to the multi-modal environment information, and identifying a current task according to the analysis result of the multi-modal environment information and the retrieved associated memory data.

[0128] The task planning module 3 is used for planning an execution strategy according to the current task.

[0129] Correspondingly, the third aspect of the embodiment of the application provides an electronic device, which comprises at least one processor and a memory connected with the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the robot control method.

[0130] Correspondingly, the fourth aspect of the embodiment of the application provides a computer readable storage medium, which stores computer instructions, and the instructions are executed by a processor to implement the robot control method.

[0131] The embodiment of the present application aims to protect a robot control method and system, the method comprising the following steps: acquiring multi-modal environment information in a detection range in real time; analyzing the multi-modal environment information, and retrieving associated memory data according to the multi-modal environment information; identifying a current task according to the analysis result of the multi-modal environment information and the retrieved associated memory data; and planning an execution strategy according to the current task; wherein the multi-modal environment information comprises environment scene information and / or environment condition information, the information type of the environment scene information comprises part or all of images, sounds and tactile sensations, and the information type of the environment condition information comprises part or all of environment temperature, environment humidity, environment brightness and environment air.

[0132] The above technical solution has the following effects:

[0133] 1. By acquiring multi-modal environment information covering images, sounds, tactile sensations and various environment conditions in real time, combining historical data and using a multi-modal information analysis model for deep analysis, the environment changes can be accurately understood, the current task can be quickly and accurately identified according to the analysis result and the associated memory data, the complex and variable scene can be adapted, and the scientificity and accuracy of the autonomous decision of the robot can be effectively improved.

[0134] 2. The identified current task is disassembled into multiple unit tasks according to the task type and object, the running parameters, initial position and initial and target parameters of the unit task of each execution unit of the robot are comprehensively considered, multiple strategies are generated and intelligently screened, the task execution is ensured to be efficient and smooth, the performance advantage of the robot is maximized, and the task completion quality and efficiency are improved.

[0135] 3. The preset voice instruction is acquired in real time, and the task execution process is adjusted accordingly, the original task continuity is ensured when no instruction is received, the instruction is responded to preferentially when received, the multi-modal information processing mechanism is matched, the high flexibility of human-computer interaction is realized, the robot is adapted to the user demand at any time, and the user experience is enhanced.

[0136] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0137] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.

[0138] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.

[0139] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.

[0140] Finally, it should be noted that the above-mentioned embodiments are merely intended to illustrate the technical solutions of the present application, but not to limit the same. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalent replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered within the scope of protection of the claims of the present application.

Claims

1. A robot control method characterized by, The method comprises the following steps: Real-time acquisition of multi-modal environmental information within a detection range; Parsing the multi-modal environmental information and retrieving associated memory data according to the multi-modal environmental information; Identifying a current task according to the parsing result of the multi-modal environmental information and the retrieved associated memory data; Planning an execution strategy according to the current task; The multi-modal environmental information comprises environmental scene information and / or environmental condition information, the information type of the environmental scene information comprises part or all of images, sounds and tactile sensations, and the information type of the environmental condition information comprises part or all of environmental temperature, environmental humidity, environmental brightness and environmental air.

2. The robot control method according to claim 1, characterized by, The parsing of the multi-modal environmental information comprises: Determining part or all of the data of semantic, timbre, sound source and decibel of sound information, parsing voice interaction demand scenarios and / or environmental sound processing trends based on a multi-modal information parsing model and historical environmental data; And / or, Determining scene change results of image information and / or tactile information based on a multi-modal information parsing model and historical environmental data; And / or, Parsing environmental change results of the environmental condition information based on a multi-modal information parsing model.

3. The robot control method according to claim 1, wherein, The associated memory data comprises an environmental data set and a task execution data set, and the environmental data groups in the environmental data set are stored in association with the task execution data groups in the task execution data set; The identification of the current task according to the parsing result of the multi-modal environmental information and the retrieved associated memory data comprises: Retrieving the environmental data groups matching the parsing result of the multi-modal environmental information in the environmental data set; Identifying a target task to be executed according to the task execution data groups associated with the retrieved environmental data groups.

4. The robot control method according to claim 3, wherein, The identification of the target task to be executed according to the task execution data groups associated with the retrieved environmental data groups comprises: In the case of retrieving multiple environmental data groups, filtering the task execution data groups currently adapted from the corresponding multiple task execution data groups according to historical task execution time and / or historical task execution frequency; and Identifying the target task to be executed according to the filtered task execution data groups.

5. The robot control method according to claim 3, wherein, The associated memory data comprises environmental data and task execution data stored in association based on a pre-training mode and environmental data and task execution data stored in history.

6. The robot control method according to claim 1, wherein, The planning of the execution strategy according to the current task comprises: Decomposing the current task into multiple unit tasks according to task types and task objects; Determining initial parameters and target parameters of each unit task; Generating multiple sets of task planning strategies according to the initial parameters and target parameters of all unit tasks; Selecting the task planning strategy currently executed by breadth-first search or depth-first search.

7. The robot control method according to claim 6, wherein, The generation of multiple sets of task planning strategies according to the initial parameters and target parameters of all unit tasks comprises: Acquiring running parameters of each execution unit of the robot and initial position data of the robot; According to the operation parameters and the initial position data, a plurality of sets of task planning strategies are generated in combination with initial parameters and target parameters of all the unit tasks.

8. The robot control method according to claim 1, wherein, The multi-modal environment information is analyzed, and associated memory data is retrieved according to the multi-modal environment information, including: determining whether the multi-modal environment information matches a preset scene triggering condition; if the multi-modal environment information does not match the preset scene triggering condition, then analyzing the multi-modal environment information and retrieving the associated memory data; if the multi-modal environment information matches the preset scene triggering condition, a task setting command corresponding to a preset scene information combination is executed, wherein the preset scene triggering condition includes at least one environmental information setting parameter associated with a corresponding task setting command.

9. The robot control method according to any one of claims 1-8, wherein, After real-time acquisition of multi-modal environment information within a detection range and before implementation of the execution strategy, the control method further includes: real-time acquisition of a preset voice instruction; if the preset voice instruction is not acquired, the execution state of the original execution step is maintained, and real-time acquisition of the preset voice instruction is continued; if the preset voice instruction is acquired, the original execution step is stopped, and a task command corresponding to the preset voice instruction is executed.

10. A robot control system, characterized by including: an information acquisition module for real-time acquisition of multi-modal environment information within a detection range, the multi-modal environment information including environmental scene information and / or environmental condition information, the information type of the environmental scene information including some or all of images, sounds, and tactile sensations, and the information type of the environmental condition information including some or all of environmental temperature, environmental humidity, environmental brightness, and environmental air; a task decision module for analyzing the multi-modal environment information and retrieving associated memory data according to the multi-modal environment information; and identifying a current task according to the analysis result of the multi-modal environment information and the retrieved associated memory data; a task planning module for planning an execution strategy according to the current task.