Robot control method and device, electronic equipment and computer readable storage medium
Patent Information
- Application Number
- CN202511103446.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-08-06
AI Technical Summary
在智能机器人中,往往配置有视觉语言模型,在利用视觉语言模型处理复杂机器人操作任务的时候由于缺乏中间的推理步骤,导致机器人在处理复杂任务的过程中表现不佳,影响任务的成功率
[0019]根据本申请提供的实施例的机器人控制方法,至少具有如下有益效果:在进行机器人控制的过程中,首先获取语音指令和初始观察图像;接着基于预训练的视觉语言模型中的因果注意力子模块、语音指令和初始观察图像进行视觉推理处理,就可以得到子目标图像;接着基于视觉语言模型中的全注意力子模块、语音指令、初始观察图像和子目标图像进行动作预测处理,就可以得到预测动作序列信息;接着对预测动作序列信息进行评估处理就可以得到动作评估信息;在动作评估信息表征预测动作序列信息满足预设条件的情况下,就可以将预测动作序列信息确定为最终动作序列信息。通过上述技术方案,在对机器人进行控制处理的过程中,还会首先生成子目标图像,后续再基于子目标图像生成预测动作序列信息,进而可以很好地提升机器人的任务执行的准确度。
Smart Images

Figure CN120839788B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to, but are not limited to, the field of robot control technology, and in particular to a robot control method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the continuous development of society and economy and the advancement of technology, intelligent robots are increasingly being applied to all aspects of people's lives. For example, they have been widely promoted and applied in the fields of smart healthcare and insurance. Intelligent robots are often equipped with visual language models. However, when using these models to handle complex robotic tasks, the lack of intermediate reasoning steps leads to poor robot performance and affects the success rate. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0004] To address the problems mentioned in the background section, this application provides a robot control method, apparatus, electronic device, and computer-readable storage medium that can improve the accuracy of robot task execution.
[0005] In a first aspect, embodiments of this application provide a robot control method, including:
[0006] Acquire voice commands and initial observation images;
[0007] Visual reasoning processing is performed based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image to obtain the sub-target image;
[0008] Based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image, action prediction processing is performed to obtain predicted action sequence information;
[0009] The predicted action sequence information is evaluated to obtain action evaluation information;
[0010] If the action evaluation information indicates that the predicted action sequence information meets preset conditions, the predicted action sequence information is determined as the final action sequence information.
[0011] Secondly, embodiments of this application also provide a robot control device, the device comprising:
[0012] The acquisition unit is used to acquire voice commands and initial observation images;
[0013] The inference unit is used to perform visual inference processing based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image to obtain the sub-target image;
[0014] The prediction unit is used to perform action prediction processing based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image to obtain predicted action sequence information;
[0015] An evaluation unit is used to evaluate and process the predicted action sequence information to obtain action evaluation information;
[0016] The determination unit is used to determine the predicted action sequence information as the final action sequence information when the action evaluation information characterizes the predicted action sequence information as meeting preset conditions.
[0017] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the robot control method described in the first aspect above.
[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for performing the robot control method described in the first aspect above.
[0019] The robot control method according to the embodiments provided in this application has at least the following beneficial effects: During robot control, voice commands and initial observation images are first acquired; then, visual reasoning processing is performed based on the causal attention submodule of a pre-trained visual language model, the voice commands, and the initial observation images to obtain sub-target images; next, action prediction processing is performed based on the full attention submodule of the visual language model, the voice commands, the initial observation images, and the sub-target images to obtain predicted action sequence information; then, the predicted action sequence information is evaluated to obtain action evaluation information; if the action evaluation information characterizes the predicted action sequence information and satisfies preset conditions, the predicted action sequence information can be determined as the final action sequence information. Through the above technical solution, during the robot control process, sub-target images are first generated, and then predicted action sequence information is generated based on the sub-target images, thereby significantly improving the accuracy of the robot's task execution. Attached Figure Description
[0020] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0021] Figure 1 This is a flowchart illustrating a robot control method provided in one embodiment of this application;
[0022] Figure 2 yes Figure 1 A schematic diagram of a specific implementation method of step S200;
[0023] Figure 3 yes Figure 1 A schematic diagram of a specific implementation of step S300;
[0024] Figure 4 yes Figure 1 A schematic diagram of a specific implementation of step S400;
[0025] Figure 5 Is it completed? Figure 1 A flowchart illustrating a specific implementation method following step S400;
[0026] Figure 6 Is it completed? Figure 1 A flowchart illustrating a specific implementation method following step S500;
[0027] Figure 7 Is it completed? Figure 6 A flowchart illustrating a specific implementation method following step S640;
[0028] Figure 8 This is a schematic diagram of a robot control device provided in one embodiment of this application;
[0029] Figure 9 This is a schematic diagram of an electronic device provided in one embodiment of this application. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0032] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0033] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0034] AI is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thought. Furthermore, artificial intelligence utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results—the theories, methods, technologies, and application systems available for use.
[0035] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0036] Artificial intelligence, or AI, is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0037] The servers involved in artificial intelligence technology can be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0038] This application provides a robot control method, device, electronic device, and computer-readable storage medium. During robot control, voice commands and initial observation images are first acquired. Then, visual reasoning processing is performed based on the causal attention submodule of a pre-trained visual language model, the voice commands, and the initial observation images to obtain sub-target images. Next, action prediction processing is performed based on the full attention submodule of the visual language model, the voice commands, the initial observation images, and the sub-target images to obtain predicted action sequence information. The predicted action sequence information is then evaluated to obtain action evaluation information. If the action evaluation information characterizes the predicted action sequence information and satisfies preset conditions, the predicted action sequence information can be determined as the final action sequence information. Through this technical solution, during robot control, sub-target images are generated first, and then predicted action sequence information is generated based on these sub-target images, thereby significantly improving the accuracy of robot task execution.
[0039] The robot control method provided in this application relates to the field of robot control technology. The robot control method provided in this application can be applied to a terminal or a server, and can also be software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0040] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0041] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0042] The embodiments of this application will be further described below with reference to the accompanying drawings.
[0043] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a robot control method provided in one embodiment of this application. The robot control method includes the following steps:
[0044] Step S100: Obtain voice commands and initial observation images.
[0045] The robot control method provided in this application first needs to acquire voice commands and initial observation images during the robot control process. Then, based on the causal attention submodule in the pre-trained visual language model, the voice commands, and the initial observation images, visual reasoning processing can be performed to obtain sub-target images. Subsequently, based on the sub-target images, predictive action sequence information generation processing can be performed to improve the accuracy of action sequence information generation.
[0046] It is worth noting that voice commands are the control instructions issued by the user to the robot, and the initial observation image is the image captured by the robot before it performs the relevant actions. After obtaining the voice commands and the initial observation image, visual reasoning can be performed based on the causal attention submodule in the pre-trained visual language model, the voice commands, and the initial observation image to obtain the corresponding sub-target image, which prepares for the subsequent generation of action sequence information and improves the accuracy of robot action sequence information generation.
[0047] For example, in the healthcare industry, robots can be used to disinfect hospital rooms. Medical staff can give the robot a voice command to "disinfect the room," at which point the robot will take pictures of the hospital room's interior to obtain initial observation images. Alternatively, in the insurance industry, robots can use optical character recognition and natural language processing technologies to extract, verify, and process data from complex documents such as claims forms and policy applications, reducing the workload of manual data entry and review. Therefore, insurance agents can give the robot a voice command to "recognize policy content," at which point the robot will take pictures of the policy content to obtain initial observation images.
[0048] It is worth noting that user permission or consent is obtained before any voice commands or initial image viewing. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain sensitive personal information from the user, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after explicitly obtaining the user's separate permission or consent is the necessary user-related data required for the normal operation of this application embodiment acquired.
[0049] Step S200: Perform visual reasoning processing based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observed image to obtain the sub-target image.
[0050] The robot control method provided in this application, after obtaining the voice command and the initial observation image, first performs visual reasoning processing based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image to obtain the sub-target image, which prepares for the generation of subsequent predicted action sequence information.
[0051] It's worth noting that visual language models are multimodal AI models that integrate computer vision and natural language processing, capable of simultaneously processing and understanding visual data such as images or videos, as well as text data. The causal attention submodule is a special attention mechanism primarily used to ensure that the model can only utilize current and previous information when processing sequential data, and cannot "see" future information. The causal attention submodule plays a crucial role in deep learning models, especially in tasks requiring sequential data processing, effectively improving model performance and generalization ability. Based on the causal attention submodule, spurious correlations can be removed, making the model more stable when dealing with out-of-distribution data. The causal attention submodule ensures that the model uses only relevant information during generation or understanding, thereby improving the accuracy and efficiency of attention.
[0052] like Figure 2 As shown, visual reasoning processing based on the causal attention submodule of a pre-trained visual language model, voice commands, and initial observed images to obtain sub-target images can include the following steps:
[0053] Step S210: Perform first speech recognition processing on the speech command to obtain first speech recognition information; and perform first encoding processing on the initial observed image to obtain first image feature information;
[0054] Step S220: Perform first semantic understanding processing on the first speech recognition information to obtain first speech text information;
[0055] Step S230: Based on the causal attention submodule, the first speech text information and the first image feature information are fused to obtain multimodal fusion information;
[0056] Step S240: Perform inference processing on the multimodal fusion information to obtain the sub-target image.
[0057] For steps S210 to S240, in the process of obtaining the sub-target image through visual reasoning processing based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image, the first speech recognition processing is performed on the voice command to obtain the first speech recognition information; and the first encoding processing is performed on the initial observation image to obtain the first image feature information; then the first semantic understanding processing is performed on the first speech recognition information to obtain the first speech text information; then the first speech text information and the first image feature information are fused based on the causal attention submodule to obtain multimodal fusion information; finally, the multimodal fusion information is inferred to obtain the corresponding sub-target image; based on the above technical solution, the generation of the sub-target image can be more accurate and faster.
[0058] It is worth noting that performing a first speech recognition process on a voice command yields the first speech recognition information. Speech recognition is a technology that converts human speech into text. The speech recognition process involves several steps: speech acquisition, speech preprocessing, feature extraction, acoustic model processing, language model processing, and decoding. Speech acquisition obtains the speech signal; preprocessing improves the quality of the speech signal; feature extraction extracts useful feature vectors; the acoustic model maps the feature vectors to the probability distribution of phonemes or sub-words; the speech model generates the most probable text sequence; and the decoding process combines the results of the acoustic and language models to generate the optimal text sequence. For example, in the medical industry, when a medical professional gives a robot the voice command "deliver internal medicine supplies to the internal medicine department," this command can be processed using speech recognition to obtain the corresponding speech recognition information. Similarly, in the insurance industry, when an insurance agent gives a robot the voice command "recognize and process the text in the claims materials," this command can also be processed using speech recognition to obtain the corresponding speech recognition information. In this application, the terms "first" and "second" are used only to distinguish different objects and to make the process of illustrating the embodiments clearer, and do not mean that the processing methods of the two are different.
[0059] It is worth noting that the initial image is first encoded to obtain the first image feature information. Image encoding is a core task in the field of computer vision. Its purpose is to transform image data into a more compact and meaningful representation for subsequent processing and analysis. Image encoding usually involves techniques such as feature extraction, dimensionality reduction, and compression. Image encoding can transform complex image data into a form that is easier to process and analyze, thus playing an important role in various computer vision tasks.
[0060] It is worth noting that after obtaining the first speech recognition information, it is necessary to perform first semantic understanding processing on the first speech recognition information to obtain the first speech-text information. Semantic understanding processing of speech recognition information is a crucial step in natural language processing, aiming to convert the text generated by speech recognition into machine-understandable and executable intents and slot values. Semantic understanding processing of speech recognition information can convert speech signals into machine-understandable and executable intents and slot values, thereby enabling various intelligent voice interaction applications. In the process of performing first semantic understanding processing on the first speech recognition information, the speech recognition information is first preprocessed, then the preprocessed text is semantically understood, and finally, context management processing is performed to obtain the corresponding speech-text information.
[0061] It is worth noting that by fusing the first speech text information with the first image feature information based on the causal attention submodule, multimodal fusion information can be obtained, which prepares for subsequent inference processing of multimodal fusion information.
[0062] Step S300: Perform action prediction processing based on the full attention submodule in the visual language model, voice command, initial observation image and sub-target image to obtain predicted action sequence information.
[0063] The robot control method provided in this application, after obtaining a sub-target image through visual reasoning processing based on the causal attention sub-module, voice command, and initial observation image in the pre-trained visual language model, can then perform action prediction processing based on the full attention sub-module, voice command, initial observation image, and sub-target image in the visual language model to obtain predicted action sequence information, thereby improving the accuracy of action sequence information generation.
[0064] It is worth noting that the full attention submodule is an attention mechanism module used to enhance the performance of neural network models. It is commonly used in computer vision and natural language processing tasks, enhancing the model's ability to perceive features by applying attention mechanisms in both channel and spatial dimensions. The full attention submodule combines channel attention and spatial attention.
[0065] like Figure 3 As shown, action prediction processing based on the full attention submodule in the visual language model, voice commands, initial observation images, and sub-target images to obtain predicted action sequence information can include the following steps:
[0066] Step S310: Perform second speech recognition processing on the speech command to obtain second speech recognition information; and perform second encoding processing on the initial observation image to obtain second image feature information; and perform third encoding processing on the sub-target image to obtain sub-target image feature information.
[0067] Step S320: Perform second semantic understanding processing on the second speech recognition information to obtain second speech text information;
[0068] Step S330: Align the second speech text information, the second image feature information, and the sub-target image feature information.
[0069] Step S340: Based on the full attention submodule, perform global attention calculation on the aligned second speech text information, second image feature information and sub-target image feature information to obtain global fusion features;
[0070] Step S350: Perform action reasoning processing based on global fusion features to obtain predicted action sequence information.
[0071] For steps S310 to S350, in the process of obtaining predicted action sequence information by performing action prediction processing based on the full attention submodule in the visual language model, voice command, initial observation image, and sub-target image, firstly, the voice command undergoes second speech recognition processing to obtain second speech recognition information; then, the initial observation image undergoes second encoding processing to obtain second image feature information; and then, the sub-target image undergoes third encoding processing to obtain sub-target image feature information; next, the second speech recognition information undergoes second semantic understanding processing to obtain second speech text information; then, the second speech text information, second image feature information, and sub-target image feature information are aligned; then, based on the full attention submodule, global attention calculation processing is performed on the aligned second speech text information, second image feature information, and sub-target image feature information to obtain global fusion features; finally, action reasoning processing can be performed based on the global fusion features to obtain predicted action sequence information. Through the above technical solution, visual thinking is performed before action. The model first generates a sub-target image representing the ideal future state based on the current initial observation image and language instructions. Then, the model uses this generated sub-target image as an additional condition to plan and output a more precise action sequence to achieve the goal. The two-stage design significantly improves the model's reasoning ability and task success rate through an explicit planning process.
[0072] It is worth noting that by aligning the second speech text information, the second image feature information, and the sub-target image feature information, the global attention calculation can be performed on the aligned second speech text information, the second image feature information, and the sub-target image feature information based on the full attention submodule to obtain global fusion features. Finally, action reasoning can be performed based on the global fusion features to obtain the corresponding predicted action sequence information. The whole process is stable and reliable.
[0073] Step S400: Evaluate the predicted action sequence information to obtain action evaluation information.
[0074] The robot control method provided in this application, after obtaining predicted action sequence information by performing action prediction processing based on the full attention submodule in the visual language model, voice commands, initial observation images, and sub-target images, can then evaluate the predicted action sequence information to obtain action evaluation information. Subsequently, if the action evaluation information characterizes the predicted action sequence information as meeting pre-defined conditions, the predicted action sequence information can be determined as the final action sequence information, further improving the accuracy of the final action sequence information determination and significantly enhancing the precision of robot control.
[0075] It is worth noting that during the evaluation and processing of predicted action sequence information, security evaluation, effectiveness evaluation, and efficiency evaluation can be performed on the predicted action sequence information respectively. Based on the above evaluation and processing, the evaluation of predicted action sequence information can be made more accurate.
[0076] like Figure 4 As shown, evaluating the predicted action sequence information to obtain action evaluation information may include the following steps:
[0077] Step S410: Perform a security assessment on the predicted action sequence information to obtain security assessment information;
[0078] Step S420: Evaluate the effectiveness of the predicted action sequence information to obtain effectiveness evaluation information;
[0079] Step S430: Efficiency evaluation is performed on the predicted action sequence information to obtain efficiency evaluation information;
[0080] Step S440: Determine action evaluation information based on safety assessment information, effectiveness assessment information, and efficiency assessment information.
[0081] For steps S410 to S440, in the process of evaluating the predicted action sequence information to obtain action evaluation information, the predicted action sequence information is first evaluated for safety to obtain safety evaluation information; then, the predicted action sequence information is evaluated for effectiveness to obtain effectiveness evaluation information; then, the predicted action sequence information is evaluated for efficiency to obtain efficiency evaluation information; finally, based on the safety evaluation information, effectiveness evaluation information, and efficiency evaluation information, the action evaluation information can be determined, making the evaluation of the predicted action sequence information more accurate.
[0082] It is worth noting that the predicted action sequence information is subjected to a security assessment to evaluate and detect whether there are any security risks; the predicted action sequence information is subjected to an effectiveness assessment to evaluate whether it can complete the corresponding control operation instructions; and the predicted action sequence information is subjected to an efficiency assessment to evaluate the action efficiency of the predicted action sequence information.
[0083] It is worth noting that after obtaining safety assessment information, effectiveness assessment information, and efficiency assessment information, motion assessment information can be determined based on these information, making the motion assessment information more accurate.
[0084] Step S500: If the motion evaluation information characterizes the predicted motion sequence information and meets the preset conditions, the predicted motion sequence information is determined as the final motion sequence information.
[0085] The robot control method provided in this application can determine the predicted action sequence information as the final action sequence information when the action evaluation information characterizes the predicted action sequence information to meet the preset conditions. Through the above technical solution, the accuracy of the generated action sequence information is further improved.
[0086] It is worth noting that the predicted action sequence information must be subject to a security assessment and meet the security requirements; the predicted action sequence information must be subject to an effectiveness assessment and meet the effectiveness requirements; and the predicted action sequence information must be subject to an efficiency assessment and meet the efficiency requirements. Only when all three of these requirements are met can the predicted action sequence information meet the preset conditions, and only then can the predicted action sequence information be determined as the final action sequence information.
[0087] like Figure 5 As shown, after evaluating the predicted action sequence information to obtain action evaluation information, the following steps may be included:
[0088] Step S610: If the motion evaluation information characterizes the predicted motion sequence information but does not meet the preset conditions, determine the evaluation difference information;
[0089] Step S620: Adjust the weight parameters of the visual language model based on the evaluation difference information.
[0090] For steps S610 to S620, after evaluating the predicted action sequence information to obtain action evaluation information, if the action evaluation information does not meet the preset conditions in representing the predicted action sequence information, evaluation difference information can be determined. Subsequently, the weight parameters of the visual language model can be adjusted based on the evaluation difference information so that the generated predicted action sequence information is adjusted, thereby further improving the accuracy of the generated predicted action sequence information.
[0091] It is worth noting that in the medical industry, evaluating the predicted action sequence information can yield action evaluation information. If the action evaluation information does not meet the preset conditions for representing the predicted action sequence information, the evaluation difference information can be determined. Finally, the weight parameters of the visual language model can be adjusted based on the evaluation difference information, so that the subsequently generated predicted action sequence information for the medical industry can be more accurate.
[0092] like Figure 6As shown, after determining the predicted action sequence information as the final action sequence information, provided that the action evaluation information characterizes the predicted action sequence information according to preset conditions, the following steps may also be included:
[0093] Step S630: The final action sequence information is marked to obtain action marking information;
[0094] Step S640: Store the action tag information in a preset action database.
[0095] For steps S630 to S640, after the predicted action sequence information is determined as the final action sequence information when the action evaluation information characterizes the predicted action sequence information, the final action sequence information can be marked to obtain action marking information. Then, the action marking information is stored in a pre-set action database. If the same voice command is given to the robot again, the corresponding final action sequence information can be extracted from the action database, and then the operation can be performed according to the final action sequence information. It is not necessary to re-analyze and process the voice command and the initial observation image, which improves the control operation efficiency of the robot and brings great convenience to the user.
[0096] For example, in the healthcare industry, if the previous voice command was "disinfect the internal medicine department," and the generated final action sequence information met the relevant requirements, then this final action sequence information could be tagged to obtain action tag information. This action tag information was then stored in the action database. If the robot subsequently received the command "disinfect the internal medicine department" again, it could directly retrieve the corresponding final action sequence information from the action database and perform the operation based on it. Similarly, in the insurance industry, if the previous voice command was "recognize and process the insurance bill," and the generated final action sequence information met the relevant requirements, then this final action sequence information could also be tagged to obtain action tag information. This action tag information was then stored in the action database. If the robot subsequently received the command "recognize and process the insurance bill," it could directly retrieve the corresponding final action sequence information from the action database and perform the operation based on it.
[0097] like Figure 7 As shown, after storing the action tag information into a preset action database, the following steps may be included:
[0098] Step S650: Obtain new voice commands;
[0099] Step S660: If the new voice command and the historical voice command are the same, extract the corresponding action marker information from the action database;
[0100] Step S670: Perform action execution processing based on the action marker information.
[0101] For steps S650 to S670, after storing the action marker information in the preset action database, a new voice command is obtained; and if the new voice command is the same as the historical voice command, the corresponding action marker information can be extracted from the action database; finally, the action execution processing can be performed according to the action marker information, which improves the robot's task execution efficiency and brings great convenience to the user.
[0102] It is worth noting that the action will only be executed directly based on the corresponding action marker information if the new voice command is the same as the historical voice command, which is suitable for repetitive operations.
[0103] In addition, such as Figure 8 As shown, one embodiment of this application also provides a robot control device 10, the device comprising:
[0104] Acquisition unit 100 is used to acquire voice commands and initial observation images;
[0105] The inference unit 200 is used to perform visual inference processing based on the causal attention submodule in the pre-trained visual language model, voice commands and initial observation images to obtain sub-target images;
[0106] The prediction unit 300 is used to perform action prediction processing based on the full attention submodule in the visual language model, the voice command, the initial observation image and the sub-target image to obtain the predicted action sequence information.
[0107] Evaluation unit 400 is used to evaluate and process the predicted action sequence information to obtain action evaluation information;
[0108] The determination unit 500 is used to determine the predicted action sequence information as the final action sequence information when the action evaluation information characterizes the predicted action sequence information as meeting preset conditions.
[0109] It should be noted that during robot control, the process begins with acquiring voice commands and initial observation images. Next, visual reasoning is performed based on the causal attention submodule of a pre-trained visual language model, the voice commands, and the initial observation images to obtain sub-target images. Then, action prediction is performed based on the full attention submodule of the visual language model, the voice commands, the initial observation images, and the sub-target images to obtain predicted action sequence information. This predicted action sequence information is then evaluated to obtain action evaluation information. If the action evaluation information characterizes the predicted action sequence information and meets preset conditions, the predicted action sequence information is determined as the final action sequence information. Through this technical solution, during robot control, sub-target images are generated first, and then predicted action sequence information is generated based on these sub-target images, thereby significantly improving the accuracy of the robot's task execution.
[0110] The specific implementation of the robot control device 10 is basically the same as the specific embodiment of the robot control method described above, and will not be repeated here.
[0111] In addition, such as Figure 9 As shown, one embodiment of this application also provides an electronic device 700, which includes: a memory 720, a processor 710, and a computer program stored on the memory 720 and executable on the processor 710.
[0112] The processor 710 and memory 720 can be connected via a bus or other means.
[0113] The non-transient software program and instructions required to implement the robot control method of the above embodiments are stored in the memory 720. When executed by the processor 710, the robot control method of each of the above embodiments is executed.
[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0115] Furthermore, one embodiment of this application provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor 710 or a controller, for example, by a processor 710 in the above-described device embodiment, such that the processor 710 performs the robot control method described in the above-described embodiment.
[0116] The above embodiments can be used in combination, and modules with the same name in different embodiments may be the same or different.
[0117] The foregoing has described specific embodiments of this application; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0119] The apparatus, device, computer-readable storage medium and method provided in the embodiments of this application are corresponding. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be described again here.
[0120] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0121] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91 SAM, Microchip PIC18F26K20, and Silicon Labs C8051 F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0122] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0123] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0128] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0129] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0130] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0131] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0132] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0133] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. The embodiments of this application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0134] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0135] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A robot control method, characterized in that, The method includes: Acquire voice commands and initial observation images; Visual reasoning processing is performed based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image to obtain the sub-target image; Based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image, action prediction processing is performed to obtain predicted action sequence information; The predicted action sequence information is evaluated to obtain action evaluation information; If the action evaluation information indicates that the predicted action sequence information meets the preset conditions, the predicted action sequence information is determined as the final action sequence information. The step of performing action prediction processing based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image to obtain predicted action sequence information includes: The voice command is subjected to a second speech recognition process to obtain second speech recognition information; the initial observation image is subjected to a second encoding process to obtain second image feature information; and the sub-target image is subjected to a third encoding process to obtain sub-target image feature information. The second speech recognition information is subjected to second semantic understanding processing to obtain second speech text information; The second voice text information, the second image feature information, and the sub-target image feature information are aligned. Based on the full attention submodule, global attention calculation is performed on the aligned second speech text information, the second image feature information, and the sub-target image feature information to obtain global fusion features; Based on the global fusion features, action reasoning is performed to obtain the predicted action sequence information.
2. The robot control method according to claim 1, characterized in that, The causal attention submodule in the pre-trained visual language model, the voice command, and the initial observed image undergo visual reasoning processing to obtain a sub-target image, including: The voice command is subjected to a first speech recognition process to obtain first speech recognition information; and the initial observed image is subjected to a first encoding process to obtain first image feature information. The first speech recognition information is subjected to first semantic understanding processing to obtain first speech text information; The first speech text information and the first image feature information are fused together based on the causal attention submodule to obtain multimodal fusion information; The multimodal fusion information is subjected to inference processing to obtain the sub-target image.
3. The robot control method according to claim 1, characterized in that, The evaluation process of the predicted action sequence information to obtain action evaluation information includes: A security assessment is performed on the predicted action sequence information to obtain security assessment information; The effectiveness of the predicted action sequence information is evaluated to obtain effectiveness evaluation information; The predicted action sequence information is evaluated for efficiency to obtain efficiency evaluation information; The action evaluation information is determined based on the security evaluation information, the effectiveness evaluation information, and the efficiency evaluation information.
4. The robot control method according to claim 1, characterized in that, After evaluating the predicted action sequence information to obtain action evaluation information, the method further includes: If the action evaluation information characterizes the predicted action sequence information but does not meet preset conditions, evaluation difference information is determined; The weight parameters of the visual language model are adjusted based on the evaluation difference information.
5. The robot control method according to claim 1, characterized in that, After determining the predicted action sequence information as the final action sequence information when the action evaluation information characterizes the predicted action sequence information as meeting preset conditions, the method further includes: The final action sequence information is then labeled to obtain action label information; The action tagging information is stored in a preset action database.
6. The robot control method according to claim 5, characterized in that, After storing the action tagging information into a preset action database, the method further includes: Get new voice commands; If the new voice command and the historical voice command are the same, the corresponding action marker information is extracted from the action database; The action is executed based on the action tag information.
7. A robot control device, characterized in that, The device includes: The acquisition unit is used to acquire voice commands and initial observation images; The inference unit is used to perform visual inference processing based on the causal attention submodule in the pre-trained visual language model, the voice command, and the initial observation image to obtain the sub-target image; The prediction unit is used to perform action prediction processing based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image to obtain predicted action sequence information; An evaluation unit is used to evaluate and process the predicted action sequence information to obtain action evaluation information; The determination unit is used to determine the predicted action sequence information as the final action sequence information when the action evaluation information characterizes the predicted action sequence information as meeting preset conditions. The step of performing action prediction processing based on the full attention submodule in the visual language model, the voice command, the initial observation image, and the sub-target image to obtain predicted action sequence information includes: The voice command is subjected to a second speech recognition process to obtain second speech recognition information; the initial observation image is subjected to a second encoding process to obtain second image feature information; and the sub-target image is subjected to a third encoding process to obtain sub-target image feature information. The second speech recognition information is subjected to second semantic understanding processing to obtain second speech text information; The second voice text information, the second image feature information, and the sub-target image feature information are aligned. Based on the full attention submodule, global attention calculation is performed on the aligned second speech text information, the second image feature information, and the sub-target image feature information to obtain global fusion features; Based on the global fusion features, action reasoning is performed to obtain the predicted action sequence information.
8. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the robot control method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the robot control method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Mechanical arm control method, device and equipment based on large visual model and storage medium
CN118143940A
Robot grabbing method and device combining vision and language instruction guidance
CN118721192A