Robot control method and device, electronic equipment and storage medium

By acquiring robot environment images and time stamps, using noise generation and prediction models to process instructions, the problem of robot instruction uncertainty is solved, and precise action execution and task completion is achieved.

CN120347757APending Publication Date: 2025-07-22PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510724476.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

During the execution of robot instructions, uncertainty caused by fuzzy expressions or environmental interference, it is difficult to restore to clear and executable instructions, affecting the task completion effect.

Method used

By obtaining the current image and time stamps of the robot environment, using noise generation rules and preset noise prediction models, instruction encoding, noise fusion and denoising processing are performed to generate clear execution instructions.

Benefits of technology

Eliminate uncertainty in robot instructions, ensure that the instructions are consistent with the actual state, and improve the accuracy of action execution and task completion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347757A_ABST
    Figure CN120347757A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a robot control method and device, electronic equipment and a storage medium, belongs to the technical field of robots, and is applied to financial and medical scenes. The method comprises the steps of obtaining a current environment image and a current time stamp of an environment where the robot is located; generating noise according to a preset noise generation rule to obtain current noise; performing instruction coding according to a preset target task instruction, the current environment image and the current time stamp to obtain a target instruction code; performing noise fusion on the target instruction code according to the current noise to obtain a current fusion code; performing noise prediction on the current fusion code according to a preset target noise prediction model to obtain target prediction noise; performing instruction denoising on the current fusion code according to the target prediction noise to obtain a target execution instruction; and controlling the robot to execute actions according to the target execution instruction. According to the embodiment of the invention, the robot instruction can be restored from an uncertain state to an executable clear instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot technology and is applied to financial and medical scenarios. In particular, it relates to a robot control method and device, an electronic device, and a storage medium. Background Art

[0002] When controlling a robot to execute an action, the robot is controlled to act by issuing commands to the robot. In the actual execution process, the commands may show uncertainty due to vague expressions or environmental interference. On the one hand, the same task objective may correspond to multiple language expressions, such as "put away the cup", "put the cup back", or "tidy up the table", etc.; on the other hand, unexpected situations may also occur during the execution process. For example, when executing the command "bring the water cup over from the table", the water cup falls and rolls due to the wind, resulting in the original command becoming invalid. Such problems will cause the robot command to deviate from the actual state. If the clear and executable target command cannot be restored from the uncertain state, it will directly affect the completion effect of the task. Therefore, how to restore the robot command from the uncertain state to a clear and executable command has become an urgent problem to be solved. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a robot control method and device, an electronic device, and a storage medium, aiming to restore the robot command from an uncertain state to a clear and executable command.

[0004] To achieve the above object, a first aspect of the embodiments of this application proposes a robot control method, and the method includes:

[0005] Obtain the current environmental image and the current time mark of the environment where the robot is located;

[0006] Generate noise according to the current time mark to obtain the current noise;

[0007] Perform command encoding according to a preset target task command, the current environmental image, and the current time mark to obtain a target command encoding;

[0008] Perform noise fusion on the target command encoding according to the current noise to obtain the current fusion encoding;

[0009] Perform noise prediction on the current fusion encoding according to a preset target noise prediction model to obtain a target predicted noise;

[0010] Perform command denoising on the current fusion encoding according to the target predicted noise to obtain a target execution command;

[0011] Control the robot to execute an action according to the target execution command.

[0012] In some embodiments, the target execution instruction includes at least two alternative instruction actions, and the alternative instruction actions are provided with alternative instruction sequence identifiers; controlling the robot to execute actions according to the target execution instruction includes:

[0013] Screen the alternative instruction actions according to the alternative instruction sequence identifiers to obtain the current target instruction action;

[0014] Control the robot to execute actions according to the current target instruction action;

[0015] After the robot executes the current target instruction action, obtain the next environment image and the next time stamp of the environment where the robot is located;

[0016] Generate noise randomly according to the next time stamp to obtain the next noise;

[0017] Generate and obtain the next fusion code according to the target task instruction, the next environment image, the next time stamp and the next noise;

[0018] Perform instruction denoising on the next fusion code through the target noise prediction model to obtain the next execution instruction;

[0019] Control the robot to execute actions according to the next execution instruction.

[0020] In some embodiments, encoding the instruction according to the preset target task instruction, the current environment image, and the current time stamp to obtain the target instruction code includes:

[0021] Perform text encoding on the target task instruction to obtain the target text code;

[0022] Perform image encoding on the current environment image to obtain the target image code;

[0023] Perform time step embedding on the current time stamp to obtain the target time code;

[0024] Concatenate the target text code, the target image code, and the target time code to obtain the target instruction code.

[0025] In some embodiments, performing image encoding on the current environment image to obtain the target image code includes:

[0026] Extract key features from the current environment image to obtain an environmental key feature vector;

[0027] Encode the environmental key feature vector to obtain the target image code.

[0028] In some embodiments, before obtaining the current environmental image of the robot and the current time stamp, it further includes pre-training the target noise prediction model, specifically including:

[0029] Obtain sample task instructions, sample environmental images, sample time stamps, and sample preset noises;

[0030] Generate an encoded sample by encoding according to the sample task instructions, the sample environmental image, the sample time stamp, and the sample preset noise to obtain a sample fusion code;

[0031] Perform noise prediction on the sample fusion code through a preset noise prediction model to obtain a sample predicted noise;

[0032] Train the preset noise prediction model according to the sample predicted noise and the sample preset noise to obtain the target noise prediction model.

[0033] In some embodiments, the training of the preset noise prediction model according to the sample predicted noise and the sample preset noise to obtain the target noise prediction model includes:

[0034] Calculate the mean square error according to the sample predicted noise and the sample preset noise to obtain a sample loss value;

[0035] Optimize the parameters of the preset noise prediction model according to the sample loss value to obtain the target noise prediction model.

[0036] In some embodiments, after training the preset noise prediction model according to the sample predicted noise and the sample preset noise to obtain the target noise prediction model, it further includes:

[0037] Freeze the model parameters of the target noise prediction model;

[0038] Splice a preset fine-tuning model into the target noise prediction model to obtain an original fine-tuning model;

[0039] Perform image enhancement on the sample environmental image to obtain a sample enhanced image;

[0040] Train the original fine-tuning model according to the sample task instructions, the sample enhanced image, the sample time stamp, and the sample preset noise to obtain a preliminary fine-tuning model;

[0041] Unfreeze the frozen model parameters in the preliminary fine-tuning model to obtain a target fine-tuning model, and replace the target noise prediction model with the target fine-tuning model.

[0042] To achieve the above object, a second aspect of the embodiments of the present application provides a robot control device, the device comprising:

[0043] A data acquisition module, configured to acquire a current environmental image and a current time stamp of the environment where the robot is located;

[0044] A noise generation module, configured to generate noise according to the current time stamp to obtain current noise;

[0045] An instruction encoding module, configured to perform instruction encoding according to a preset target task instruction, the current environmental image, and the current time stamp to obtain a target instruction encoding;

[0046] A noise fusion module, configured to perform noise fusion on the target instruction encoding according to the current noise to obtain a current fusion encoding;

[0047] A noise prediction module, configured to perform noise prediction on the current fusion encoding according to a preset target noise prediction model to obtain a target predicted noise;

[0048] An instruction denoising module, configured to perform instruction denoising on the current fusion encoding according to the target predicted noise to obtain a target execution instruction;

[0049] An execution action module, configured to control the robot to execute an action according to the target execution instruction.

[0050] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method described in the first aspect when executing the computer program.

[0051] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, and the computer program implementing the method described in the first aspect when executed by a processor.

[0052] A robot control method and device, an electronic device, and a storage medium proposed in this application obtain the current environmental image and the current time mark of the environment where the robot is located, and after encoding the instructions based on the preset target task instructions, the current environmental image, and the current time mark, then generate the current noise using the preset noise generation rule, and fuse the current noise into the target instruction encoding to obtain the current fusion encoding; then predict the target prediction noise in the current fusion encoding based on the preset target noise prediction model, and use the target prediction noise to denoise the current fusion encoding to obtain a clear and executable target execution instruction, and finally control the robot to execute actions according to the target execution instruction. In this way, the embodiments of this application eliminate the uncertainty in the robot instructions, make the robot instructions consistent with the actual state, avoid execution errors caused by fuzzy instructions or environmental interference, thus solving the problem that the robot instructions are uncertain and difficult to execute precisely, and improving the action execution accuracy and task completion effect of the robot. Description of the Drawings

[0053] Figure 1 is the flowchart of the robot control method provided by the embodiment of this application;

[0054] Figure 2 is the flowchart of the robot control method provided by another embodiment of this application;

[0055] Figure 3 is Figure 2 the flowchart of step S204 in

[0056] Figure 4 is the flowchart of the robot control method provided by another embodiment of this application;

[0057] Figure 5 is Figure 1 the flowchart of step S103 in

[0058] Figure 6 is Figure 5 the flowchart of step S502 in

[0059] Figure 7 is Figure 1 the flowchart of step S107 in

[0060] Figure 8 is the structural schematic diagram of the robot control device provided by the embodiment of this application;

[0061] Figure 9 is the hardware structural schematic diagram of the electronic device provided by the embodiment of this application. Detailed Embodiments

[0062] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] It should be noted that although functional module division is carried out in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the sequence in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0065] First, some terms involved in the present application are analyzed:

[0066] Noise Schedule: Noise Schedule refers to a predefined scheduling strategy in the Diffusion Model where the noise intensity varies with the time step, used to control the degree of gradually adding Gaussian noise to the data during the forward diffusion process. It is an important parameter mechanism in generative modeling and probabilistic modeling in artificial intelligence. Noise Schedule is usually given in the form of a sequence of noise variances or standard deviations corresponding to a series of time steps, used to simulate the process of the original sample gradually evolving into pure noise. Its design directly affects the smoothness of the diffusion process, training stability, and the quality and efficiency of subsequent reverse denoising generation. In multiple technical fields such as image generation, speech synthesis, medical reconstruction, and 3D modeling, a reasonable Noise Schedule can improve the model's approximation ability to the data distribution and enhance the diversity and detail fidelity of the samples generated by the model. Common types of Noise Schedule include linear, cosine, exponential, etc. forms, and researchers can also customize and optimize according to task requirements. As the core control parameter of the diffusion probability model, Noise Schedule is an important basic link to achieve high-quality generation results, improve the model training performance and steady-state convergence ability.

[0067] DINOv2: DINOv2 is a large-scale visual representation model constructed based on a self-supervised learning mechanism and is an important research achievement in the field of computer vision and representation learning in artificial intelligence. Based on the idea of contrastive learning, this model automatically learns the structural semantic features in images through a data training method without artificial labels to construct a general and high-quality visual feature representation. DINOv2 uses Vision Transformer (ViT) as the backbone network in its model structure and introduces a knowledge distillation and multi-perspective alignment mechanism, enabling the model to have stronger discriminative and generalization abilities at different scales and semantic levels. This model is widely applied in downstream tasks such as image classification, object detection, semantic segmentation, and image retrieval, especially showing significant advantages in scenarios with scarce data annotation or cross-domain transfer learning. As a pre-training framework for obtaining visual semantic representations without supervision, DINOv2 is one of the key technical paths to promote the generalization, adaptability, and cross-task application of visual understanding.

[0068] Qformer: Qformer is a neural network structure for multi-modal alignment and interaction modeling and is one of the key technologies in the field of multi-modal learning and large model pre-training in artificial intelligence. Qformer is used to establish an efficient information transfer channel between the visual encoder and the language large model. Its core mechanism is to introduce a set of learnable query vectors (Query Tokens) in the Transformer structure. By interacting with visual features, it realizes the extraction and recombination of key semantic information in images. Qformer not only has the global modeling ability of Transformer but also can achieve semantic bridging between images and texts while maintaining computational efficiency. It is widely applied in tasks such as image-text question answering, image caption generation, visual language reasoning, and multi-modal retrieval. In the construction of multi-modal large models, as an intermediate converter for visual feature compression and abstraction, Qformer significantly improves the generalization ability and cross-modal alignment quality of the model under low computing power conditions and is one of the core components to promote the modularization, lightweight, and efficiency of general visual language model structures.

[0069] When controlling a robot to execute actions, the robot's actions are controlled by issuing instructions to the robot. During the actual execution process, the instructions may exhibit uncertainty due to vague expressions or environmental interference. On the one hand, the same task objective may correspond to multiple language expressions, such as "put away the cup", "put the cup back", or "tidy up the table", etc.; on the other hand, unexpected situations may also occur during the execution process. For example, when executing the instruction "bring the water cup over from the table", the water cup falls and rolls due to the wind, resulting in the invalidation of the original instruction. Such problems will cause a deviation between the robot's instructions and the actual state. If the clear and executable target instruction cannot be restored from the uncertain state, it will directly affect the completion effect of the task. Therefore, how to restore the robot's instructions from the uncertain state to a clear and executable instruction has become an urgent problem to be solved.

[0070] Based on this, the embodiments of the present application provide a robot control method, device, electronic device, and storage medium, aiming to restore the robot's instructions from an uncertain state to a clear and executable instruction.

[0071] The robot control method, device, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the robot control method in the embodiments of the present application is described.

[0072] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0073] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0074] The robot control method provided by the embodiments of this application relates to the field of robot technology and is applied to financial and medical scenarios. The robot control method provided by the embodiments of this application can be applied to a terminal, or to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the robot control method, etc., but is not limited to the above forms.

[0075] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0076] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0077] Figure 1 is an optional flowchart of the robot control method provided by the embodiments of this application. Figure 1 The method in may include but is not limited to steps S101 to S107.

[0078] Step S101, obtain the current environmental image and the current time mark of the environment where the robot is located;

[0079] Step S102, generate noise according to the current time mark to obtain the current noise;

[0080] Step S103, perform instruction encoding according to the preset target task instruction, the current environmental image and the current time mark to obtain the target instruction encoding;

[0081] Step S104, perform noise fusion on the target instruction encoding according to the current noise to obtain the current fusion encoding;

[0082] Step S105, perform noise prediction on the current fusion encoding according to the preset target noise prediction model to obtain the target predicted noise;

[0083] Step S106, perform instruction denoising on the current fusion encoding according to the target predicted noise to obtain the target execution instruction;

[0084] Step S107, control the robot to execute actions according to the target execution instruction.

[0085] In steps S101 to S107 illustrated in the embodiments of the present application, by obtaining the current environmental image and the current time mark of the environment where the robot is located, and performing instruction encoding according to the preset target task instruction, the current environmental image and the current time mark, then generating the current noise by using the current time mark, and fusing the current noise into the target instruction encoding to obtain the current fusion encoding; subsequently, predicting the target predicted noise in the current fusion encoding based on the preset target noise prediction model, and then using the target predicted noise to perform denoising processing on the current fusion encoding to obtain a clear and executable target execution instruction, and finally controlling the robot to execute actions according to the target execution instruction. In this way, the embodiments of the present application achieve the elimination of the uncertainty existing in the robot instructions, making the robot instructions consistent with the actual state, avoiding execution errors caused by ambiguous instructions or environmental interference, thus solving the problem that the robot instructions have uncertainty and are difficult to execute precisely, and improving the action execution accuracy and task completion effect of the robot.

[0086] In step S101 of some embodiments, the current environmental image is the image information obtained in real time by a camera installed on the robot, and this image information includes visual information such as the spatial layout of the environment where the robot is located, the positions of obstacles, and the positions, appearances, and states of target objects, and is used to reflect the actual situation of the environment where the robot is currently located.

[0087] The current time marker is used to clearly indicate the degree of adding noise to the target instruction encoding, that is, which step in the multi-step denoising process the current fused encoding with added noise is in. For example, the initial state corresponds to the maximum noise intensity, and as the denoising steps progress, the noise intensity gradually decreases. Therefore, when adding noise to the target instruction encoding to form the current fused encoding in step S103, the current time marker is added simultaneously, so that in the subsequent noise prediction and instruction denoising processes, the target noise prediction model can accurately predict the noise components contained in the current fused encoding according to this clear noise intensity level, and gradually and specifically remove the corresponding noise, thereby finally obtaining a target execution instruction with a clear denoising effect and stable quality, ensuring that the robot can accurately execute actions.

[0088] In step S102 of some embodiments, the noise generation is determined based on the denoising progress corresponding to the current time marker. Specifically, the current time marker represents the current stage of the denoising process, and each stage corresponds to a preset specific noise intensity level. The noise generation process uses the current time marker to look up and determine the corresponding noise intensity level from the preset noise intensity rules, and thus generates the current noise composed of random numbers matching this stage. The current noise is represented as a random vector with a specific intensity and is used for subsequent fusion with the target instruction encoding. For example, when the current time marker is the first execution of a task, the corresponding noise level intensity is 500, and noise is added 500 times according to 500 to obtain the current noise.

[0089] Please refer to Figure 2 , in some embodiments, before step S101, the robot control method may further include but is not limited to steps S201 to S204:

[0090] Step S201, obtaining a sample task instruction, a sample environment image, a sample time marker, and a sample preset noise;

[0091] Step S202, performing encoding generation according to the sample task instruction, the sample environment image, the sample time marker, and the sample preset noise to obtain a sample fused encoding;

[0092] Step S203, performing noise prediction on the sample fused encoding through a preset noise prediction model to obtain a sample predicted noise;

[0093] Step S204, training the preset noise prediction model according to the sample predicted noise and the sample preset noise to obtain a target noise prediction model.

[0094] In the steps S201 to S204 illustrated in the embodiments of the present application, by obtaining a sample task instruction, a sample environment image, a sample time mark, and a sample preset noise, and encoding according to the above data to generate a sample fusion code, and then predicting the noise in the sample fusion code through a preset noise prediction model to obtain a sample predicted noise; subsequently, according to the difference between the sample predicted noise and the sample preset noise, the preset noise prediction model is trained to optimize and obtain a target noise prediction model. In this way, the embodiments of the present application enable the target noise prediction model to accurately learn the noise feature distribution of different task instructions under various environment images and noise levels, significantly improve the prediction accuracy of the noise components in the fusion code, and thus effectively achieve the denoising process of the task instruction in the actual environment.

[0095] In step S201 of some embodiments, the sample task instruction is, the sample environment image is, and the sample time mark is.

[0096] It should be noted that the sample preset noise is the noise constructed according to the sample task instruction, the sample environment image, and the sample time mark, and corresponds to the sample task instruction, the sample environment image, and the sample time mark.

[0097] For example, assume that the sample task instruction is a 10×M matrix after encoding. After scrambling the sample task instruction (for example, the same task objective may correspond to multiple language expressions, such as "put away the cup", "put the cup back", or "tidy up the table"), and then encoding, a 10×M matrix is also obtained. Subtracting the two matrices can obtain a 10×M noise matrix.

[0098] Similarly, for the noise matrix corresponding to the assumed picture is 32×M, and the time mark is a 1×M noise matrix.

[0099] At the same time, for example, the sample time mark is the first execution of the task, and the corresponding noise level intensity is 500. According to 500, noise is added 500 times to obtain a 1×M noise matrix.

[0100] Finally, all the noise matrices are concatenated to obtain the sample preset noise, that is, a (10 + 32 + 1 + 1)×M matrix, which is used to simulate the task instruction noise, the environment image noise, and the time mark noise existing in the actual execution process, as well as the artificially added noise.

[0101] That is to say, before step S201, it further includes:

[0102] Data augmentation is performed according to the sample task instruction, the sample environment image, and the sample time mark to obtain an enhanced task instruction, an enhanced environment image, and an enhanced time mark.

[0103] According to the enhancement task instruction, the environment image, the enhancement time mark and the sample task instruction, the sample environment image and the sample time mark are enhanced to generate noise and obtain preliminary noise.

[0104] Noise generation is performed based on the sample time stamps to obtain random noise.

[0105] The preliminary noise and random noise are concatenated together to obtain the sample preset noise.

[0106] In some embodiments, the principle of step S202 is the same as that of steps S501 to S504 below, and will not be repeated here.

[0107] In some embodiments, step S203 has the same principle as step S105 below, and will not be repeated here.

[0108] See also Figure 3 In some embodiments, step S204 may include but is not limited to steps S301 to S302:

[0109] Step S301, calculating the mean square error based on the sample prediction noise and the sample preset noise to obtain the sample loss value;

[0110] Step S302, optimizing the parameters of the preset noise prediction model according to the sample loss value to obtain a target noise prediction model.

[0111] In steps S301 to S302 shown in the embodiment of the present application, by calculating the mean square error between the sample prediction noise and the sample preset noise, a sample loss value that can accurately reflect the accuracy of the noise prediction is obtained, and the parameters of the preset noise prediction model are optimized to achieve the target noise prediction model to accurately predict the noise components in the fusion coding, effectively improve the accuracy of noise prediction and the quality of instruction denoising, and ensure that the robot can accurately execute task instructions.

[0112] In step S301 of some embodiments, the sample loss value is calculated by squaring the difference between the sample prediction noise output by the preset noise prediction model and the pre-set sample preset noise, and averaging the results of all the calculated square differences to obtain the sample loss value.

[0113] In step S302 of some embodiments, parameter optimization is implemented by taking the sample loss value as the optimization target, performing back propagation optimization on each network parameter in the preset noise prediction model, and continuously adjusting the network parameters through optimization algorithms such as gradient descent to gradually reduce the sample loss value, so that the output of the preset noise prediction model gradually approaches the actual sample preset noise, and finally obtains a target noise prediction model with higher prediction accuracy.

[0114] Please refer to Figure 4 In some embodiments, after step S204, the robot control method may further include, but is not limited to, steps S401 to S405:

[0115] Step S401, freeze the model parameters of the target noise prediction model;

[0116] Step S402, splice the preset fine-tuning model into the target noise prediction model to obtain the original fine-tuning model;

[0117] Step S403, perform image enhancement on the sample environmental image to obtain a sample enhanced image;

[0118] Step S404, train the original fine-tuning model according to the sample task instruction, the sample enhanced image, the sample time stamp, and the sample preset noise to obtain a preliminary fine-tuning model;

[0119] Step S405, unfreeze the frozen model parameters in the preliminary fine-tuning model to obtain a target fine-tuning model, and replace the target noise prediction model with the target fine-tuning model.

[0120] Steps S401 to S405 shown in the embodiments of the present application first fix the network parameters of the target noise prediction model, splice the preset fine-tuning model, and use the sample environmental image after image enhancement to train the spliced original fine-tuning model, so as to obtain a preliminary fine-tuning model. Subsequently, the frozen network parameters in the preliminary fine-tuning model are unfrozen to obtain the final target fine-tuning model, which replaces the original target noise prediction model to achieve fine adjustment of the model parameters, further improving the adaptability of the target fine-tuning model to actual environmental changes and the accuracy of noise prediction, and ensuring that the robot can execute tasks stably and reliably in a complex and changeable environment.

[0121] It should be noted that steps S401 to S405 are used in a new environment. For example, when a user brings the robot home and uses it after obtaining the robot, fine-tuning training can be quickly performed through steps S401 to S405 instead of full-scale learning, so as to improve the adaptability of the target fine-tuning model to actual environmental changes and the accuracy of noise prediction, and ensure that the robot can execute tasks stably and reliably in a complex and changeable environment.

[0122] In step S401 of some embodiments, freezing the model parameters means setting the network parameters in the target noise prediction model not to participate in training updates, that is, in subsequent training processes, these frozen network parameters remain unchanged.

[0123] In step S402 of some embodiments, the preset fine-tuning model is a pre-constructed network structure, specifically obtained by splicing and connecting with the target noise prediction model in a network layer form to obtain the original fine-tuning model, thereby forming an overall model with an additional trainable network structure.

[0124] In step S403 of some embodiments, image enhancement refers to performing various data enhancement operations on the sample environmental image, including but not limited to brightness adjustment, contrast enhancement, random rotation, scale change, or random cropping, etc., to generate a more abundant sample enhanced image and improve the generalization performance of the model under different visual conditions.

[0125] The principles of step S404 and steps S201 to S204 in some embodiments are the same and will not be elaborated here.

[0126] Please refer to Figure 5 , in some embodiments, step S103 includes but is not limited to steps S501 to S504:

[0127] Step S501, perform text encoding on the target task instruction to obtain the target text encoding;

[0128] Step S502, perform image encoding on the current environmental image to obtain the target image encoding;

[0129] Step S503, perform time step embedding on the current time stamp to obtain the target time encoding;

[0130] Step S504, splice the target text encoding, the target image encoding, and the target time encoding to obtain the target instruction encoding.

[0131] Steps S501 to S504 illustrated in the embodiments of the present application, by respectively performing encoding processing on the target task instruction, the current environmental image, and the current time stamp to extract the key feature representations in their respective modalities, and then splicing and integrating the above feature representations, thereby constructing the target instruction encoding that fuses language, vision, and time information.

[0132] In step S501 of some embodiments, the text encoding includes:

[0133] Obtain the corresponding action data according to the target task data.

[0134] Perform instruction encoding on the target task instruction through a preset text encoder.

[0135] Perform action encoding on the action data through a preset action encoder.

[0136] Splice the encoding of the target task instruction and the encoding of the action data together to obtain the target text encoding.

[0137] Instruction encoding refers to inputting the target task instruction into a text encoder constructed based on a language modeling mechanism, parsing the lexical structure and semantic meaning of the task statement, and converting it into a vector representation with a fixed-length dimension, that is, instruction encoding. In one embodiment, the text encoder is CLIP.

[0138] Action encoding refers to inputting action data into an action encoder, where the action encoder is a NoiseSchedule network.

[0139] The reason for introducing action data is that it is difficult to fully express the specific action details and dynamic execution characteristics involved in the process of a robot executing a task only relying on the text information of the target task instruction. Especially in the scenario of task descriptions with ambiguous or polysemous expressions, the instruction itself cannot precisely define the corresponding operation sequence. As a historical execution behavior record matching the target task instruction, action data can supplement the task intention at the semantic level and provide specific motion trajectories or control elements at the operation level. By encoding the action data and splicing it with the instruction encoding, a target text encoding with both language semantics and motion characteristics can be constructed, thereby improving the subsequent model's understanding ability and generation accuracy of task instructions, and enhancing the expression integrity and discrimination ability of instruction encoding in a multi-modal environment.

[0140] Please refer to Figure 6 , in some embodiments, step S502 includes but is not limited to steps S601 to S602:

[0141] Step S601, extracting key features from the current environmental image to obtain an environmental key feature vector;

[0142] Step S602, encoding the environmental key feature vector to obtain a target image encoding.

[0143] Steps S601 to S602 illustrated in the embodiments of the present application, by first extracting the key visual information in the current environmental image and then performing feature encoding on the extraction result to obtain a more recognizable and structured target image encoding, helps to ensure that the image features have a clear environmental orientation during the fusion process and improves the model's recognition ability of the environmental state.

[0144] In step S601 of some embodiments, the key feature extraction process is to perform forward processing on the current environmental image based on a trained image perception model, identify the operationally significant regions, the positions, morphological features, and their relative layouts of the target objects from the image, and compress this information into a feature vector as the environmental key feature vector. In one embodiment, the image perception model is DINOv2.

[0145] In step S602 of some embodiments, the encoding process is to input the environmental key feature vector into the image encoding module for feature reconstruction and structural compression, so as to generate a target image encoding with a fixed dimension. The generated target image encoding can represent the core state information of the current environment and is used for fusion with other modality data. In one embodiment, the image encoding module is Qformer.

[0146] In step S503 of some embodiments, the time step embedding process is to use the current time stamp as the input signal, and through the sine-cosine embedding function, obtain a target time encoding with the ability to express temporal information. This time encoding is used to identify the corresponding denoising stage in the current noise fusion process, so as to provide indication information on the denoising intensity for subsequent noise prediction.

[0147] In step S504 of some embodiments, the concatenation process is to concatenate the target text encoding, the target image encoding, and the target time encoding in sequence to obtain the target instruction encoding.

[0148] In step S104 of some embodiments, the current fusion encoding is obtained by concatenating the current noise and the target instruction encoding. Specifically, the target instruction encoding is generated by encoding the preset target task instruction, the current environmental image, and the current time stamp, and the current noise is composed of a random vector determined by the current time stamp. Both the target instruction encoding and the current noise are represented in vector form, and by directly concatenating the two along the feature dimension, the current fusion encoding containing the target task instruction feature, the environmental visual feature, the time feature, and the specific intensity noise feature is formed.

[0149] In step S105 of some embodiments, the target noise prediction model is constructed using a Transformer network. Specifically, this Transformer network includes multiple sequentially stacked network layers, and feature extraction and prediction analysis are performed on the input current fusion encoding through the multi-head attention mechanism. After receiving the current fusion encoding containing the task instruction, the environmental image, the time stamp, and the noise feature, the target noise prediction model gradually predicts the noise components existing in the current fusion encoding through each network layer in the model and outputs the corresponding target predicted noise in the current fusion encoding.

[0150] In step S106 of some embodiments, the instruction denoising process performs denoising processing on the current fusion encoding according to the target predicted noise, so as to recover the target instruction encoding from the current fusion encoding, and further restore the target execution instruction according to the target instruction encoding. Specifically, the current fusion encoding is an encoding vector formed by splicing and fusing the target instruction encoding and the current noise. Assuming that the dimension of the current fusion encoding is M×N, where the current noise in the current fusion encoding is a 1×N vector. The dimension of the target predicted noise should also be M×N. The first (M - 1)×N is the noise prediction for commands, pictures, and time steps, which is used to simulate the possible noise in the environment and instructions. The last 1×N is the prediction of the artificially added noise. After subtraction, the noise existing in the environment and instructions, as well as the artificially added noise, are removed together to obtain a subtraction vector. Then, the subtraction vector is input into a preset decoder to convert the subtraction vector that combines environmental information and task instructions into a clear instruction.

[0151] Please refer to Figure 7 , in some embodiments, the target execution instruction includes at least two alternative instruction actions, and the alternative instruction actions are provided with alternative instruction sequence identifiers. Step S107 may include but is not limited to steps S701 to S707:

[0152] Step S701, screen the alternative instruction actions according to the alternative instruction sequence identifiers to obtain the current target instruction action;

[0153] Step S702, control the robot to execute the action according to the current target instruction action;

[0154] Step S703, after the robot executes the current target instruction action, obtain the next environmental image and the next time mark of the environment where the robot is located;

[0155] Step S704, randomly generate noise according to the next time mark to obtain the next noise;

[0156] Step S705, generate encoding according to the target task instruction, the next environmental image, the next time mark, and the next noise to obtain the next fusion encoding;

[0157] Step S706, perform instruction denoising on the next fusion encoding through the target noise prediction model to obtain the next execution instruction;

[0158] Step S707, control the robot to execute the action according to the next execution instruction.

[0159] In the steps S701 to S707 illustrated in the embodiments of the present application, by sequentially screening multiple alternative instruction actions and controlling the robot to perform actual operations based on the instruction actions, after the actions are executed, the updated environmental image and time stamp are synchronously obtained, and a new fusion code is regenerated in combination with the newly generated noise and denoised, so as to obtain a new execution instruction, and the next control is performed accordingly, realizing a closed-loop process of instruction - perception - execution. Thus, by iteratively repeating the above process, the embodiments of the present application enable the robot to continuously receive, update, and adjust task instructions under the condition of dynamic environmental changes, ensuring the coherence and accuracy of action execution.

[0160] In step S701 of some embodiments, the alternative instruction actions refer to a set of possible instruction action sequences generated according to the current fusion code, and the alternative instruction sequence identifier is the action sequence information used to indicate the priority selection and actual execution in the current execution cycle. Screening the alternative instruction actions based on the alternative instruction sequence identifier includes: sequentially screening to obtain the first N actions and determining the target instruction action to be executed in the current stage.

[0161] The principles of steps S702 to S707 in some embodiments are similar to those of steps S101 to S107, and will not be elaborated here.

[0162] It should be noted that the reason for continuously refreshing the instructions and controlling the robot to perform actions through the next execution instruction is that the environment in which the robot is actually executing has dynamic changes. For example, the position offset of the operation object, the appearance of obstacles, or the change of lighting conditions may all cause the original task instructions to no longer be fully applicable. By re - obtaining the environmental image and time information after executing a certain number of instruction actions and updating the fusion code and execution instructions accordingly, it can ensure that each round of instructions is highly consistent with the current environmental state, thereby effectively improving the stability and accuracy of the robot's task execution and enhancing its adaptability to dynamic environments.

[0163] For example: after generating 16 actions, the robot first performs the first few (such as the first 4), then looks at the newly taken picture (the environment may have changed), updates the data, generates a new 16 actions, and repeats this process (3 times per second, at a frequency of 3 Hz) until the task is completed (such as "pour the coffee beans into the bowl").

[0164] Please refer to Figure 8 , the embodiments of the present application further provide a robot control device, which can implement the above - mentioned robot control method. The device includes:

[0165] An acquisition data module 801, configured to acquire the current environmental image and the current time stamp of the environment where the robot is located;

[0166] A noise generation module 802 is configured to generate noise according to a current time stamp to obtain current noise;

[0167] An instruction encoding module 803 is configured to perform instruction encoding according to a preset target task instruction, a current environment image, and a current time stamp to obtain a target instruction encoding;

[0168] A noise fusion module 804 is configured to perform noise fusion on the target instruction encoding according to the current noise to obtain a current fusion encoding;

[0169] A noise prediction module 805 is configured to perform noise prediction on the current fusion encoding according to a preset target noise prediction model to obtain a target predicted noise;

[0170] An instruction denoising module 806 is configured to perform instruction denoising on the current fusion encoding according to the target predicted noise to obtain a target execution instruction;

[0171] An execution action module 807 is configured to control a robot to execute an action according to the target execution instruction.

[0172] The specific implementation manner of this robot control device is basically the same as the specific embodiments of the above-mentioned robot control method, and will not be elaborated herein.

[0173] An embodiment of this application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned robot control method is implemented. The electronic device may be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0174] Please refer to Figure 9 , Figure 9 , which shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0175] A processor 901, which may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of this application;

[0176] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the robot control method of the embodiments of this application;

[0177] The input / output interface 903 is used to implement information input and output;

[0178] The communication interface 904 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0179] The bus 905 transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0180] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0181] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned robot control method is implemented.

[0182] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0183] The robot control method, robot control device, electronic device, and storage medium provided by the embodiments of the present application obtain the current environmental image and the current time mark of the environment where the robot is located, perform instruction encoding based on the preset target task instruction, the current environmental image, and the current time mark, then generate the current noise using the preset noise generation rule, and fuse the current noise into the target instruction encoding to obtain the current fusion encoding. Subsequently, based on the preset target noise prediction model, predict the target prediction noise in the current fusion encoding, and then use the target prediction noise to denoise the current fusion encoding to obtain a clear and executable target execution instruction. Finally, control the robot to execute actions according to the target execution instruction. In this way, the embodiments of the present application eliminate the uncertainty in the robot instructions, make the robot instructions consistent with the actual state, avoid execution errors caused by fuzzy instructions or environmental interference, thereby solving the problem that the robot instructions have uncertainty and are difficult to execute precisely, and improving the action execution accuracy and task completion effect of the robot.

[0184] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0185] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0188] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0189] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0190] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0191] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, can exist separately physically for each unit, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0194] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A robot control method, characterized in that, The method includes: Obtaining the current environmental image and the current time mark of the environment where the robot is located; Generating noise according to the current time mark to obtain the current noise; Performing instruction encoding according to a preset target task instruction, the current environmental image, and the current time mark to obtain a target instruction encoding; Performing noise fusion on the target instruction encoding according to the current noise to obtain a current fusion encoding; Performing noise prediction on the current fusion encoding according to a preset target noise prediction model to obtain a target predicted noise; Performing instruction denoising on the current fusion encoding according to the target predicted noise to obtain a target execution instruction; Controlling the robot to perform actions according to the target execution instruction.

2. The method according to claim 1, wherein The target execution instruction includes at least two alternative instruction actions, and the alternative instruction actions are provided with alternative instruction sequence identifiers; the controlling the robot to perform actions according to the target execution instruction includes: Screening the alternative instruction actions according to the alternative instruction sequence identifiers to obtain a current target instruction action; Controlling the robot to perform actions according to the current target instruction action; After the robot executes the current target instruction action, obtaining the next environmental image and the next time mark of the environment where the robot is located; Randomly generating noise according to the next time mark to obtain the next noise; Performing encoding generation according to the target task instruction, the next environmental image, the next time mark, and the next noise to obtain a next fusion encoding; Performing instruction denoising on the next fusion encoding through the target noise prediction model to obtain a next execution instruction; Controlling the robot to perform actions according to the next execution instruction.

3. The method according to claim 1, wherein The performing instruction encoding according to a preset target task instruction, the current environmental image, and the current time mark to obtain a target instruction encoding includes: Performing text encoding on the target task instruction to obtain a target text encoding; Performing image encoding on the current environmental image to obtain a target image encoding; Performing time step embedding on the current time mark to obtain a target time encoding; Concatenating the target text encoding, the target image encoding, and the target time encoding to obtain the target instruction encoding.

4. The method according to claim 3, characterized in that, The performing image encoding on the current environmental image to obtain a target image encoding includes: Performing key feature extraction on the current environmental image to obtain an environmental key feature vector; Encoding the environmental key feature vector to obtain the target image encoding.

5. The method according to any one of claims 1 to 4, characterized in that, Before obtaining the current environmental image and the current time mark of the environment where the robot is located, it further includes pre-training the target noise prediction model, specifically including: Obtaining a sample task instruction, a sample environmental image, a sample time mark, and a sample preset noise; Performing encoding generation according to the sample task instruction, the sample environmental image, the sample time mark, and the sample preset noise to obtain a sample fusion encoding; Performing noise prediction on the sample fusion encoding through a preset noise prediction model to obtain a sample predicted noise; Train the preset noise prediction model based on the predicted noise of the sample and the preset noise of the sample to obtain the target noise prediction model.

6. The method according to claim 5, wherein The training of the preset noise prediction model based on the predicted noise of the sample and the preset noise of the sample to obtain the target noise prediction model includes: Calculate the mean square error based on the predicted noise of the sample and the preset noise of the sample to obtain the sample loss value; Optimize the parameters of the preset noise prediction model according to the sample loss value to obtain the target noise prediction model.

7. The method according to claim 5, wherein After training the preset noise prediction model based on the predicted noise of the sample and the preset noise of the sample to obtain the target noise prediction model, it further includes: Freeze the model parameters of the target noise prediction model; Concatenate a preset fine-tuning model to the target noise prediction model to obtain an original fine-tuning model; Enhance the sample environment image to obtain a sample enhanced image; Train the original fine-tuning model based on the sample task instruction, the sample enhanced image, the sample time stamp, and the preset noise of the sample to obtain a preliminary fine-tuning model; Unfreeze the frozen model parameters in the preliminary fine-tuning model to obtain a target fine-tuning model, and replace the target noise prediction model with the target fine-tuning model.

8. A robot control device, characterized in that, The device includes: A data acquisition module for acquiring the current environment image and the current time stamp of the environment where the robot is located; A noise generation module for generating current noise according to the current time stamp; An instruction encoding module for encoding an instruction according to a preset target task instruction, the current environment image, and the current time stamp to obtain a target instruction encoding; A noise fusion module for fusing noise to the target instruction encoding according to the current noise to obtain a current fused encoding; A noise prediction module for predicting noise of the current fused encoding according to a preset target noise prediction model to obtain a target predicted noise; An instruction denoising module for denoising the current fused encoding according to the target predicted noise to obtain a target execution instruction; An action execution module for controlling the robot to execute an action according to the target execution instruction.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the robot control method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the robot control method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Semantic constraint instant distillation humanoid robot control method based on large language model and related equipment

    CN121340233A

  • A semantic constraint instant distillation humanoid robot control method based on a large language model and related equipment

    CN121340233B

  • Medical robot multi-mode interaction control method and system based on environment perception

    CN122033948A