Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

230 results about "Action model" patented technology

Robot control method, system and equipment based on multi-modal large model and medium

The invention relates to the technical field of robot control, and discloses a robot control method, system, equipment and medium based on a multi-modal large model, and the method comprises the steps: collecting the multi-source modal data of a scene where an operation task is located, and carrying out the processing through a machine learning model, obtaining a multi-modal feature, and carrying out the position coding and Transform fusion processing, multi-modal fusion features are obtained, the multi-modal fusion features and the constructed job task knowledge base are input into a large language model to decompose a target job task, a human-in-the-loop mechanism is introduced to optimize a decomposition result, and a sub-task sequence is obtained; according to a subtask type in the subtask sequence, processing the subtask sequence through a visual language action model or a reinforcement learning model, and generating a motion instruction to enable the robot to start an execution process of the target operation task; live-line work tasks are processed through the multi-modal large models LLM, VLA and the like, and the work efficiency of the autonomous distribution network live-line work robot is improved.
Owner:WENZHOU ELECTRIC POWER BUREAU +2

Bipedal action model for humanoid robot

The present disclosure provides a system for generating motor control commands for a humanoid robot, comprising an alpha model with over 1 billion parameters that processes visual observations and language instructions at a first frequency to generate contextual embeddings, and a beta model operating at a higher second frequency. The beta model includes an embodiment-specific state encoder projecting robot state information into a shared embedding space, a diffusion transformer module generating denoised action sequences through iterative flow-matching that cross-attends to the alpha model's contextual embeddings, and an embodiment-specific action decoder converting denoised sequences into motor control commands. The beta model generates action chunks comprising future action sequences over a predetermined time horizon in a single inference step, with the complete system having less than 5 billion parameters.
Owner:FIGURE AI INC

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot comprising a torso having an alpha model deployed on a first GPU, and wherein said alpha model includes a first number of parameters and is configured to receive a natural language command from a human and generate processed data, a beta model deployed on a second GPU, and wherein said beta model includes a second number of parameters and is configured to receive the processed data from the alpha model and provide output data used to control an extent of the left wrist, and wherein the first number of parameters is larger than the second number of parameters, and a unified training framework is used to jointly train the alpha model and the beta model.
Owner:FIGURE AI INC

Three-dimensional live-action model generation method for city updating

The invention relates to a three-dimensional live-action model generation method for city updating. The method comprises the following steps: collecting multi-source city image information, and obtaining semantic data according to the multi-source city image information; performing city modeling visualization according to the multi-source city image information to generate a city model; and matching the semantic data with the city model to obtain a three-dimensional real scene model containing the semantic data. According to the method, city information can be obtained more comprehensively by performing total factor accurate collection through a multi-source three-dimensional monitoring technology, the accuracy of a real scene model is improved, semantic data is matched with a city model, a three-dimensional real scene model containing the semantic data is obtained, multiple materials and textures of the model are combined, and the accuracy of the real scene model is improved. And the model is lightened, so that the purpose of reducing the resource loading time is achieved.
Owner:GUANGDONG URBAN & RURAL PLANNING & DESIGN INST

Human-computer interaction method and system based on vision-language-action model

The invention discloses a human-computer interaction method and system based on a vision-language-action model, and belongs to the field of human-computer interaction. According to the method, an anchoring ring strategy is adopted to collect human teaching data to finely adjust the VLA model, and then the trained model is applied to an actual human-computer interaction scene. In the data acquisition stage, a first operator guides a master robot to execute task actions, and a slave robot synchronously moves and interacts with a second operator, and returns to a predefined initial position after each interaction; a teaching sample is formed by recording a robot state, an environment image and an instruction text, and a high-quality data set is generated through data enhancement. In the application stage, the real-time robot state, the environment image and the instruction text serve as input, an action instruction is generated through the VLA model, and the robot is driven to complete a cooperation task. According to the method, the data utilization efficiency and the model generalization ability are remarkably improved, the difference between simulation and the real environment is effectively overcome, and efficient, safe and natural man-machine cooperation is achieved.
Owner:ZHEJIANG UNIV

Method, device, equipment and product for controlling robot to execute task

The invention relates to a method, a device, equipment and a program product for controlling a robot to execute tasks. The method includes acquiring a user input for a robot and an image captured by a camera of the robot. The method includes generating a sequence of actions for the robot based on the user input and the image by a visual language action model trained by reinforcement learning for a world model. In addition, the method further comprises the step of controlling the robot to execute the task corresponding to the user input by executing the action sequence.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Interactive interface task automation utilizing generative artificial intelligence (AI) action models improved with retrieval-augmented generation (RAG)

PendingUS20260037318A1Mathematical modelsResource allocationRequest - actionEngineering
This disclosure describes a framework for performing user-requested tasks automatically across an interactive interface using various types of machine learning models. Specifically, this disclosure outlines and describes a task execution system that utilizes a generative artificial intelligence (AI) action model and retrieval-augmented generation (RAG) to complete user-requested actions across an interactive interface. The task execution system solves many of the current limitations of LAMs by using a generative AI action model to determine a session plan, which includes a set of actions for accomplishing stages of the actionable task across the interactive interface, obtaining visual context information of each interactive interface segment, integrates RAG results to improve the accuracy of both the session plan and individual actions, and self-corrects when faced with unexpected obstacles.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Robot action generation method and device

The invention provides a robot action generation method and device, and the method comprises the steps: responding to a target task received by a robot, and obtaining image data and text description information associated with the target task; on the basis of the image data and the text description information, potential action information and fused visual representation information are determined by utilizing a large language model obtained by supervised training provided based on a potential action model; compressing and splicing the image data, the potential action information and the fused visual representation information to obtain control sequence information; and de-noising and splicing the state information and the control sequence information corresponding to the robot to generate action sequence information, so that the robot executes a corresponding action based on the action sequence information. By means of the method, the smoothness and coherence of actions generated by the robot are improved.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

Experiment task execution method and device

The invention provides an experiment task execution method and device.The method comprises the steps that task information of a target experiment task is divided based on a visual language model, a subtask sequence is obtained, and the subtask sequence is formed by arranging multiple subtasks from front to back according to the execution sequence; processing the text information corresponding to the current sub-task and the experiment image before the current sub-task is executed through the visual language model from the first sub-task in the sub-task sequence to obtain a visual prompt image, and executing the visual prompt image through a visual language action model. And processing based on the text information, the experiment image and the visual prompt image, guiding a robot to execute the current sub-task until the last sub-task in the sub-task sequence is completed, and determining that the target experiment task is completed. According to the method, the success rate, the operation safety and the regulation compliance of the experiment task are effectively improved, and the method has high universality and safety.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Robot control method based on visual language action model and related equipment thereof

The invention provides a robot control method based on a visual language action model and related equipment thereof, and relates to the technical field of intelligent robots, and the method comprises the steps: obtaining visual observation information, a target language instruction and robot body state information of a target robot; reasoning the visual observation information, the target language instruction and the robot body state information through a pre-trained visual language action model to obtain a predicted low-dimensional potential vector; reconstructing the predicted low-dimensional potential vector into a robot action instruction through a pre-trained high-dimensional action reconstruction module; wherein the dimension of the robot action instruction is matched with the action space dimension of the target robot; and the target robot is driven to execute the high-dimensional robot action instruction. According to the method, the limitation of the existing visual language action model on the action output dimension can be solved, so that the high-degree-of-freedom robot is effectively controlled.
Owner:PAXINI TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal interaction method and interaction system applied to intelligent robot

The invention discloses a multi-modal interaction method and interaction system applied to an intelligent robot. The method comprises the following steps: collecting a scene image and processing the scene image into a three-dimensional point cloud and a two-dimensional texture feature; the Gemini Robotics-ER model is used for extracting features, and the vision-language-action model is used for analyzing a language instruction into a sequence capable of being recognized by a machine; and fusing the features to generate an interactive decision matrix, planning a trajectory, calculating kinetic parameters, driving the robot to execute actions and feeding back in real time. The system comprises a multispectral visual information acquisition and preprocessing unit, a Gemini Robotics-ER model processing unit, a natural language instruction analysis unit, a vision-language-action cooperative processing unit, a trajectory planning and dynamics calculation unit and a motion control and feedback unit, and all the units work cooperatively. According to the method and the system, through multi-modal fusion and closed-loop control, interaction accuracy and real-time performance are improved, and industrial scene requirements are met.
Owner:ZHENGXIN (SUZHOU) TECHNOLOGY CO LTD

Action model learning system and method based on embedded AI

The invention belongs to the technical field of artificial intelligence, embedded systems and action confrontation command, and discloses an action model learning system and method based on embedded AI. The system comprises an information collection module used for collecting historical action information and carrying out data preprocessing, and the historical action information comprises action environment data, confrontation situation information and decision data; the feature extraction and model construction module is used for extracting features from the preprocessed historical action information by using a deep learning algorithm to obtain action dynamic change features and enemy action mode features, and matching the action dynamic change features and the enemy action mode features with corresponding decision data results to form a training set; training the embedded AI model by using the training set to generate an action model; according to the invention, dynamic data from actions can be processed and learned in real time, so that the action model is continuously optimized, and the command decision-making capability is improved.
Owner:CHINESE PEOPLES LIBERATION ARMY AVIATION COLLEGE

Method for adjusting inertial sensing range and sensitivity

ActiveCN120831132AMeasurement devicesProject completionControl engineering
The invention provides an inertial sensing range and sensitivity adjusting method, which belongs to the technical field of parameter adjustment, and comprises the following steps: after an item opening instruction is obtained, obtaining basic information of a moving body, and calling an adaptive sensing range according to the basic information; different detection nodes are set in the whole process before a project is started, the feedback degree of a moving body is recorded in sequence when execution is carried out to each node, values of corresponding positions of an internal curve are increased or decreased according to the feedback degree, and corresponding sensitivity is dynamically set according to the internal curve. The feedback speed and external performance of the moving body are judged in the project execution process, and the follow-up content and links of the project are adjusted in combination with the action model library till the project is completed or closed. The inertial sensing range and sensitivity adjusting method provided by the invention can carry out specific dynamic adjustment according to the state of the moving body and the actual situation, has strong pertinence and ensures smooth completion of a project.
Owner:MT MICROSYST

Method for training action expert model, robot control method and electronic equipment

The embodiment of the invention provides a method for training an action expert model, a robot control method and electronic equipment, and relates to the technical field of artificial intelligence and robots. A first action expert model used for predicting motion direction information when the robot executes the to-be-processed task and a second action expert model used for predicting action detail information when the robot executes the to-be-processed task are trained; the training process of the first action expert model and the second action expert model is layered; moreover, the second action expert model can be trained by means of the partial training data of the first action expert model, so that the first action expert model can be utilized to provide directional guidance for the second action expert model to generate an executable action sequence, and thus, the operation efficiency is improved. The accuracy of generating an executable action sequence by a visual language action model can be improved, and then the accuracy of executing an operation task by a robot is improved.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot comprising a torso having an alpha model deployed on a first GPU, and wherein said alpha model includes a first number of parameters and is configured to receive a natural language command from a human and generate processed data, a beta model deployed on a second GPU, and wherein said beta model includes a second number of parameters and is configured to receive the processed data from the alpha model and provide output data used to control an extent of the left wrist, and wherein the first number of parameters is larger than the second number of parameters, and a unified training framework is used to jointly train the alpha model and the beta model.
Owner:FIGURE AI INC

Visual chain-of-thought reasoning for robot vision-language-action models

Apparatuses, systems, and techniques are disclosed for controlling a robot to execute a task. In at least one embodiment, a current image of the robot in an environment and a text describing the task are obtained. A future image of the robot in the environment is predicted based on the current image and the text. Subsequently, one or more actions are predicted based on the current image, the future image, and the text. The one or more actions can move the robot from a first state corresponding to the current image to a second state corresponding to the future image. The robot executes the sequence of actions to move in the environment.
Owner:NVIDIA CORP

Lightweight method and system for highway engineering three-dimensional live-action model and medium

The invention discloses a lightweight method and system for a highway engineering three-dimensional live-action model and a medium, and relates to the technical field of three-dimensional live-action modeling. The method comprises the following steps: firstly, carrying out engineering semantic segmentation on an input three-dimensional grid model, and extracting design parameters in a BIM / GIS (Basic Information Model / Geographic Information System); then, constructing a multi-level constraint model fusing semantic constraints, engineering precision constraints and design parameter constraints; and finally, carrying out iterative simplification under the guidance and limitation of the multi-level constraint model by adopting an improved quadratic error measurement algorithm. The method solves the problems of semantic information loss, out-of-control local precision and disjunction with design intention when a traditional pure geometric simplification method is applied to highway engineering, and can intelligently reserve key engineering characteristics, control geometric errors and fit a design form while ensuring that the data size of the model is greatly reduced. And a high-quality lightweight model suitable for professional analysis and application is generated.
Owner:SICHUAN HIGHWAY PLANNING SURVEY DESIGN AND RESEARCH INSTITUTE LTD

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot system comprising a mechanical structure with at least 30 degrees of freedom across torso, arms, and legs, actuators driving the degrees of freedom, sensors including cameras and proprioceptive sensors, and a computing system implementing a hierarchical bipedal action model (BAM). The BAM includes: a Delta model processing sensor data and user input to generate latent representations at a first frequency; a Gamma model receiving latent representations to generate human task actions at a higher second frequency; a Beta model translating task actions into joint configurations at a higher third frequency; and an Alpha model converting joint configurations into actuator control signals at a higher fourth frequency.
Owner:FIGURE AI INC

Multi-sensor fusion-based badminton player action posture analysis system and method

PendingCN121838273AImage enhancementImage analysisCentre of pressureSimulation
The invention discloses a badminton player action posture analysis system and method based on multi-sensor fusion, and relates to the technical field of athletic training auxiliary systems.The system comprises a wearable sensing subsystem, an environment sensing subsystem and a central processing and feedback subsystem; the wearable sensing subsystem is arranged on an inertia measurement unit and a pressure sensing insole of a body to collect movement inertia and plantar pressure data; the environment perception subsystem collects global videos through multiple cameras. The central processing and feedback subsystem performs synchronous processing on multi-source data, adopts a hierarchical fusion algorithm, restrains inertia integral drift by utilizing a plantar contact state, reconstructs a three-dimensional skeleton posture and a motion trail in combination with visual key points and skeleton restraint, drives a digital twin model, compares with a standard motion model, and performs three-dimensional motion control on the three-dimensional skeleton posture and the motion trail. Quantitative evaluation results such as joint angle deviation, time sequence difference and pressure center track are generated, and visual feedback is performed through a display device and an augmented reality terminal for badminton training evaluation and technical deviation correction.
Owner:GUIZHOU UNIV

Emotion service robot system based on multi-mode mental theory

The invention relates to the technical field of artificial intelligence, intelligent and robot control, and discloses an emotion service robot system based on a multi-mode mental theory, which comprises a hierarchical mental reasoning module, a connection module and an action semantic execution module, the hierarchical mental reasoning module is used for generating a decision text containing a high-level strategy according to the multi-modal environment information collected by the robot; the connection module is used for constructing a semantic instruction according to the decision text generated by the hierarchical mental reasoning module; and the action semantic execution module is used for generating a control action of the robot according to the semantic instruction constructed by the connection module and the real-time image acquired by the robot. According to the method, a control architecture fusing a vision-language-action model and a robot center view angle mental reasoning hierarchy is constructed, so that the robot can carry out multi-order belief reasoning and implicit target inference from the view angle of the robot. According to the system, the robot is no longer a pure instruction follower and has an active service capability.
Owner:JILIN UNIVERSITY

Model training method, storage medium, electronic equipment and program product

The invention provides a model training method, a storage medium, electronic equipment and a program product, and relates to the technical field of deep learning. The method is used for training a visual language action model, and the model comprises a visual language module and an action expert module. The model training method comprises the following steps: acquiring action training data, wherein the action training data comprises an environment image, a task instruction and a real action sequence; a visual language module is used for processing the environment image and the task instruction to obtain a first fusion feature vector, and the first fusion feature vector comprises visual features and language features; applying an attention mask to the target visual features in the first fusion feature vector by using an action expert module to obtain a second fusion feature vector, and generating a predicted action sequence based on the second fusion feature vector; and adjusting parameters of the visual language action model based on the difference between the predicted action sequence and the real action sequence.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Vision-language-action model training method and system utilizing inference data closed-loop optimization

The invention discloses a vision-language-action model training method and system utilizing inference data closed-loop optimization, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, performing initial training on a vision-language-action model by utilizing a training data set; deploying the trained model in a task environment to execute a task, monitoring a task execution result in real time, and capturing and structurally recording failure track data of the current task when task execution failure is recognized; then, based on the captured failure trajectory data, generating a negative prompt text for guiding the model to avoid repeated error behaviors; and finally, combining failure trajectory data with the generated negative prompt text, retraining the vision-language-action model, and inhibiting the model from generating action output similar to the error action sequence. According to the method, the data utilization rate is greatly improved, the data dependence and acquisition cost are effectively reduced, and the complete closed loop of the VLA model training and reasoning process is realized.
Owner:ZHEJIANG UNIV

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot system comprising a mechanical structure including a torso, two arms, and two legs providing at least 30 degrees of freedom, actuators coupled to the degrees of freedom, a sensor suite comprising at least one camera and proprioceptive sensors including joint encoders and an inertial measurement unit, a computing system comprising at least one processor and memory storing instructions which, when executed, implement a hierarchical bipedal action model including a Beta model configured to receive multimodal input data and generate a token sequence indicative of task intent and environmental state, and an Alpha model configured to condition on the token sequence and current robot pose data to output continuous action chunks comprising sequences of future target joint states over a finite horizon, and a low-level controller configured to convert the continuous action chunks into actuator control signals for execution.
Owner:FIGURE AI INC

Training and use of a bipedal action model for humanoid robot

The present disclosure provides a method for controlling a humanoid robot using a hierarchical bipedal action model (BAM), the method comprising obtaining a base controller by training in simulation with reinforcement learning, instantiating an initial BAM including a Gamma model configured to generate intermediate goals, a Beta model configured to translate the intermediate goals into task-space actions, and an Alpha model configured to translate the task-space actions and robot state into motor commands, deploying the initial BAM such that at least the Alpha model executes on-board the humanoid robot, causing the humanoid robot to perform an initial task and logging sensor and control data to form a first dataset, based on the first dataset, training at least one policy of the BAM to generate a refined BAM, and deploying the refined BAM to control the humanoid robot autonomously.
Owner:FIGURE AI INC

Robot operation task control method, device and equipment based on visual language action model, robot and medium

The invention provides a robot operation task control method, device and equipment based on a visual language action model, a robot and a medium, and relates to the technical field of sensors and robots. The method comprises the following steps: predicting weights between visual angles of corresponding visual image data based on global feature vectors corresponding to the visual image data of a plurality of visual angles of a robot; according to the plurality of local feature vectors corresponding to each piece of visual image data, predicting the weight in the visual angle of the corresponding image area; fusing the feature vectors of the reserved areas in the visual images in combination with the inter-view-angle weight and the in-view-angle weight of each image area to obtain a fused feature vector; and inputting the fusion feature vector into a robot operation task control strategy network based on a visual language action model to obtain an operation task control instruction. The motion execution precision can be improved, the computing power burden is reduced, the computing efficiency is improved, and the task success rate of complex and fine tasks is improved.
Owner:BEIJING ZHUJI POWER TECHNOLOGY CO LTD

Intelligent driving cooperative control terminal based on visual language action model

The invention discloses an intelligent driving cooperative control terminal based on a visual language action model, which comprises an instruction optimization processing unit, an execution management analysis unit and a dynamic adaptive cooperative processing unit, relates to the technical field of intelligent cooperative control, and solves the problem of how to realize strategy dynamic switching based on real-time scene classification. In order to solve the technical problem of improving the self-adaptive capability of the system to complex traffic scenes, according to the invention, through a dynamic self-adaptive cooperative processing unit, coherent actions are disassembled and then conflict analysis is carried out, the grading standards of emergency actions and non-emergency actions are determined, and a pause-recovery mechanism is adopted for the non-emergency actions, so that the self-adaptive capability of the system to the complex traffic scenes is improved. The method comprises the following steps: solving conflicts for emergency actions through parameter adjustment or equivalent security substitution, improving action execution coherence and security, realizing scene fine-grained classification based on a multi-modal fusion model, triggering strategy switching through scene confidence, and starting corresponding strategies for a composite scene after confidence of all sub-scenes reaches the standard.
Owner:SUZHOU AUTOMOBILE RES INST OF TSINGHUA UNIV (WUJIANG) +1

Target injection type fine tuning method for visual language action model

The invention discloses a visual language action model-oriented target injection type fine tuning method, which comprises the following steps of: firstly, constructing any existing visual language action model, introducing a condition image generation model, and generating a target image with consistent semantics and vision according to an initial observation image and a task target instruction; secondly, target image features are injected into observation input through zero-initialization convolution, parameters are gradually increased from zero, it is ensured that interference noise is not introduced in the initial stage of fine adjustment to destroy a pre-training strategy, and in the training process, along with gradual optimization of the parameters, target image feature information is gradually fused into model representation, and the target image feature information is obtained; therefore, the understanding ability and the execution performance of the task target are improved. According to the method, through a lightweight target image injection mechanism and an efficient fine adjustment process, the performance of the model on various reference tasks can be remarkably improved in few training rounds, and the problem that an existing visual language action model cannot systematically introduce target image guidance is effectively solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Aircraft maintenance simulation model training method and aircraft maintenance simulation method

The invention provides a training method of an aircraft maintenance simulation model and an aircraft maintenance simulation method, relates to the technical field of aircraft maintenance simulation, and aims to predict future evolution of aircraft maintenance so as to improve authenticity of aircraft maintenance simulation. The method comprises the following steps: acquiring training data related to maintenance simulation; training a maintenance simulation model based on the training data; the aircraft maintenance simulation model comprises a video word segmentation device, a multi-modal input encoder, a multi-modal token sequence and a multi-modal output encoder, wherein the video word segmentation device is used for encoding videos related to aircraft maintenance simulation into a video token sequence; the multi-modal input encoder is used for encoding multi-modal input data related to aircraft maintenance into a multi-modal token sequence; the potential action model is used for determining a potential action representation of a maintenance action based on the video token sequence and the multi-modal token sequence, and the dynamic prediction model is used for predicting a prediction token at the next moment based on the video token sequence, the multi-modal token sequence and the potential action representation so as to simulate a maintenance scene at the next moment.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD

Multi-robot collaboration method based on visual language model and related equipment thereof

The invention provides a multi-robot cooperation method based on a visual language model and related equipment thereof, the method is applied to a central scheduling server, the central scheduling server is in communication connection with a plurality of robots, and the method comprises the following steps: obtaining a target cooperation task, and decomposing the target cooperation task into ordered subtask sets; circularly executing the following steps until the subtask set is completed: acquiring real-time multi-modal observation data of the plurality of robots; according to the real-time multi-modal observation data, determining a sub-task to be executed currently, and generating a corresponding structured sub-task instruction for each robot; and issuing the sub-task instruction to the corresponding robot, so that the robot generates a joint control instruction according to the sub-task instruction and the real-time multi-modal observation data through a pre-configured visual language action model to control the robot to complete a corresponding action. According to the method, the high-level instruction can be mapped into the coordinated action of the multiple robots, so that the collaborative operation efficiency is improved.
Owner:PAXINI TECHNOLOGY (SHENZHEN) CO LTD

Robot control method and device, storage medium and robot

The invention discloses a robot control method and device, a storage medium and a robot, and relates to the technical field of robots. Determining a task instruction, a scene image collected for a preset scene and the state of a robot body; and based on the scene image, determining a target three-dimensional space feature corresponding to the preset scene. And processing the task instruction, the target three-dimensional space characteristics and the state of the robot body through a visual language action model, determining an action instruction to be executed by the robot, and controlling the robot to execute the action instruction so as to complete a task corresponding to the task instruction. As the target three-dimensional space features imply information such as geometrical shapes of the objects in the space, relative position relations between the objects and spatial layout, surrounding environment perception is performed by using the target three-dimensional space features on the premise of not increasing hardware cost, the precision of the robot for sensing the surrounding environment can be improved, and the robot experience is improved. And thus, the determined action instruction is more accurate, and the task execution success rate of the robot can be improved.
Owner:SHENZHEN SWEET POTATO ROBOT CO LTD +1