Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

225 results about "Robot action" patented technology

Action generation method and device based on multi-modal pre-training, robot and medium

The invention relates to the technical field of artificial intelligence, can be applied to the field of medical health and the field of financial transaction, discloses an action generation method and device based on multi-modal pre-training, a robot and a medium, and is applied to a high-frequency action generation scene of an intelligent surgical robot or an intelligent customer service and wealth management scene. The method comprises the following steps: acquiring a language instruction, a visual image and robot body sensing data; performing multi-modal feature alignment and cross-modal feature fusion on the language instruction, the visual image and the robot body sensing data through a pre-trained visual language model to generate a fused joint feature vector; generating a target continuous control instruction based on the fused joint features through a pre-trained target action model by adopting a flow matching technology; and generating continuous actions of the robot based on the target continuous control instruction. According to the method, the robot action generation efficiency and the cross-platform adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Task processing method and device based on expert sub-path, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a task processing method, device, equipment and medium based on an expert sub-path, comprising: setting a shared attention layer, an understanding expert sub-path, a control expert sub-path and a router module in a neural network, the method comprises the following steps: constructing a robot action trajectory data set, a visual text pair data set and a structured language template set, training based on the robot action trajectory data set and the structured language template set to obtain an intermediate training model, and training based on the intermediate training model and the visual text pair data set to generate a unified training model; and when target task input is received, the router module selects a corresponding expert sub-path, and a result is output through the unified training model. According to the method, staged training is combined with an expert mixing mechanism, catastrophic forgetting is avoided, semantic understanding and robot control keep stable coexistence in a unified model, and the accuracy of task execution is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Annotation model for humanoid robot data

The present disclosure provides a method for generating annotation data for robotic training using a hierarchical transformer-based model with multiple layers. The transformer-based model includes Alpha models generating low-level control outputs and Beta models generating high-level control outputs. The method receives multimodal input data comprising visual sensor data and natural language instructions, processes this data through the hierarchical transformer-based model to generate annotations at different abstraction levels, wherein Beta models create semantic annotations describing task objectives and Alpha models generate motor command annotations specifying robotic actions, and stores these annotations with the input data to create annotated training data for robotic control systems.
Owner:FIGURE AI INC

Intelligent robot with body and robot motion control system and method

The invention discloses an intelligent robot with a body and a robot motion control system and method.The intelligent robot with the body is provided with an execution component, a moving module and the robot motion control system, and the robot motion control system comprises a multi-modal input encoder, a multi-modal output encoder and a multi-modal output encoder, the multi-modal feature fusion module is used for collecting and processing multi-modal data to obtain multi-modal features and fusing the multi-modal features to obtain a multi-modal feature token sequence; the action expert network is used for mapping the fused multi-modal feature token sequence into an abstract robot action sequence block; and the execution controller is used for converting the abstract robot action sequence block into a control instruction which can be executed by bottom hardware. According to the method, direct mapping from environment perception and semantic understanding to movement control and execution of specified task actions can be realized, so that the task understanding, response speed and execution completion degree of the robot in a complex environment are improved.
Owner:HEFEI INNOVATION RES INST BEIHANG UNIV +1

System and method for robot planning using large language models

A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface configured to receive a plurality of multimodal inputs each specifying instructions for performing a task in a different modality including audio, video, and a text modality. The controller also comprises a multimodal large language model, an action sequence decoder, and a controller. The multimodal LLM includes a multimodal LLM encoder and an LLM decoder. The multimodal LLM encoder is trained with machine learning to transform the multimodal instructions into encodings and the LLM decoder is configured to decode the encodings into a sequence of robotic instructions. The action sequence decoder is trained with machine learning to transform the sequence of robotic instructions into a sequence of actions using a library of robotic skills. The controller is configured to control a robot according to the sequence of actions.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Zero-sample cross-domain imitation learning method based on three-dimensional semantic point cloud and related equipment

The invention discloses a zero-sample cross-domain imitation learning method based on three-dimensional semantic point cloud and related equipment, and the method comprises the steps: constructing a training scene in a physical simulation engine, randomly generating an operation object pose, and enabling a robot to act to complete an operation task; an original image is observed, semantic segmentation privilege information is obtained by utilizing a simulation environment, an operation object is segmented, and semantic point cloud of the operation object is obtained by combining a depth map generated by the simulation environment; synchronously recording robot joint angles, and constructing a training data set; inputting the semantic point cloud of the operation object and the training data set into an imitation learning model, performing action blocking, extracting features, performing feature splicing, generating a predicted action sequence, and completing model training; deploying the trained model to a real environment, obtaining the semantic point cloud of an operation object, reading the current joint angle of the robot, sending the angle to the model for reasoning, outputting a robot action sequence, and completing a specified task. According to the method, the generalization ability of the strategy can be improved while the operation precision is ensured.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Multi-modal data processing method supporting multi-task switching and related equipment

The invention provides a multi-modal data processing method supporting multi-task switching and related equipment, the method is applied to the technical field of robots, and the method comprises the following steps: acquiring an environment image, a historical task instruction, a current task instruction, a robot state and robot action noise of an environment where a robot is located at present; generating fusion features and multi-task condition features by using the environment image, the historical task instruction and the current task instruction; robot state features are extracted from the robot states; and the robot action noise, the fusion features, the multi-task condition features and the robot state features are input into a preset action expert model for action decision making, so that an optimal action sequence of the robot is obtained. According to the scheme, the optimal action sequence of the robot is obtained according to the historical task instruction, the current task instruction and the environment image, the problem of complexity during task switching of the robot is effectively solved, and the accuracy of robot actions output during task switching is improved.
Owner:BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Robot action generation method and device

The invention provides a robot action generation method and device, and the method comprises the steps: responding to a target task received by a robot, and obtaining image data and text description information associated with the target task; on the basis of the image data and the text description information, potential action information and fused visual representation information are determined by utilizing a large language model obtained by supervised training provided based on a potential action model; compressing and splicing the image data, the potential action information and the fused visual representation information to obtain control sequence information; and de-noising and splicing the state information and the control sequence information corresponding to the robot to generate action sequence information, so that the robot executes a corresponding action based on the action sequence information. By means of the method, the smoothness and coherence of actions generated by the robot are improved.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

Robot control method based on visual language action model and related equipment thereof

The invention provides a robot control method based on a visual language action model and related equipment thereof, and relates to the technical field of intelligent robots, and the method comprises the steps: obtaining visual observation information, a target language instruction and robot body state information of a target robot; reasoning the visual observation information, the target language instruction and the robot body state information through a pre-trained visual language action model to obtain a predicted low-dimensional potential vector; reconstructing the predicted low-dimensional potential vector into a robot action instruction through a pre-trained high-dimensional action reconstruction module; wherein the dimension of the robot action instruction is matched with the action space dimension of the target robot; and the target robot is driven to execute the high-dimensional robot action instruction. According to the method, the limitation of the existing visual language action model on the action output dimension can be solved, so that the high-degree-of-freedom robot is effectively controlled.
Owner:PAXINI TECHNOLOGY (SHENZHEN) CO LTD

Robot skill migration method and system based on multi-view object trajectory prediction

The invention relates to the technical field of robot skill migration, in particular to a robot skill migration method and system based on multi-view object trajectory prediction.The method comprises the steps that a first skill migration model is constructed, and the first skill migration model comprises a skill perception coding module A1 and a multi-view trajectory prediction diffusion module B1 which are connected in sequence; a second skill migration model is constructed based on the trained first skill migration model, and the second skill migration model comprises a skill perception coding module A2 and a multi-view trajectory prediction diffusion module B2 which are connected in sequence; inputting a to-be-tested robot multi-view image and a to-be-tested language instruction into the second skill migration model to obtain a multi-view target object prediction trajectory; and according to the multi-view target object prediction trajectory, an executable robot action is generated, and the action in the to-be-tested language instruction is completed. According to the method, robot skill migration can be efficiently and robustly realized.
Owner:SHANDONG UNIV

Personal intelligent multi-source data quality evaluation and verification method, device, medium and product

ActiveCN121188440AData setPhysical security
The embodiment of the invention relates to the technical field of information, and discloses an intelligent multi-source data quality evaluation and verification method and device, a medium and a product, and the method comprises the steps: converting intelligent multi-source data into standardized format data; performing quality evaluation on the standardized format data, wherein the quality evaluation comprises at least one of integrity check, availability check and consistency check; loading a robot URDF model corresponding to the standardized format data, driving the URDF model to move according to joint data, and synchronously playing corresponding visual data for realizing linkage playback of model actions and visual pictures; analyzing linkage playback through the first large model and / or the second large model; wherein the first large model is used for judging the physical safety of the robot action, and the second large model is used for judging the high-level semantic correctness of the robot action and task annotation, so that a high-quality and reliable self-service intelligent data set is constructed.
Owner:SHANGHAI COOPERS TECHNOLOGY CO LTD

Robot action reasoning method and system based on Gaussian action field

The invention relates to a robot motion reasoning method and system based on a Gaussian motion field, and the method comprises the steps: inputting a sparse and uncalibrated multi-view RGB image, extracting mixed scene features through a visual Transform backbone network, constructing a dynamic Gaussian motion field, endowing each Gaussian unit with a learnable motion attribute, and carrying out the reasoning of the motion of a robot. And realizing synchronous modeling of scene geometry and motion evolution. The system further comprises a multi-modal query module, a point cloud registration module, a diffusion model optimization module and a closed-loop control module which are respectively used for current and future scene reconstruction, mechanical arm tail end clamping jaw action estimation, action sequence optimization and real-time feedback and model updating in the execution process. According to the method, through a unified space-time modeling and closed-loop control mechanism, the robot operation problem in dynamic, shielding and uncalibrated environments is effectively solved, and the method is suitable for various application scenes such as industrial automation and service robots.
Owner:TSINGHUA UNIVERSITY

Health monitoring system of old-age care robot

The invention discloses a health monitoring system of an old-age care robot, which relates to the technical field of health monitoring, and comprises the steps of constructing an indoor multi-modal data pool, generating an encrypted feature vector, establishing a steady-state physiological behavior baseline according to a historical encrypted feature vector, and outputting a health risk level and an intervention instruction set. According to the invention, all-time and all-scene health monitoring is realized by integrating a multi-mode sensor, old people do not need to wear equipment, feature compression and homomorphic encryption technologies are adopted, data security is ensured, a threshold value can be updated in real time, and the method is suitable for being applied to the health monitoring of the old people. The health state of the old people is accurately monitored, health risk prediction and personalized intervention are carried out in combination with analysis of behaviors such as falling and static abnormity, the prompt mode is dynamically adjusted according to the ability of the old people, the data processing efficiency is optimized through the edge cloud collaborative architecture, the data privacy and stability are guaranteed, and the life quality and safety are improved.
Owner:NANJING XIAOZHUANG UNIV

Robot action prediction method and device, computer equipment and storage medium

The invention relates to the technical field of robots, and discloses a robot action prediction method and device, computer equipment and a storage medium, and the method comprises the steps: generating an input observation sequence according to a position code and a current feature vector, and predicting an action sequence through Transform and the input observation sequence; and exponential decay weighted average processing is carried out on the predicted action sequence, and a target action is determined. Through the above mode, the multi-modal observation data is converted into the unified feature vector through the feature extraction network, key information in the observation data is reserved, the time sequence dependency relationship and context information in the observation data are fully captured by using the Transform decoder, and the robot action sequence is accurately predicted. And exponential decay weighted average processing is performed on the predicted action sequence, so that the action sequence is further smoothed, the prediction instability is reduced, the finally determined target action is more accurate, and the performance and success rate of the robot during task execution are improved.
Owner:XIAN YOUIBOT ROBOTICS TECHNOLOGY CO LTD

Robot action generation method and system thereof, medium, equipment and program product

The embodiment of the invention provides a robot action generation method and system, a medium, equipment and a program product. The method comprises the following steps: acquiring visual input data of a robot and corresponding language instruction data; obtaining an intermediate action representation by using a pre-constructed cross-modal shared semantic space, a corresponding relationship between the shared semantic representation and the intermediate action representation, the visual input data and the language instruction data; wherein the shared semantic representation is obtained on the basis of visual input data and / or language instruction data by using a cross-modal shared semantic space; and action parameters are generated according to the intermediate action representation so as to drive the robot to execute corresponding actions. Due to the fact that the cross-modal shared semantic space can enable semantic information from different modalities to be measured under the unified scale, and the corresponding relation can serve as middle bridging representation from semantics to middle action representation, the ability of a robot to understand and execute common language instructions in diversified real scenes is improved.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Robot action redirection method, storage medium, electronic equipment and product

The invention provides a robot action redirection method, a storage medium, electronic equipment and a product, and relates to the field of robot control. The method comprises the steps that motion data of a to-be-simulated object and pre-configured joint offset parameters are obtained, wherein the joint offset parameters are used for describing rotation transformation data from a joint coordinate system of the to-be-simulated object to a joint coordinate system of a robot; based on the joint offset parameters, joint offset mapping from the to-be-simulated object to the robot is established; and based on the joint offset mapping, converting the motion data of the to-be-simulated object into motion data of the robot so as to drive the robot to execute corresponding actions. According to the method, the difference between the to-be-simulated object and the robot in joint coordinate system rotation transformation can be effectively eliminated, the physical feasibility and motion naturalness of robot actions are ensured, and the interaction requirements in a complex interaction scene are met.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Robot autonomous operation optimization method and system based on virtual simulation and multi-hypothesis planning

The invention relates to a robot autonomous operation optimization method and system based on virtual simulation and multi-hypothesis planning. The method comprises the steps that a virtual simulation platform scene with the same configuration as a real experiment scene is built; robot training is carried out in the virtual simulation platform scene, and an autonomous operation strategy is generated; constructing a plurality of hypothetical states based on the current state of the robot, and obtaining a virtual track action sequence of each hypothetical state based on an autonomous operation strategy; sequentially verifying whether the virtual track action sequence corresponding to each virtual state is feasible or not according to a time sequence, and if the virtual track action sequence is feasible, taking the virtual track action sequence as an optimal virtual track action sequence; and aligning the real robot state with the assumed state corresponding to the optimal virtual track action sequence, and generating a real robot action sequence based on the current state and the optimal virtual track action sequence by adopting a cooperative action integration mechanism. Compared with the prior art, the method has the advantages that the difference between virtuality and reality of the model is reduced, and the execution success rate is increased.
Owner:SHANGHAI TONGJI INDEPENDENT INTELLIGENT UNMANNED SYSTEMS RESEARCH INSTITUTE +1

System and Method for Interactive Robot Action Replanning Using Large Language Models

A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface to receive multimodal inputs specifying instructions for performing a task in audio, video, and a text modality. The controller transforms the multimodal instructions into encodings using a large language model (LLM) encoder and decodes the encodings into a first sequence of robotic instructions and a robot action description of the actions using an LLM decoder. Human feedback input is received corresponding to at least one action in the first sequence of actions and the controller encodes the feedback input with the robot action description. The controller feeds the encoded data along with multimodal features generated from the encodings into the LLM decoder to generate a corrected sequence of actions. The controller is configured to control a robot according to the corrected sequence of actions.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Surgical robot multi-module collaborative real-time decision planning system

A multi-module collaborative real-time decision planning system for a surgical robot comprises eight modules including a data acquisition and processing module, a system modeling module, a preoperative planning module, a neural network module, a decision planning module, an execution control module, an autonomous obstacle avoidance module and an autonomous evaluation module. The system obtains robot state and operation environment information in real time through multiple sensors, and constructs an operation environment and motion model; the inverse kinematics is solved by using a neural network, and the multi-solution problem of the inverse kinematics is solved; the preoperative planning module identifies key points and generates an initial path, the decision planning module dynamically adjusts the path in combination with joint constraints, and the execution control module controls the robot to act in a closed-loop mode according to instructions and corrects the relation between force and current in real time; the autonomous obstacle avoidance module monitors the current of a driving motor to realize collision detection and path adjustment; and the autonomous evaluation module continuously evaluates the operation quality through a reward function and feeds back optimization. The system can significantly improve the real-time performance, precision and safety of the surgical robot in a complex environment, and is suitable for the fields of medical treatment, industry, space exploration and the like.
Owner:BEIJING RES INST OF PRECISE MECHATRONICS CONTROLS

Robot action planning method and device, terminal and medium

The invention provides a robot action planning method and device, a terminal and a medium. The method comprises the steps that a user instruction, a scene perception image set, sensor information and a safety threshold table are acquired; generating an action draft according to the user instruction and the scene perception image set; performing virtual simulation on the scene perception image set and the sensor information according to a dynamic virtual twinning method to obtain virtual twinning state information; performing processing through the action draft and the virtual twinning state information, determining a risk vector of current iteration, generating an iterative action draft based on the risk vector of the current iteration and a safety threshold table, and obtaining iterative virtual twinning state information; and determining an iterative risk vector according to the iterative action draft and the iterative virtual twin state information, and obtaining a robot action instruction. On the premise of not depending on a pre-established and perfect global physical model, accurate physical consequence prediction is carried out on locally and dynamically changing interaction scenes.
Owner:CHONGQING VEHICLE TEST & RES INST CO LTD

Hierarchical robot operation strategy generation method, device and equipment

The invention discloses a hierarchical robot operation strategy generation method, device and equipment, and the generation method comprises the steps: collecting a hand-eye camera image and an external camera image of a robot, and obtaining a natural language instruction; taking a hand-eye camera image and an external camera image as image observation, and inputting the images into a cyclic consistency variational auto-encoder to obtain semantic features and spatial features; generating a new-view-angle hand-eye camera image, performing multi-view-angle fusion with image observation, and updating semantic features; inputting the semantic features and the natural language instruction into a basic skill discriminator to generate a robot skill category; and inputting the spatial features and the robot skill categories into an action generator to generate robot actions. The invention relates to the technical field of robot intelligent control and artificial intelligence, can significantly improve the success rate and reliability of the robot executing long-time-sequence and multi-task operation in a complex and changing environment, and has a wide application prospect.
Owner:NAT UNIV OF DEFENSE TECH

Data generation method and device for smart operation, equipment and storage medium

The invention provides a data generation method, device and equipment for dexterous operation and a storage medium, and the method comprises the steps: obtaining at least one piece of demonstration data generated when a robot executes a dexterous demonstration task, the dexterous demonstration task comprises the initial pose of at least one target object, the demonstration data comprises a first robot action sequence and first observation data corresponding to the first robot action sequence; according to the first observation data, the corresponding first robot action sequence is divided into a motion section and a skill section corresponding to each target object; responding to multiple times of pose adjustment of the target object, performing pose adaptation on the motion section and the skill section corresponding to the target object, and generating a plurality of second robot action sequences; and executing the second robot action sequence in the simulation environment, collecting corresponding second observation data, and pairing the second observation data with the second robot action sequence to generate training data.
Owner:北京中科慧灵机器人技术有限公司

Robot action sequence generation method and device, equipment and medium

The invention provides a robot action sequence generation method and device, equipment and a medium, and the method comprises the steps: inputting real-time multi-modal observation data and a current joint state into a pre-trained action prediction model, and generating an initial action prediction sequence containing a plurality of future time steps; constructing a space-time heterogeneous guide mask matrix by utilizing the current tail end linear speed and the execution state of a tail end operation part in the target robot; and correcting the initial action prediction sequence by using the space-time heterogeneous guide mask matrix and the historical action sequence queue to obtain a target action sequence, and sending the target action sequence to a controller of the target robot. By means of the method and device, the problems that in the track splicing process of a traditional method, a mechanical arm shakes, and an end effector is switched in an uncertain state are effectively solved, and the stability and response precision of target robot control are improved.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

Humanoid robot remote control method based on virtual reality technology

The invention provides a humanoid robot remote control method based on a virtual reality technology, and relates to the technical field of robot motion control, and the method comprises the steps: extracting the displacement features and rotation variation features of the motion of an arm tail end according to a state information set of the arm tail end in virtual reality, and forming a feature set; the robot simulates arm end actions in the virtual reality environment according to the state information set, obtains a simulation time difference according to the simulation process, marks and trains the feature set according to the simulation time difference to obtain a prediction time difference, obtains the state information set and the prediction time difference in real time, and compares the state information set and the prediction time difference to form action delay; and feedback is carried out to the virtual reality environment, so that a natural person in the virtual reality environment can know the delay amount of the current action at the robot end in real time, the action rhythm of the robot can be accurately controlled, the action of the robot better fits the robot, and the control accuracy is further improved.
Owner:XINJIANG UNIV OF SCI & TECH

Robot action sequence generation method based on hierarchical attention mechanism

The invention discloses a robot action sequence generation method based on a hierarchical attention mechanism, and the method comprises the steps: 1, inputting a plurality of robot operation images into a CNN layer, and extracting an image feature sequence; projecting the multiple pieces of joint position information into a joint state vector sequence through a linear layer of the CNN layer; taking a to-be-executed action sequence as a prediction target; taking the joint state vector sequence, the image feature sequence and the prediction target as input data of an encoder; 2, mapping the joint state vector sequence into a style variable by adopting an encoder, and adding randomly initialized learnable vector marks at the front ends of the image feature sequence, the joint state vector sequence and the action sequence to be executed; and 3, based on the sequence sparse attention, the trained decoder is adopted to predict a future action sequence of the mechanical arm of the robot. The invention provides a new solution for improving the fine operation capability of the robot.
Owner:XIAN UNIV OF TECH

Family robot action generation method based on Re-Plan principle

The invention provides a home robot action generation method based on the Re-Plan principle, and belongs to the technical field of robots. According to the invention, a dynamic gating fusion network is designed to improve the multi-mode sensing efficiency of the home service robot, and a joint space-time security optimization mechanism is introduced to improve the quality of multi-mode feature fusion; on the basis of a HomeAct-1. 2k data set in combination with a curriculum learning strategy, lightweight fine tuning of a Qwen-VL-Chat model is completed, an action generation framework based on the Re-Plan principle is designed on the basis of the model, and the task execution capacity of the home service robot in complex tasks is effectively improved.
Owner:HUNAN UNIVERSITY SUZHOU INSTITUTE

Robot control method and device, controller, storage medium and program product

The invention discloses a robot control method and device, a controller, a storage medium and a computer program product. The robot control method and device are applied to a scene in which a robot is jointly controlled based on safety equipment and N control ends. The control method of the robot comprises the following steps: establishing communication connection with at least one control end which needs to control the robot in the N control ends; recording one control end which currently needs to control the robot in the at least one control end which establishes the communication connection as a current control end; acquiring an operation command sent by the current control end, and acquiring security information sent by the security equipment; and based on the operation command sent by the current control end and the safety information sent by the safety equipment, cooperatively controlling the robot to act. According to the scheme, the multi-terminal equipment, the safety equipment and the robot body are cooperatively controlled, so that the multi-terminal equipment simultaneously controls the cooperative work of the robot, and the programming efficiency and accuracy are improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI

Multi-modal flow matching-based motion prediction method and device for robot with body

The invention provides a multi-modal flow matching-based motion prediction method and device for a body-equipped robot, and relates to the technical field of intelligent robots, and the method comprises the steps: obtaining an instruction text, an image feature set collected by a robot, and a depth image set corresponding to the image features of all moments and positions; performing feature splicing processing and feature refining processing on the image feature set to obtain an image sequence feature set, and performing feature fusion processing on the depth image set based on the image sequence feature set to determine a target visual feature; text features in the instruction text are fused with target visual features, text visual modal feature information is determined, and based on the text visual modal feature information, motion of the robot mechanical arm is predicted, and motion pose prediction features are determined. According to the method, the accuracy of motion prediction of the body-equipped robot can be remarkably improved.
Owner:SHENZHEN SHIHE ROBOTIC TECH CO LTD

Robot action generation method integrating multi-layer feature bridging and world knowledge prediction

The invention discloses a robot action generation method integrating multi-layer feature bridging and world knowledge prediction. The method comprises the following steps: extracting multi-layer middle layer visual features and action query hidden variables by utilizing a pre-trained visual language model; an initialization strategy based on robot body sensing state guidance is adopted, real-time pose priori is injected for action query, and traditional all-zero initialization is replaced; a spatial perception vector gating mechanism is introduced into the bridging attention module, and fine-grained selection of specific image region features is achieved; meanwhile, a world knowledge prediction task is integrated, and physical common knowledge is enhanced through explicit modeling environment depth, semantics and a dynamic region; and finally, replacing a traditional L1 regression action head with a diffusion model architecture, and generating an optimal action sequence under multi-modal distribution through an iterative denoising process. According to the method, the operation precision of the robot in a long-range complex task is improved, and the problem of control failure caused by an action'averaging effect 'is solved while the light weight of the model is kept.
Owner:NANJING UNIV

Robot action control method and system based on deep learning

The invention provides a robot action control method and system based on deep learning, and belongs to the technical field of artificial intelligence and robot control crossing. The method comprises the steps that S1, a robot action generation model is created based on a music encoder, an action encoder, an action generator, an action discriminator and a fusion layer; s2, acquiring a large amount of historical music and historical action videos to construct a data set; s3, training a robot action generation model based on the data set; s4, deploying a robot action generation model passing the test; and S5, collecting real-time music, preprocessing the real-time music, inputting the preprocessed real-time music into the deployed robot action generation model to obtain a command action, performing coordinate mapping on the command action to obtain a robot control instruction, and controlling the robot based on the robot control instruction. The method has the advantages that the anthropomorphic degree of robot action display is greatly improved, and action distortion is reduced.
Owner:XIAMEN UNIV