Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

146 results about "Action model" patented technology

Bipedal action model for humanoid robot

The present disclosure provides a system for generating motor control commands for a humanoid robot, comprising an alpha model with over 1 billion parameters that processes visual observations and language instructions at a first frequency to generate contextual embeddings, and a beta model operating at a higher second frequency. The beta model includes an embodiment-specific state encoder projecting robot state information into a shared embedding space, a diffusion transformer module generating denoised action sequences through iterative flow-matching that cross-attends to the alpha model's contextual embeddings, and an embodiment-specific action decoder converting denoised sequences into motor control commands. The beta model generates action chunks comprising future action sequences over a predetermined time horizon in a single inference step, with the complete system having less than 5 billion parameters.
Owner:FIGURE AI INC

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot comprising a torso having an alpha model deployed on a first GPU, and wherein said alpha model includes a first number of parameters and is configured to receive a natural language command from a human and generate processed data, a beta model deployed on a second GPU, and wherein said beta model includes a second number of parameters and is configured to receive the processed data from the alpha model and provide output data used to control an extent of the left wrist, and wherein the first number of parameters is larger than the second number of parameters, and a unified training framework is used to jointly train the alpha model and the beta model.
Owner:FIGURE AI INC

Interactive interface task automation utilizing generative artificial intelligence (AI) action models improved with retrieval-augmented generation (RAG)

PendingUS20260037318A1Mathematical modelsResource allocationRequest - actionEngineering
This disclosure describes a framework for performing user-requested tasks automatically across an interactive interface using various types of machine learning models. Specifically, this disclosure outlines and describes a task execution system that utilizes a generative artificial intelligence (AI) action model and retrieval-augmented generation (RAG) to complete user-requested actions across an interactive interface. The task execution system solves many of the current limitations of LAMs by using a generative AI action model to determine a session plan, which includes a set of actions for accomplishing stages of the actionable task across the interactive interface, obtaining visual context information of each interactive interface segment, integrates RAG results to improve the accuracy of both the session plan and individual actions, and self-corrects when faced with unexpected obstacles.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot comprising a torso having an alpha model deployed on a first GPU, and wherein said alpha model includes a first number of parameters and is configured to receive a natural language command from a human and generate processed data, a beta model deployed on a second GPU, and wherein said beta model includes a second number of parameters and is configured to receive the processed data from the alpha model and provide output data used to control an extent of the left wrist, and wherein the first number of parameters is larger than the second number of parameters, and a unified training framework is used to jointly train the alpha model and the beta model.
Owner:FIGURE AI INC

Visual chain-of-thought reasoning for robot vision-language-action models

Apparatuses, systems, and techniques are disclosed for controlling a robot to execute a task. In at least one embodiment, a current image of the robot in an environment and a text describing the task are obtained. A future image of the robot in the environment is predicted based on the current image and the text. Subsequently, one or more actions are predicted based on the current image, the future image, and the text. The one or more actions can move the robot from a first state corresponding to the current image to a second state corresponding to the future image. The robot executes the sequence of actions to move in the environment.
Owner:NVIDIA CORP

Lightweight method and system for highway engineering three-dimensional live-action model and medium

The invention discloses a lightweight method and system for a highway engineering three-dimensional live-action model and a medium, and relates to the technical field of three-dimensional live-action modeling. The method comprises the following steps: firstly, carrying out engineering semantic segmentation on an input three-dimensional grid model, and extracting design parameters in a BIM / GIS (Basic Information Model / Geographic Information System); then, constructing a multi-level constraint model fusing semantic constraints, engineering precision constraints and design parameter constraints; and finally, carrying out iterative simplification under the guidance and limitation of the multi-level constraint model by adopting an improved quadratic error measurement algorithm. The method solves the problems of semantic information loss, out-of-control local precision and disjunction with design intention when a traditional pure geometric simplification method is applied to highway engineering, and can intelligently reserve key engineering characteristics, control geometric errors and fit a design form while ensuring that the data size of the model is greatly reduced. And a high-quality lightweight model suitable for professional analysis and application is generated.
Owner:SICHUAN HIGHWAY PLANNING SURVEY DESIGN AND RESEARCH INSTITUTE LTD

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot system comprising a mechanical structure with at least 30 degrees of freedom across torso, arms, and legs, actuators driving the degrees of freedom, sensors including cameras and proprioceptive sensors, and a computing system implementing a hierarchical bipedal action model (BAM). The BAM includes: a Delta model processing sensor data and user input to generate latent representations at a first frequency; a Gamma model receiving latent representations to generate human task actions at a higher second frequency; a Beta model translating task actions into joint configurations at a higher third frequency; and an Alpha model converting joint configurations into actuator control signals at a higher fourth frequency.
Owner:FIGURE AI INC

Multi-sensor fusion-based badminton player action posture analysis system and method

PendingCN121838273AImage enhancementImage analysisCentre of pressureSimulation
The invention discloses a badminton player action posture analysis system and method based on multi-sensor fusion, and relates to the technical field of athletic training auxiliary systems.The system comprises a wearable sensing subsystem, an environment sensing subsystem and a central processing and feedback subsystem; the wearable sensing subsystem is arranged on an inertia measurement unit and a pressure sensing insole of a body to collect movement inertia and plantar pressure data; the environment perception subsystem collects global videos through multiple cameras. The central processing and feedback subsystem performs synchronous processing on multi-source data, adopts a hierarchical fusion algorithm, restrains inertia integral drift by utilizing a plantar contact state, reconstructs a three-dimensional skeleton posture and a motion trail in combination with visual key points and skeleton restraint, drives a digital twin model, compares with a standard motion model, and performs three-dimensional motion control on the three-dimensional skeleton posture and the motion trail. Quantitative evaluation results such as joint angle deviation, time sequence difference and pressure center track are generated, and visual feedback is performed through a display device and an augmented reality terminal for badminton training evaluation and technical deviation correction.
Owner:GUIZHOU UNIV

Emotion service robot system based on multi-mode mental theory

The invention relates to the technical field of artificial intelligence, intelligent and robot control, and discloses an emotion service robot system based on a multi-mode mental theory, which comprises a hierarchical mental reasoning module, a connection module and an action semantic execution module, the hierarchical mental reasoning module is used for generating a decision text containing a high-level strategy according to the multi-modal environment information collected by the robot; the connection module is used for constructing a semantic instruction according to the decision text generated by the hierarchical mental reasoning module; and the action semantic execution module is used for generating a control action of the robot according to the semantic instruction constructed by the connection module and the real-time image acquired by the robot. According to the method, a control architecture fusing a vision-language-action model and a robot center view angle mental reasoning hierarchy is constructed, so that the robot can carry out multi-order belief reasoning and implicit target inference from the view angle of the robot. According to the system, the robot is no longer a pure instruction follower and has an active service capability.
Owner:JILIN UNIVERSITY

Vision-language-action model training method and system utilizing inference data closed-loop optimization

The invention discloses a vision-language-action model training method and system utilizing inference data closed-loop optimization, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, performing initial training on a vision-language-action model by utilizing a training data set; deploying the trained model in a task environment to execute a task, monitoring a task execution result in real time, and capturing and structurally recording failure track data of the current task when task execution failure is recognized; then, based on the captured failure trajectory data, generating a negative prompt text for guiding the model to avoid repeated error behaviors; and finally, combining failure trajectory data with the generated negative prompt text, retraining the vision-language-action model, and inhibiting the model from generating action output similar to the error action sequence. According to the method, the data utilization rate is greatly improved, the data dependence and acquisition cost are effectively reduced, and the complete closed loop of the VLA model training and reasoning process is realized.
Owner:ZHEJIANG UNIV

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot system comprising a mechanical structure including a torso, two arms, and two legs providing at least 30 degrees of freedom, actuators coupled to the degrees of freedom, a sensor suite comprising at least one camera and proprioceptive sensors including joint encoders and an inertial measurement unit, a computing system comprising at least one processor and memory storing instructions which, when executed, implement a hierarchical bipedal action model including a Beta model configured to receive multimodal input data and generate a token sequence indicative of task intent and environmental state, and an Alpha model configured to condition on the token sequence and current robot pose data to output continuous action chunks comprising sequences of future target joint states over a finite horizon, and a low-level controller configured to convert the continuous action chunks into actuator control signals for execution.
Owner:FIGURE AI INC

Training and use of a bipedal action model for humanoid robot

The present disclosure provides a method for controlling a humanoid robot using a hierarchical bipedal action model (BAM), the method comprising obtaining a base controller by training in simulation with reinforcement learning, instantiating an initial BAM including a Gamma model configured to generate intermediate goals, a Beta model configured to translate the intermediate goals into task-space actions, and an Alpha model configured to translate the task-space actions and robot state into motor commands, deploying the initial BAM such that at least the Alpha model executes on-board the humanoid robot, causing the humanoid robot to perform an initial task and logging sensor and control data to form a first dataset, based on the first dataset, training at least one policy of the BAM to generate a refined BAM, and deploying the refined BAM to control the humanoid robot autonomously.
Owner:FIGURE AI INC

Robot operation task control method, device and equipment based on visual language action model, robot and medium

The invention provides a robot operation task control method, device and equipment based on a visual language action model, a robot and a medium, and relates to the technical field of sensors and robots. The method comprises the following steps: predicting weights between visual angles of corresponding visual image data based on global feature vectors corresponding to the visual image data of a plurality of visual angles of a robot; according to the plurality of local feature vectors corresponding to each piece of visual image data, predicting the weight in the visual angle of the corresponding image area; fusing the feature vectors of the reserved areas in the visual images in combination with the inter-view-angle weight and the in-view-angle weight of each image area to obtain a fused feature vector; and inputting the fusion feature vector into a robot operation task control strategy network based on a visual language action model to obtain an operation task control instruction. The motion execution precision can be improved, the computing power burden is reduced, the computing efficiency is improved, and the task success rate of complex and fine tasks is improved.
Owner:BEIJING ZHUJI POWER TECHNOLOGY CO LTD

Target injection type fine tuning method for visual language action model

The invention discloses a visual language action model-oriented target injection type fine tuning method, which comprises the following steps of: firstly, constructing any existing visual language action model, introducing a condition image generation model, and generating a target image with consistent semantics and vision according to an initial observation image and a task target instruction; secondly, target image features are injected into observation input through zero-initialization convolution, parameters are gradually increased from zero, it is ensured that interference noise is not introduced in the initial stage of fine adjustment to destroy a pre-training strategy, and in the training process, along with gradual optimization of the parameters, target image feature information is gradually fused into model representation, and the target image feature information is obtained; therefore, the understanding ability and the execution performance of the task target are improved. According to the method, through a lightweight target image injection mechanism and an efficient fine adjustment process, the performance of the model on various reference tasks can be remarkably improved in few training rounds, and the problem that an existing visual language action model cannot systematically introduce target image guidance is effectively solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Aircraft maintenance simulation model training method and aircraft maintenance simulation method

The invention provides a training method of an aircraft maintenance simulation model and an aircraft maintenance simulation method, relates to the technical field of aircraft maintenance simulation, and aims to predict future evolution of aircraft maintenance so as to improve authenticity of aircraft maintenance simulation. The method comprises the following steps: acquiring training data related to maintenance simulation; training a maintenance simulation model based on the training data; the aircraft maintenance simulation model comprises a video word segmentation device, a multi-modal input encoder, a multi-modal token sequence and a multi-modal output encoder, wherein the video word segmentation device is used for encoding videos related to aircraft maintenance simulation into a video token sequence; the multi-modal input encoder is used for encoding multi-modal input data related to aircraft maintenance into a multi-modal token sequence; the potential action model is used for determining a potential action representation of a maintenance action based on the video token sequence and the multi-modal token sequence, and the dynamic prediction model is used for predicting a prediction token at the next moment based on the video token sequence, the multi-modal token sequence and the potential action representation so as to simulate a maintenance scene at the next moment.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD

Multi-robot collaboration method based on visual language model and related equipment thereof

The invention provides a multi-robot cooperation method based on a visual language model and related equipment thereof, the method is applied to a central scheduling server, the central scheduling server is in communication connection with a plurality of robots, and the method comprises the following steps: obtaining a target cooperation task, and decomposing the target cooperation task into ordered subtask sets; circularly executing the following steps until the subtask set is completed: acquiring real-time multi-modal observation data of the plurality of robots; according to the real-time multi-modal observation data, determining a sub-task to be executed currently, and generating a corresponding structured sub-task instruction for each robot; and issuing the sub-task instruction to the corresponding robot, so that the robot generates a joint control instruction according to the sub-task instruction and the real-time multi-modal observation data through a pre-configured visual language action model to control the robot to complete a corresponding action. According to the method, the high-level instruction can be mapped into the coordinated action of the multiple robots, so that the collaborative operation efficiency is improved.
Owner:PAXINI TECHNOLOGY (SHENZHEN) CO LTD

Robot control method and device, storage medium and robot

The invention discloses a robot control method and device, a storage medium and a robot, and relates to the technical field of robots. Determining a task instruction, a scene image collected for a preset scene and the state of a robot body; and based on the scene image, determining a target three-dimensional space feature corresponding to the preset scene. And processing the task instruction, the target three-dimensional space characteristics and the state of the robot body through a visual language action model, determining an action instruction to be executed by the robot, and controlling the robot to execute the action instruction so as to complete a task corresponding to the task instruction. As the target three-dimensional space features imply information such as geometrical shapes of the objects in the space, relative position relations between the objects and spatial layout, surrounding environment perception is performed by using the target three-dimensional space features on the premise of not increasing hardware cost, the precision of the robot for sensing the surrounding environment can be improved, and the robot experience is improved. And thus, the determined action instruction is more accurate, and the task execution success rate of the robot can be improved.
Owner:SHENZHEN SWEET POTATO ROBOT CO LTD +1

Intelligent agent action sequence generation method based on visual language action model

The invention discloses an agent action sequence generation method based on a visual language action model, belongs to the technical field of artificial intelligence, and can solve the problems that an existing method is highly dependent on an additional observation mode and an auxiliary module, needs additional training and fine adjustment, and is limited in deployment flexibility and expandability. The method comprises the following steps: S1, determining an output hidden state of a feedforward network of a visual language action model according to observation information and a language instruction of an intelligent agent; s2, determining the probability distribution of each action mark corresponding to the current feed-forward network, and determining the uncertainty of the current feed-forward network; s3, determining observation characteristics to be injected according to the uncertainty of the current feed-forward network, and updating the output hidden state of the next feed-forward network by using the observation characteristics; and S4, taking the next feed-forward network as the current feed-forward network, repeating the steps S2 and S3, and generating an agent action sequence according to the output hidden state of the last feed-forward network. The method is used for generating the agent action sequence.
Owner:ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD +1

Bipedal action model for humanoid robot

The present disclosure provides a humanoid robot system comprising a mechanical structure with at least 30 degrees of freedom across torso, arms, and legs, actuators driving the degrees of freedom, sensors including cameras and proprioceptive sensors, and a computing system implementing a hierarchical bipedal action model (BAM). The BAM includes: a Delta model processing sensor data and user input to generate latent representations at a first frequency; a Gamma model receiving latent representations to generate human task actions at a higher second frequency; a Beta model translating task actions into joint configurations at a higher third frequency; and an Alpha model converting joint configurations into actuator control signals at a higher fourth frequency.
Owner:FIGURE AI INC

Visual and tactile language action large model training method, visual and tactile language action large model and robot

The invention provides a training method of a visual sense and tactile sense language action large model, the visual sense and tactile sense language action large model and a robot, and relates to the technical field of industrial robots. The method comprises the following steps: training a bridging module between a task planning system and an action execution system; training the task planning system; and finally, performing end-to-end supervision fine tuning on the visual and tactile language action large model. Task planning and an action execution function are decoupled, and a tactile encoder and a three-dimensional point cloud encoder are introduced in stages, so that the model can fuse tactile information in a planning stage to understand physical interaction, and point cloud data can be utilized to realize accurate positioning in an execution stage; by adopting the three-stage training method, the training stability and convergence efficiency of a multi-mode and multi-task model are effectively guaranteed, the operation precision, success rate and environmental adaptability of a robot in industrial assembly tasks are fundamentally improved, and the practical application of a visual language action model in industrial flexible production is promoted.
Owner:ZHEJIANG SHIYUE TECHNOLOGY CO LTD

Automatic driving vision-language-action model and training method thereof

The invention discloses an automatic driving vision-language-action model and a training method thereof. The model comprises a vision encoder module, a text encoder module, a cross-modal adapter module, a large language model module, a subtask module and a generative planner module. The method comprises the following steps: injecting domain knowledge; pre-training a generative planner module; aligning the training vision-text; and finely adjusting the small sample based on a draft chain thinking mechanism. The visual encoder and the cross-modal adapter carry out multi-scale encoding on each view angle of an input image, dynamic and static target and time sequence consistency modeling is completed in an aerial view space at the same time, then a language model is injected, and the generative planner generates track distribution with global semantic prompt and local dynamic state as conditions. And the scene understanding and reasoning capability of the model can be improved. According to the method, the adaptive capacity and the real-time performance of the model to a complex driving environment are enhanced by guiding the automatic driving vision-language-action model to think in multiple steps and outputting the abstract.
Owner:DALIAN UNIV OF TECH

Air-ground collaborative navigation method based on hierarchical visual language action model

The invention provides an air-ground collaborative navigation method based on a hierarchical visual language action model, which comprises the following steps that: a ground robot and an air unmanned aerial vehicle respectively acquire relevant data and send the relevant data to a central processing system for processing; a global path planning module of the central processing system generates a coarse-grained global path key point sequence, namely a sparse path point sequence, based on the processed and aligned aerial top view image, and then a local path planning module of the central processing system generates dense local path points between adjacent sparse path points based on the ground view image; and a capability perception evaluation module of the central processing system evaluates the trafficability of each output candidate path according to the physical capability of the ground robot so as to provide a basis for final path selection. The method effectively solves the problems that a traditional air-ground collaborative framework lacks quantitative evaluation and closed-loop feedback mechanisms for physical capabilities of ground robots and does not make full use of global environment information to carry out robot passage planning and the like.
Owner:NANJING AGRICULTURAL UNIVERSITY

Adding Voice or Chat User interface to graphical user interface (gui)-based virtualized applications and desktops using large language and large action models

Methods and systems for enhanced remote desktop interfaces are described. A computing system may train, using historical or live information, a LAM to execute, within a remote desktop application, textual actions with their parameters if any. A user declarative request (voice or chat) may be interpreted by a LLM to match a specific action (and potentially ask for the corresponding parameters in a conversational way). Subsequently, from the textual action and its parameters, the LAM may execute the action within a remote desktop application and report the result to the user via voice or chat.
Owner:CITRIX SYSTEMS INC

Training method of security event reproduction model and security event reproduction method

The invention provides a training method of a security event reoccurrence model and a security event reoccurrence method, relates to the technical field of security event reoccurrence, and aims to perform anti-fact reasoning through the security event reoccurrence model so as to improve the accuracy of security event reoccurrence. The method comprises the steps of obtaining training data; training a security event reproduction model based on the training data; wherein the security event reproduction model comprises a video word segmentation device, a multi-modal event injection module, a potential action model and a dynamic prediction model; wherein the video word segmentation device is used for encoding a video related to a security event into a video token sequence; the multi-modal event injection module is used for encoding multi-modal data related to a security event into a multi-modal event token sequence; the potential action model is used for coding into action representation based on a video token sequence and a multi-modal event token sequence; the dynamic prediction model is used for outputting a prediction token sequence and a key point probability sequence of the security event based on the video token sequence and the action representation.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD

Hydraulic support control system and method based on action learning and reproduction

The invention discloses a hydraulic support control system and method based on action learning and reproduction. The system comprises a man-machine interaction layer, a support controller, a signal acquisition module, an electromagnetic valve driving circuit, a sensor group and an electromagnetic valve group. The method comprises the following steps that S1, an operator manually controls a hydraulic support to complete a hydraulic support moving action process, and a core processor synchronously records monitoring data of a sensor set and opening and closing state data of an electromagnetic valve set in the process, stores the monitoring data and the opening and closing state data as original operation data and stores the original operation data in a storage; s2, the core processor generates a descending, shifting and ascending action model capable of being executed by the support controller according to the original operation data, and the descending, shifting and ascending action model is stored in a memory; s3, the core processor analyzes the descending, shifting and ascending motion model and outputs control instructions to the electromagnetic valve driving circuit one by one, and meanwhile the sensor set feeds monitoring signals back to the core processor; and after all actions in the descending, shifting and ascending action model are completed, the process is ended.
Owner:ZHENGZHOU HENGDA INTELLIGENT CONTROL TECHNOLOGY CO LTD

Robot control method and device based on image-text interleaving instruction, equipment and medium

The invention relates to the technical field of artificial intelligence, provides a robot control method and device based on an image-text interleaving instruction, equipment and a medium, is applied to financial and medical health care service scenes, and can construct an initial model which takes an image-text mixed format as an input data format and takes a visual language model as a backbone model. The limitation that a traditional vision-language-action model can only process image observation and plain text instructions is broken through; performing language-action pre-training on the initial model based on a potential action distribution function and a flow matching mechanism, so that the model can learn potential representation of action intention from a language; instruction tuning training is carried out based on a hybrid expert mechanism to ensure that the model is flexibly switched between language reasoning and action prediction, so that the model can accurately reasone from a complex and natural image-text instruction and carry out action planning; and generating a predicted action sequence by using the robot action generation model, so as to accurately control the robot based on the image-text interlacing instruction.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method of controlling robot based on simplified map and natural-language command and computer program therefor

An embodiment of the present disclosure provides a method of controlling a robot based on a simplified map and a natural-language command. The method may include acquiring an image of a space, acquiring a simplified map of the space, the simplified map including a current position of a robot and a destination position of the robot and not including at least some elements describing the space, acquiring a natural-language command specifying compliance requirements that the robot is to comply with during an operation of the robot, generating input tokens based on at least one of the image, the simplified map, and the natural-language command, and generating waypoint tokens by processing the input tokens with a trained vision-language-action model.
Owner:MAUM AI INC

Robot skill knowledge characterization method based on Granger causality test

The invention provides a robot skill knowledge characterization method based on Granger causality test. The method comprises the following steps: collecting original video stream data in a robot operation process; standard operation steps and sudden abnormal conditions of the robot are coded into a standard operation state and a random event state respectively, stage variables are formed, executable operation after the standard operation or the random event occurs is coded into an executable operation state, and action variables are formed; modeling the sequential relationship between the historical stage variable and the current action variable by using a VAR model to obtain a VAR-based stage-action model; performing Granger causality test on variables in the stage-action model based on the VAR to confirm the causality between the variables; and according to the result of the Granger causal test, constructing a causal relationship network in the robot operation process, and forming a causal rule device. According to the method, the real dependency relationship in the time series data can be deeply mined, and a skill characterization framework based on a causal mechanism is established.
Owner:UNIV OF SCI & TECH BEIJING

Tabular policy action models for clinical applications

In some aspects, the present disclosure provides a computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to perform a computer task by a user expressing to the computer-implemented system, and wherein the computer-implemented system is configured to perform the task by generating and executing a computer-executable program.
Owner:TORTUS AI LTD