Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

84 results about "Robot learning" patented technology

Robot learning is a research field at the intersection of machine learning and robotics. It studies techniques allowing a robot to acquire novel skills or adapt to its environment through learning algorithms. The embodiment of the robot, situated in a physical embedding, provides at the same time specific difficulties (e.g. high-dimensionality, real time constraints for collecting data and learning) and opportunities for guiding the learning process (e.g. sensorimotor synergies, motor primitives).

Visual servo double-arm robot migration simulation learning method from simulation to reality

The invention discloses a visual servo double-arm robot migration simulation learning method from simulation to reality, which comprises the following steps: S110, constructing a virtual simulation scene according to a real task scene, and establishing a relation between the virtual simulation scene and the real task scene; s120, the collected teaching mechanical arm and executing mechanical arm operation data are replayed and optimized in the virtual simulation scene, so that a data set used for final training is constructed; s130, a deep imitation learning network is designed and achieved, input of the deep imitation learning network comprises visual data obtained in the real task scene and state data of all joint motors of the teaching mechanical arm, and output of the deep imitation learning network is predicted states of all joint motors of the execution mechanical arm at the next moment; and S140, migrating the fully trained and converged deep imitation learning network from a virtual simulation scene to a double-arm robot in a real task scene. According to the invention, the exploration efficiency and generalization ability of the two-arm robot learning network are improved.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Robot continuous imitation learning method and system based on double diffusion model

The invention belongs to the technical field of robot learning, and particularly discloses a robot continuous imitation learning method and system based on a double diffusion model. Comprising the following steps: generating old task data in a submerged space by utilizing a pre-trained latent diffusion model, generating a model through guidance of a task identifier in combination with conditional diffusion, generating an image and state data related to an old task, modeling to construct a diffusion model, realizing multi-modal mapping from a state to an action, and generating a diversified action sequence; high-quality images and state data are generated through a pre-trained variational auto-encoder and a conditional diffusion model, and the consistency of the generated data in time and the accuracy of global semantics are ensured through time sequence continuity constraint and semantic consistency constraint; and performing feedback optimization training on the latent diffusion model. According to the method, the loss weight is dynamically adjusted, and the model is ensured to show good generalization ability and stability in a complex task in combination with collaborative training of an imitation learning strategy and a generator.
Owner:HUAZHONG UNIV OF SCI & TECH

Robot control method and device based on artificial intelligence, equipment and medium

The invention relates to an artificial intelligence technology, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a robot control method, device and equipment based on artificial intelligence and a medium. Generating a planning path instruction of the robot according to the decomposition task; executing the path planning instruction in the simulation environment, and outputting operation track information and environment state change data of the robot; the simulation execution result is verified, and if verification succeeds, the natural language task instruction, the decomposition task, the operation track information and the success label serve as multi-modal information to be stored; training the initial diffusion strategy model based on the multi-modal information to generate a robot decision model; and controlling the robot based on the robot decision model. Through the automatic and large-scale data generation process, the robot learning efficiency and robustness are remarkably improved, and a natural language instruction can be responded to execute multiple tasks.
Owner:PING AN TECH (BEIJING) CO LTD

Cross-modal perception driven compliance control system for robot with body

The invention relates to the technical field of robot control, in particular to a cross-modal perceptual driving body robot compliance control system which comprises the steps that a sensor is adopted to synchronously collect environment information, multi-source data bias is eliminated through a space-time alignment algorithm, a radial basis function neural network is adopted to analyze multi-modal fusion features, and a multi-modal model is obtained; human operation intention probability distribution is extracted to decompose a task into a path planning layer and a motion control layer, a collision-free trajectory is generated through an RRT algorithm, a high-fidelity physical engine is adopted to construct a virtual interaction scene, and robot learning results are shared through federal learning. According to the method, the problems of inaccurate perception, incoordination between intention recognition and interaction control, difficulty in control strategy verification, slow new task adaptation and difficulty in multi-robot learning result sharing caused by multi-source data deviation and large multi-modal semantic difference of the body robot in a complex environment are solved.
Owner:CHANGCHUN UNIV OF TECH

Robot training system and method based on plot memory

The invention discloses a robot training system and method based on plot memory in the technical field of artificial intelligence and robot learning, and solves the problems that an existing robot training method lacks an effective memory scheduling system, an empirical value quantification mechanism is not intelligent, and the multi-modal information fusion capability is insufficient. The system comprises a sensing module, a multi-modal unified memory encoder, a progressive memory scheduling system, a multi-dimensional value quantitative evaluation mechanism, an intelligent experience hierarchical scheduler, a multi-modal unified code retriever, a strategy generation module, an execution module and an anomaly detection module. The multi-modal unified memory encoder adopts a hierarchical dimensionality reduction Transform architecture, and fuses an RGB image, a depth image, force sensor data and joint angle information into a 576-dimensional unified feature vector; the progressive memory scheduling system comprises a working memory structure, a short-term memory structure and a long-term memory structure. The multi-dimensional value quantitative evaluation mechanism carries out quantitative scoring on experience based on reward evaluation, novelty evaluation and uncertainty evaluation.
Owner:SHANGHAI MODUAN TECHNOLOGY CO LTD

Logistics robot learning system based on artificial intelligence algorithm

The invention relates to the technical field of logistics robots, discloses a logistics robot learning system based on an artificial intelligence algorithm, and aims to solve the problem that a traditional logistics robot system lacks autonomous learning and environment change adaptation capabilities. The system integrates a path planning module, a task execution strategy generation module, a strategy adjustment module, a real-time monitoring module and a task configuration module, and can perceive a logistics environment in real time and plan an optimal path. And according to a path planning result, the system generates a specific task execution strategy, and continuously optimizes the strategy in a task execution process to cope with behavior differences. The real-time monitoring module ensures that the robot behavior is consistent with the planning result, and if deviation is found, strategy adjustment is triggered. The task configuration module configures targeted tasks according to robot performance and work requirements. According to the system, the autonomous learning and adaptive capacity of the logistics robot are improved, and the efficiency and safety of logistics operation are effectively improved.
Owner:SHANDONG VOCATIONAL COLLEGE OF LIGHT IND

Transform-DQN-based navigation method and device for mobile robot

The invention relates to the technical field of robot path planning, in particular to a navigation method and device of a mobile robot based on Transform-DQN, and can solve the problems that an existing algorithm is slow in algorithm convergence time, poor in path planning strategy performance and the like in a complex dynamic environment to a certain extent. According to the method, environment information and state information are obtained through multi-sensor fusion sensing environment; obtaining an expected action of the current mobile robot by using an optimal strategy obtained by a DQN algorithm; and controlling the movement of the mobile robot according to the current expected action. According to the technical scheme, the reward function considering multiple factors is set to interact with the mobile robot, so that the accuracy of the algorithm is improved; in the training process, an attenuation mode of an adjustable greedy factor is set, exploration and learning of the mobile robot in environments with different complexity degrees are balanced, a Transform model is introduced into an experience playback mechanism, the long-term dependency relationship between experiences is captured, the learning effect of the robot is enhanced, and the training efficiency is improved.
Owner:CHANGZHOU UNIV

Robot imitation learning method based on diffusion model

The invention discloses a robot imitation learning method based on a diffusion model, and belongs to the technical field of robot learning and body intelligence. Comprising the following steps: carrying out standardization processing on input image observation data and a robot state, and carrying out image enhancement by adopting gray scale transformation, random erasure and Gaussian blur; based on the enhanced image and the robot state, training a diffusion model to generate an action sequence, including forward diffusion, conditional feature fusion, noise prediction and joint loss optimization; in real-time control, features are extracted through an image enhancement network, sampling is accelerated through DDIM to generate an action sequence, noise scheduling coefficients are dynamically adjusted based on visual feature differences, and closed-loop optimization is achieved. According to the method, the image enhancement technology and the diffusion strategy are deeply fused, joint optimization of visual features and action sequences is achieved, the action generation accuracy and robustness of the robot under the complex visual interference and small sample conditions are remarkably improved, and the success rate is improved by 18% compared with a base line under 40 sample sizes.
Owner:NANCHANG UNIV +1

Humanoid robot falling recovery control method based on multi-stage course learning

According to the humanoid robot tumble recovery control method based on multi-stage curriculum learning, robot tumble crawling actions are decomposed into a plurality of key frames, frame-by-frame staged learning is carried out, meanwhile, a mixed internal model and a reinforcement learning framework are introduced, the mixed internal model obtains a current observation value of a robot in real time, and through speed estimation, the robot tumble recovery control method based on multi-stage curriculum learning is realized. According to the method, feature information is obtained, the robustness of a strategy is enhanced, the difference between simulation and reality is reduced, a reinforced semester framework mixing internal optimization and near-end strategy optimization is combined, and through interaction with a simulation environment, robot learning is rapidly and stably transited from one key frame to the next key frame, learning of complex human shape motion behaviors is achieved, and the learning efficiency is improved. According to the application, the man-machine interaction safety performance is optimized, so that the robot can reliably operate in scenes needing close man-machine cooperation, such as medical care and home service, and a technical foundation is laid for constructing a more intelligent and safer next-generation service robot system.
Owner:SONGYAN POWER (BEIJING) TECHNOLOGY CO LTD

Systems and methods for robot learning and controlling a robot

A method may include receiving input data identifying an operation to be completed in an environment. The method may also include identifying, using an artificial intelligence (AI) model, a series of tasks to be performed by robots to complete the operation based on the input data. In addition, the method may include identifying a subset of the robots to perform the series of tasks based on capabilities to be used to perform the series of tasks. Further, the method may include causing the subset of the robots to autonomously perform the series of tasks to complete the operation.
Owner:COLLABORATIVE ROBOTICS

Lifelong robot learning for mobile robots

A method for improving a mobile robot configured to perform a task in an environment using an operational program is disclosed. Data is received, the data being recorded by the mobile robot using one or more sensors when the mobile robot is navigated in the environment to perform the task. A database and / or model associated with the environment is updated to include the recorded data therein. The operational procedure of the mobile robot can be modified based on the database and / or the model to generate a modified operational procedure for performing the task in the environment, the modified operational procedure improving performance of the mobile robot. Furthermore, based on the database and / or the model, suggestions for improving the performance of the mobile robot in performing the task in the environment can be determined and displayed to a user for consideration.
Owner:ROBERT BOSCH GMBH

Devices, systems, and methods for transferring physical skills to robots

A robotic training system enabling intuitive skill transfer through human demonstration using paired devices, allowing robots to perform complex manipulation tasks traditionally requiring extensive manual programming or expensive hardware setups. Leader robotic devices configured for human manipulation and follower robotic devices replicate the leader's movements. Force sensors, torque sensors, or force-torque sensors measure and process force data, torque data, or force-torque data for recording training demonstrations for artificial intelligence (AI) models. Demonstration devices include user control interfaces, multiple workspace viewpoints, motion modification for haptic feedback, and interchangeable end effector tools. A control system enables automatic transitions between position and force control based on sensed interactions, while incorporating learned policies from demonstrations with visual and force data. Methods for robot programming leverage demonstration data, generating natural language descriptions of actions and human-readable narratives of robot programs. The system provides a comprehensive, low-cost solution for robot learning from human demonstration.
Owner:STANDARD BOTS CO

A humanoid robot control method, system, storage medium, and program product

The present invention provides a humanoid robot control method, system, storage medium and program product, belonging to the field of computer vision. The method includes: preprocessing expert action data and processing the expert action data into expert data equivalent to the skeletal structure of the target robot; building a robot with a humanoid structure in a simulation environment, configuring the joint parameters of the robot, and each joint degree of freedom is controlled by an independent physical control module; constructing a policy representation method for the robot, including a state space, an action space, a reward function, and a multi-frame control method; initializing the robot; minimizing the difference between the robot action and the expert action in each frame, maximizing the reward function, and driving the robot to learn. The present invention can assist the learning process of the humanoid robot, enabling the robot to be anthropomorphic while completing tasks and improving the training speed.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Control Method, Device, and Medium of a Debate Robot Based on Speech Processing

The present disclosure provides a control method, device, medium and speech processing-based debate robot system for a speech processing-based debate robot. The control method for the speech processing-based debate robot includes: obtaining a debate material training set; inputting the debate material training set into a preset training model to control the debate robot to learn the debate materials, where the preset training model includes a low-rank adapter for optimizing the preset training model; inputting a preset debate scenario, the debate duration corresponding to the preset debate scenario, and the debater role of the debate robot into the debate robot; and controlling the debate robot to output target debate language within the debate duration corresponding to the preset debate scenario according to the preset debate scenario and the speech information of the debate scenario. Through the present disclosure, a robot debater can be generated to adapt to various debate scenarios and break through time limitations to accompany humans in debate training or debate competitions.
Owner:TIANHUA COLLEGE OF SHANGHAI NORMAL UNIV

Method for determining development parameters of a natural gas hydrate reservoir and related apparatus

The application provides a method for determining development parameters of a natural gas hydrate reservoir and related equipment. In the process of particle swarm intelligent optimization, a robot learning algorithm is used for data mining and training to obtain a training model. The potential particles are determined based on the training model. The global optimal objective function is determined based on the combination of the potential particles and the updated development parameters. In the case where the convergence condition is reached based on the development parameter combination corresponding to the global optimal objective function value, the development parameter combination corresponding to the global optimal objective function value is determined as the target development parameter combination. The calculation process of the intelligent optimization algorithm can be improved by the machine learning algorithm, and the optimization time of the development parameters of the natural gas hydrate reservoir is shortened.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Wearable robot data collection system with human-machine operation interface

A data collection system that performs data collection of human-driven robot actions for robot learning. The data collection system includes: i) a wearable computation subsystem that is worn by a human data collector and that controls the data collection process and ii) a human-machine operation interface subsystem that allows the human data collector to use the human-machine operation interface to operate an attached robotic gripper to perform one or more actions. A user interface subsystem receives instructions from the wearable computation subsystem that direct the human data collector to perform the one or more actions using the human-machine operation interface subsystem. A visual sensing subsystem includes one or more cameras that collect raw visual data related to the pose and movement of the robotic gripper while performing the one or more actions. A data collection subsystem receives collected data related to the one or more actions.
Owner:ACUMINO

A blockchain-based machine learning method, device and medium

The application discloses a kind of robot learning methods, equipment and medium based on blockchain, method includes: determining the blockchain platform created based on blockchain framework in advance;Determine operation instruction, and write operation instruction in blockchain platform;Receive the operation request sent by robot;According to operation request, and the operation instruction stored in blockchain platform, determine the current operation instruction corresponding to the robot;Send current operation instruction to robot, to make robot execute corresponding operation mode based on current operation instruction.The whole process of robot learning and obtaining operation instruction is stored in blockchain platform.As blockchain platform is distributed storage, data tampering of single node will not take effect, which ensures the authenticity and credibility of data on the blockchain platform.When robot encounters new working scenario, it only needs to obtain corresponding operation instruction in blockchain platform to work, without artificial re-editing, which improves work efficiency.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

A quadruped robot low-noise gait control method and system based on a soft landing reward function

The present application relates to a kind of soft landing reward function-based quadruped robot low-noise gait control method and system, the method includes: constructing quadruped robot motion control environment, and defining foot contact determination standard;Design soft landing reward function based on vertical contact force, for the negative punishment of large contact force at the moment of landing, guide quadruped robot to learn low contact force landing mode;Collect the real-time contact force of quadruped robot foot and historical contact state, with soft landing reward function as core optimization target, using reinforcement learning algorithm, in combination with gait stability constraint, iterative training is carried out to strategy network, and low-noise gait strategy network is obtained, for optimizing output gait control signal, to control the gait action of quadruped robot accordingly.Compared with prior art, the present application can consider the noise control and gait stability of robot by accurately determining the moment of landing, designing contact force negative reward and combining reinforcement learning to actively optimize gait strategy.
Owner:FUDAN UNIVERSITY

Multi-robot assembly line method and system based on temporal logic control strategy

The present invention discloses a multi-robot assembly line method and system based on a temporal logic control strategy, comprising: expressing the robot's task specification based on a control strategy synthesized by temporal logic of a parity check game, constructing a reward automaton with a potential energy function according to the acceptance condition of the synthesized strategy to assign a reward value to the robot's behavior; decomposing the comprehensive strategy of a robot group consisting of a generalized reactive specification of rank 1 into a reward automaton for each robot, and expanding the reward automaton with potential energy on the MDP; proposing a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy; the present invention captures the temporal attributes of the task based on temporal logic, decomposes the generated comprehensive strategy into multiple individual reward automata to guide the robots to learn the optimal strategy, and proposes a reward shaping algorithm based on value iteration to improve the efficiency of the robot group in learning the optimal assembly strategy and avoid falling into the problem of falling into local optimality.
Owner:CHANGZHOU UNIV

Robot, learning data collection device, learning data collection method, and computer program product

The invention relates to a robot, a learning data collection device, a learning data collection method, and a computer program product. A robot capable of collecting learning data for recognizing a speaker from speech includes: an external stimulus detection unit that detects an external stimulus; a data collection unit that, when the external stimulus detection unit detects, as the external stimulus, speech data indicating the detected speech, as the learning data, stores the speech data in a storage unit, in a case where the speech satisfying a given collection condition is detected by the external stimulus detection unit as the external stimulus; and a condition changing unit that, when the external stimulus detection unit detects the external stimulus that satisfies a specific condition, changes the collection condition for a period from the detection of the external stimulus that satisfies the specific condition until a prescribed time elapses, and changes the collection condition for a period from the detection of the external stimulus that satisfies the specific condition to the detection of the external stimulus that satisfies the specific condition. The external stimulus detection unit is configured to detect the external stimulus such that the speech detected by the external stimulus detection unit is easy to satisfy the collection condition.
Owner:CASIO COMPUTER CO LTD

A method for legged robot reinforcement learning control based on constrained markov decision process and manifold variance compression

PendingCN122626199AAlgorithmLegged robot
The application belongs to the technical field of intelligent control of legged robots, and particularly relates to a legged robot reinforcement learning control method based on a constrained Markov decision process and manifold variance compression, which comprises collecting multi-source state data of the legged robot and preprocessing the same; constructing a strategy-value network and a danger assessment network; proposing state-dependent manifold variance compression and performing action sampling; constructing a Lagrange model of the constrained Markov decision process; updating the strategy network and performing alternating optimization of the Lagrange multiplier. The application can adaptively compress the exploration noise of high-risk joints in the strategy training process, and simultaneously realize automatic balancing of multiple physical constraints through a dynamic Lagrange multiplier, so as to fundamentally improve the running safety and robustness of the legged robot under extreme postures and complex working conditions while ensuring the learning efficiency of the legged robot, and provide reliable protection for real physical deployment.
Owner:GUANGDONG UNIV OF TECH

ROBOT LEARNING DEVICE WITH SYMBOL PROGRAMMING FUNCTION

Robot teaching device (10) for generating an operating program (16) for a robot (20) by arranging command symbols (60 to 64) expressing operating commands for the robot (20), comprising a marker indicator (43) which, when an operating command contains multiple position data, displays multiple markers (69) relating to identifiers (68) of the position data in conjunction with a single command symbol (61) on it.
Owner:FANUC LTD

Robot precision landing planning method and device based on prior terrain model

The embodiment of the application provides a robot precise landing planning method and device based on a priori terrain model, which carries out digital modeling on a preset terrain in a simulation environment, constructs a digital map containing terrain geometry and semantic information, and clearly calibrates the center coordinates of each ideal landing area as a target landing point. In the training process, the global coordinates of the robot foot end are tracked in real time, and the three-dimensional deviation of the robot foot end from the nearest target landing point is determined through map mapping, and then a reward function with the deviation size as the evaluation index is designed to drive the robot to learn a high-precision landing strategy. The present application introduces the priori terrain target point as a guide, so that the reward signal has a clear geometric meaning and can directly and efficiently drive the strategy to converge to the precise landing behavior, greatly improving the motion accuracy and reliability on the structured terrain.
Owner:HANGZHOU YUNSHENCHU TECH CO LTD

An artificial intelligence-based robot control method, device, equipment and medium

The application relates to an artificial intelligence technology, which can be applied to a medical health, financial technology and other business system platform, and discloses a robot control method, device, equipment and medium based on artificial intelligence, which comprises the following steps: decomposing a natural language task instruction; generating a planning path instruction of a robot according to the decomposed task; executing the planning path instruction in a simulation environment, and outputting operation track information of the robot and environment state change data; verifying the simulation execution result, and if the verification is successful, storing the natural language task instruction, the decomposed task, the operation track information and a success label as multi-modal information; training an initial diffusion strategy model based on the multi-modal information, generating a robot decision model; and controlling the robot based on the robot decision model. The application significantly improves the learning efficiency and robustness of the robot through an automatic and large-scale data generation process, and can respond to natural language instructions to execute multi-task.
Owner:PING AN TECH (BEIJING) CO LTD

Embodied intelligence training corpus generation method and system based on adversarial data governance, terminal and medium

PendingCN122332956AData streamData segment
This invention discloses a method, system, terminal, and medium for generating embodied intelligence training corpus based on adversarial data governance, relating to the field of robot learning technology. The key technical points are: acquiring multimodal data streams generated by robots during human-machine collaborative interaction; performing time alignment processing on the multimodal data streams to obtain aligned standardized data streams; identifying high-value data segments from the standardized data streams using a preset value evaluation function; and performing coordinate regularization processing on the identified high-value data segments to generate training corpus. This invention introduces an adversarial operator role into the human-machine collaborative loop, systematically injecting multidimensional perturbations and actively generating high-difficulty edge scenario data containing instability-recovery logic. This results in training corpus that not only has extremely high information density, covering failure-recovery manifolds that are difficult to collect using traditional data, but also possesses strong generalization characteristics.
Owner:TIANFU JIANGXI LAB

Llarva: vision-action instruction tuning for enhanced robot learning

PendingUS20260249456A1SimulationComputer vision
A robotic device includes a robot having an end-effector, and a large modality model (LMM) pre-trained on vision-language tasks and fine-tuned on image-visual trace pairs. A method of predicting a next sequence of actions for a robot using a large modality model (LMM) includes receiving, at the LMM, an image input and a language input, using the LMM to predict a next action sequence for a robot having an end-effector, and using the LMM to produce predicted visual traces of the end-effector.
Owner:RGT UNIV OF CALIFORNIA

Robot learning through retrieval and self improvement

Implementations are provided for an interactive machine learning methodology that allows non-expert users to use natural language to teach new skills, particularly to robots, through language grounding and understanding. In various implementations, a plurality of natural language summaries may be retrieved. Each of the natural language summaries may describe details of robotic performance of a task, and may include, or be usable to retrieve, a corresponding set of reference modulation values. A set of modulation values corresponding to a natural language request may be generated based on the plurality of natural language summaries. The natural language request may specify one or more constraints on robotic performance of the task. A robot control signal may be generated based on the generated set of modulation values.
Owner:GDM HOLDING LLC

Method and apparatus for training palletizing robot based on reinforcement learning using attention mechanism of vision transformer

A robot training apparatus according to an embodiment may perform the operations of: acquiring state information including a state of a pallet and the size of an object to be loaded onto the pallet; converting the state information into a patch of a predefined size in order to input same into a vision transformer model; on the basis of an actor network of a reinforcement learning model including a policy function for determining a loading position of the object in the direction that maximizes an expected value of the final loading rate of the pallet, determining, from the patch, a position to load the object onto the pallet; on the basis of a critic network of the reinforcement learning model including a value function for deriving the expected value according to the determination, deriving an expected value according to the determination; and updating parameters of the policy function and the value function in the direction that minimizes a loss of a loss function calculated on the basis of the determination and the expected value.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION