Robot control method based on multi-modal large model
Through the robot control method of multimodal large model, combined with multimodal sensors and deep learning algorithms, the problem of insufficient single modal data in the existing technology is solved, and the robot can be efficiently adapted and decided in complex environments is achieved, and execution efficiency and security are improved.
Patent Information
- Application Number
- CN202510206165.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing robot control methods rely on single modal data, resulting in the inability to fully understand environmental diversity in complex and dynamic environments, resulting in insufficient security, poor adaptability and inefficient execution.
The robot control method of multimodal large model is adopted. By installing a multimodal sensor group, visual, auditory and tactile information is collected in real time, combined with long-term and short-term memory networks and dynamic evolution algorithms, the behavioral strategy is dynamically updated and its evolutionary adjustment volume is optimized.
It realizes the rapid adaptation and efficient decision-making of robots in complex environments, improves the efficiency and success rate of autonomous tasks, and avoids unnecessary mistakes through risk assessment mechanisms.
Smart Images

Figure CN119681909B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control, and in particular to a robot control method based on a multi-modal large model. Background Art
[0002] With the rapid development of artificial intelligence and robotics, the application scope of robots has been continuously expanded, gradually shifting from traditional single-task execution to more complex and dynamic multi-task environments. This shift not only improves the operational efficiency of robots, but also enables robots to play a role in a wider range of application scenarios. In these human-machine interaction scenarios, robots face rapidly changing environments and complex user needs, so it is particularly important to have the ability to integrate multiple sensory inputs.
[0003] The existing technology has the following deficiencies: the existing robot control methods mostly rely on the collection and processing of single-modal data, and have certain limitations in environmental perception and decision execution. Especially in complex and dynamic scenarios, robots cannot fully understand the diversity of the environment, resulting in insufficient safety, poor adaptability and low execution efficiency. When robots rely only on visual information to implement behavioral strategies, they ignore important information conveyed by audio or contact information, thereby increasing the risk of collision or reducing reaction speed. Existing intelligent decision-making lacks flexibility and adaptability when dealing with complex environmental changes. It relies on a pre-set model architecture and lacks the ability to update in real time for new situations, making it difficult for robots to make adjustments when facing emergencies or abnormal environments, which can easily lead to failure or inefficiency.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention
[0005] The purpose of the present invention is to provide a robot control method based on a multi-modal large model to solve the problems in the above-mentioned background technology.
[0006] In order to achieve the above object, the present invention provides the following technical solution: a robot control method based on a multi-modal large model, comprising the following steps:
[0007] Step 1: Collect the robot's perception data in real time by installing a multimodal sensor group, and build a three-dimensional posture model of the robot itself to obtain the posture change data of the multimodal sensor group in real time, where the perception data specifically includes environmental video data, environmental audio data, and contact status data;
[0008] Step 2: Processing environmental video data, environmental audio data, contact state data and posture change data, and extracting features of corresponding data respectively, wherein the features of corresponding data specifically include environmental color information features, environmental object edge features, spectrum features, time series features and posture features;
[0009] Step 3: Receive and process the environmental color information features, environmental object edge features, spectrum features, timing features, and posture features in real time through the long short-term memory network, and dynamically update the hidden state and memory state to generate a predictable robot behavior strategy;
[0010] Step 4: Optimize the evolution adjustment of the robot's behavior strategy through a dynamic evolutionary algorithm, and calculate the risk assessment value based on completion time, energy consumption, and number of failures to achieve the evolution of the optimal behavior strategy;
[0011] Step 5: By creating a Q-value table and using the ε-greedy strategy, the Q-value of the robot's optimal behavior strategy is updated based on the immediate reward. The feedback loop helps the robot adjust and optimize its behavior strategy in real time.
[0012] Preferably, in step one, the multimodal sensor group specifically includes an RGB camera for collecting environmental video data in real time, a microphone array for collecting environmental audio data in real time, and a tactile sensor for collecting contact status data in real time. The specific steps of installing the multimodal sensor group to collect the robot's perception data in real time include installing the RGB camera to the robot's head position, collecting environmental video data in real time and extracting static key frames of the environmental video data through inter-frame differences, installing the microphone array to the robot's chest position as the robot's multi-audio input point, collecting environmental audio data in real time and recording timestamps and multi-audio input point positions, installing the tactile sensor to the robot's hand position, and collecting contact status data in real time. The robot collects and records the state data of the touch screen and timestamps the touch screen. The robot coordinates the relative center positions of the RGB camera, microphone array and tactile sensor to maintain the spatial synchronization of the multimodal sensor group. The robot's own master clock source is selected as the time reference of the RGB camera, microphone array and tactile sensor. The robot's own posture 3D model is constructed through the inertial measurement unit of the robot's head, chest and hand positions. The RGB camera, microphone array and tactile sensor are used as the robot's feature points and mapped to the posture 3D model. The initial position of the robot's feature points is defined. According to the collaborative positioning of the multimodal sensor group, the GPS positioning API is called to obtain the posture change data of the robot's feature points in real time.
[0013] Preferably, in the step 2, the specific steps of processing the environmental video data, the environmental audio data, the contact state data and the posture change data include converting the static key frame of the environmental video data into a grayscale image and fixing the size to 224x224 pixels, adjusting the brightness and contrast of the static key frame by grayscale stretching, highlighting the target details in the environmental video data by sharpening and edge enhancement, and the specific steps of extracting the features of the corresponding data include extracting the environmental color information features of the static key frame by color space conversion, extracting the environmental object edge features of the static key frame by Canny edge detection, separating the multiple audio input points of the microphone array based on beamforming, dividing the environmental audio stream data into 20s frames, applying a Hamming window to each frame to reduce the edge effect and extracting the spectral features of the environmental audio stream data by short-time Fourier transform, converting the spectral features to the Mel frequency domain by a Mel filter group, dynamically compressing the range of the Mel frequency domain by logarithmic processing, and converting it into cepstrum coefficients through discrete cosine transform, extracting the time series features of the contact state data, and extracting the posture features of the posture change data.
[0014] Preferably, in the step three, a predictable behavior strategy is defined according to the mapping relationship between the environmental video data, the environmental audio data, the contact state data and the posture change data and the robot's behavior strategy, and the environmental color information features, the environmental object edge features, the spectrum features, the timing features and the posture features are received as vector inputs through the long short-term memory network at each time step of the real-time acquisition data; the specific steps of dynamically updating the hidden state and the memory state to generate a predictable robot behavior strategy include initializing the hidden state and the memory state of the long short-term memory network to all zero vectors, linearly changing the current vector input and the hidden state of the previous time step by using the sigmoid function through the input gate, and outputting an information activation value indicating the importance of the feature, and updating the hidden state and the memory state of the long short-term memory network by t The anh function linearly changes the current vector input and the hidden state of the previous time step, and outputs a new candidate memory vector. The sigmoid function is used through the forget gate to linearly change the current vector input and the hidden state of the previous time step, and outputs a value indicating the proportion of information that needs to be forgotten. The information activation value indicating the importance of the feature and the value indicating the proportion of information that needs to be forgotten are combined and the memory unit is updated according to the candidate memory vector. The sigmoid function is used through the output gate to linearly change the current vector input and the hidden state of the previous time step, and outputs an information activation value indicating the importance of the output. The current memory unit state is transformed through the tanh function and multiplied by the information activation value of the output gate to update and output the robot's behavior strategy in the current time step.
[0015] Preferably, in step 4, the optimal behavior strategy is evolved through a dynamic evolution algorithm, and its specific formula is:
[0016]
[0017] in, represents the robot's optimal behavior strategy, represents the robot's behavior strategy in the current time step, It represents the evolution adjustment of the behavior strategy, records the completion time, energy consumption and number of failures of the robot to complete the behavior strategy, and calculates the risk assessment value of the behavior strategy. The specific formula is:
[0018]
[0019] in, represents the risk assessment value of the behavior strategy, Indicates the completion time. represents the energy consumption of the behavioral strategy, represents the number of failures of the behavior strategy, represents the risk weight of the completion time, represents the risk weight of energy consumption, Represents the risk weight of the number of failures.
[0020] Preferably, in step 5, a Q value table representing the expected benefits of the robot adopting the optimal behavior strategy under the current data is created, the Q values corresponding to all data-behavior strategies are initialized to 0, and the optimal behavior strategy is selected by the ε-greedy strategy with a probability feedback of 1-ε. If the optimal behavior strategy is successfully completed, a positive Q value reward is given, and if the optimal behavior strategy fails or has an error, a negative Q value reward is given. The Q value is updated according to the reward, and the specific formula is:
[0021]
[0022] in, represents the Q value of the behavior strategy a adopted by the robot under the updated current data s, represents the Q value of the behavior strategy a adopted by the robot under the current data s, represents the learning rate that controls the Q value update speed, It represents the immediate reward obtained by the robot after adopting the optimal behavior strategy. represents the discount factor for evaluating the impact of future rewards, It represents the maximum Q value of all possible behavior strategies a that the robot may take under the current data s. If the current Q value is greater than 20, the weight of the corresponding optimal behavior strategy is increased by 10% to encourage the optimal behavior strategy to be selected again under similar data. If the current Q value is less than -20, the weight of the corresponding optimal behavior strategy is reduced by 10% to reduce the probability of the optimal behavior strategy being executed again under similar data. Repeat the iteration until the strategy converges.
[0023] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0024] 1. By installing multimodal sensors on the robot's head, chest and hands to achieve comprehensive perception of the environment, it can not only collect rich visual, auditory and tactile information in real time, but also maintain spatial synchronization between sensors through collaborative positioning technology to ensure data accuracy and consistency. By performing inter-frame difference processing on the environmental video data collected by the RGB camera, extracting static key frames and performing image processing, it can more clearly identify and analyze the target objects and their characteristics in the environment. Audio data is collected in real time through the multiple audio input points of the microphone array, and the spectral features are extracted using short-time Fourier transform, which enhances the robot's ability to understand environmental sounds and multiple sensory information. The integration of information enables the robot to consider more contextual information when making decisions, significantly improving its intelligent decision-making capabilities. By processing multimodal data through long short-term memory networks, the robot can generate predictable behavioral strategies at each time step, allowing the robot to quickly adapt and respond reasonably in complex and changing environments, thereby improving its efficiency and success rate in autonomous execution of tasks. By optimizing behavioral strategies through dynamic evolutionary algorithms, the robot can continuously learn and adapt to new environmental challenges. The risk assessment mechanism that records indicators such as completion time, energy consumption, and number of failures enables the robot to fully consider potential risks when making decisions, thereby avoiding unnecessary mistakes.
[0025] 2. By introducing the Q-value table and ε-greedy strategy, an effective learning and self-optimization mechanism is constructed, so that the robot can self-adjust and optimize in a constantly changing environment. By initializing the Q-value corresponding to all data-behavior strategies to 0, the robot continuously updates the Q-value through the feedback mechanism during the execution of the task, forming an expected benefit evaluation for each behavior strategy, so that the robot can effectively balance between exploration and utilization, ensuring that it can adjust its behavior strategy in time when facing new situations. When the robot successfully completes a behavior strategy, it will receive a positive reward, thereby increasing the Q-value of the strategy; conversely, if the strategy execution fails or an error occurs, a negative reward will be given, reducing the corresponding Q-value, which not only promotes the robot to strengthen the successful strategy, but also effectively inhibits the re-execution of bad strategies. For the dynamic adjustment mechanism of the current Q-value, when the Q-value is greater than 20, the weight of the corresponding strategy is increased to encourage the robot to choose the strategy again in similar situations; and when the Q-value is less than -20, the weight of the strategy is reduced, reducing its execution probability, further enhancing the robot's learning ability, enabling it to quickly find the optimal behavior strategy when facing a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0027] Figure 1 The present invention is a flow chart of the method of robot control based on multi-modal large model. DETAILED DESCRIPTION
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of the present disclosure will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art.
[0029] The present invention provides Figure 1 The robot control method based on the multimodal large model shown includes the following steps:
[0030] Step 1: Collect the robot's perception data in real time by installing a multimodal sensor group, and build a three-dimensional posture model of the robot itself to obtain the posture change data of the multimodal sensor group in real time, where the perception data specifically includes environmental video data, environmental audio data, and contact status data;
[0031] Install an RGB camera to the robot's head position, collect environmental video data in real time and extract static key frames of the environmental video data through inter-frame difference, install a microphone array to the robot's chest position as the robot's multi-audio input point, collect environmental audio data in real time and record timestamps and multi-audio input point positions, install a tactile sensor to the robot's hand position, collect contact state data in real time and record timestamps, collaboratively locate the relative center positions of the RGB camera, microphone array and tactile sensor to maintain spatial synchronization of the multi-modal sensor group, select the robot's own master clock source as the time reference for the RGB camera, microphone array and tactile sensor, build the robot's own posture 3D model through the inertial measurement unit at the robot's head, chest and hand positions, use the RGB camera, microphone array and tactile sensor as robot feature points and map them to the posture 3D model, define the initial positions of the robot's feature points, and obtain the posture change data of the robot's feature points in real time based on the collaborative positioning of the multi-modal sensor group and call the GPS positioning API.
[0032] Step 2: Processing environmental video data, environmental audio data, contact state data and posture change data, and extracting features of corresponding data respectively, wherein the features of corresponding data specifically include environmental color information features, environmental object edge features, spectrum features, time series features and posture features;
[0033] The static key frames of the environmental video data are converted into grayscale images and fixed in size to 224x224 pixels. The brightness and contrast of the static key frames are adjusted by grayscale stretching. Sharpening and edge enhancement are used to highlight the target details in the environmental video data. The environmental color information features of the static key frames are extracted by color space conversion. The edge features of the environmental objects in the static key frames are extracted by Canny edge detection. The multiple audio input points of the microphone array are separated based on beamforming. The environmental audio stream data is divided into 20s frames. A Hamming window is applied to each frame to reduce the edge effect and the spectral features of the environmental audio stream data are extracted by short-time Fourier transform. The spectral features are converted to the Mel frequency domain by a Mel filter bank. The range of the Mel frequency domain is dynamically compressed by logarithmic processing and converted into cepstrum coefficients through discrete cosine transform. The temporal features of the contact state data are extracted, and the posture features of the posture change data are extracted.
[0034] Step 3: Receive and process the environmental color information features, environmental object edge features, spectrum features, timing features, and posture features in real time through the long short-term memory network, and dynamically update the hidden state and memory state to generate a predictable robot behavior strategy;
[0035] According to the mapping relationship between the environmental video data, environmental audio data, contact state data and posture change data and the robot's behavior strategy, a predictable behavior strategy is defined, and the environmental color information features, environmental object edge features, spectrum features, timing features and posture features are received as vector inputs through the long short-term memory network at each time step of the real-time acquisition data. The hidden state and memory state of the long short-term memory network are initialized to all zero vectors. The sigmoid function is used through the input gate to linearly change the current vector input and the hidden state of the previous time step, and the information activation value representing the importance of the feature is output. The current vector input and the hidden state of the previous time step are converted into a linear vector using the tanh function. Perform linear changes and output new candidate memory vectors. Use the sigmoid function through the forget gate to linearly change the current vector input and the hidden state of the previous time step, and output the value indicating the proportion of information that needs to be forgotten. Combine the information activation value indicating the importance of the feature and the value indicating the proportion of information that needs to be forgotten and update the memory unit according to the candidate memory vector. Use the sigmoid function through the output gate to linearly change the current vector input and the hidden state of the previous time step, and output the information activation value indicating the importance of the output. Transform the current memory unit state through the tanh function and multiply it with the information activation value of the output gate to update and output the robot's behavior strategy in the current time step.
[0036] The mapping relationship between environmental video data, environmental audio data and the robot's behavioral strategy includes forward / backward, turning, and defining predictable behavioral strategies includes adjusting speed and direction, turning left, turning right or turning around. The mapping relationship between contact state data and the robot's behavioral strategy includes dynamic grasping and releasing objects. Defining predictable behavioral strategies includes adjusting the grasping action and grasping force of the robot's hand and enabling the release action. The mapping relationship between environmental audio data and the robot's behavioral strategy includes command recognition and active interaction. Defining predictable behavioral strategies includes executing command behavior strategies and active communication actions.
[0037] Step 4: Optimize the evolution adjustment of the robot's behavior strategy through a dynamic evolutionary algorithm, and calculate the risk assessment value based on completion time, energy consumption, and number of failures to achieve the evolution of the optimal behavior strategy;
[0038] The optimal behavior strategy is evolved through a dynamic evolutionary algorithm, and its specific formula is:
[0039]
[0040] in, represents the robot's optimal behavior strategy, represents the robot's behavior strategy in the current time step, It represents the evolution adjustment of the behavior strategy, records the completion time, energy consumption and number of failures of the robot to complete the behavior strategy, and calculates the risk assessment value of the behavior strategy. The specific formula is:
[0041]
[0042] in, represents the risk assessment value of the behavior strategy, Indicates the completion time. represents the energy consumption of the behavioral strategy, represents the number of failures of the behavior strategy, represents the risk weight of the completion time, represents the risk weight of energy consumption, Represents the risk weight of the number of failures.
[0043] Step 5: Create a Q-value table and use the ε-greedy strategy to update the Q-value of the robot's optimal behavior strategy based on the immediate reward. The feedback loop helps the robot adjust and optimize its behavior strategy in real time.
[0044] Create a Q value table representing the expected benefits of the robot taking the optimal behavior strategy under the current data, initialize the Q values corresponding to all data-behavior strategies to 0, select the optimal behavior strategy through the ε-greedy strategy with a probability feedback of 1-ε, and give a positive Q value reward if the optimal behavior strategy is successfully completed. If the optimal behavior strategy fails or has an error, a negative Q value reward is given. Update the Q value according to the reward. The specific formula is:
[0045]
[0046] in, represents the Q value of the behavior strategy a adopted by the robot under the updated current data s, represents the Q value of the behavior strategy a adopted by the robot under the current data s, represents the learning rate that controls the Q value update speed, It represents the immediate reward obtained by the robot after adopting the optimal behavior strategy. represents the discount factor for evaluating the impact of future rewards, It represents the maximum Q value of all possible behavior strategies a that the robot may take under the current data s. If the current Q value is greater than 20, the weight of the corresponding optimal behavior strategy is increased by 10% to encourage the optimal behavior strategy to be selected again under similar data. If the current Q value is less than -20, the weight of the corresponding optimal behavior strategy is reduced by 10% to reduce the probability of the optimal behavior strategy being executed again under similar data. Repeat the iteration until the strategy converges.
[0047] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A robot control method based on a multimodal large model, characterized in that: The following steps are involved: Step 1: Collect the robot's perception data in real time by installing a multimodal sensor group, and build a three-dimensional posture model of the robot itself to obtain the posture change data of the multimodal sensor group in real time, where the perception data specifically includes environmental video data, environmental audio data, and contact status data; The RGB camera, microphone array and tactile sensor are used as robot feature points and mapped to the posture 3D model, the initial position of the robot feature points is defined, and the GPS positioning API is called based on the collaborative positioning of the multimodal sensor group; Step 2: Processing environmental video data, environmental audio data, contact state data and posture change data, and extracting features of corresponding data respectively, wherein the features of corresponding data specifically include environmental color information features, environmental object edge features, spectrum features, time series features and posture features; Step 3: Receive and process the environmental color information features, environmental object edge features, spectrum features, timing features, and posture features in real time through the long short-term memory network, and dynamically update the hidden state and memory state to generate a predictable robot behavior strategy; Step 4: Optimize the evolution adjustment of the robot's behavior strategy through a dynamic evolutionary algorithm, and calculate the risk assessment value based on completion time, energy consumption, and number of failures to achieve the evolution of the optimal behavior strategy; The specific formula of the dynamic evolutionary algorithm is: ; in, represents the robot's optimal behavior strategy, represents the robot's behavior strategy in the current time step, represents the evolutionary adjustment of the behavioral strategy; The specific formula for calculating the risk assessment value is: ; in, represents the risk assessment value of the behavior strategy, Indicates the completion time. represents the energy consumption of the behavioral strategy, represents the number of failures of the behavior strategy, represents the risk weight of the completion time, represents the risk weight of energy consumption, The risk weight representing the number of failures; Step 5: By creating a Q-value table and using the ε-greedy strategy, the Q-value of the robot's optimal behavior strategy is updated based on the immediate reward. The feedback loop helps the robot adjust and optimize its behavior strategy in real time.
2. The robot control method based on multimodal large model according to claim 1, characterized in that: In the step 1, the multimodal sensor group specifically includes an RGB camera for collecting environmental video data in real time, a microphone array for collecting environmental audio data in real time, and a tactile sensor for collecting contact state data in real time.
3. The robot control method based on multimodal large model according to claim 1, characterized in that: In the step 2, the specific steps of processing the environmental video data, environmental audio data, contact status data and posture change data include: using grayscale stretching to adjust the brightness and contrast of static key frames, using sharpening and edge enhancement to highlight the target details in the environmental video data, separating the multiple audio input points of the microphone array based on beamforming, dividing the environmental audio stream data into 20s frames, applying a Hamming window to each frame to reduce edge effects, using logarithmic processing to dynamically compress the range of the Mel frequency domain, and converting it into cepstral coefficients through discrete cosine transformation.
4. The robot control method based on multimodal large model according to claim 1, characterized in that: In the step 2, the specific steps of extracting the features of the corresponding data include: extracting the environmental color information features of the static key frame through color space conversion, extracting the edge features of the environmental objects of the static key frame by using Canny edge detection, extracting the spectrum features of the environmental audio stream data by short-time Fourier transform, converting the spectrum features to the Mel frequency domain by using the Mel filter group, extracting the time series features of the contact state data, and extracting the posture features of the posture change data.
5. The robot control method based on multimodal large model according to claim 1, characterized in that: In the step three, the specific steps of generating a predictable robot behavior strategy are: receiving environmental color information features, environmental object edge features, spectrum features, timing features, and posture features as vector inputs, initializing the hidden state and memory state of the long short-term memory network to all zero vectors, using the sigmoid function through the input gate to linearly change the current vector input and the hidden state of the previous time step, and outputting an information activation value representing the importance of the feature, using the tanh function to linearly change the current vector input and the hidden state of the previous time step, and outputting a new candidate memory vector, using the sigmoid function through the forget gate to linearly change the current vector input and the hidden state of the previous time step, and outputting a value representing the proportion of information that needs to be forgotten, combining the information activation value representing the importance of the feature and the information proportion value representing the need to be forgotten and updating the memory unit according to the candidate memory vector, using the sigmoid function through the output gate to linearly change the current vector input and the hidden state of the previous time step, and outputting an information activation value representing the importance of the output, transforming the current memory unit state through the tanh function, and multiplying it with the information activation value of the output gate to update and output the robot's behavior strategy in the current time step.
6. The robot control method based on multimodal large model according to claim 1, characterized in that: In step 5, the specific formula for updating the Q value of the robot to adopt the optimal behavior strategy is: ; in, represents the Q value of the behavior strategy a adopted by the robot under the updated current data s, represents the Q value of the behavior strategy a adopted by the robot under the current data s, represents the learning rate that controls the Q value update speed, It represents the immediate reward obtained by the robot after adopting the optimal behavior strategy. represents the discount factor for evaluating the impact of future rewards, It represents the maximum Q value of all possible behavior strategies a that the robot may take under the current data s.
Citation Information
Patent Citations
Robot control system and method, storage medium, controller and robot
CN118927246A
Intelligent robot operation management system based on electric power field
CN119148635A
Unknown dynamic environment mobile robot path planning method based on reinforcement learning
CN119413174A