Robot following method, device, electronic device and storage medium
By obtaining and processing robot following information in real time, using the speed control model of reinforcement learning training, the problem of susceptible to interference during the robot following process is solved, and the stable target follow-up and obstacle avoidance effects are achieved.
Patent Information
- Application Number
- CN202210647762.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-06-08
AI Technical Summary
Existing robot following methods are susceptible to disturbances in the environment, which leads to easy follow-up or missing targets, reducing follow-up stability.
By obtaining the motion following information of the robot following the target object in real time, including position, obstacle relationship and its own state, the preset speed control model is used for reinforcement learning training, output speed control signals, and control the robot to maintain a preset distance from the target and avoid collisions.
Improves the stability of robot following, ensuring that the robot remains within a preset distance from the target and avoids collision with obstacles.
Smart Images

Figure CN114815851B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a robot following method, device, electronic equipment and storage medium. Background Art
[0002] With the continuous emergence of various types of robots, such as logistics robots, patrol robots, etc., robots are usually required to follow to perform tasks.
[0003] Robots usually follow in relatively complex environments, which contain complex factors such as static and dynamic obstacles, targets and non-targets. The current solutions mainly use lidar data as input or visual information for navigation and obstacle avoidance. The following process is easily affected by interference from the environment, so it is easy to follow the wrong target or lose the target, which reduces the following stability. Summary of the Invention
[0004] The present invention provides a robot following method, device, electronic device and storage medium, which can realize acceleration and deceleration according to the target distance and angle and keep a certain distance from the target to complete variable speed following.
[0005] According to one aspect of the present invention, there is provided a robot following method, the method comprising:
[0006] Determine motion following information when the robot follows the target object; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state;
[0007] The motion following information is input into a preset speed control model, and a speed control signal for the robot is output; the speed control model is obtained by reinforcement learning training according to a preset reward function;
[0008] The robot is controlled to continue to follow the target object according to the speed control signal, so that the robot and the target object are kept within a preset distance range and the robot is prevented from colliding with surrounding obstacles.
[0009] According to another aspect of the present invention, there is provided a robot following device, the device comprising:
[0010] A motion following determination module is used to determine the motion following information of the robot when following the target object; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state;
[0011] A speed signal acquisition module is used to input the motion tracking information into a preset speed control model and output a speed control signal for the robot; the speed control model is obtained by reinforcement learning training according to a preset reward function;
[0012] The motion following control module is used to control the robot to continue to follow the target object according to the speed control signal, so that the robot and the target object remain within a preset distance range and prevent the robot from colliding with surrounding obstacles.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the robot following method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the robot following method according to any embodiment of the present invention when executed.
[0018] The technical solution of the embodiment of the present invention obtains motion following information of the robot in real time when following the movement of the target object; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state, and the motion following information is input into a preset speed control model to output a speed control signal for the robot; the speed control model is obtained by reinforcement learning training according to a preset reward function, and can control the robot to continue to follow the target object according to the speed control signal, and keep the robot and the target object within a preset distance range and prevent the robot from colliding with surrounding obstacles, thereby solving the problem that the following process is easily affected by interference objects in the environment, and is prone to following the wrong target or losing the target, thereby improving the following stability.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 is a flowchart of a robot following method provided according to an embodiment of the present invention;
[0022] Figure 2 is a schematic diagram of a learning process of an autoencoder in dynamic vision to which an embodiment of the present invention is applicable;
[0023] Figure 3 is a schematic diagram of a process for using an autoencoder in dynamic vision to which an embodiment of the present invention is applicable;
[0024] Figure 4 This is a schematic structural diagram of a robot following device provided according to an embodiment of the present invention;
[0025] Figure 5 It is a structural diagram of an electronic device for implementing the robot following method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "target", "current", "next", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0029] Currently, most path planning processes use raw camera image information as input, enabling machines to learn functions such as navigation and obstacle avoidance. A smaller portion uses LiDAR data as input. Navigation and obstacle avoidance in a mapped environment require pre-mapped environments, making them suitable only for indoor environments and not for outdoor navigation and obstacle avoidance. Directly using visual image modeling makes it difficult to transfer from simulated training environments to real-world environments. LiDAR is expensive, and its imaging performance is significantly affected by weather and lighting conditions. LiDAR operates on optical detection principles, and lasers can pass through transparent glass, resulting in a certain probability of missed detections, which can affect the robot's navigation and obstacle avoidance capabilities in real-world environments.
[0030] The robot following method, device, electronic device and storage medium provided in this application are described in detail below through various embodiments and optional solutions.
[0031] Figure 1 A flowchart of a robot following method is provided for an embodiment of the present invention. This embodiment is applicable to controlling a robot to follow a target object in real time. The method can be executed by a robot following device, which can be implemented in the form of hardware and / or software. The robot following device can be configured in any electronic device with network communication function. Figure 1 As shown, the robot following method in this embodiment may include the following steps:
[0032] S110 : Determine motion following information when the robot follows the target object.
[0033] Among them, motion tracking information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion status.
[0034] The robot can be a logistics robot, patrol robot, or household pet robot, and the target object can be a user who has control authority over the robot and wants to use it, or a user who needs the robot's assistance to perform a task. When the target object is moving (such as walking or running), the robot needs to follow the target object so that the robot can reach the target object as quickly as possible when the target object needs the robot to perform a task.
[0035] The robot's own motion state may include the robot's motion position, motion speed, motion posture, motion trajectory trend, etc. Surrounding obstacles may include but are not limited to static objects and moving objects around the robot.
[0036] In an optional solution of this embodiment, determining the motion following information when the robot follows the target object may include steps A1-A2:
[0037] Step A1: Obtain a depth map of the scene ahead of the robot, and reconstruct the depth map of the scene ahead through a preset encoder to obtain a latent vector feature of the scene ahead.
[0038] The robot is equipped with a depth camera, which captures the scene in front of it during movement, generating a depth image of the scene, including any obstacles in front of it. A depth image, also known as a range image, is an image whose pixel values are the distance (depth) from the depth camera to each point in the scene.
[0039] Step A2: Determine the motion relationship between the robot and surrounding obstacles based on the latent vector features of the front scene and the robot's motion state features.
[0040] The preset encoder can be based on a variational autoencoder or other encoder trained using depth images acquired from depth cameras in both real and simulated environments. The reconstruction function of the variational autoencoder automatically pre-processes the image to obtain the required latent vector, saving time in image processing. Furthermore, by simultaneously modeling both simulated and real depth maps to generate the preset variational autoencoder, compatibility between simulation and reality is achieved, facilitating the transition from simulation to reality.
[0041] In an optional manner of this embodiment, reconstructing the front scene depth map by a preset encoder to obtain the front scene latent vector feature may include steps B1-B3:
[0042] Step B1: Divide the front scene depth map into a ground area depth map and a non-ground area depth map.
[0043] Step B2: reconstruct the ground area depth map through the encoder corresponding to the ground area to obtain the latent vector features of the front ground area.
[0044] Step B3: reconstruct the non-ground area depth map through the encoder corresponding to the non-ground area to obtain the latent vector feature of the front non-ground area.
[0045] See also Figure 2Since the depth map contains ground and non-ground areas, their depth distributions vary greatly. It is difficult to balance the areas with large depth differences when acquiring latent vectors. Therefore, we can model the ground and non-ground areas separately and construct variational autoencoders corresponding to the ground and non-ground areas, so as to reconstruct the depth maps of the ground and non-ground areas respectively.
[0046] Optionally, a training scene depth map (including real scenes and simulated scenes) is obtained. The training scene depth map can be first divided into a training ground area depth map and a training non-ground area depth map. The training ground area depth map is reconstructed using a variational autoencoder, and the reconstructed latent vector features are used for decoding to obtain a preprocessed training ground area depth map. By comparing the preprocessed training ground area depth map with the pre-processed training ground area depth map, the parameters of the variational autoencoder are updated and adjusted to obtain a variational autoencoder corresponding to the appropriate ground area. Similarly, the same training method is adopted for the training non-ground area depth map to obtain a variational autoencoder corresponding to the non-ground area.
[0047] See also Figure 3 When reconstructing the depth map of the scene ahead, a region dividing line between the ground area and the non-ground area can be determined, and the depth map of the scene ahead can be divided into a ground area depth map and a non-ground area depth map using the region dividing line. The region dividing line can be predetermined based on the orientation of a depth camera mounted on the robot, the height of the depth camera, and the angle between the orientation of the depth camera and the ground. The region dividing line can be a fuzzy dividing line determined by identifying the depth values of pixels in the depth map of the scene ahead and performing statistical analysis on the depth values.
[0048] In an optional manner of this embodiment, determining the motion relationship between the robot and surrounding obstacles based on the latent vector features of the front scene and the robot's motion state features may include the following process:
[0049] The latent vector features of the ground area, the latent vector features of the non-ground area, and the robot's motion state features at the current moment are input into the pre-trained sequence model, and the latent vector used to characterize the motion relationship between the robot and the surrounding obstacles at the next moment is output.
[0050] A sequence model is used to model the ground area latent vector, non-ground area latent vector, and robot's motion trajectory generated during depth map reconstruction. This allows the model to estimate the position of obstacles around the robot in the future and output a latent vector representing the motion relationship between the robot and the obstacles at the next moment. The sequence model can be trained using a sequence model structure such as an LSTM.
[0051] Through dynamic visual obstacle avoidance, during the variable speed following process, the depth map is encoded into a latent vector describing the external environment, and a sequence model is used to learn the relative motion relationship between the robot and obstacles in the environment from the environmental latent vector and the robot's motion state and motion trajectory for subsequent variable zoom following.
[0052] In another optional solution of this embodiment, determining the motion following information when the robot follows the target object may include steps C1-C2:
[0053] Step C1: Determine the distance and angle of the target object relative to the robot in the robot coordinate system by using visual positioning or ultra-wideband technology.
[0054] Step C2: Determine whether there are obstacles within a preset distance range around the robot using an obstacle detection sensor.
[0055] Among them, obstacle detection sensors include but are not limited to ultrasonic sensors, lidar sensors, etc., which are used to detect whether there are obstacles around the robot that require the robot to avoid them urgently.
[0056] This approach eliminates the need for pre-built maps. Vision positioning or ultra-wideband (UWB) positioning devices can obtain the target object's distance and angle relative to the robot's coordinate system in real time, allowing the robot to accelerate and decelerate based on the target's distance and angle, maintaining a certain distance from the target and enabling variable speed following. Nearby emergency obstacle avoidance: Ultrasonic and other sensors detect obstacles within a certain range. When an obstacle appears within a safe range, emergency avoidance is performed. Multimodal data input ensures stability.
[0057] S120. Input the motion following information into a preset speed control model to output a speed control signal for the robot; wherein the preset speed control model is obtained by reinforcement learning training according to a preset reward function.
[0058] The inputs to reinforcement learning include: latent vectors generated by the sequence model that represent the motion relationship between stationary objects, moving objects, and the robot in the environment; the robot's motion state (including but not limited to linear velocity, angular velocity, etc.); the position of the target object in the robot's body coordinate system (including distance P and angle θ), where P represents the distance and θ represents the angle with the positive direction of the x-axis; and ultrasonic and other sensors that detect whether there are obstacles within a safe distance.
[0059] The output of reinforcement learning includes control signals for the robot, such as speed commands. When this is a speed command, the control signal includes the robot's linear velocity in the x-axis and y-axis directions, as well as its angular velocity in the yaw direction. The speed commands sent by the robot can be as follows: control_speed_x = delta_x * speed_x, control_speed_y = delta_y * speed_y, control_w = delta_w * speed_w. Delta can be adjusted based on the robot's performance and effectiveness in the real environment.
[0060] The preset speed control model is trained through reinforcement learning using a preset reward function. For example, it can be implemented using a reinforcement learning algorithm based on the actor-critic framework, such as PPO, TRPO, DDPG, or A3C. The preset reward function also includes instructions for maintaining the robot's distance from the target object within a preset range, ensuring smooth speed changes during the robot's speed change, and controlling the angle between the robot's preset direction and the target object's preset direction within a preset range. The reward function includes, but is not limited to, maintaining a certain distance between the robot and the target, avoiding collisions, ensuring smooth motion, and minimizing the angle between the target and the robot.
[0061] S130 , controlling the robot to continue following the target object according to the speed control signal, keeping the robot and the target object within a preset distance range, and preventing the robot from colliding with surrounding obstacles.
[0062] Reinforcement learning through reward functions enables the robot to maintain a certain distance from a moving target and automatically and smoothly accelerate and decelerate according to the distance. Furthermore, during the following process, the robot uses the dynamic vision module to proactively react to obstacles in the environment and plan an appropriate route.
[0063] According to the technical solution of the embodiment of the present invention, the motion following information of the robot following the target object is obtained in real time; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and the surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state, and the motion following information is input into a preset speed control model to output a speed control signal for the robot; the speed control model is obtained by reinforcement learning training according to a preset reward function, and can control the robot to continue to follow the target object according to the speed control signal, and keep the robot and the target object within a preset distance range and prevent the robot from colliding with surrounding obstacles, thereby solving the problem that the following process is easily affected by interference objects in the environment, and is prone to following the wrong target or losing the target, thereby improving the following stability.
[0064] Figure 4 The present invention provides a structural block diagram of a robot following device. This embodiment is applicable to controlling a robot to follow a target object in real time. The robot following device can be implemented in the form of hardware and / or software. The robot following device can be configured in any electronic device with network communication function. Figure 4 As shown, the robot following device in this embodiment may include the following: a motion following determination module 410, a speed signal acquisition module 420 and a motion following control module 430. Among them:
[0065] A motion following determination module 410 is configured to determine motion following information when the robot follows a target object; the motion following information is configured to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, the presence of obstacles around the robot, and the robot's own motion state;
[0066] The speed signal acquisition module 420 is used to input the motion tracking information into a preset speed control model and output a speed control signal for the robot; the speed control model is obtained by reinforcement learning training according to a preset reward function;
[0067] The motion following control module 430 is used to control the robot to continue to follow the target object according to the speed control signal, so that the robot and the target object remain within a preset distance range and prevent the robot from colliding with surrounding obstacles.
[0068] Based on the above embodiment, optionally, determining the motion following information when the robot follows the target object includes:
[0069] Obtain a depth map of the scene ahead of the robot, and reconstruct the depth map of the scene ahead using a preset variational autoencoder to obtain a latent vector feature of the scene ahead;
[0070] The motion relationship between the robot and surrounding obstacles is determined based on the latent vector features of the front scene and the robot's motion state features.
[0071] Based on the above embodiment, optionally, reconstructing the front scene depth map by a preset encoder to obtain the front scene latent vector feature includes:
[0072] Splitting the front scene depth map into a ground area depth map and a non-ground area depth map;
[0073] The ground area depth map is reconstructed through the encoder corresponding to the ground area to obtain the hidden vector features of the front ground area;
[0074] The non-ground area depth map is reconstructed through the encoder corresponding to the non-ground area to obtain the latent vector features of the front non-ground area.
[0075] Based on the above embodiment, optionally, determining the motion relationship between the robot and surrounding obstacles based on the latent vector features of the front scene and the robot's motion state features includes:
[0076] The latent vector features of the ground area, the latent vector features of the non-ground area, and the robot's motion state features at the current moment are input into the pre-trained sequence model, and the latent vector used to characterize the motion relationship between the robot and the surrounding obstacles at the next moment is output.
[0077] Based on the above embodiment, optionally, determining the motion following information when the robot follows the target object includes:
[0078] The distance and angle of the target object relative to the robot in the robot coordinate system are determined by visual positioning or ultra-wideband technology; the obstacle detection sensor is used to determine whether there are obstacles within a preset distance range around the robot.
[0079] Based on the above embodiment, optionally, the preset reward function also includes a function for indicating that the distance between the robot and the target object remains within a preset distance range, that the robot changes speed smoothly during speed change, and that the angle between the preset mark direction of the robot and the preset mark direction of the target object is controlled within a preset range.
[0080] Based on the above embodiment, optionally, the obstacle detection sensor is an ultrasonic sensor and a radar sensor.
[0081] The robot following device provided in the embodiment of the present invention can execute the robot following method provided in any embodiment of the present invention, and has the corresponding functions and beneficial effects of executing the robot following method. For detailed process, please refer to the relevant operations of the robot following method in the above embodiment.
[0082] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0083] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0084] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0085] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the robot following method.
[0086] In some embodiments, method XXX can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the robot following method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the robot following method in any other suitable manner (e.g., via firmware).
[0087] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0089] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0091] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0092] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0093] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0094] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A robot following method, characterized in that: The method comprises: Determine motion following information when the robot follows the target object; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state; The motion following information is input into a preset speed control model, and a speed control signal for the robot is output; the speed control model is obtained by reinforcement learning training according to a preset reward function; Controlling the robot to continue following the target object according to the speed control signal, so that the robot and the target object remain within a preset distance range and preventing the robot from colliding with surrounding obstacles; The determination of the motion following information when the robot follows the target object includes: Obtain a depth map of the scene ahead of the robot, and reconstruct the depth map of the scene ahead using a preset encoder to obtain a latent vector feature of the scene ahead; Determine the motion relationship between the robot and surrounding obstacles based on the latent vector features of the front scene and the robot's motion state features; The process of reconstructing the front scene depth map by a preset encoder to obtain the front scene latent vector feature includes: Splitting the front scene depth map into a ground area depth map and a non-ground area depth map; The ground area depth map is reconstructed through the encoder corresponding to the ground area to obtain the hidden vector features of the front ground area; The non-ground area depth map is reconstructed through the encoder corresponding to the non-ground area to obtain the hidden vector features of the non-ground area in front; The motion relationship between the robot and surrounding obstacles is determined based on the latent vector features of the front scene and the robot's motion state features, including: The latent vector features of the ground area, the latent vector features of the non-ground area, and the robot's motion state features at the current moment are input into the pre-trained sequence model, and the latent vector used to characterize the motion relationship between the robot and the surrounding obstacles at the next moment is output.
2. The method according to claim 1, characterized in that Determine the motion following information when the robot follows the target object, including: The distance and angle of the target object relative to the robot in the robot coordinate system are determined by visual positioning or ultra-wideband technology; the obstacle detection sensor is used to determine whether there are obstacles within a preset distance range around the robot.
3. The method according to claim 1, characterized in that The preset reward function also includes a function for indicating that the distance between the robot and the target object remains within a preset distance range, that the robot changes speed smoothly during speed change, and that the angle between the preset mark direction of the robot and the preset mark direction of the target object is controlled within a preset range.
4. The method according to claim 2, characterized in that The obstacle detection sensors are ultrasonic sensors and radar sensors.
5. A robot following device, characterized in that: The device comprises: A motion following determination module is used to determine the motion following information of the robot when following the target object; the motion following information is used to describe the position of the target object in the robot coordinate system, the relative motion relationship between the robot and surrounding obstacles, whether there are obstacles around the robot, and the robot's own motion state; A speed signal acquisition module is used to input the motion tracking information into a preset speed control model and output a speed control signal for the robot; the speed control model is obtained by reinforcement learning training according to a preset reward function; a motion following control module, configured to control the robot to continue following the target object according to the speed control signal, so as to keep the robot and the target object within a preset distance range and prevent the robot from colliding with surrounding obstacles; The determination of the motion following information when the robot follows the target object includes: Obtain a depth map of the scene ahead of the robot, and reconstruct the depth map of the scene ahead using a preset encoder to obtain a latent vector feature of the scene ahead; Determine the motion relationship between the robot and surrounding obstacles based on the latent vector features of the front scene and the robot's motion state features; The process of reconstructing the front scene depth map by a preset encoder to obtain the front scene latent vector feature includes: Splitting the front scene depth map into a ground area depth map and a non-ground area depth map; The ground area depth map is reconstructed through the encoder corresponding to the ground area to obtain the hidden vector features of the front ground area; The non-ground area depth map is reconstructed through the encoder corresponding to the non-ground area to obtain the hidden vector features of the non-ground area in front; The motion relationship between the robot and surrounding obstacles is determined based on the latent vector features of the front scene and the robot's motion state features, including: The latent vector features of the ground area, the latent vector features of the non-ground area, and the robot's motion state features at the current moment are input into the pre-trained sequence model, and the latent vector used to characterize the motion relationship between the robot and the surrounding obstacles at the next moment is output.
6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the robot following method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the robot following method according to any one of claims 1 to 4 when executed.
Citation Information
Patent Citations
Target following and dynamic obstacle avoidance control method for speed difference slip steering vehicle
CN110989576A
Pedestrian accompanying control method and device of robot, mobile robot and medium
CN113467462A
Automatic driving decision-making method fusing multi-source data and synthesizing multi-dimensional indexes
CN113743469A