Robot guiding method based on target detection

By mounting a camera and object detection model on the robot, combined with path planning and guidance control modules, the problem of low reliability of robot navigation in dynamic environments is solved, and efficient and safe robot guidance is achieved.

CN120044956APending Publication Date: 2025-05-27HUBEI NORMAL UNIV

Patent Information

Application Number
CN202510265784.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to achieve autonomous navigation and efficient guidance of robots in dynamic and complex environments, especially when target positions and environmental obstacles change rapidly, resulting in a reduced reliability of the robot guidance pace.

Method used

By loading a camera on the robot to obtain image data, using the target detection model to identify and locate target objects in the environment, combining the robot's current position information, using the path planning method to generate obstacle avoidance paths, and adjust the robot's motion trajectory in real time through the guidance control module.

Benefits of technology

It improves the navigation safety and adaptability of robots in complex environments, ensures that robots can quickly respond to environmental changes, avoid collisions, and improve task execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044956A_ABST
    Figure CN120044956A_ABST
Patent Text Reader

Abstract

The invention discloses a robot guiding method based on target detection, and the method comprises the following steps: S1, obtaining image data in a scene through a camera carried by a robot, and carrying out the preprocessing of the obtained image data; s2, inputting the preprocessed image data into a target detection model, and identifying and positioning a target object in the environment; s3, according to the target position output by the target detection model, combining the current position information of the robot, and using a path planning method to generate an obstacle avoidance path; and S4, controlling the robot to move according to the obstacle avoidance path through a guide control module, and adjusting the motion track of the robot in real time. The target object in the environment is identified and positioned through the target detection model, the obstacle avoidance path is generated by using a path planning method according to the target position and the current position information of the robot, the action of the robot is controlled, the movement track is adjusted in real time, and the pace guiding reliability of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot guidance, and particularly to a robot guidance method based on object detection. Background Art

[0002] With the rapid development of artificial intelligence and robot technology, more and more robots are applied in various industries, including industrial production, smart home, medical health and other fields. In these applications, how to achieve autonomous navigation and efficient guidance of robots has become a key issue.

[0003] Traditional robot guidance methods usually rely on pre-determined path planning to achieve robot positioning and navigation. However, with the complexity and dynamic changes of the environment, these traditional methods often face some challenges. For example, fixed landmarks may be lost, sensors may be interfered, or the robot needs to work in a more dynamic and unpredictable environment.

[0004] In this context, object detection technology provides new possibilities for robot guidance. By real-time identifying and tracking specific objects in the environment, object detection can help robots to conduct effective guidance in unknown or dynamic environments. Different from traditional methods, the guidance method based on object detection can react in real-time according to the changes in the environment, improving the flexibility and adaptability of the robot.

[0005] In the prior art, Chinese Patent No. CN119273097A discloses "An Asset Change Management Method, Product, Device and Medium for a Data Center". After obtaining the task information input by the current user, the robot matches the target task to be executed from each task to be executed stored in advance, and then guides the staff to quickly find the corresponding positions of the cabinet and the current target device in the computer room and perform corresponding operations on the target device. However, the above solution needs to store each task to be executed in advance. The robot needs to match the target task to be executed from each task to be executed stored in advance and then guide the staff. However, in a dynamic environment, there will be many unknown obstacles affecting the robot, reducing the reliability of the robot's guiding steps.

[0006] Therefore, it is urgent to design a robot guidance method based on object detection to solve the problems existing in the above prior art. Summary of the Invention

[0007] The purpose of the present invention is to provide a robot guidance method based on object detection, which uses an object detection model to identify and locate target objects in the environment. At the same time, according to the target position and combined with the current position information of the robot, a path planning method is used to generate an obstacle avoidance path, control the robot's movement, and adjust the movement trajectory in real-time, so as to improve the reliability of the robot's guiding steps.

[0008] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect of the present invention, a robot guidance method based on object detection is provided, and the method includes the following steps: S1: Use the camera carried by the robot to obtain image data in the scene and preprocess the obtained image data; S2: Input the preprocessed image data into the object detection model to identify and locate the target objects in the environment; S3: According to the target positions output by the object detection model, combined with the current position information of the robot, use a path planning method to generate an obstacle avoidance path; S4: Control the movement of the robot through the guidance control module according to the obstacle avoidance path generated by the path planning method and adjust the movement trajectory of the robot in real time.

[0009] As an embodiment of the present application, step S1 specifically includes: S11: Select an RGB camera suitable for color image acquisition in a standard environment; S12: Use the RGB camera carried by the robot to capture image data in real time and output a continuous video stream. The RGB camera captures 30 frames of images per second, and the image data is captured continuously at the frame rate. The size formula of each frame of image is expressed as follows:

[0010] where the pixels of each frame of image are , is the width, is the height, and 3 represents the three RGB channels; Each frame of image is represented as a matrix, and the matrix formula is expressed as follows:

[0011] where represents the frame image at the th time in the image, is the timestamp; S13: Use an image enhancement technology to improve the brightness distribution of the image and make the pixel value distribution of the image uniform, and the formula is expressed as follows:

[0012] where represents the video frame after enhancement processing; represents the maximum value of all pixel values of the video frame; represents the minimum value of all pixel values of the video frame, is the amplification factor, which is used to adjust the overall brightness and contrast of the enhanced image; is the reduction factor, which is used to adjust the dynamic range of the image; S14: Normalize the image, scale the pixel values from [0, 255] to [0, 1], and the formula is as follows: .

[0013] As an embodiment of the present application, the step S2 specifically includes: S21: Input the preprocessed frame image into the target detection model. The size of the frame image is , where is the width, is the height, and is the channel; S22: Use a convolutional neural network as the backbone network to extract the features of the frame image. Through the convolutional outputs of different levels, a series of feature maps from low level to high level are obtained, and the feature maps from different layers are fused to form multi-scale feature maps; S23: Design a multi-scale feature position encoding, fuse the spatial position information of the multi-scale feature maps extracted by the convolutional neural network, and perform operations with the corresponding position encoding and send them to the encoder for processing; S24: The encoder receives the multi-scale feature maps containing the position encoding, introduces a multi-scale self-attention module, and combines the feed-forward neural network to stack multiple levels of self-attention modules; S25: Through the decoder, the convolutional target query vector interacts with the multi-scale feature maps to perform target classification, processes targets of different sizes and scales, and outputs the category and bounding box of each target; and design a loss function to optimize the detection accuracy of the target detection model for targets of different scales.

[0014] As an embodiment of the present application, the step S22 specifically includes: S221: Use a 1x1 convolutional kernel to gradually perform downsampling convolution on the frame image, and perform pooling operations on each layer of the image to extract low-level features; S222: Use a 3x3 convolutional kernel to gradually perform upsampling convolution on the frame image that has passed through the 1x1 convolutional kernel, and perform pooling operations on each layer of the image to restore high-level features; S223: Concatenate the upsampled high-level features and the low-level features to form multi-scale feature maps, and the formula is as follows:

[0015]

[0016]

[0017] Among them, is the input frame image, is the feature map obtained after 1x1 convolution and linear transformation, is the feature map obtained after 3x3 convolution and linear transformation, is and the multi-scale feature map obtained by concatenation, and are learnable weight matrices.

[0018] As an embodiment of the present application, the step S23 specifically includes: S231: Generate corresponding position encodings for feature maps of each scale. The position encodings are generated by sine and cosine functions, and the calculation formula for the position encodings is as follows:

[0019]

[0020] Among them, represents the horizontal and vertical coordinates of the position; represents the index of the dimension; represents the total dimension of the position encoding, indicates that the even position encoding is calculated using the cosine function, indicates that the odd position encoding is calculated using the sine function; S232: Use element-wise addition to fuse the feature map of each scale with the spatial position encoding corresponding to that scale to generate a feature map containing spatial position information and image content information. The calculation formula is as follows:

[0021] Among them, is the feature map containing the position encoding of the th scale, is the position encoding of the feature map of the th scale; is the feature map of the th scale in the multi-scale feature map th.

[0022] As an embodiment of the present application, the step S24 specifically includes: S241: For the multi-scale feature maps containing position encoding received by the encoder, the feature maps of each scale are respectively input into the corresponding self-attention module. In each self-attention module, first, the input feature maps are linearly mapped into three matrices respectively, and the formula is expressed as follows:

[0023]

[0024]

[0025] Among them, is the feature map containing position encoding of the th scale, , , are the linear transformation matrices of query, key, and value of the th scale respectively, , , are the query, key, and value matrices respectively; Then, calculate the similarity scores between the query and the key, and obtain the attention weights through the scaled dot-product attention mechanism. The calculation formula is as follows:

[0026] Among them, is the transposed matrix, is the scaling factor; S242: Perform layer normalization on the output features obtained through multi-scale self-attention calculation, input the features after layer normalization into the feed-forward neural network, perform layer normalization operation on the output of the feed-forward neural network again, add the results of the two layer normalizations to obtain the output of the current-level self-attention module, and stack the structure composed of the above self-attention module and the feed-forward neural network to form multiple levels of self-attention modules.

[0027] As an embodiment of the present application, the step S25 specifically includes: S251: Initialize a group of learnable target query vectors, initialize them through the convolution kernel, and set their dimensions to be the same as the number of channels of the input feature map; S252: Use the multi-head cross-attention mechanism to directly interact the target query vectors initialized by the convolution kernel with the multi-scale feature maps. During this process, the target query vectors are adaptively adjusted according to the content of the input feature maps; S253: Concatenate the context features obtained through the multi-head cross-attention mechanism with the target query vectors, and then input the concatenated result into the fully connected layer. After calculation by the fully connected layer, directly predict the class scores and bounding boxes of the targets; S254: A loss function designed includes a classification loss function and a bounding box regression loss function, and the loss function formula The expression is as follows:

[0028]

[0029]

[0030] in, is the classification loss function, is the bounding box regression loss function, is the total number of categories; is the balancing factor of the category; is the true category label; is the probability of the class predicted by the model; is the weight of the entropy regularization term, , , , They are the horizontal coordinate, vertical coordinate, height, and width of the center point of the prediction box. , , , They are the horizontal coordinate, vertical coordinate, height, and width of the center point of the real frame.

[0031] As an embodiment of the present application, step S3 specifically includes: S31: Obtain the target object output by the target detection model, automatically calculate the frame set in the target detection through the positioning device of the robot, and merge it with the current position data of the robot; S32: Calculate a path using a path planning method according to the position of the target object and environmental obstacle information.

[0032] As an embodiment of the present application, step S32 specifically includes: S321: Define a dynamic window, the range of speed and steering angle that the robot can choose at the current moment, including the range of linear speed that the robot can choose and the range of angular speed that the robot can choose. The specific calculation formula is as follows:

[0033]

[0034]

[0035]

[0036] in, is the current linear velocity of the robot; is the maximum acceleration of the robot; is the current angular velocity of the robot; is the maximum angular velocity of the robot; and are the minimum and maximum limits of speed respectively; and are the minimum and maximum limits of angular velocity, respectively; S322: According to the dynamic window, the path planning method generates a series of possible trajectories based on the current state and the dynamic window. These trajectories represent the possible movements of the robot starting from the current state in the future. Each trajectory records the position and speed changes of the robot. The specific calculation formula is as follows:

[0037]

[0038] in, and is the current position coordinate of the robot; and is the position coordinate of the target; is a weight coefficient that controls the influence of target proximity on the total cost; and It is at the moment Changes in the robot's linear and angular velocities; is a weight coefficient that controls the influence of smoothness on the total cost; It is a weight coefficient that controls the influence of obstacle avoidance cost in the total cost. The total cost is finally obtained through the weight coefficient. The optimal path is selected in the path planning process according to the sum of the total cost.

[0039] As an embodiment of the present application, step S4 specifically includes: S41: Obtain the current position of the robot in real time through the robot's built-in positioning device, including its coordinates and orientation; S42: Obtaining an obstacle avoidance path from the path planning method, the guidance control module generates a control instruction for the robot movement and sends it to the robot's motion control system; S43: Each time after the robot's motion control system executes a control instruction, the robot's own positioning device acquires the robot's new position in real time, and the guidance control module continuously updates the control instruction according to the new position error.

[0040] The beneficial effects of the present invention are: (1) The present invention first obtains images through a camera and passes them into a target detection model, and uses a convolutional neural network to extract multi-scale feature maps; low-level features are downsampled through 1x1 convolution, and high-level features are upsampled through 3x3 convolution, and the two are spliced ​​to form a multi-scale feature map; the feature map of each scale is combined with the position encoding, and further processed through a self-attention mechanism and a feedforward neural network. Finally, the convolution target query vector interacts with the multi-scale feature map to perform target classification and bounding box regression. The model optimizes the target detection accuracy through classification loss and bounding box regression loss.

[0041] (2) The present invention combines the target position output by the target detection model with the current position information of the robot and uses a path planning method to generate an obstacle avoidance path. During the path planning process, the target object position output by the target detection model is firstly fused with the current position data of the robot to calculate the target frame position. The path planning method performs path calculation based on the target position and environmental obstacle information. The dynamic window is defined to determine the speed and angle range that the robot can select, taking into account the limitations of the current speed, maximum acceleration and angular velocity. Then, the method generates a series of possible trajectories and performs weighted calculation based on the target proximity, smoothness and obstacle avoidance cost. The trajectory with the smallest total cost is selected as the optimal path to ensure that the robot approaches the target smoothly while avoiding obstacles. This method can take into account the position of the target object and environmental obstacles in real time, thereby generating an optimal path to avoid collisions and significantly improve the navigation safety of the robot in complex environments.

[0042] (3) The present invention obtains the current position of the robot, including its coordinates and orientation, in real time through the robot's built-in positioning device, obtains the obstacle avoidance path from the path planning method, generates control instructions for the robot's movement through the guidance control module and sends them to the robot's motion control system; each time the robot's motion control system executes the control instruction, it obtains the robot's new position and continuously updates the control instruction based on the new position error, which can quickly respond to environmental changes and ensure that the robot always moves along a safe and efficient path; this real-time adjustment capability significantly improves the robot's adaptability and flexibility in dynamic scenarios.

[0043] (4) The present invention can reduce the time and energy consumption of the robot in the process of searching for targets and avoiding obstacles through accurate target detection and efficient path planning; the robot can reach the target location and complete the task more quickly, thereby improving the efficiency of task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic flow chart of a robot guidance method based on target detection provided in an embodiment of the present invention; Figure 2Flowchart for obtaining scenario information of a robot guidance method based on object detection provided in an embodiment of the present invention; Figure 3 Flowchart for object detection of a robot guidance method based on object detection provided in an embodiment of the present invention; Figure 4 Flowchart for path planning method of a robot guidance method based on object detection provided in an embodiment of the present invention. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0047] In the present invention, unless otherwise clearly defined and limited, terms such as "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0048] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that satisfies both A and B at the same time. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0049] Reference Figures 1 to 4 , the first aspect of the present invention provides a robot guiding method based on object detection, and the method includes the following steps: S1: Use the camera mounted on the robot to obtain image data in the scene, and preprocess the obtained image data; S2: Input the preprocessed image data into the object detection model to identify and locate the target objects in the environment; S3: According to the target positions output by the object detection model, combined with the current position information of the robot, use the path planning method to generate an obstacle avoidance path; S4: Control the robot's movement through the guiding control module according to the obstacle avoidance path generated by the path planning method, and adjust the robot's movement trajectory in real time.

[0050] Specifically, the present invention uses the camera mounted on the robot to obtain image data in the scene, preprocesses the obtained image data and then inputs it into a designed object detection model to identify and locate the target objects in the environment. At the same time, according to the target positions output by the object detection model, combined with the current position information of the robot, use the path planning method to generate an obstacle avoidance path, and control the robot's movement through the guiding control module according to the path planning result, and adjust the robot's movement trajectory in real time to improve the reliability of the robot guiding steps.

[0051] As Figure 2 shown, as an embodiment of the present application, the step S1 specifically includes: S11: Select an RGB camera suitable for color image acquisition in a standard environment; S12: Use the RGB camera mounted on the robot to capture image data in real time and output a continuous video stream. The RGB camera captures 30 frames of images per second, and the image data is captured continuously at the frame rate. The size formula of each frame of image is expressed as follows:

[0052] where the pixels of each frame of image are , is the width, is the height, and 3 represents the three RGB channels; Each frame of image is represented as a matrix, and the matrix formula is expressed as follows:

[0053] where, represents the frame image at the th time in the image, is the timestamp; S13: Preprocess the collected image data to improve the effect of subsequent object detection. Use an image enhancement technique to improve the brightness distribution of the image. This operation makes the pixel value distribution of the image uniform and enhances the details of the image. The formula is as follows:

[0054] Where, represents the video frame after enhancement processing; after each frame of image undergoes this processing, its brightness and contrast will be adjusted according to the amplification factor and reduction factor; represents the maximum value of all pixel values of the video frame; represents the minimum value of all pixel values of the video frame, is the amplification factor, which is used to adjust the overall brightness and contrast of the enhanced image; is the reduction factor, which is used to adjust the dynamic range of the image; S14: Then perform normalization processing on the image, scaling the pixel values from [0, 255] to [0, 1]. The formula is as follows: .

[0055] Specifically, in the robot vision system of the present invention, a camera is used to collect color images in a standard environment, and image data is captured in real time through the camera, outputting a continuous video stream; each frame of image is represented in matrix form, where the size of each frame of image consists of the image width, height, and three RGB channels; 30 frames of images are collected per second, and the image data is continuously captured at the frame rate; in order to improve the effect of subsequent object detection, the image is enhanced, the brightness distribution is adjusted, and the enhancement formula adjusts the image brightness and contrast according to the maximum and minimum values of the image, using the amplification factor and reduction factor; then, the image is normalized for subsequent processing and analysis.

[0056] As Figure 3 shown, as an embodiment of the present application, the step S2 specifically includes: S21: Input the preprocessed frame image into the object detection model. The size of the frame image is , is the width, is the height, is the channel; S22: Use a convolutional neural network as the backbone network to extract the features of the frame image. Through the convolutional outputs of different levels, a series of feature maps from low level to high level are obtained, and the feature maps from different layers are fused to form multi-scale feature maps; S23: Design a multi-scale feature position encoding, fuse the spatial position information of the multi-scale feature maps extracted by the convolutional neural network, and perform operations with the corresponding position encoding and send it to the encoder for processing; S24: The encoder receives the multi-scale feature maps containing the position encoding, introduces a multi-scale self-attention module, and stacks the feed-forward neural network to form multiple levels of self-attention modules; S25: Through the decoder, the convolutional target query vector interacts with the multi-scale feature maps to perform target classification, processes targets of different sizes and scales, and outputs the category and bounding box of each target; and design a loss function to optimize the detection accuracy of the target detection model for targets of different scales.

[0057] As an embodiment of the present application, the step S22 specifically includes: S221: Use a 1x1 convolutional kernel to gradually perform downsampling convolution on the frame image, and perform pooling operations on each layer of the image. The output of each layer represents the features of the input image at different scales, and low-level features are extracted; S222: Use a 3x3 convolutional kernel to gradually perform upsampling convolution on the frame image that has passed through the 1x1 convolutional kernel, and perform pooling operations on each layer of the image to restore high-level features; S223: Concatenate the upsampled high-level features and the low-level features to form a multi-scale feature map. The formula is expressed as follows:

[0058]

[0059]

[0060] Among them, is the input frame image, is the feature map obtained after 1x1 convolution and linear transformation, is the feature map obtained after 3x3 convolution and linear transformation, is and the multi-scale feature map obtained by concatenation, and are learnable weight matrices for processing the feature maps.

[0061] As an embodiment of the present application, the step S23 specifically includes: S231: Generate corresponding position encodings for the feature maps of each scale. The position encodings are generated by sine and cosine functions. The position encoding calculation formula is as follows:

[0062]

[0063] Among them, represents the horizontal and vertical coordinates of the position; represents the index of the dimension; represents the total dimension of the position encoding, indicating that the even position encoding is calculated using the cosine function, indicating that the odd position encoding is calculated using the sine function; S232: Use element-wise addition to fuse the feature map of each scale with the spatial position encoding corresponding to that scale to generate a feature map containing spatial position information and image content information. The calculation formula is as follows:

[0064] Among them, is the th scale's feature map containing position encoding, is the th scale's position encoding of the feature map; is the th scale's feature map in the multi-scale feature map Each scale's feature map is combined with the corresponding spatial position encoding.

[0065] As an embodiment of the present application, the step S24 specifically includes: S241: For the multi-scale feature map containing position encoding received by the encoder, input the feature map of each scale into the corresponding self-attention module respectively. In each self-attention module, first linearly map the input feature map into three matrices respectively. The formula is expressed as follows:

[0066]

[0067]

[0068] Among them, is the th scale's feature map containing position encoding, , , are respectively the th scale's linear transformation matrices of query, key, and value, , , are respectively the query, key, and value matrices; Then, calculate the similarity score between the query and the key, and obtain the attention weights through the scaled dot-product attention mechanism. The calculation formula is as follows:

[0069] Among them, is the transposed matrix, is the scaling factor, Multiply the output attention weights by the value vectors to obtain the final output; S242: Perform layer normalization on the output features obtained through multi-scale self-attention calculation. Input the features after layer normalization into the feed-forward neural network, and perform layer normalization on the output of the feed-forward neural network again. Add the results of the two layer normalizations to obtain the output of the self-attention module at the current level. Stack the structure composed of the above self-attention module and the feed-forward neural network to form multiple levels of self-attention modules.

[0070] As an embodiment of the present application, the step S25 specifically includes: S251: Initialize a group of learnable target query vectors, initialize them through convolutional kernels, and set their dimensions to be the same as the number of channels of the input feature map; S252: Use the multi-head cross-attention mechanism to enable the target query vectors initialized by the convolutional kernels to directly interact with the multi-scale feature maps. During this process, the target query vectors are adaptively adjusted according to the content of the input feature maps; S253: Concatenate the context features obtained through the multi-head cross-attention mechanism with the target query vectors, and then input the concatenated results into the fully connected layer. After calculation by the fully connected layer, directly predict the class scores and bounding boxes of the targets; S254: Design a loss function including a classification loss function and a bounding box regression loss function to optimize the performance of the object detection model. The final loss function is the weighted sum of the classification loss and the bounding box regression loss. The loss function formula is expressed as follows:

[0071]

[0072]

[0073] Among them, is the classification loss function, is the bounding box regression loss function, is the total number of classes; is the class balance factor; is the true class label; is the probability of the class predicted by the model; is the weight of the entropy regularization term, , , , are respectively the abscissa, ordinate, height, and width of the center point of the predicted bounding box, , , , are respectively the abscissa, ordinate, height, and width of the center point representing the ground truth bounding box.

[0074] Specifically, in object detection, the acquired image data is input into the object detection model to identify and locate the target objects in the image; the object detection model analyzes the input image, identifies the target objects, and outputs their category and location information, and realizes the location through the image annotation area. Each image annotation area consists of the center coordinates, width, and height of the target object, and each target object is assigned a category label; there are multiple target objects in the image, and the object detection model will detect each target separately and assign a unique identifier; in each frame, the model will detect multiple target objects, generate multiple bounding boxes and category labels, and ensure the matching and association of the target objects between multiple frames by calculating the overlap degree of the target bounding boxes between different frames; specifically, first, the image is obtained through the camera and input into the object detection model, and the convolutional neural network is used to extract the multi-scale feature map. The low-level features are downsampled by 1x1 convolution, and the high-level features are upsampled by 3x3 convolution. The two are concatenated to form the multi-scale feature map. Each scale of the feature map is combined with the position encoding and further processed through the self-attention mechanism and the feed-forward neural network. Finally, the target query vector interacts with the multi-scale feature map for target classification and bounding box regression, and the model optimizes the object detection accuracy through the classification loss and the bounding box regression loss.

[0075] As Figure 4 shown, as an embodiment of the present application, step S3 specifically includes: S31: Obtain the target objects output by the object detection model, automatically calculate the set bounding boxes in the object detection through the positioning device of the robot, and fuse them with the current position data of the robot; S32: Perform path calculation using a path planning method according to the position of the target object and the environmental obstacle information.

[0076] As an embodiment of the present application, step S32 specifically includes: S321: Define a dynamic window, which is the range of the speed and steering angle that the robot can select at the current moment, including the linear velocity range that the robot can select and the angular velocity range that the robot can select. The specific calculation formula is as follows:

[0077]

[0078]

[0079]

[0080] Among them, is the current linear velocity of the robot; is the maximum acceleration of the robot; is the current angular velocity of the robot; is the maximum angular velocity of the robot; and are the minimum and maximum limits of the velocity respectively; and are the minimum and maximum limits of the angular velocity respectively; S322: According to the dynamic window, the path planning method generates a series of possible trajectories based on the current state and the dynamic window. These trajectories represent the possible movements of the robot starting from the current state within a certain period of time in the future. Each trajectory records the position and velocity changes of the robot. The specific calculation formula is as follows:

[0081]

[0082] Among them, and are the current position coordinates of the robot; and are the position coordinates of the target; is a weight coefficient that controls the influence degree of the target proximity in the total cost; and are the changes in the linear velocity and angular velocity of the robot at time ; is a weight coefficient that controls the influence degree of the smoothness in the total cost; is a weight coefficient that controls the influence degree of the obstacle avoidance cost in the total cost. The total cost is finally obtained through the weighted coefficients. The optimal path is selected according to the sum of the total costs during the path planning process.

[0083] Specifically, in the path planning process, the target object position output by the target detection model is first fused with the robot's current position data to calculate the target frame position; the path planning method performs path calculation based on the target position and environmental obstacle information; by defining a dynamic window, the robot's selectable speed and angle range are determined, taking into account the limitations of the current speed, maximum acceleration, and angular velocity; then, the method generates a series of possible trajectories and performs weighted calculations based on target proximity, smoothness, and obstacle avoidance costs, selecting the trajectory with the smallest total cost as the optimal path to ensure that the robot approaches the target smoothly while avoiding obstacles.

[0084] As an embodiment of the present application, step S4 specifically includes: S41: Obtain the current position of the robot in real time through the robot's built-in positioning device, including its coordinates and orientation; S42: Obtaining an obstacle avoidance path from the path planning method, the guidance control module generates a control instruction for the robot movement and sends it to the robot's motion control system; S43: Each time after the robot's motion control system executes the control instruction, the robot's new position is acquired in real time through the robot's own positioning device, and the guidance control module continuously updates the control instruction according to the new position error.

[0085] The present invention can reduce the time and energy consumption of the robot in the process of finding the target and avoiding obstacles through accurate target detection and efficient path planning. The robot can reach the target position and complete the task faster, thereby improving the efficiency of task execution.

[0086] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A robot guidance method based on target detection, characterized in that: The method comprises the following steps: S1: using the camera carried by the robot to obtain image data in the scene, and preprocessing the obtained image data; S2: Input the preprocessed image data into the target detection model to identify and locate the target objects in the environment; S3: Generate an obstacle avoidance path using a path planning method based on the target position output by the target detection model and the current position information of the robot; S4: The robot is controlled by guiding the control module according to the obstacle avoidance path generated by the path planning method, and the movement trajectory of the robot is adjusted in real time.

2. A robot guidance method based on target detection according to claim 1, characterized in that: The step S1 specifically includes: S11: Select an RGB camera suitable for color image acquisition in a standard environment; S12: The RGB camera carried by the robot captures image data in real time and outputs a continuous video stream. The RGB camera collects 30 frames of images per second, and the image data is continuously captured at the frame rate. The size formula of each frame of the image is expressed as follows: The pixels of each frame are , is the width, is the height, 3 represents the three RGB channels; Each frame of the image is represented as a matrix, and the matrix formula is expressed as follows: in, Indicates the image The frame image of time, is the timestamp; S13: Use an image enhancement technology to adjust the brightness distribution of the image so that the pixel value distribution of the image is uniform. The formula is as follows: in, Represents the enhanced video frame; Indicates the maximum value of all pixel values ​​in the video frame; Indicates the minimum value of all pixel values ​​in the video frame. is the magnification factor, which is used to adjust the overall brightness and contrast of the enhanced image; is the reduction factor, used to adjust the dynamic range of the image; S14: Normalize the image and scale the pixel value from [0,255] to [0,1]. The formula is as follows: 。 3. A robot guidance method based on target detection according to claim 1, characterized in that: The step S2 specifically includes: S21: The pre-processed frame image Input to the target detection model, the frame image The size is , is the width, is the height, For the channel; S22: Use a convolutional neural network as the backbone network to extract the features of the frame image. Through the convolution outputs at different levels, a series of feature maps from low to high layers are obtained. The feature maps from different layers are fused to form a multi-scale feature map. S23: Design a multi-scale feature position encoding, fuse the multi-scale feature map extracted by the convolutional neural network with the spatial position information, and perform operations with the corresponding position encoding to send it to the encoder for processing; S24: The encoder receives a multi-scale feature map containing position encoding, introduces a multi-scale self-attention module, and combines the feedforward neural network stack to form a multi-level self-attention module; S25: Through the decoder, the convolutional target query vector interacts with the multi-scale feature map to perform target classification, process targets of different sizes and scales, and output the category and bounding box of each target; and design a loss function to optimize the detection accuracy of the target detection model for targets of different scales.

4. A robot guidance method based on target detection according to claim 3, characterized in that: The step S22 specifically includes: S221: Use a 1x1 convolution kernel to gradually downsample the frame image, and perform a pooling operation on each layer of the image to extract low-level features; S222: Use a 3x3 convolution kernel to gradually upsample the frame image that has passed through the 1x1 convolution kernel, and perform a pooling operation on each layer of the image to restore high-level features; S223: The upsampled high-level features are concatenated with the low-level features to form a multi-scale feature map. The formula is as follows: in, is the input frame image, yes The feature map obtained by 1x1 convolution and linear transformation, yes The feature map obtained by 3x3 convolution and linear transformation, yes and The multi-scale feature map obtained by splicing, and is a learnable weight matrix.

5. A robot guidance method based on target detection according to claim 4, characterized in that: The step S23 specifically includes: S231: Generate corresponding position codes for feature maps of each scale. The position codes are generated by sine and cosine functions. The position code calculation formula is as follows: in, The horizontal and vertical coordinates representing the position; Indicates the index of the dimension; represents the total dimension of the position encoding, Indicates that the even position encoding is calculated using the cosine function, Indicates that the odd position code is calculated using a sine function; S232: The feature map of each scale is fused with the spatial position code corresponding to the scale by element-by-element addition to generate a feature map containing spatial position information and image content information. The calculation formula is as follows: in, It is The feature map containing position encoding at each scale, It is Position encoding of feature maps at each scale; It is a multi-scale feature map Middle Feature maps of different scales.

6. A robot guidance method based on target detection according to claim 5, characterized in that: The step S24 specifically includes: S241: For the multi-scale feature map containing position encoding received by the encoder, the feature map of each scale is input into the corresponding self-attention module. In each self-attention module, the input feature map is first linearly mapped into three matrices, and the formula is expressed as follows: in, It is The feature map containing position encoding at each scale, , , They are The linear transformation matrix of query, key, and value at each scale, , , They are query, key, and value matrices respectively; Then the similarity score between the query and the key is calculated, and the attention weight is obtained by the scaled dot product attention mechanism. The calculation formula is as follows: in, is the transposed matrix, is the scaling factor; S242: Perform layer normalization on the output features obtained through multi-scale self-attention calculation, input the layer-normalized features into the feedforward neural network, perform layer normalization on the output of the feedforward neural network again, add the results of the two layer normalizations to obtain the output of the self-attention module at the current level, stack the structure composed of the above self-attention module and the feedforward neural network to form self-attention modules at multiple levels.

7. A robot guidance method based on target detection according to claim 6, characterized in that: The step S25 specifically includes: S251: Initialize a set of learnable target query vectors, initialize them through convolution kernels, and set their dimensions to be consistent with the number of channels of the input feature map; S252: Use a multi-head cross attention mechanism to make the target query vector initialized by the convolution kernel directly interact with the multi-scale feature map. In this process, the target query vector is adaptively adjusted according to the content of the input feature map; S253: Concatenate the context features obtained through the multi-head cross attention mechanism with the target query vector, and then input the concatenation result into the fully connected layer. After calculation by the fully connected layer, the category score and bounding box of the target are directly predicted; S254: A loss function designed includes a classification loss function and a bounding box regression loss function, and the loss function formula The expression is as follows: in, is the classification loss function, is the bounding box regression loss function, is the total number of categories; is the balancing factor of the category; is the true category label; is the probability of the class predicted by the model; is the weight of the entropy regularization term, , , , They are the horizontal coordinate, vertical coordinate, height, and width of the center point of the prediction box. , , , They are the horizontal coordinate, vertical coordinate, height, and width of the center point of the real frame.

8. The robot guidance method based on target detection according to claim 1, characterized in that: The step S3 specifically includes: S31: Obtain the target position output by the target detection model, automatically calculate the frame set in the target detection through the positioning device of the robot, and merge it with the current position data of the robot; S32: Calculate a path using a path planning method according to the position of the target object and environmental obstacle information.

9. The robot guidance method based on target detection according to claim 8, characterized in that: The step S32 specifically includes: S321: Define a dynamic window, the range of speed and steering angle that the robot can choose at the current moment, including the range of linear speed that the robot can choose and the range of angular speed that the robot can choose. The specific calculation formula is as follows: in, is the current linear velocity of the robot; is the maximum acceleration of the robot; is the current angular velocity of the robot; is the maximum angular velocity of the robot; and are the minimum and maximum limits of speed respectively; and are the minimum and maximum limits of angular velocity, respectively; S322: According to the dynamic window, the path planning method generates a series of trajectories based on the current state and the dynamic window. These trajectories represent the movement of the robot from the current state in the future. Each trajectory records the position and speed changes of the robot. The specific calculation formula is as follows: in, and is the current position coordinate of the robot; and is the position coordinate of the target; is a weight coefficient that controls the influence of target proximity on the total cost; and It is at the moment Changes in the robot's linear and angular velocities; is a weight coefficient that controls the influence of smoothness on the total cost; It is a weight coefficient that controls the influence of obstacle avoidance cost in the total cost. The total cost is finally obtained through the weight coefficient. The optimal path is selected in the path planning process according to the sum of the total cost.

10. The robot guidance method based on target detection according to claim 1, characterized in that: The step S4 specifically includes: S41: Obtain the current position of the robot in real time through the robot's own positioning device, including its coordinates and orientation; S42: Obtaining an obstacle avoidance path from the path planning method, the guidance control module generates a control instruction for the robot movement and sends it to the robot's motion control system; S43: Each time after the robot's motion control system executes a control instruction, the robot's own positioning device acquires the robot's new position in real time, and the guidance control module continuously updates the control instruction according to the new position error.

Citation Information

Patent Citations

  • Asset change management method of data center, product, equipment and medium

    CN119273097A

  • Scene danger level assessment method

    CN116109913A

  • Target object recognition model construction method and device and computer equipment

    CN116740487A

  • Security check contraband detection method based on multi-scale attention and data enhancement

    CN116883933A

  • Handwritten mathematical formula identification method

    CN117542064A

Cited By

  • Robot motion control method, electronic equipment and storage medium

    CN121670637A