Mobile robot field positioning method, system, device and medium

CN118334121BActive Publication Date: 2026-09-25XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410477889.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2026-09-25
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

[0006]有鉴于此,本公开实施例提供一种移动机器人野外定位方法、系统、设备及介质,至少部分解决现有技术中存在模型训练效率和定位性能较差的问题

Benefits of technology

[0019]本公开实施例的有益效果为:通过本公开的方案,通过充分利用参数计算过程中的梯度方差信息,启发式调整每次迭代中的参数更新,使得每次迭代中具有更小方差的梯度分量被赋予更高的权重,而具有更大方差的梯度分量被赋予更小的权重,进而加快模型参数的收敛效率,缩短训练过程,提高了定位性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118334121B_ABST
    Figure CN118334121B_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure provides a mobile robot field positioning method, system, device and medium, which belongs to the technical field of measurement, and specifically comprises: shooting an image of a reference object as training data; recording labels corresponding to the training data to form a training set in combination with the training data; constructing an image detection model and inputting the training set for training; inputting an image shot by a mobile robot into the trained image detection model to obtain a probability of existence of a reference object in the image; when the probability of existence of the reference object exceeds a first threshold value, marking the position of the reference object in the image, the position being marked by an anchor box, and calculating the proportion of the area of the anchor box to the area of the image, when the proportion exceeds a second threshold value, positioning the mobile robot near the reference object; and according to the category of the reference object, searching for pre-stored coordinate information corresponding to the reference object to obtain a position estimation of the mobile robot. Through the scheme of the present disclosure, the model training efficiency and positioning performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of measurement technology, and in particular to a method, system, device and medium for positioning a mobile robot in the field. Background Technology

[0002] Currently, localization is a fundamental and unavoidable problem in the field of mobile robotics. One of the prerequisites for a mobile robot to successfully complete a task is its ability to perceive its own position and attitude within the environment. Utilizing the robot's attitude information allows for higher-level planning and control of the task at hand, highlighting the significant importance of localization research in mobile robotics. Various methods exist for mobile robot localization, with the most widely used including positioning using BeiDou or GPS satellite signals and calculating positioning based on the robot's onboard encoders and inertial sensing units. However, the former is limited by its environment, only usable in environments with BeiDou or GPS signal coverage, and unusable when satellite signals are interfered with or cut off. The latter, on the other hand, is affected by factors such as sensor cumulative errors and slippage, resulting in significant positioning errors during large-scale movements.

[0003] Compared to traditional navigation sensors, visual sensors offer advantages such as lower cost, lighter weight, and lower power consumption. They also acquire richer information, have shorter sampling periods, and do not rely on external satellite signals. Image-based landmark localization is a typical visual localization method. This method uses existing objects in the environment with known locations as reference points, or artificially designed objects with specific characteristics, as reference points. It employs pattern recognition to detect reference points in images captured during the robot's movement, simultaneously determining the presence and location of the reference point within the image. When a reference point is detected and the anchor frame marking the reference point's location occupies a proportion exceeding a certain threshold, the robot is considered to be near the reference point. The robot's position in the world coordinate system is then calculated based on the reference point's location within the image and its position in the world coordinate system. Therefore, effectively detecting reference points and determining their locations in images has become a research hotspot in the field of visual navigation for mobile robots.

[0004] Existing pattern recognition methods widely employ deep learning models based on convolutional neural networks as reference object detection models. After training the model a certain number of times using sufficient image data, the model can determine whether an object to be detected exists in the input image and determine the location of the object within the image. However, since calculating the parameters of the model (i.e., the training process) involves a large computational burden from calculating all sample data, random sampling is often used to iteratively calculate the model parameters. However, the randomness introduced by sampling results in slow convergence speed and long computation cycles, making this technique inefficient during the training phase.

[0005] It is evident that there is an urgent need for a mobile robot field localization method with high model training efficiency and strong localization performance. Summary of the Invention

[0006] In view of this, the present disclosure provides a method, system, device and medium for mobile robot field positioning, which at least partially solves the problems of poor model training efficiency and positioning performance in the prior art.

[0007] In a first aspect, embodiments of this disclosure provide a method for mobile robot field positioning, including: Step 1: Use the camera carried by the mobile robot to take multiple images of reference objects of different categories with known location information from different angles and positions as training data. Step 2: Mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine them with the training data to form a training set; Step 3: Construct an image detection model based on the YOLOv5 convolutional neural network and train it using the training set as input; Step 4: Run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image; Step 5: When the probability of the existence of a reference object exceeds the first threshold, the location of the reference object is marked in the image. This location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When this proportion exceeds the second threshold, the mobile robot is positioned near the reference object. Step 6: Retrieve the pre-stored coordinate information corresponding to the reference object based on its category to obtain the position estimate of the mobile robot.

[0008] According to one specific implementation of this disclosure, the image detection model includes an input end, a backbone network, a neck network, and an output end; The input terminal is used to implement Mosaic data augmentation, adaptive anchor box calculation and adaptive image scaling to obtain multiple channels of input feature maps of the same size, wherein the feature maps in different channels are random flipping and scaling operations on the original image; The backbone network includes a Focus structure and a CSP structure. The Focus structure is used to divide the input feature map into multiple sub-maps and arrange them in different channels. The CSP structure is used to divide the output feature map of the Focus structure into two parts, one part is processed by the sub-network, and the other part is processed by the next layer. The neck network includes a feature pyramid and a path aggregation network; The output terminal is used to determine the position and size of the anchor frame, thereby marking the position of the reference object in the image, and simultaneously outputting the probability of detecting the reference object.

[0009] According to a specific implementation of an embodiment of this disclosure, step 3 specifically includes: Step 3.1, Initialize step size parameters and its benchmark value Momentum coefficient and its benchmark value Second-order central moment coefficients Momentum decay rate Momentum variance coefficient and the initial moving weighted average momentum vector Initial moving weighted average second central moment Initial moving weighted average momentum variance Initial parameters of the image detection model Number of iterations ; Step 3.2: Extract a predetermined number of samples from the training data as input to the image detection model to obtain the output of the image detection model, namely: the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. Then, calculate the CIOU loss function at the t-th iteration. The output of the image detection model includes the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. The expression of the CIOU loss function is:

[0010] in, This represents the area of ​​the overlapping region between the anchor boxes output by the image detection model and the anchor boxes in the label. Let represent the area of ​​the union of the anchor boxes output by the image detection model and the anchor boxes in the labels; d is the distance between the center points of the anchor boxes output by the image detection model and the anchor boxes in the labels; and c is the diagonal distance of the minimum bounding rectangle of the anchor boxes output by the image detection model and the anchor boxes in the labels. , These are the width and height of the anchor box in the label, respectively. , These are the width and height of the anchor box output by the image detection model, respectively; Step 3.3: Calculate the gradient vector of the CIOU loss function with respect to the parameters of the image detection model using the error backpropagation method.

[0011] in, This indicates the operator for obtaining partial derivatives. The CIOU loss function is a function of the image detection model parameters. Represents the gradient vector; Step 3.4, calculate the moving weighted average momentum vector: ; Step 3.5, calculate the moving weighted average momentum variance: ; Step 3.6, calculate the second-order central moments of the moving weighted average: ; Step 3.7, calculate the step size rotation coefficient:

[0012] in, Indicates from Extract the minimum value from each dimension of the vector. Indicates from Extract the maximum value from each dimension of the vector. Representing dimensions and A vector that is identical and whose values ​​in all dimensions are 1; Step 3.8, calculate the normalized rotation coefficients:

[0013] In this context, division between vectors represents an element-wise division operation between two vectors. This indicates the search for the L2 norm of the orientation quantity. This indicates that a vector is multiplied element by element. Step 3.9, update the parameters of the image detection model: ; Step 3.10, update the step size parameter and momentum coefficient: , ; Step 3.11, Increment the number of iterations: ; Step 3.12, if If yes, proceed to step 3.2; otherwise, training ends. The parameters of the image detection model are given, and the trained image detection model is obtained, where T is the preset number of iterations.

[0014] Secondly, embodiments of this disclosure provide a mobile robot field positioning system, comprising: The data acquisition module is used to take multiple images of reference objects of different categories with known location information from different angles and positions using the camera carried by the mobile robot itself as training data; The data processing module is used to mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine them with the training data to form a training set. The training module is used to build an image detection model based on the YOLOv5 convolutional neural network and input the training set for training. The imaging module is used to run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image. The determination module is used to mark the location of the reference object in the image when the probability of the existence of the reference object exceeds the first threshold. The location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When the proportion exceeds the second threshold, the mobile robot is positioned near the reference object. The localization module is used to retrieve the pre-stored coordinate information corresponding to the reference object based on the category of the reference object, and obtain the position estimate of the mobile robot.

[0015] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: At least one processor; and, The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the mobile robot field positioning method in the first aspect or any implementation thereof.

[0016] Fourthly, embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the mobile robot field positioning method in the first aspect or any implementation thereof.

[0017] Fifthly, embodiments of this disclosure also provide a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the mobile robot field positioning method in the first aspect or any implementation thereof.

[0018] The mobile robot field localization scheme in this embodiment includes: Step 1, using the camera carried by the mobile robot to take multiple images of reference objects of different categories with known location information from different angles and positions as training data; Step 2, marking the outer contour of the reference object with anchor boxes in the training data, and recording the pixel coordinates of the upper left and lower right corners of the anchor boxes as labels corresponding to the training data, and combining the training data to form a training set; Step 3, constructing an image detection model based on a YOLOv5 convolutional neural network and inputting the training set for training; Step 4, running the mobile robot, inputting the images taken by the mobile robot into the trained image detection model, and obtaining the probability of the presence of reference objects in the image; Step 5, when the probability of the presence of a reference object exceeds a first threshold, marking the location of the reference object in the image, which is marked by an anchor box, and calculating the proportion of the area of ​​the anchor box to the area of ​​the image, when the proportion exceeds a second threshold, positioning the mobile robot near the reference object; Step 6, retrieving the pre-stored coordinate information corresponding to the reference object according to its category, and obtaining the position estimate of the mobile robot.

[0019] The beneficial effects of the embodiments of this disclosure are as follows: By fully utilizing the gradient variance information in the parameter calculation process, the parameter updates in each iteration are heuristically adjusted, so that the gradient components with smaller variance are given higher weights in each iteration, while the gradient components with larger variance are given smaller weights, thereby accelerating the convergence efficiency of model parameters, shortening the training process, and improving localization performance. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating a mobile robot field positioning method provided in an embodiment of this disclosure; Figure 2 A schematic diagram of mobile robot positioning for an embodiment of the present disclosure; Figure 3A schematic diagram of the CIOU loss function descent curve during model training provided in this embodiment of the disclosure; Figure 4 This is a schematic diagram of the structure of a mobile robot field positioning system provided in an embodiment of the present disclosure; Figure 5 A schematic diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0022] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0023] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0024] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0025] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0026] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0027] This disclosure provides a method for mobile robot localization in the field, which can be applied to the robot localization process in a measurement or detection scenario.

[0028] See Figure 1 This is a flowchart illustrating a mobile robot field positioning method provided in an embodiment of this disclosure. Figure 1 As shown, the method mainly includes the following steps: Step 1: Use the camera carried by the mobile robot to take multiple images of reference objects of different categories with known location information from different angles and positions as training data. In practice, for mobile robots equipped with cameras, multiple images of reference objects with known location information of different categories can be taken from different angles and positions by the mobile robot's own camera during its movement in the field, which can then be used as training data.

[0029] Step 2: Mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine them with the training data to form a training set; Step 3: Construct an image detection model based on the YOLOv5 convolutional neural network and train it using the training set as input; Based on the above embodiments, the image detection model includes an input end, a backbone network, a neck network, and an output end; The input terminal is used to implement Mosaic data augmentation, adaptive anchor box calculation and adaptive image scaling to obtain multiple channels of input feature maps of the same size, wherein the feature maps in different channels are random flipping and scaling operations on the original image; The backbone network includes a Focus structure and a CSP structure. The Focus structure is used to divide the input feature map into multiple sub-maps and arrange them in different channels. The CSP structure is used to divide the output feature map of the Focus structure into two parts, one part is processed by the sub-network, and the other part is processed by the next layer. The neck network includes a feature pyramid and a path aggregation network; The output terminal is used to determine the position and size of the anchor frame, thereby marking the position of the reference object in the image, and simultaneously outputting the probability of detecting the reference object.

[0030] Furthermore, step 3 specifically includes: Step 3.1, Initialize step size parameters and its benchmark value Momentum coefficient and its benchmark value Second-order central moment coefficients Momentum decay rate Momentum variance coefficient and the initial moving weighted average momentum vector Initial moving weighted average second central moment Initial moving weighted average momentum variance Initial parameters of the image detection model Number of iterations ; Step 3.2: Extract a predetermined number of samples from the training data as input to the image detection model to obtain the output of the image detection model, namely: the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. Then, calculate the CIOU loss function at the t-th iteration. The output of the image detection model includes the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. The expression of the CIOU loss function is:

[0031] in, This represents the area of ​​the overlapping region between the anchor boxes output by the image detection model and the anchor boxes in the label. Let represent the area of ​​the union of the anchor boxes output by the image detection model and the anchor boxes in the labels; d is the distance between the center points of the anchor boxes output by the image detection model and the anchor boxes in the labels; and c is the diagonal distance of the minimum bounding rectangle of the anchor boxes output by the image detection model and the anchor boxes in the labels. , These are the width and height of the anchor box in the label, respectively. , These are the width and height of the anchor box output by the image detection model, respectively; Step 3.3: Calculate the gradient vector of the CIOU loss function with respect to the parameters of the image detection model using the error backpropagation method.

[0032] in, This indicates the operator for obtaining partial derivatives. The CIOU loss function is a function of the image detection model parameters. Represents the gradient vector; Step 3.4, calculate the moving weighted average momentum vector: ; Step 3.5, calculate the moving weighted average momentum variance: ; Step 3.6, calculate the second-order central moments of the moving weighted average: ; Step 3.7, calculate the step size rotation coefficient:

[0033] in, Indicates from Extract the minimum value from each dimension of the vector. Indicates from Extract the maximum value from each dimension of the vector. Representing dimensions and A vector that is identical and whose values ​​in all dimensions are 1; Step 3.8, calculate the normalized rotation coefficients:

[0034] In this context, division between vectors represents an element-wise division operation between two vectors. This indicates the search for the L2 norm of the orientation quantity. This indicates that a vector is multiplied element by element. Step 3.9, update the parameters of the image detection model: ; Step 3.10, update the step size parameter and momentum coefficient: , ; Step 3.11, Increment the number of iterations: ; Step 3.12, if If yes, proceed to step 3.2; otherwise, training ends. The parameters of the image detection model are given, and the trained image detection model is obtained, where T is the preset number of iterations.

[0035] In practical implementation, the YOLOv5 convolutional neural network model can be used as the image detection model. Its input is an image captured by a camera carried by a mobile robot, and its output is the probability and category of a reference object present in the image. The model consists of four parts: an input layer, a backbone network, a neck network, and an output layer. The input layer implements Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling to obtain multiple channels of input feature maps of the same size. The feature maps within different channels are obtained by randomly flipping and scaling the original image. The backbone network includes a Focus structure and a Cross Stage Partial (CSP) structure. The Focus structure divides the input feature map into multiple sub-maps arranged in different channels. The CSP structure divides the output feature map of the Focus structure into two parts: one part is processed by a sub-network, and the other part is processed by the next layer. The neck network includes a feature pyramid and a path aggregation network. The output layer displays the position and size of the anchor boxes, thus marking the location of the reference object in the image, and simultaneously outputs the probability of detecting the reference object. Then, using pre-captured images containing specified reference objects as training data, a special adaptive momentum method was developed to optimize the model's parameters. This method calculates the moving weighted average momentum vector and its variance, as well as the moving weighted average of the second-order central moments of the model parameters during previous iterations. Combining the above statistical information, the parameter update rule for each iteration is calculated. It has advantages such as adaptive adjustment of the iteration step size, fast convergence speed, and simple parameter calculation process.

[0036] Step 4: Run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image; In practice, after the image detection model is trained, when a mobile robot is used for field measurement and localization is required, the mobile robot can be run and the images captured by the mobile robot can be input into the trained image detection model to obtain the probability of the presence of reference objects in the image, so as to carry out subsequent operation procedures.

[0037] Step 5: When the probability of the existence of a reference object exceeds the first threshold, the location of the reference object is marked in the image. This location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When this proportion exceeds the second threshold, the mobile robot is positioned near the reference object. For example, when the probability of detecting reference object A, B, or C in an image exceeds a certain threshold, the location of the reference object is marked in the image. This location is marked by an anchor frame. Based on the position of reference object A, B, or C in the image marked by the anchor frame, the pixel coordinates of the upper left and lower right corners of the anchor frame are obtained. If the anchor frame occupies more than 10% of the image area, the robot is considered to be near the reference object. Of course, in practical applications, the specific values ​​of the first and second thresholds can be adjusted adaptively, which will not be elaborated here.

[0038] Step 6: Retrieve the pre-stored coordinate information corresponding to the reference object based on its category to obtain the position estimate of the mobile robot.

[0039] For example, when it is determined that the mobile robot is near reference object A, the position estimate of the mobile robot can be obtained by retrieving the pre-stored coordinate information corresponding to reference object A according to the category of reference object A.

[0040] The mobile robot field localization method provided in this embodiment establishes a YOLOv5-based image detection model based on images captured during the mobile robot's operation. Then, it uses pre-collected images of known-location reference objects as a training dataset. The YOLOv5 model parameters are iteratively calculated using the image detection model parameter calculation method, enabling rapid model training. It fully utilizes the gradient variance information during parameter calculation, heuristically adjusting parameter updates in each iteration. This assigns higher weights to gradient components with smaller variance and lower weights to gradient components with larger variance, thereby accelerating model parameter convergence and shortening the training process. After parameter iteration is complete, images captured during the mobile robot's movement are used as input to the image detection model. The probability of a reference object being present in the image is calculated. If this probability exceeds a certain threshold, the reference object is marked using anchor boxes. The mobile robot's orientation relative to the reference object is calculated based on the anchor box position, and the mobile robot is positioned near the reference object. Simultaneously, the position coordinates of the detected reference object are retrieved, thus obtaining the mobile robot's position coordinates and achieving high-precision mobile robot localization.

[0041] The method disclosed herein will be further explained below with reference to a specific embodiment. Addressing the localization problem of mobile robots on campus, an image detection model based on YOLOv5 is established using images captured during the robot's operation as the basis for judgment. Then, images of known reference objects are used as a training dataset. The parameters of the YOLOv5 model are iteratively calculated using the image detection model parameter calculation method described above, achieving rapid model training. After the parameter iterative calculation is complete, images captured during the robot's movement are used as input to the image detection model. The probability of the presence of a reference object in the image is calculated. If this probability is greater than a certain threshold, the reference object in the image is marked using anchor boxes. The orientation of the mobile robot relative to the reference object is calculated based on the position of the anchor boxes, and the mobile robot is positioned near the reference object. Simultaneously, the position coordinates corresponding to the detected reference object are retrieved, thereby obtaining the position coordinates of the mobile robot and achieving mobile robot localization.

[0042] Taking the localization of a mobile robot on campus as an example, we select the landmark buildings "South Gate of Democracy Building", "West Gate of Democracy Building", and "Peace Building" as reference points, and the global coordinates of these buildings are known. The following steps are performed to achieve image-based localization of the mobile robot: Images of the "South Gate of the Democracy Building," "West Gate of the Democracy Building," and "Peace Building" were taken from different angles and positions using a camera carried by the mobile robot, and used as a training dataset.

[0043] In the training dataset, anchor boxes are used to mark the outer contours of "South Gate of Democracy Building", "West Gate of Democracy Building" and "Peace Building", and the pixel coordinates of the upper left and lower right corners of the anchor boxes are recorded as the corresponding labels for the training data.

[0044] The YOLOv5 convolutional neural network model is used as the image detection model. Its input is the image taken by the camera carried by the mobile robot, and the output is the probability and category of the presence of "South Gate of Democracy Building", "West Gate of Democracy Building" and "Peace Building" in the image.

[0045] Initialize step size parameters: and its benchmark values: Momentum coefficient: and its benchmark values: Second-order central moment coefficients: Momentum decay rate: Momentum variance coefficient: And the initial moving weighted average momentum vector: Initial moving weighted average second-order central moments: Initial moving weighted average momentum variance: Randomly initialize the initial parameters of the image detection model. Initialize the number of iterations .

[0046] Thirty-two samples are extracted from the training dataset as input to the image detection model. After obtaining the model's output, the CIOU loss function is calculated, which is the probability of the existence of "South Gate of Democracy Building", "West Gate of Democracy Building", and "Peace Building" and the position of the anchor box used to label "South Gate of Democracy Building", "West Gate of Democracy Building", and "Peace Building" in the image. Then, the CIOU loss function at the t-th iteration is calculated, as shown in the following expression:

[0047] In the formula, This represents the area of ​​the overlapping region between the anchor boxes output by the model and the anchor boxes in the label. d represents the area of ​​the region corresponding to the union of the anchor boxes output by the model and the anchor boxes in the label; d is the distance between the center points of the anchor boxes output by the model and the anchor boxes in the label; and c is the diagonal distance of the minimum bounding rectangle of the anchor boxes output by the model and the anchor boxes in the label. , These are the width and height of the anchor box in the label, respectively. , These represent the width and height of the anchor frame output by the model, respectively.

[0048] The CIOU loss function is calculated using the error backpropagation method. The gradient vector is denoted as: In the formula, The symbol represents the operator for taking partial derivatives. CIOU loss function is the CIOU loss function with respect to the parameters of the image detection model. This represents the gradient vector.

[0049] Calculate the moving weighted average momentum vector: .

[0050] Calculate the moving weighted average momentum variance: .

[0051] Calculate the second central moment of the moving weighted average: .

[0052] Calculate the step size rotation coefficient: In the formula, Indicates from Extract the minimum value from each dimension of the vector. Indicates from Extract the maximum value from each dimension of the vector. Representing dimensions and A vector that is identical and whose values ​​are all 1 in each dimension.

[0053] Calculate the normalized rotation coefficients: In the formula, division between vectors represents an element-wise division operation between two vectors. This indicates the search for the L2 norm of the orientation quantity. This indicates that a vector is multiplied element by element.

[0054] Update the parameters of the image detection model: .

[0055] Update step size and momentum coefficient: , .

[0056] Increasing the number of iterations: .

[0057] If t < 1000, then jump to 5; otherwise, select... These are the parameters of the image detection model.

[0058] Run the mobile robot and input the images captured by the mobile robot into the image detection model to obtain the probability of detecting "South Gate of Democracy Building", "West Gate of Democracy Building" and "Peace Building" in the image.

[0059] When the probability of detecting "South Gate of Democracy Building", "West Gate of Democracy Building" or "Peace Building" in the image in 16 exceeds a certain threshold, the location of the reference object is marked in the image, and the location is marked by an anchor frame.

[0060] Based on the position of the "South Gate of the Democracy Building," "West Gate of the Democracy Building," or "Peace Building" marked by the anchor frame in the image, the pixel coordinates of the upper left and lower right corners of the anchor frame are obtained. If the anchor frame occupies more than 10% of the image area, the robot is considered to be near the reference object. In this example, the mobile robot identified the "South Gate of the Democracy Building," marked it with an anchor frame, and measured the pixel coordinates of the upper left and lower right corners of the anchor frame as follows: , Based on the pixel values ​​of the image. The calculated ratio of the anchor frame area to the image area is 15.32%, which is greater than the threshold of 10%. Therefore, the mobile robot is located near the Democracy Building, with its latitude and longitude as 28.170507°N and 112.930566°E.

[0061] During model parameter calculation, the change in the CIOU loss function reflects the change in the model's recognition accuracy during training. To demonstrate the effectiveness of the YOLOv5 model parameter calculation method described in this invention, the method is compared with the Adam parameter calculation method. Both methods are used to train a YOLOv5 model, and the curves showing the change in the CIOU loss function of the corresponding model during the iteration process are plotted. Figure 2 The curves showing a faster decline indicate that the corresponding model parameter calculation method is more efficient. In the last iteration, the CIOU loss function value obtained using the described model parameter calculation method was 0.02989, while the CIOU loss function value obtained using the Adam method was 0.03431. Therefore, the model trained using the described method has a smaller error.

[0062] For a corresponding method embodiment, see [link to relevant documentation]. Figure 4 This disclosure also provides a mobile robot field positioning system 40, comprising: The data acquisition module 401 is used to take multiple images of reference objects of different categories with known location information from different angles and positions using the camera carried by the mobile robot itself as training data. The data processing module 402 is used to mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine the training data to form a training set. Training module 403 is used to build an image detection model based on YOLOv5 convolutional neural network and input the training set for training; The imaging module 404 is used to run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image. The determination module 405 is used to mark the location of the reference object in the image when the probability of the existence of the reference object exceeds the first threshold. The location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When the proportion exceeds the second threshold, the mobile robot is positioned near the reference object. The positioning module 406 is used to retrieve the pre-stored coordinate information corresponding to the reference object based on the category of the reference object, and obtain the position estimate of the mobile robot.

[0063] Figure 4 The system shown can execute the contents of the above method embodiments. For the parts not described in detail in this embodiment, please refer to the contents recorded in the above method embodiments, and they will not be repeated here.

[0064] See Figure 5This disclosure also provides an electronic device 50, which includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed, enables the at least one processor to perform the mobile robot field positioning method described in the foregoing method embodiments.

[0065] This disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the mobile robot field positioning method in the foregoing method embodiments.

[0066] This disclosure also provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the mobile robot field positioning method in the foregoing method embodiments.

[0067] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device 50 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0068] like Figure 5 As shown, electronic device 50 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 50. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0069] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 50 to communicate wirelessly or wiredly with other devices to exchange data. Although an electronic device 50 with various devices is shown in the figure, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0070] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0071] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0072] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0073] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the relevant steps of the above-described method embodiments.

[0074] Alternatively, the aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the relevant steps of the above method embodiments.

[0075] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0077] The units described in the embodiments of this disclosure can be implemented in software or in hardware.

[0078] It should be understood that the various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof.

[0079] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for field positioning of a mobile robot, characterized in that, include: Step 1: Use the camera carried by the mobile robot to take multiple images of reference objects of different categories with known location information from different angles and positions as training data. Step 2: Mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine them with the training data to form a training set; Step 3: Construct an image detection model based on the YOLOv5 convolutional neural network and train it using the training set as input; Step 3 specifically includes: Step 3.1, Initialize step size parameters and its benchmark value Momentum coefficient and its benchmark value Second-order central moment coefficients Momentum decay rate Momentum variance coefficient and the initial moving weighted average momentum vector Initial moving weighted average second central moment Initial moving weighted average momentum variance Initial parameters of the image detection model Number of iterations ; Step 3.2: Extract a predetermined number of samples from the training data as input to the image detection model to obtain the output of the image detection model, namely: the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. Then, calculate the CIOU loss function at the t-th iteration. The output of the image detection model includes the probability of the presence of a reference object and the position of the anchor box used to mark the reference object in the image. The expression of the CIOU loss function is: in, This represents the area of ​​the overlapping region between the anchor boxes output by the image detection model and the anchor boxes in the label. Let represent the area of ​​the union of the anchor boxes output by the image detection model and the anchor boxes in the labels; d is the distance between the center points of the anchor boxes output by the image detection model and the anchor boxes in the labels; and c is the diagonal distance of the minimum bounding rectangle of the anchor boxes output by the image detection model and the anchor boxes in the labels. , These are the width and height of the anchor box in the label, respectively. , These are the width and height of the anchor box output by the image detection model, respectively; Step 3.3: Calculate the gradient vector of the CIOU loss function with respect to the parameters of the image detection model using the error backpropagation method. in, This indicates the operator for obtaining partial derivatives. The CIOU loss function is a function of the image detection model parameters. Represents the gradient vector; Step 3.4, calculate the moving weighted average momentum vector: ; Step 3.5, calculate the moving weighted average momentum variance: ; Step 3.6, calculate the second-order central moments of the moving weighted average: ; Step 3.7, calculate the step size rotation coefficient: in, Indicates from Extract the minimum value from each dimension of the vector. Indicates from Extract the maximum value from each dimension of the vector. Representing dimensions and A vector that is identical and whose values ​​in all dimensions are 1; Step 3.8, calculate the normalized rotation coefficients: In this context, division between vectors represents an element-wise division operation between two vectors. This indicates the search for the L2 norm of the orientation quantity. This indicates that a vector is multiplied element by element. Step 3.9, update the parameters of the image detection model: ; Step 3.10, update the step size parameter and momentum coefficient: , ; Step 3.11, Increment the number of iterations: ; Step 3.12, if If yes, proceed to step 3.2; otherwise, training ends. The parameters of the image detection model are given, and the trained image detection model is obtained, where T is the preset number of iterations; Step 4: Run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image; Step 5: When the probability of the existence of a reference object exceeds the first threshold, the location of the reference object is marked in the image. This location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When this proportion exceeds the second threshold, the mobile robot is positioned near the reference object. Step 6: Retrieve the pre-stored coordinate information corresponding to the reference object based on its category to obtain the position estimate of the mobile robot.

2. The method according to claim 1, characterized in that... The image detection model includes an input terminal, a backbone network, a neck network, and an output terminal. The input terminal is used to implement Mosaic data augmentation, adaptive anchor box calculation and adaptive image scaling to obtain multiple channels of input feature maps of the same size, wherein the feature maps in different channels are random flipping and scaling operations on the original image; The backbone network includes a Focus structure and a CSP structure. The Focus structure is used to divide the input feature map into multiple sub-maps and arrange them in different channels. The CSP structure is used to divide the output feature map of the Focus structure into two parts, one part is processed by the sub-network, and the other part is processed by the next layer. The neck network includes a feature pyramid and a path aggregation network; The output terminal is used to determine the position and size of the anchor frame, thereby marking the position of the reference object in the image, and simultaneously outputting the probability of detecting the reference object.

3. A mobile robot field positioning system, used to execute the mobile robot field positioning method according to any one of claims 1 to 2, characterized in that, include: The data acquisition module is used to take multiple images of reference objects of different categories with known location information from different angles and positions using the camera carried by the mobile robot itself as training data; The data processing module is used to mark the outer contour of the reference object with anchor boxes in the training data, and record the pixel coordinates of the upper left and lower right corners of the anchor boxes as the corresponding labels of the training data, and combine them with the training data to form a training set. The training module is used to build an image detection model based on the YOLOv5 convolutional neural network and input the training set for training. The imaging module is used to run the mobile robot and input the images captured by the mobile robot into the trained image detection model to obtain the probability that a reference object exists in the image. The determination module is used to mark the location of the reference object in the image when the probability of the existence of the reference object exceeds the first threshold. The location is marked by an anchor frame, and the proportion of the area of ​​the anchor frame to the area of ​​the image is calculated. When the proportion exceeds the second threshold, the mobile robot is positioned near the reference object. The localization module is used to retrieve the pre-stored coordinate information corresponding to the reference object based on the category of the reference object, and obtain the position estimate of the mobile robot.

4. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed, enables the at least one processor to perform the mobile robot field positioning method according to any one of claims 1-2.

5. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the mobile robot field positioning method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN112132107A

  • Three-dimensional target detection method and device and computer storage medium

    CN115909269A