Robot positioning method based on combination of improved YOLOv8 and depth map

By combining the improved YOLOv8 image segmentation model with depth map technology, the problem of insufficient accuracy and reliability of traditional positioning methods in complex environments is solved, and higher positioning accuracy and robustness are achieved.

CN120070835APending Publication Date: 2025-05-30SHENYANG JIANZHU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411936097.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional robot positioning methods are difficult to meet the requirements of positioning accuracy and reliability in indoors or complex environments where GPS signals are limited, and visual positioning methods have positioning accuracy fluctuations and instability problems in lighting changes, viewing angle problems and complex environments.

Method used

The improved YOLOv8 image segmentation model is combined with the depth map technology to extract feature information in the environment through image segmentation processing, and use the spatial depth information provided by the depth map to determine the position information of the robot.

Benefits of technology

It improves the accuracy and robustness of robot positioning, can effectively reduce errors in complex environments, improve positioning accuracy, and perform relatively stable in environments with dynamic changes or uneven light.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070835A_ABST
    Figure CN120070835A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot positioning, and provides a robot positioning method based on combination of improved YOLOv8 and a depth map, and the method comprises the steps: obtaining a to-be-processed image, carrying out the image segmentation processing of the to-be-processed image based on a trained YOLOv8 segmentation model, and obtaining a segmentation result corresponding to the to-be-processed image; wherein the to-be-processed image is shot by a robot, and the segmentation result comprises information of a plurality of feature objects in the to-be-processed image; determining a depth map corresponding to the to-be-processed image, and determining target image information based on the multiple pieces of feature object information in the to-be-processed image and the depth map; and determining a reference feature object based on the target image information, and determining position information of the robot based on the target image information and the reference feature object. According to the embodiment, the improved YOLOv8 image segmentation model and the depth map technology are combined, the positioning precision of the robot is improved, meanwhile, the used feature objects are all road signs which can be understood by people for positioning, and high robustness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of robot positioning, and more particularly, to a robot positioning method based on the combination of improved YOLOv8 and depth maps. Background Art

[0002] With the rapid development of robot technology, the demand for robot navigation and positioning in various complex environments is increasing. In robot navigation, accurate positioning information can help the robot determine its own position, so as to plan an effective travel path. Traditional robot positioning methods, such as GPS-based positioning and odometer calculation, perform well in open and well-signaled environments, but in indoor scenarios where GPS signals are limited, their positioning accuracy and reliability often cannot meet the actual needs.

[0003] In related technologies, the wide application of deep learning technology in the field of image recognition and processing provides a new direction for vision-based robot positioning methods. Vision sensors can capture rich information in the environment and process it in combination with deep learning models to infer the position of the robot; vision-based positioning methods can overcome the limitations of traditional positioning technologies in GPS signal-limited areas and show good results especially in indoor or complex environments. However, relying solely on visual information still faces many challenges: First, factors such as environmental light changes, perspective problems, and object occlusion will have an adverse impact on the quality of the image, resulting in fluctuations in positioning accuracy; in addition, vision sensors may not provide sufficient spatial information in some scenarios, especially when the environment is too complex or lacks obvious features, the positioning results may be unstable or have large errors. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a robot positioning method based on the combination of improved YOLOv8 and depth maps. By combining the improved YOLOv8 image segmentation model with depth map technology, the positioning accuracy of the robot is improved, and at the same time, the feature objects used for positioning are all road signs that can be understood by people, with strong robustness.

[0005] An embodiment of the present disclosure provides a robot positioning method based on the combination of improved YOLOv8 and depth maps, including:

[0006] Obtain an image to be processed, and perform image segmentation processing on the image to be processed based on a trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed; wherein, the image to be processed is captured by a robot, and the segmentation result includes information of multiple feature objects in the image to be processed;

[0007] Determine the depth map corresponding to the image to be processed, and determine the target image information based on the information of multiple feature objects in the image to be processed and the depth map;

[0008] Determine the reference feature object based on the target image information, and determine the position information of the robot based on the target image information and the reference feature object.

[0009] In some possible embodiments, before performing image segmentation processing on the image to be processed based on the trained YOLOv8 segmentation model, it includes:

[0010] Construct an initial YOLOv8 segmentation model; wherein, the initial YOLOv8 segmentation model includes a feature extraction module, an upsampling module, an attention extraction module, and an image segmentation module;

[0011] Replace the upsampling module in the initial YOLOv8 segmentation model with a dynamic upsampling module, and replace the attention extraction module with a Triplet attention extraction module to obtain the YOLOv8 segmentation model to be trained;

[0012] Obtain a label image set corresponding to the robot; wherein, the label image set includes multiple environmental label images, and each environmental label image includes multiple feature objects;

[0013] Train the YOLOv8 segmentation model to be trained based on the label image set to obtain the trained YOLOv8 segmentation model.

[0014] In some possible embodiments, the training of the YOLOv8 segmentation model to be trained based on the label image set includes:

[0015] Extract features from each environmental label image in the label image set based on the feature extraction module to obtain an original feature map corresponding to each environmental label image;

[0016] Perform dynamic upsampling processing on each environmental label image in the label image set based on the dynamic upsampling model to obtain a high-resolution environmental label image corresponding to each environmental label image;

[0017] Extract global features from the high-resolution environmental label image corresponding to each environmental label image based on the Triplet attention extraction module to obtain an attention weight map corresponding to each high-resolution environmental label image;

[0018] Based on the image segmentation module, determine the segmentation result corresponding to each environmental label image according to the original feature map corresponding to each environmental label image and the attention weight map corresponding to each high-resolution environmental label image respectively;

[0019] Calculate the loss value between the segmentation result corresponding to each environmental label image and the environmental label sample image corresponding to each environmental label image based on a preset loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value;

[0020] Repeat the above steps until the training result meets the preset requirements to obtain the trained YOLOv8 segmentation model.

[0021] In some possible embodiments, the determining the target image information based on the multiple feature object information and the depth map in the image to be processed includes:

[0022] Register the multiple feature object information in the image to be processed with the depth map, and obtain the target image information based on the registration result.

[0023] Before determining the depth map corresponding to the image to be processed in some possible embodiments, it includes:

[0024] Determine the target calibration object in the environment corresponding to the robot, and determine the scale factor based on the size information of the target calibration object and the distance between the robot and the target calibration object.

[0025] In some possible embodiments, the determining the reference feature object based on the target image information includes:

[0026] Calculate the distances between the robot and each feature object in the target image information respectively based on the scale factor and the target image information, and determine the first reference feature object and the second reference feature object based on the calculation results.

[0027] In some possible embodiments, the determining the position information of the robot based on the target image information and the reference feature object includes:

[0028] Determine the depth information of the first reference feature object and the second reference feature object respectively according to the target image information; and determine the distance between the first reference feature object and the robot according to the scale factor and the depth information of the first reference feature object; determine the distance between the second reference feature object and the robot according to the scale factor and the depth information of the second reference feature object;

[0029] Based on the distance between the first reference feature and the robot and the distance between the second reference feature and the robot, the position information of the robot is determined based on the cosine theorem.

[0030] An embodiment of the present disclosure provides a robot positioning device based on the combination of improved YOLOv8 and depth map, including:

[0031] An image segmentation module, configured to obtain a to-be-processed image, and perform image segmentation processing on the to-be-processed image based on a trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the to-be-processed image; wherein, the to-be-processed image is captured by a robot, and the segmentation result includes information of multiple features in the to-be-processed image;

[0032] An image determination module, configured to determine a depth map corresponding to the to-be-processed image, and determine target image information based on the information of multiple features in the to-be-processed image and the depth map;

[0033] A position determination module, configured to determine a reference feature based on the target image information, and determine the position information of the robot based on the target image information and the reference feature.

[0034] In some possible embodiments, the device further includes:

[0035] A model construction module, configured to construct an initial YOLOv8 segmentation model; wherein, the initial YOLOv8 segmentation model includes a feature extraction module, an upsampling module, an attention extraction module, and an image segmentation module;

[0036] A model determination module, configured to replace the upsampling module in the initial YOLOv8 segmentation model with a dynamic upsampling module, and replace the attention extraction module with a Triplet attention extraction module to obtain a YOLOv8 segmentation model to be trained;

[0037] A data acquisition module, configured to acquire a label image set corresponding to the robot; wherein, the label image set includes multiple environmental label images, and each environmental label image includes multiple features;

[0038] A model training module, configured to train the YOLOv8 segmentation model to be trained based on the label image set to obtain the trained YOLOv8 segmentation model.

[0039] In some possible embodiments, the model training module is specifically configured to:

[0040] Based on the feature extraction module, feature extraction is respectively performed on each of the environmental label images in the label image set to obtain an original feature map corresponding to each of the environmental label images;

[0041] Based on the dynamic upsampling model, dynamic upsampling processing is respectively performed on each of the environmental label images in the label image set to obtain a high-resolution environmental label image corresponding to each of the environmental label images;

[0042] Based on the Triplet attention extraction module, global feature extraction is respectively performed on the high-resolution environmental label image corresponding to each of the environmental label images to obtain an attention weight map corresponding to each of the high-resolution environmental label images;

[0043] Based on the image segmentation module, the segmentation result corresponding to each of the environmental label images is respectively determined according to the original feature map corresponding to each of the environmental label images and the attention weight map corresponding to each of the high-resolution environmental label images;

[0044] Based on a preset loss function, the loss value between the segmentation result corresponding to each of the environmental label images and the environmental label sample image corresponding to each of the environmental label images is calculated, and the YOLOv8 segmentation model to be trained is adjusted based on the loss value;

[0045] Repeat the above steps until the training result meets the preset requirements to obtain the trained YOLOv8 segmentation model.

[0046] In some possible embodiments, the image determination module is specifically configured to:

[0047] Register the information of multiple feature objects in the image to be processed with the depth map, and obtain the target image information based on the registration result.

[0048] In some possible embodiments, the image determination module is further configured to:

[0049] Determine the target calibration object in the environment corresponding to the robot, and determine the scale factor based on the size information of the target calibration object and the distance between the robot and the target calibration object.

[0050] In some possible embodiments, the position determination module is specifically configured to:

[0051] Based on the scale factor and the target image information, calculate the distances between the robot and each feature object in the target image information respectively, and determine the first reference feature object and the second reference feature object based on the calculation result.

[0052] In some possible embodiments, the position determination module is specifically configured to:

[0053] Determine the depth information of the first reference feature and the second reference feature respectively according to the target image information; and determine the distance between the first reference feature and the robot according to the scale factor and the depth information of the first reference feature; determine the distance between the second reference feature and the robot according to the scale factor and the depth information of the second reference feature;

[0054] Based on the distance between the first reference feature and the robot and the distance between the second reference feature and the robot, determine the position information of the robot based on the cosine theorem.

[0055] An embodiment of the present disclosure provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the robot positioning method based on the combination of the improved YOLOv8 and the depth map as described in any of the above possible implementation manners is executed.

[0056] An embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the robot positioning method based on the combination of the improved YOLOv8 and the depth map as described in any of the above possible implementation manners is implemented.

[0057] In the robot positioning method based on the combination of the improved YOLOv8 and the depth map provided in the embodiments of the present disclosure, by combining the improved YOLOv8 image segmentation model and the depth map technology, key feature information in the environment can be accurately extracted and analyzed. At the same time, using the spatial depth information provided by the depth map, the accuracy and robustness of robot positioning are improved. The present disclosure can effectively reduce errors in complex environments and improve positioning accuracy. At the same time, the method provided by the present disclosure performs relatively stably in dynamically changing or uneven illumination environments. In addition, the positioning method combined with depth information enhances the perception ability of obstacles and target objects, enabling the robot to better adapt to complex working scenarios when performing tasks, thereby improving the efficiency of its autonomous navigation and decision-making.

[0058] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be cited in the embodiments will be briefly introduced below. The accompanying drawings here are incorporated into the specification and form a part of this specification. These accompanying drawings show the embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0060] Figure 1 Shows the flowchart of a robot positioning method based on the combination of improved YOLOv8 and depth map provided by the embodiments of the present disclosure;

[0061] Figure 2 Shows the flowchart of a method for determining a YOLOv8 segmentation model provided by the embodiments of the present disclosure;

[0062] Figure 3 Shows the flowchart of a method for training a YOLOv8 segmentation model provided by the embodiments of the present disclosure;

[0063] Figure 4 Shows the structural schematic diagram of a robot positioning device based on the combination of improved YOLOv8 and depth map provided by the embodiments of the present disclosure;

[0064] Figure 5 Shows the structural schematic diagram of another robot positioning device based on the combination of improved YOLOv8 and depth map provided by the embodiments of the present disclosure;

[0065] Figure 6 Shows the structural schematic diagram of a computer device provided by the embodiments of the present disclosure. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure to be protected, but only represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0067] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures.

[0068] The term "and / or" in this document merely describes an associated relationship and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the term "at least one" in this document means any one of multiple items or any combination of at least two of multiple items. For example, including at least one of A, B, and C can represent selecting any one or more elements from the set composed of A, B, and C.

[0069] With the rapid development of robotics, the need for robots to navigate and position in various complex environments has become increasingly prominent. During the robot navigation process, accurate positioning information is crucial for the robot to determine its own position and plan its travel path. Traditional positioning methods, such as GPS-based positioning and odometer estimation, although perform well in open areas and environments with good signals, in scenarios where GPS signals are restricted, especially in indoor environments, the positioning accuracy and reliability of these methods often cannot meet the actual application requirements.

[0070] In this context, the extensive application of deep learning technology in the field of image recognition and processing has provided a new solution for vision-based robot positioning. Visual sensors can capture rich information in the environment and, combined with deep learning models, process the images to infer the position of the robot. Compared with traditional positioning technologies, vision-based positioning methods can effectively overcome the limitations of GPS signal-restricted areas and show better effects especially in indoor or other complex environments. For example, camera-based visual positioning can not only provide rich feature information in space but also fuse with other sensor data to improve the accuracy and reliability of positioning.

[0071] It has been found through research that positioning methods relying on visual information still face many challenges. First, factors such as changes in environmental lighting, perspective deviation, and object occlusion will affect the image quality, thereby causing fluctuations in positioning accuracy. Especially in low-light, dynamic environments, or complex scenes, the degradation of image quality may lead to an increase in positioning errors. Second, in some specific scenarios, such as when the environment is too complex or lacks significant features, visual sensors may not be able to provide sufficient spatial information, resulting in unstable positioning results and even large errors. In addition, single visual information processing often cannot cope with the continuously changing dynamic features in the environment, so it is necessary to fuse multiple sensor technologies, such as lidar, IMU, etc., to enhance the robustness and stability of the positioning system.

[0072] Based on the above research, an embodiment of the present disclosure provides a robot positioning method based on the combination of improved YOLOv8 and depth maps. Specifically, first, an image to be processed is obtained, and the image to be processed is subjected to image segmentation processing based on a trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed; wherein, the image to be processed is captured by a robot, and the segmentation result includes information of multiple feature objects in the image to be processed; then, a depth map corresponding to the image to be processed is determined, and target image information is determined based on the information of multiple feature objects in the image to be processed and the depth map; finally, a reference feature object is determined based on the target image information, and the position information of the robot is determined based on the target image information and the reference feature object.

[0073] In the embodiment of the present disclosure, by combining the improved YOLOv8 image segmentation model with the depth map technology, the key feature object information in the environment can be accurately extracted and analyzed. At the same time, by using the spatial depth information provided by the depth map, the accuracy and robustness of robot positioning are improved. The present disclosure can effectively reduce errors in complex environments and improve the positioning accuracy. At the same time, the method provided by the present disclosure performs relatively stably in environments with dynamic changes or uneven illumination. In addition, the positioning method combined with depth information enhances the perception ability of obstacles and target objects, enabling the robot to better adapt to complex working scenarios when performing tasks, thereby improving the efficiency of its autonomous navigation and decision-making.

[0074] To facilitate the understanding of this embodiment, the execution subject of the robot positioning method based on the combination of improved YOLOv8 and depth maps provided by the embodiment of the present disclosure is introduced in detail first. The execution subject of the robot positioning method based on the combination of improved YOLOv8 and depth maps provided by the embodiment of the present disclosure is a computer device. This computer device can be a terminal device or a server. Among them, the terminal device can also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method can also be applied to an implementation environment composed of a computer device and a server.

[0075] The following will describe in detail the robot positioning method based on the combination of improved YOLOv8 and depth maps provided by the embodiments of the present application with reference to the accompanying drawings. Refer to Figure 1 As shown, it is a flowchart of a robot positioning method based on the combination of improved YOLOv8 and depth maps provided by an embodiment of the present disclosure. The method includes the following S101 to S103:

[0076] S101. Obtain the image to be processed, and perform image segmentation processing on the image to be processed based on the trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed.

[0077] It can be understood that the image to be processed is usually captured by a camera, lidar or other vision sensors equipped on the robot, and its image form can be a static image (such as a still photo) or a dynamic video frame (from a real-time video stream). The visual information captured by these images is about the environment where the robot is located, and can include data such as objects, structures, and colors in the scene. Then, based on the trained YOLOv8 segmentation model, image segmentation processing is performed on the image to be processed to obtain a segmentation result. The segmentation result contains information about multiple feature objects in the image to be processed, specifically including the category of the feature object, the boundary of each feature object in the image (i.e., the edge contour), and the specific position information of the object, such as coordinates and area size, etc. Here, image segmentation processing refers to dividing an image into multiple meaningful regions, and each region usually corresponds to a different object or feature. For example, in an image of an indoor environment, the segmentation result may show objects such as "sofa", "TV", "floor", and "ceiling", and assign a pixel-level region to each object.

[0078] Specifically, the YOLOv8 (You Only Look Once version 8) image segmentation model is the latest version based on the YOLO series, focusing on real-time image detection and segmentation tasks. YOLOv8 combines the capabilities of object detection and image segmentation, and can perform object detection (locating objects) and semantic segmentation (pixel-level object segmentation) simultaneously in a single inference process. Compared with the earlier versions of YOLO, YOLOv8 has made significant improvements in speed and accuracy, can handle more complex image segmentation tasks, and has optimized the adaptability to different types of scenes and objects. Referring to Figure 2 As shown, the YOLOv8 segmentation model is obtained through the following steps S201 - S204:

[0079] S201. Construct an initial YOLOv8 segmentation model; wherein, the initial YOLOv8 segmentation model includes a feature extraction module, an upsampling module, an attention extraction module, and an image segmentation module.

[0080] It is understandable that the initial YOLOv8 segmentation model mainly involves four core modules: a feature extraction module, an upsampling module, an attention extraction module, and an image segmentation module. The feature extraction module extracts meaningful features such as edges, textures, and other high-level features from the input image through multiple convolutional layers, laying the foundation for object recognition; the upsampling module restores the low-resolution feature map to a higher resolution to ensure that subsequent image segmentation can finely process the classification of each pixel, usually achieved through transposed convolution or interpolation methods; the attention extraction module enhances the model's attention to important regions by weighting the feature map, suppressing irrelevant or interfering information, and improving the detection and segmentation accuracy; the image segmentation module is responsible for classifying each pixel point in the image into different categories to achieve accurate object segmentation.

[0081] S202, replace the upsampling module in the initial YOLOv8 segmentation model with a dynamic upsampling module, and replace the attention extraction module with a Triplet attention extraction module to obtain the YOLOv8 segmentation model to be trained.

[0082] Exemplarily, in a resource-constrained environment, the initial YOLOv8 segmentation model may have problems such as long training time and high resource consumption. Therefore, to improve the training efficiency, the present disclosure replaces the upsampling module in the original model with a dynamic upsampling module. Different from the static upsampling rate, the dynamic upsampling module can automatically adjust the sampling rate according to the changes in real-time data, adapt to the changes in data distribution, thereby improving the performance of the model and avoiding the limitations that may be brought by the static sampling strategy.

[0083] Specifically, the present disclosure selects the Dysample module as the new dynamic upsampling module. The Dysample module is lighter and more efficient than traditional upsampling methods (such as bilinear interpolation and nearest neighbor interpolation). It can not only reduce the computational load but also avoid time-consuming dynamic convolution and additional sub-networks, showing excellent performance in detection and segmentation tasks. At the same time, the Dysample module performs upsampling through a simplified point sampling method, which can better save computational resources, and is also relatively simple to implement in Pytorch and does not require a customized CUDA package. Compared with the kernel-based dynamic upsampling method, the Dysample module has fewer parameters, lower floating-point operation counts (FLOP), lower GPU memory requirements, and lower latency, thus broadening its application scenarios and enhancing the practical value of the model.

[0084] Exemplarily, the attention mechanism enables the model to selectively focus on relatively important parts of the input data by mimicking human attention, thereby improving the performance of the task. Through dynamic weight allocation and global calculation, it plays an important role in capturing long-range dependencies, dynamic weight allocation, multi-modal data fusion, and enhancing model interpretability. In this disclosure, the attention extraction module in the original model is replaced with a Triplet attention extraction module, which enhances the model's ability to focus on different dimensions of the feature map by considering information interaction in both spatial and channel dimensions, thereby improving the model's feature representation ability and overall performance. The input of this module is a tensor with the shape of C×H×W. The whole module consists of three parallel branches, and each branch performs the same operations, namely Z-pool operation, convolution operation, and applying the Sigmoid function to compress the output value between 0 and 1. The Z-pool operation is a pooling operation that performs average pooling and max pooling on the input tensor in the channel dimension and concatenates the results in the channel dimension. The convolution operation uses a 7×7 convolution, batch normalization, and the Sigmoid function to generate attention weights. Finally, after dimension permutation again, the weights are applied to the original input tensor through a 1×1 convolution to achieve the purpose of feature reweighting.

[0085] S203. Obtain a set of labeled images corresponding to the robot.

[0086] It can be understood that the set of labeled images refers to a dataset containing multiple images with manual annotations. In these images, each pixel has a clear label indicating which category or object it belongs to. The set of labeled images includes multiple environmental labeled images, and each environmental labeled image includes multiple feature objects. Feature objects refer to objects that need to be segmented or detected in the image; these feature objects can be common objects in daily life, such as tables, chairs, and cabinets, etc.

[0087] S204. Train the YOLOv8 segmentation model to be trained based on the set of labeled images to obtain the trained YOLOv8 segmentation model.

[0088] It can be understood that after obtaining the set of labeled images, the YOLOv8 segmentation model to be trained can be trained according to the set of labeled images. As shown in Figure 3 , it specifically includes the following steps S2041 to S2047:

[0089] S2041. Respectively perform feature extraction on each environmental labeled image in the set of labeled images based on the feature extraction module to obtain an original feature map corresponding to each environmental labeled image.

[0090] It is understandable that the feature extraction module extracts features from each environmental label image in the label image set, analyzes basic information such as pixels, colors, and shapes in the image, and extracts a feature map that can represent the image content. In the YOLOv8 segmentation model, the feature extraction module usually uses deep learning structures such as convolutional neural networks (CNNs) to capture local and global information of the image

[0091] S2042, Based on the dynamic upsampling model, perform dynamic upsampling processing on each of the environmental label images in the label image set to obtain a high-resolution environmental label image corresponding to each of the environmental label images.

[0092] Exemplarily, each environmental label image is processed by the dynamic upsampling model, thereby improving the resolution of the image, enabling the model to more meticulously identify and process small objects or complex details in the image. The design of the dynamic upsampling model can flexibly adjust the parameters in the upsampling process according to the characteristics of the input image to ensure that the output high-resolution image can improve the accuracy of subsequent segmentation.

[0093] S2043, Based on the Triplet attention extraction module, perform global feature extraction on the high-resolution environmental label image corresponding to each of the environmental label images to obtain an attention weight map corresponding to each of the high-resolution environmental label images.

[0094] Specifically, using the Triplet attention extraction module to perform global feature extraction on the high-resolution environmental label image can concentrate more computing resources on important regions or key features when processing complex images, thereby enhancing the segmentation effect. The Triplet attention module specifically extracts the global features of the image through three key elements (usually query, key, and value) to obtain an attention weight map corresponding to each high-resolution image. This weight map reflects the importance of each region in the image and can guide the model to more precisely process the regions of interest during segmentation.

[0095] S2044, Based on the image segmentation module, determine the segmentation result corresponding to each of the environmental label images according to the original feature map corresponding to each of the environmental label images and the attention weight map corresponding to each of the high-resolution environmental label images.

[0096] Here, the YOLOv8 segmentation model combines the original feature map and the attention weight map for final image segmentation. The image segmentation module utilizes the information in the feature map and the weights provided by the attention mechanism to generate the segmentation result. The YOLOv8 segmentation model is based on multi-layer convolutional operations in deep learning. By processing and refining the image layer by layer, it finally completes the segmentation of objects in the image. In this way, it is ensured that the model can accurately identify the target regions in each labeled image and effectively separate them from the background or other regions.

[0097] S2045, calculate the loss value between the segmentation result corresponding to each environment labeled image and the environment label sample image corresponding to each environment labeled image based on a preset loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value.

[0098] Exemplarily, after the image segmentation result is determined, it needs to be compared with the ground truth labeled image to evaluate the accuracy of the segmentation. To measure the segmentation quality, a preset loss function is used to calculate the difference between the segmentation result of each image and the ground truth environment labeled image, so as to find out the errors existing in the model during the training process, and guide the model adjustment by calculating the loss value. The smaller the loss value, the closer the segmentation result of the model is to the ground truth label, otherwise it needs to be further optimized. Based on these loss values, the weights and parameters of the model are adjusted accordingly, so that the model can gradually improve its segmentation accuracy.

[0099] Here, the preset loss function can be a cross-entropy loss function, Dice loss, IoU loss, etc., which are not specifically limited here.

[0100] S2046, determine whether the training result meets the preset requirements; if so, execute step S2047; if not, return to step S2041.

[0101] It can be understood that after completing one training iteration, it is necessary to determine whether the training result meets the preset accuracy requirements. Here, if the training result does not meet the requirements (that is, the value of the loss function is still large or the segmentation accuracy is not high), then return to the feature extraction step (S2041) to re-train and optimize. In this way, it is ensured that the model is continuously strengthened and improved during the training process until it reaches a high performance standard.

[0102] S2047, obtain the trained YOLOv8 segmentation model.

[0103] Here, if the training result meets the preset requirements, it indicates that the segmentation ability of the model is accurate enough, and the training can be ended and the trained YOLOv8 segmentation model can be output.

[0104] S102. Determine the depth map corresponding to the image to be processed, and determine the target image information based on the information of multiple feature objects in the image to be processed and the depth map.

[0105] It can be understood that a depth map is an image, and each pixel value in it represents the relative distance from the corresponding point in the image to the camera or sensor. When calculating the depth map corresponding to an image, a monocular depth estimation method is usually adopted. The monocular depth estimation method estimates the depth information of each pixel from a single image, that is, infers the distance from each pixel to the robot camera, and uses a two-dimensional image to infer the depth relationship in three-dimensional space. The depth map predicted by the monocular depth estimation provides the depth information of the objects in the scene, which is crucial for understanding the three-dimensional space structure. It is widely used in fields such as autonomous driving, augmented reality, and robot navigation.

[0106] Exemplarily, the present disclosure uses the Lite-Mono monocular depth estimation network to predict the depth map. Lite-Mono is a lightweight monocular depth estimation network. The Lite-Mono network combines a CNN with a Transformer module. The feature map extracted by the CNN is input into the Transformer encoder to capture global context information. Through the self-attention mechanism, the Transformer can effectively capture the long-range dependence relationships in the feature map. Compared with traditional depth estimation methods, the advantage of Lite-Mono lies in its high efficiency and easy deployment, especially suitable for application scenarios with low requirements for computing resources or requiring fast response, and can reduce the computational complexity and model parameters while ensuring relatively high accuracy.

[0107] Specifically, after image preprocessing the image to be processed, multi-level features of the input image are extracted through a feature extraction module, and then depth decoding is performed. The size of the feature map is upsampled to the original size through upsampling. Subsequently, skip connections are performed, and the intermediate layer feature map in the backbone network is concatenated or added to the feature map of the decoder. The number of channels is compressed through a convolutional layer, and the encoded feature map is decoded into a depth map. Secondly, the depth map is post-processed, filtered and optimized to reduce noise and smooth the depth map to improve the details and accuracy of the depth map. Finally, the final depth map is obtained.

[0108] In some other embodiments, the depth map can also be predicted by methods such as stereo vision and deep learning, which are not specifically limited herein.

[0109] It can be understood that after obtaining the depth map corresponding to the image to be processed, the information of multiple feature objects in the image to be processed can be registered with the depth map, and the visual elements in the two-dimensional image are mapped into three-dimensional space. Thus, based on these spatial coordinate relationships, the target image information can be obtained more accurately.

[0110] Here, the registration process generally involves the extraction and matching of image features, as well as the calculation of spatial transformation based on these features. Using computer vision algorithms such as SIFT, SURF, or other deep learning feature extraction methods, key points and their descriptors can be effectively identified from the image, and then corresponding matching points can be found on the depth map.

[0111] Exemplarily, the present disclosure draws segmentation masks corresponding to multiple feature object information onto the depth map to obtain target image information, so as to enhance the information expression ability of the depth map. As an efficient object detection and segmentation model, YOLOv8 can quickly and accurately identify multiple objects in the image and generate corresponding segmentation masks. Overlaying these masks onto the depth map means that not only can the two-dimensional boundary information of the object be obtained, but also its position and depth relationship in the three-dimensional space can be known.

[0112] It can be understood that the depth information in the depth map output by the Lite-Mono network is relative depth, which only provides the distance relationship between objects in the scene, and lacks specific physical size or actual distance information, which may cause limitations in practical applications. To convert the relative depth into absolute depth with practical value, a reference system, that is, a target calibration object, needs to be introduced. After determining the target calibration object in the environment corresponding to the robot, it can be realized through various sensors integrated on the robot, such as lidar (LiDAR), ultrasonic sensors, or vision-based ranging methods (such as stereo vision), to measure the actual distance between the robot and the calibration object. When the distance between the robot and the target calibration object is obtained, the obtained distance information is combined with the known size of the calibration object to calculate the scale factor. Here, the scale factor is actually a proportional coefficient, which converts the depth value in the relative depth map into a specific depth value corresponding to the actual size.

[0113] Specifically, the relative depth value of the area where the calibration object is located is extracted from the depth map; then, using the actual size of the calibration object and the measured distance from the robot to the calibration object, the scale factor is deduced through geometric relationships. In this way, through the scale factor and the depth value of each pixel point in the depth map, the relative depth value of each pixel point can be converted into an absolute depth value.

[0114] S103, determine a reference feature object based on the target image information, and determine the position information of the robot based on the target image information and the reference feature object.

[0115] Here, after obtaining the target image information, the distance between the robot and each feature in the target image information can be calculated to determine the reference features. Specifically, by calculating the depth information and scale factor of each feature, the distances between these features and the robot can be determined. The depth information reflects the distance of the feature from the camera, while the scale factor is used to adjust or calibrate the measurement value to ensure the accuracy of the distance. Based on these distance data, the two features closest to the robot can be determined as the first reference feature and the second reference feature respectively, which are used to calculate the position information of the robot.

[0116] Exemplarily, after determining the first reference feature and the second reference feature, the calculation process of the robot position information may include the following steps (a) - (b):

[0117] (a) Determine the depth information of the first reference feature and the second reference feature respectively according to the target image information; and determine the distance between the first reference feature and the robot according to the scale factor and the depth information of the first reference feature; determine the distance between the second reference feature and the robot according to the scale factor and the depth information of the second reference feature;

[0118] (b) Based on the distance between the first reference feature and the robot and the distance between the second reference feature and the robot, determine the position information of the robot based on the cosine theorem.

[0119] It can be understood that the target image information includes the depth information of each feature. Therefore, after determining the first reference feature and the second reference feature, the depth information corresponding to the first reference feature and the second reference feature can be determined according to the target image information, and then the actual distances between the two reference features and the robot can be calculated according to the depth information and the scale factor.

[0120] Specifically, after calculating the distance between the first reference feature and the robot and the distance between the second reference feature and the robot, the position information of the robot is further determined by the cosine theorem. The cosine theorem can calculate the relative position of the third point when the distance between two points and the angle between the two points are known. In the present disclosure, the first reference feature and the second reference feature provide the positions of two known points, and the position of the robot can be calculated by the distances between these two features and the robot and the relative angle between them.

[0121] The robot positioning method based on the combination of improved YOLOv8 and depth map provided by the embodiments of the present disclosure can accurately extract and analyze the information of key feature objects in the environment by combining the improved YOLOv8 image segmentation model and depth map technology. At the same time, by using the spatial depth information provided by the depth map, the accuracy and robustness of robot positioning are improved. The present disclosure can effectively reduce errors and improve positioning accuracy in complex environments. At the same time, the method provided by the present disclosure performs relatively stably in environments with dynamic changes or uneven illumination. In addition, the positioning method combined with depth information enhances the perception ability of obstacles and target objects, enabling the robot to better adapt to complex working scenarios when performing tasks, thereby improving the efficiency of its autonomous navigation and decision-making.

[0122] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0123] Based on the same inventive concept, the embodiments of the present disclosure also provide a robot positioning device based on the combination of improved YOLOv8 and depth map corresponding to the robot positioning method based on the combination of improved YOLOv8 and depth map. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above-mentioned robot positioning method based on the combination of improved YOLOv8 and depth map in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0124] Referring to Figure 4 As shown, it is a schematic diagram of a robot positioning device 400 based on the combination of improved YOLOv8 and depth map provided by the embodiments of the present disclosure. The device includes:

[0125] An image segmentation module 401, configured to obtain an image to be processed, and perform image segmentation processing on the image to be processed based on a trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed; wherein, the image to be processed is captured by a robot, and the segmentation result includes information of multiple feature objects in the image to be processed;

[0126] An image determination module 402, configured to determine a depth map corresponding to the image to be processed, and determine target image information based on the information of multiple feature objects in the image to be processed and the depth map;

[0127] A position determination module 403, configured to determine a reference feature object based on the target image information, and determine the position information of the robot based on the target image information and the reference feature object.

[0128] In some possible embodiments, referring to Figure 5As shown, the device further includes:

[0129] A model construction module 404 for constructing an initial YOLOv8 segmentation model; wherein, the initial YOLOv8 segmentation model includes a feature extraction module, an upsampling module, an attention extraction module, and an image segmentation module;

[0130] A model determination module 405 for replacing the upsampling module in the initial YOLOv8 segmentation model with a dynamic upsampling module and replacing the attention extraction module with a Triplet attention extraction module to obtain a YOLOv8 segmentation model to be trained;

[0131] A data acquisition module 406 for acquiring a set of labeled images corresponding to the robot; wherein, the set of labeled images includes multiple environmental labeled images, and each of the environmental labeled images includes multiple feature objects;

[0132] A model training module 407 for training the YOLOv8 segmentation model to be trained based on the set of labeled images to obtain the trained YOLOv8 segmentation model.

[0133] In some possible embodiments, the model training module 407 is specifically configured to:

[0134] Extract features from each of the environmental labeled images in the set of labeled images based on the feature extraction module to obtain an original feature map corresponding to each of the environmental labeled images;

[0135] Perform dynamic upsampling processing on each of the environmental labeled images in the set of labeled images based on the dynamic upsampling model to obtain a high-resolution environmental labeled image corresponding to each of the environmental labeled images;

[0136] Extract global features from the high-resolution environmental labeled image corresponding to each of the environmental labeled images based on the Triplet attention extraction module to obtain an attention weight map corresponding to each of the high-resolution environmental labeled images;

[0137] Based on the image segmentation module, determine a segmentation result corresponding to each of the environmental labeled images according to the original feature map corresponding to each of the environmental labeled images and the attention weight map corresponding to each of the high-resolution environmental labeled images;

[0138] Calculate a loss value between the segmentation result corresponding to each of the environmental labeled images and the environmental labeled sample image corresponding to each of the environmental labeled images based on a preset loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value;

[0139] Repeat the above steps until the training result meets the preset requirements, and obtain the trained YOLOv8 segmentation model.

[0140] In some possible embodiments, the image determination module 402 is specifically configured to:

[0141] Register the information of multiple feature objects in the image to be processed with the depth map, and obtain the target image information based on the registration result.

[0142] In some possible embodiments, the image determination module 402 is further configured to:

[0143] Determine the target calibration object in the environment corresponding to the robot, and determine the scale factor based on the size information of the target calibration object and the distance between the robot and the target calibration object.

[0144] In some possible embodiments, the position determination module 403 is specifically configured to:

[0145] Calculate the distances between the robot and each feature object in the target image information respectively based on the scale factor and the target image information, and determine the first reference feature object and the second reference feature object based on the calculation results.

[0146] In some possible embodiments, the position determination module 403 is specifically configured to:

[0147] Determine the depth information of the first reference feature object and the second reference feature object respectively according to the target image information; and determine the distance between the first reference feature object and the robot according to the scale factor and the depth information of the first reference feature object; determine the distance between the second reference feature object and the robot according to the scale factor and the depth information of the second reference feature object;

[0148] Based on the distances between the first reference feature object and the robot and between the second reference feature object and the robot, determine the position information of the robot based on the cosine theorem.

[0149] Based on the same inventive concept, an embodiment of the present disclosure also provides a computer device. Refer to Figure 6 As shown, it is a schematic structural diagram of a computer device 600 provided by an embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. Among them, the memory 602 is used to store execution instructions, including an internal memory 6021 and an external memory 6022; the internal memory 6021 here is also called the main memory, which is used to temporarily store the operation data in the processor 601 and the data exchanged with the external memory 6022 such as a hard disk, and the processor 601 exchanges data with the external memory 6022 through the internal memory 6021.

[0150] In the embodiments of the present application, the memory 602 is specifically configured to store the application program code for executing the solution of the present application, and is controlled by the processor 601 to execute. That is, when the computer device 600 runs, the processor 601 communicates with the memory 602 through the bus 603, so that the processor 601 executes the application program code stored in the memory 602, and further executes the method described in any of the foregoing embodiments.

[0151] Among them, the memory 602 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0152] The processor 601 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0153] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the computer device 600. In other embodiments of the present application, the computer device 600 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange different components. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0154] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the robot positioning method based on the combination of improved YOLOv8 and depth map described in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.

[0155] Embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the robot positioning method based on the combination of improved YOLOv8 and depth map described in the above method embodiments. For details, please refer to the above method embodiments and will not be elaborated here.

[0156] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0157] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0158] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0160] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0161] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A robot positioning method based on improved YOLOv8 combined with depth map, characterized in that: include: Acquire an image to be processed, and perform image segmentation processing on the image to be processed based on the trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed; wherein the image to be processed is obtained by photographing a robot, and the segmentation result includes information of multiple feature objects in the image to be processed; Determine a depth map corresponding to the image to be processed, and determine target image information based on a plurality of feature object information in the image to be processed and the depth map; A reference feature is determined based on the target image information, and position information of the robot is determined based on the target image information and the reference feature.

2. The method according to claim 1, characterized in that Before performing image segmentation processing on the image to be processed based on the trained YOLOv8 segmentation model, the method includes: Constructing an initial YOLOv8 segmentation model; wherein the initial YOLOv8 segmentation model includes a feature extraction module, an upsampling module, an attention extraction module and an image segmentation module; The upsampling module in the initial YOLOv8 segmentation model is replaced with a dynamic upsampling module, and the attention extraction module is replaced with a Triplet attention extraction module to obtain a YOLOv8 segmentation model to be trained; Acquire a label image set corresponding to the robot; wherein the label image set includes a plurality of environment label images, and each of the environment label images includes a plurality of feature objects; The YOLOv8 segmentation model to be trained is trained based on the labeled image set to obtain the trained YOLOv8 segmentation model.

3. The method according to claim 2, characterized in that The training of the YOLOv8 segmentation model to be trained based on the label image set includes: Based on the feature extraction module, feature extraction is performed on each of the environment label images in the label image set to obtain an original feature map corresponding to each of the environment label images; Based on the dynamic upsampling model, each of the environment label images in the label image set is dynamically upsampled to obtain a high-resolution environment label image corresponding to each of the environment label images; Based on the Triplet attention extraction module, global feature extraction is performed on the high-resolution environment label image corresponding to each of the environment label images to obtain an attention weight map corresponding to each of the high-resolution environment label images; Determine, based on the image segmentation module, a segmentation result corresponding to each of the environment label images according to the original feature map corresponding to each of the environment label images and the attention weight map corresponding to each of the high-resolution environment label images; Calculate the loss value between the segmentation result corresponding to each of the environment label images and the environment label sample image corresponding to each of the environment label images based on a preset loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value; Repeat the above steps until the training results meet the preset requirements to obtain the trained YOLOv8 segmentation model.

4. The method according to claim 1, characterized in that: The determining target image information based on the plurality of feature object information in the image to be processed and the depth map comprises: The plurality of feature object information in the image to be processed is registered with the depth map, and the target image information is obtained based on the registration result.

5. The method according to claim 1, characterized in that Before determining the depth map corresponding to the image to be processed, the method includes: A target calibration object in an environment corresponding to the robot is determined, and a scale factor is determined based on size information of the target calibration object and a distance between the robot and the target calibration object.

6. The method according to claim 5, characterized in that The determining of the reference feature based on the target image information comprises: The distances between the robot and each feature in the target image information are calculated based on the scale factor and the target image information, and the first reference feature and the second reference feature are determined based on the calculation results.

7. The method according to claim 6, characterized in that The determining the position information of the robot based on the target image information and the reference feature object comprises: Determine the depth information of the first reference feature and the second reference feature respectively according to the target image information; and determine the distance between the first reference feature and the robot according to the scale factor and the depth information of the first reference feature; and determine the distance between the second reference feature and the robot according to the scale factor and the depth information of the second reference feature; According to the distance between the first reference feature and the robot and the distance between the second reference feature and the robot, the position information of the robot is determined based on the law of cosines.

8. A robot positioning device based on improved YOLOv8 combined with a depth map, characterized in that: include: An image segmentation module is used to obtain an image to be processed, and perform image segmentation processing on the image to be processed based on a trained YOLOv8 segmentation model to obtain a segmentation result corresponding to the image to be processed; wherein the image to be processed is obtained by photographing a robot, and the segmentation result includes information of multiple feature objects in the image to be processed; An image determination module, used to determine a depth map corresponding to the image to be processed, and determine target image information based on a plurality of feature object information in the image to be processed and the depth map; A position determination module is used to determine a reference feature based on the target image information, and to determine the position information of the robot based on the target image information and the reference feature.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.