Target detection and positioning method and system for indoor target search task

By improving the YOLOv8 model and combining the binocular camera to obtain depth information, the problem of overlapping small object detection and occlusion in complex environments is solved, and more efficient target search and positioning effects are achieved.

CN120198631APending Publication Date: 2025-06-24CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149138.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In a real search environment, the target may be very small and there are problems such as target occlusion and overlap, making the target search more difficult and easily lead to false detection or missed detection. The YOLO series algorithms have limitations in dealing with overlapping situations of small object detection and occlusion.

Method used

Improve the YOLOv8 object detection model, and improve the robustness and detection accuracy of the model by building the CSPSPPF-S module and introducing SPD-Conv convolution and CA attention mechanism. Combined with the binocular camera to obtain depth information, match the target detection results to achieve more accurate positioning.

Benefits of technology

When dealing with complex scenarios and small targets, the improved model performance is more stable, improving the efficiency of target search, and realizing object detection and positioning in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198631A_ABST
    Figure CN120198631A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor target search task-oriented target detection and positioning method and system, and the method comprises the following steps: constructing an image data set of indoor common articles, and dividing the image data set into a training set and a test set according to a proportion; based on the improved YOLOv8 target detection model, constructing a target detection network; training the improved target detection network by adopting the training set according to preset parameters; images acquired by the binocular camera are transmitted into the trained target detection network for prediction, and the category of a target and the coordinates of a frame are obtained; acquiring depth information of the image based on a binocular camera, and matching the predicted frame center coordinate with the depth information to obtain the category and coordinate of target detection; a system interface is designed based on PyQt5, and target detection and positioning visualization and data storage are achieved. According to the invention, the detection precision of small targets can be improved, the omission ratio is reduced, and the indoor target search efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an object detection and positioning method and system for indoor target search tasks. Background Art

[0002] Robot target search is a key research topic in the field of robotics. Its process involves multiple steps, which are complex and challenging. However, in a real search environment, the target may be very small, and there are problems such as target occlusion and overlap, which make target search more difficult and prone to false detection or missed detection.

[0003] Object detection is a key task in target search. By automatically identifying and locating target objects, it can significantly improve search efficiency and accuracy. The YOLO series of algorithms are a class of algorithms widely concerned in the field of object detection. These methods transform the object detection task into a regression problem and directly predict the category and location of the target from the image, thus possessing high speed and real-time performance. However, the YOLO series of algorithms still have certain limitations in the accuracy of object detection, especially in dealing with small object detection and occlusion and overlap situations.

[0004] In the target search task, it is not enough to only detect the target object, but also to locate the target. The binocular vision system obtains images from two cameras at different angles and calculates the depth information of the object using the parallax. The stereo vision system has been applied in aspects such as robot navigation and driverless driving. Summary of the Invention

[0005] Object of the Invention: The object of the present invention is to provide an object detection and positioning method and system for indoor target search tasks, which are more stable in dealing with complex scenarios and small targets and improve the efficiency of target search.

[0006] Technical Solution: An object detection and positioning method for indoor target search tasks includes the following steps:

[0007] S1, construct an image dataset of common indoor items, and divide the image dataset into a training set and a test set according to a ratio;

[0008] S2, based on the improved YOLOv8 object detection model, construct an object detection network;

[0009] S3, use the training set to train the improved object detection network according to preset parameters;

[0010] S4, input the image obtained by the binocular camera into the trained object detection network for prediction to obtain the category of the target and the coordinates of the box;

[0011] S5. Based on the depth information of the images obtained by the binocular camera, match the predicted box center coordinates with the depth information to obtain the category and coordinates of the target detection;

[0012] S6. Design the system interface based on PyQt5 to achieve the visualization and data storage of target detection and positioning.

[0013] Furthermore, constructing an image dataset of common indoor items includes the following steps:

[0014] Extract all indoor item datasets from the COCO2017 dataset;

[0015] Convert the JSON format of the COCO dataset into the TXT format suitable for the training of the YOLO series algorithms;

[0016] Collect indoor images on the network based on the crawler, use the annotation tool to annotate the items in the images, and annotate them in the YOLO format;

[0017] Fuse the processed COCO dataset with the self-built dataset to expand the indoor target detection dataset.

[0018] Furthermore, improving the YOLOv8 target detection model includes:

[0019] Use the constructed CSPSPPF-S module to replace the SPPF module in the original network structure;

[0020] Introduce the SPD-Conv convolution in the backbone structure to replace the convolution block with a stride of 2;

[0021] Fuse the CA attention mechanism into the C2f module to construct the C2f-CA module.

[0022] Furthermore, using the constructed CSPSPPF-S module to replace the SPPF module in the original network structure includes:

[0023] Combine the CSP technology and retain the pooling structure in the SPPF module to construct the CSPSPPF-S module. First, divide the input into two branches;

[0024] The first branch consists of a convolutional layer, the parameter-free attention mechanism SimAM, the SPPF module, and a convolutional layer;

[0025] The second branch consists of a convolutional layer and the parameter-free attention mechanism SimAM;

[0026] After fusing the two branches, output through another convolutional layer.

[0027] Furthermore, introducing the SPD-Conv convolution in the backbone structure to replace the convolution block with a stride of 2 includes:

[0028] The input feature map is split into four identical sub-feature maps in the H and W dimensions, and the split sub-feature maps have the same number of channels as the original feature map;

[0029] The split sub-feature maps are concatenated in the channel dimension;

[0030] The concatenated new feature map passes through a convolutional layer to form an SPD-Conv convolutional module;

[0031] Finally, the SPD-Conv convolutional block replaces 5 convolutional layers with a stride of 2 in the YOLOv8 backbone network.

[0032] Furthermore, the CA attention mechanism is integrated into the C2f module to construct the C2f-CA module. The specific implementation is as follows: the CA attention mechanism is incorporated into the original YOLOv8 C2f module and then the output features are obtained; the three C2f modules in the neck of YOLOv8 are replaced with the improved C2f-CA modules.

[0033] Furthermore, training the improved object detection network based on the collected dataset includes:

[0034] Set the number of training epochs to 300, the learning rate to 0.01, the weight decay to 0.0005, freeze the training for the first 50 epochs, the batch size for frozen training is 16, and the batch size after unfreezing is 8;

[0035] Save the weight file of the object detection network at the end of training.

[0036] Furthermore, based on the binocular camera to obtain the depth information of the image, matching the predicted box center coordinates with the depth information to obtain the category and coordinates of object detection includes:

[0037] Calibrate the binocular camera based on the Zhang Zhengyou calibration method to obtain the parameters of the binocular camera;

[0038] Use the calibrated parameters to correct the images obtained by the binocular camera. After passing the corrected images through the stereo matching algorithm to obtain the disparity map, at the same time, the obtained images are input into the trained object detection model to obtain the predicted object categories and box coordinates;

[0039] Pass the disparity map through the projection matrix Q in the parameters to obtain a mapping map. Each pixel has three channels, which store the three-dimensional point coordinates (x, y, z) of this pixel position in the camera coordinate system;

[0040] Traverse all predicted object categories, then calculate the center coordinates of the object box, match the coordinates with the (x, y) of the mapping map, and the obtained 3D coordinates (x, y, z) are used as the coordinates of the object based on the camera coordinate system.

[0041] An object detection and positioning system for indoor object search tasks, which is used to execute the object detection and positioning methods of any one of the above, designs the system interface based on PyQt5, and includes a sensor unit, a CPU / GPU processing unit, an output unit, and a storage unit; the sensor unit uses a binocular camera to obtain image data; the image data is transmitted to the CPU / GPU processing unit through USB, executes object detection and stereo matching algorithms, displays the processed object detection and positioning results through the output unit, and at the same time stores the results in the MySQL database of the storage unit.

[0042] Compared with the prior art, the remarkable effects of the present invention are as follows:

[0043] 1. The present invention improves the YOLOv8 algorithm, replaces the SPPF module in the original network structure with the CSPSPPF-S module, introduces the backbone structure of SPD-Conv convolution, and replaces the C2f module in the original network structure with C2f-CA. The improved model effectively improves the robustness and detection accuracy, especially shows more stability when dealing with complex scenes and small targets, and improves the efficiency of object search;

[0044] 2. The present invention integrates a binocular camera, calibrates the binocular camera, obtains the depth information of the image through stereo matching, matches the category coordinates after object detection with the depth information to obtain the object category and the coordinates in the camera coordinate system, and realizes object detection and positioning in complex environments; it has very high application value in the fields of home service, logistics distribution, medical care, and search and rescue;

[0045] 3. The present invention designs a system interface, integrates binocular stereo vision, object detection algorithms, and databases, realizes the integration of object detection and positioning and data storage, and facilitates visual operation by users. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flow chart of the present invention;

[0047] Figure 2 is a schematic diagram of the improved YOLOv8 model;

[0048] Figure 3 is a schematic diagram of the structure of the CSPSPPF-S module;

[0049] Figure 4 is a schematic diagram of the structure of the C2f-CA module;

[0050] Figure 5 is a schematic diagram of the structure of the positioning system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The present invention will be further described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0052] As Figure 1 shown is the flow chart of the present invention, which specifically includes the following steps:

[0053] Step 1, construct an image dataset of common indoor items;

[0054] It includes the following steps:

[0055] Step 11, extract all indoor item datasets from the COCO2017 dataset;

[0056] Step 12, convert the JSON format of the COCO dataset into the TXT format suitable for training YOLO series algorithms;

[0057] Step 13, collect indoor pictures on the network based on a crawler, use a labeling tool to label the items in the images, and label them in the YOLO format;

[0058] Step 14, fuse the processed COCO dataset with the self-built dataset to expand the indoor object detection dataset, and obtain an image dataset of common indoor items; and divide the image dataset into a training set and a test set according to 7:3.

[0059] Step 2, improve the YOLOv8 object detection model and construct an object detection network;

[0060] As Figure 2 shown, it is a diagram of improving the YOLOv8 object detection model. Constructing an object detection network includes the following steps:

[0061] Step 21, use the constructed CSPSPPF-S module (Cross Stage Partial Scalable Packet Processing Framework-SimAM, cross-stage partial simple attention spatial pyramid pooling fast module) to replace the SPPF module (Scalable Packet Processing Framework, spatial pyramid pooling fast version) of the original network structure;

[0062] It includes the following steps:

[0063] Step 211, combine the CSP (Cross Stage Partial, cross-stage partial module) technology and retain the pooling structure in the SPPF module to construct the CSPSPPF-S module. As Figure 3 shown, first divide the input into two branches;

[0064] The first branch consists of a convolutional layer, a parameter-free attention mechanism SimAM (Simple Attention Module), an SPPF module, and a convolutional layer;

[0065] The second branch consists of a convolutional layer and a parameter-free attention mechanism SimAM;

[0066] Step 212: After fusing the two branches, pass through a convolutional layer for output.

[0067] Step 22: Introduce an SPD-Conv (Space-to-Depth Convolution) convolution in the backbone structure to replace the convolution block with a stride of 2;

[0068] It includes the following steps:

[0069] Step 221: Split the input feature map into four identical sub-feature maps in the H and W dimensions. The split sub-feature maps have the same number of channels as the original feature map;

[0070] Step 222: Concatenate the split sub-feature maps in the channel dimension;

[0071] Step 223: Pass the concatenated new feature map through a convolutional layer to form an SPD-Conv convolution module;

[0072] Step 224: Finally, replace 5 convolution layers with a stride of 2 in the YOLOv8 backbone network with SPD-Conv convolution blocks.

[0073] Step 23: Incorporate the CA (Coordinate Attention) attention mechanism into the C2f (CSP Bottleneck with 2 convolutions) module to construct the C2f-CA module;

[0074] It includes the following steps:

[0075] Step 231: Output features after integrating the CA attention mechanism into the original YOLOv8 C2f module;

[0076] Step 232: Replace the three C2f modules in the neck of YOLOv8 with C2f-CA modules. The structure of the C2f-CA module is as Figure 4 shown.

[0077] Step 3: Train the improved object detection network based on the constructed training set according to preset parameters;

[0078] It includes the following steps:

[0079] Step 31: Set the number of training rounds to 300, the learning rate to 0.01, the weight decay to 0.0005, freeze the training for the first 50 rounds, set the batch_size for the frozen training to 16, and the batch_size after unfreezing to 8;

[0080] Step 32: Save the model weight file after the training ends.

[0081] Step 4: Input the images obtained by the binocular camera into the trained object detection network for prediction to obtain the category of the object and the coordinates of the bounding box;

[0082] Step 5: Based on the depth information of the images obtained by the binocular camera, match the center coordinates of the predicted bounding box with the depth information to obtain the category and coordinates of the object detection;

[0083] It includes the following steps:

[0084] Step 51: Calibrate the binocular camera based on the Zhang Zhengyou calibration method to obtain the parameters of the binocular camera;

[0085] Step 52: Use the calibrated parameters to correct the images obtained by the binocular camera. After correcting the images, obtain the disparity map through the SGBM stereo matching algorithm. At the same time, input the obtained images into the trained object detection model to get the predicted object categories and bounding box coordinates;

[0086] Step 53: Pass the disparity map through the projection matrix Q (the parameter matrix obtained after calibrating the binocular camera) in the parameters to obtain a mapping map. Each pixel has three channels, which store the three-dimensional point coordinates (x, y, z) of this pixel position in the camera coordinate system;

[0087] Step 54: Traverse all the predicted object categories, then calculate the center coordinates of the object bounding box, match the coordinates with the (x, y) of the mapping map, and the obtained 3D coordinates (x, y, z) are approximately the coordinates of the object based on the camera coordinate system.

[0088] Step 6: Design the system interface based on PyQt5 to realize the visual use and data storage of object detection and positioning;

[0089] Use PyQt5 to design the system interface and store the data in the MySQL database.

[0090] Such as Figure 5As shown in the figure, a target detection and positioning system for indoor target search tasks includes a sensor unit, a processing unit, an output unit, and a storage unit. The sensor unit uses a binocular camera to obtain image data. The image data is transmitted to the CPU / GPU processing unit through USB, where target detection and stereo matching algorithms are executed. The processed target detection and positioning results are displayed through the output unit, and at the same time, the results are stored in the MySQL database of the storage unit.

[0091] In summary, the present invention discloses a method and system for target detection and positioning for indoor target search tasks. The method includes: constructing an image dataset of common indoor items; improving the YOLOv8 target detection model to construct a target detection network; training the improved target detection network based on the constructed dataset according to preset parameters; inputting the images obtained by the binocular camera into the trained target detection network for prediction to obtain the category of the target and the coordinates of the box; based on the depth information of the images obtained by the binocular camera, matching the predicted box center coordinates with the depth information to obtain the category and coordinates of the target detection; designing the system interface based on PyQt5, using the MySQL database to store data, and realizing the visual use and data storage of target detection and positioning.

[0092] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A target detection and positioning method for indoor target search tasks, characterized in that: The steps include: S1, build an image dataset of common indoor objects and divide the image dataset into a training set and a test set in proportion; S2, builds a target detection network based on the improved YOLOv8 target detection model; S3, using the training set to train the improved target detection network according to preset parameters; S4, passing the image acquired by the binocular camera into the trained target detection network for prediction, and obtaining the target category and the coordinates of the box; S5, obtaining the depth information of the image based on the binocular camera, matching the predicted frame center coordinates with the depth information, and obtaining the category and coordinates of the target detection; S6,designs the system interface based on PyQt5 to realize the visualization and data storage of target detection and positioning.

2. The target detection and positioning method for indoor target search tasks according to claim 1 is characterized in that: The steps to construct an image dataset of common indoor objects include the following: Extract all indoor object datasets from the COCO2017 dataset; Convert the JSON format of the COCO dataset into the TXT format suitable for YOLO series algorithm training; Collect indoor pictures from the Internet using crawlers, and use annotation tools to annotate objects in the images in YOLO format; The processed COCO dataset is fused with the self-built dataset to expand the indoor object detection dataset.

3. The target detection and positioning method and system for indoor target search tasks according to claim 1, characterized in that: Improvements to the YOLOv8 target detection model include: Use the constructed CSPSPPF-S module to replace the SPPF module of the original network structure; Introduce SPD-Conv convolution in the backbone structure to replace the convolution block with a step size of 2; The CA attention mechanism is integrated into the C2f module to construct the C2f-CA module.

4. The target detection and positioning method for indoor target search tasks according to claim 3 is characterized in that: The SPPF modules of the original network structure are replaced by the constructed CSPSPPF-S modules, including: Combining CSP technology and retaining the pooling structure in the SPPF module to build the CSPSPPF-S module, first divide the input into two branches; The first branch consists of a convolutional layer, a parameter-free attention mechanism SimAM, an SPPF module, and a convolutional layer; The second branch consists of a convolutional layer and a parameter-free attention mechanism SimAM; The two branches are fused and then output through a convolutional layer.

5. The target detection and positioning method for indoor target search tasks according to claim 3 is characterized in that: The SPD-Conv convolution is introduced in the backbone structure to replace the convolution block with a step size of 2, including: The input feature map is split into four identical sub-feature maps in the H and W dimensions. The split sub-feature maps have the same channels as the original feature map. Concatenate the split sub-feature maps in the channel dimension; The concatenated new feature map passes through a convolution layer to form the SPD-Conv convolution module; Finally, the SPD-Conv convolution block replaces the five convolutional layers with a stride of 2 of the YOLOv8 backbone network.

6. The target detection and positioning method for indoor target search tasks according to claim 3 is characterized in that: The CA attention mechanism is integrated into the C2f module to construct the C2f-CA module. The specific implementation is as follows: the CA attention mechanism is integrated into the original C2f module of YOLOv8 and the features are output; the three C2f modules at the neck of YOLOv8 are replaced with the improved C2f-CA module.

7. The target detection and positioning method for indoor target search tasks according to claim 1, characterized in that: The improved object detection network trained based on the collected dataset includes: Set the number of training rounds to 300, the learning rate to 0.01, the weight decay to 0.0005, the frozen training to the first 50 rounds, the batch_size of the frozen training to 16, and the batch_size after thawing to 8; Save the target detection network weight file after training.

8. The target detection and positioning method for indoor target search tasks according to claim 1, characterized in that: Based on the depth information of the image obtained by the binocular camera, the predicted frame center coordinates are matched with the depth information to obtain the category and coordinates of the target detection, including: The binocular camera is calibrated based on Zhang Zhengyou's calibration method to obtain the parameters of the binocular camera; The image acquired by the binocular camera is corrected using the calibrated parameters. The disparity map is obtained by applying the stereo matching algorithm to the corrected image. At the same time, the acquired image is passed into the trained target detection model to obtain the predicted target category and frame coordinates. The disparity map is passed through the projection matrix Q in the parameters to obtain a mapping map. Each pixel has three channels, which store the three-dimensional point coordinates (x, y, z) of the pixel position in the camera coordinate system. Traverse all predicted target categories, then calculate the center coordinates of the target box, match the coordinates with the (x, y) of the mapping image, and use the matched 3D coordinates (x, y, z) as the coordinates of the target based on the camera coordinate system.

9. A target detection and positioning system for indoor target search tasks, used to execute the target detection and positioning method according to any one of claims 1 to 8, characterized in that: The system interface is designed based on PyQt5, including a sensor unit, a CPU / GPU processing unit, an output unit and a storage unit. The sensor unit uses a binocular camera to acquire image data. The image data is transmitted to the CPU / GPU processing unit via USB to execute target detection and stereo matching algorithms. The processed target detection and positioning results are displayed through the output unit, and the results are stored in the MySQL database of the storage unit.