A target detection and positioning method for a city scene, a UAV and a storage medium
By equipping drones with edge intelligent devices and using a lightweight YOLOv8-seg model and interactive control information to optimize image segmentation, the problems of low target detection accuracy and real-time performance in urban scenes are solved, and high-precision target segmentation and positioning in complex urban environments are achieved.
Patent Information
- Application Number
- CN202410945300.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-07-15
AI Technical Summary
In the existing technology, the accuracy and real-time performance of target object detection in urban scenes are low, especially in complex urban environments, where target segmentation is inaccurate and positioning accuracy is not high, and the back-end server has difficulty meeting real-time requirements when processing complex scenes.
Equipped with edge intelligent devices, the drone uses image acquisition equipment to obtain images, performs image instance segmentation and target object detection through the lightweight YOLOv8-seg model, receives interactive control information to perform image segmentation correction, and combines morphological algorithms to optimize segmentation edges to achieve target detection and positioning.
It improves the accuracy and real-time performance of target detection in urban scenarios, realizes the real-time segmentation and positioning of typical targets such as urban roads, vehicles, and buildings, and meets the high-precision and efficient computing requirements in complex scenarios.
Smart Images

Figure CN119360258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart cities, and in particular to a target detection and positioning method, a drone, and a storage medium for urban scenes. Background Art
[0002] In the field of smart cities, drones usually use their onboard visible light cameras to capture images and videos, and then use back-end servers to store and analyze the captured images and videos, thereby achieving target detection and positioning in the images and videos.
[0003] However, existing visible light cameras typically rely on monocular cameras, resulting in inaccurate object segmentation and low positioning accuracy. Furthermore, when back-end servers perform semantic segmentation and image recognition on images and videos, they often struggle with accurate and real-time object detection and positioning due to the complex pixel classification and high information processing requirements in urban scenes.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a target detection and positioning method, a drone and a storage medium for urban scenes, so as to solve the problems in the prior art of complex urban scenes and low accuracy and real-time performance of target object detection.
[0006] The technical solutions of the present invention are as follows:
[0007] In a first aspect, this embodiment discloses a method for target detection and positioning in an urban scene, wherein the method is applied to an edge intelligent device carried by a drone, wherein the drone is provided with an image acquisition device, and the method includes:
[0008] Acquire a drone image of an urban scene from the image acquisition device;
[0009] Performing image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs;
[0010] Receive interactive control information, perform image segmentation correction on each candidate ROI according to the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image and segmentation correction parameters corresponding to the target object categories.
[0011] Optionally, the step of acquiring a drone image in an urban scene includes:
[0012] Through the communication connection with the image acquisition device, video image information is obtained from the image acquisition device, and the format of the video image information is converted into a recognizable data format using a coding and decoding algorithm to obtain the drone image.
[0013] Optionally, the edge intelligent device performs image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs, including:
[0014] Performing lightweight processing on the model accelerator and / or the target detection model to obtain a lightweight accelerator and a lightweight target detection model after lightweight processing;
[0015] A lightweight target detection model is used to perform image instance segmentation and target object detection on the drone image, and the lightweight accelerator is used to control the process of image instance segmentation and target object detection to obtain multiple candidate ROIs.
[0016] Optionally, the model accelerator is a TensorRT network, and the lightweight target detection model is a lightweight YOLOv8-seg model.
[0017] Optionally, the step of receiving interactive control information, performing image segmentation correction on each candidate ROI according to the interactive control information, and obtaining target detection and positioning results includes:
[0018] receiving interactive control information, and obtaining target object categories and segmentation correction parameters corresponding to each target object category contained in the interactive control information;
[0019] An image segmentation algorithm is used to further segment the target object region in each candidate ROI based on the segmentation correction parameter.
[0020] Optionally, the step of further segmenting the region where the target object belonging to the target object category is located in each candidate ROI using an image segmentation algorithm based on the segmentation correction parameter includes:
[0021] Obtain the detection target object area corresponding to the target object category in each candidate ROI, as well as the detection score corresponding to each detection target object;
[0022] The segmentation edge position of the area where the target object is located is determined according to the segmentation correction parameters and the detection score, and the candidate ROI is further segmented based on the determined segmentation edge position.
[0023] Optionally, the step of confirming the segmentation edge position of the area where the target object is located according to the segmentation correction parameter and the detection score includes:
[0024] According to the segmentation correction parameter corresponding to the target object category and the detection score corresponding to the detected target object, the target objects to be segmented whose detection scores are lower than the segmentation correction parameter of the corresponding category are screened out;
[0025] The segmentation edge position of the region where the target object to be segmented is located is determined according to the score correction parameter.
[0026] Optionally, after the step of further segmenting the region where the target object belonging to the target object category is located in each candidate ROI using an image segmentation algorithm based on the segmentation correction parameter, the step further includes:
[0027] The morphological algorithm is used to fill and segment the detection target image after image segmentation.
[0028] In a second aspect, this embodiment discloses a drone, comprising: an edge intelligent device and an image acquisition device;
[0029] The image acquisition device is used to capture drone images in urban scenes;
[0030] The edge intelligent device is used to obtain the drone image, perform image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs, and receive interactive control information, perform image segmentation correction on each candidate ROI according to the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image, and segmentation correction parameters corresponding to the target object categories.
[0031] In a third aspect, this embodiment discloses a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the target detection and positioning method for urban scenes.
[0032] This embodiment provides a method for detecting and positioning targets in urban scenes, a drone, and a storage medium. These methods are applied to an edge intelligent device carried by a drone. The method obtains drone images of urban scenes from the image acquisition device, performs image instance segmentation and target object detection on the drone images, and obtains multiple candidate ROIs. The method receives interactive control information and performs image segmentation correction on each candidate ROI based on the interactive control information to obtain target detection and positioning results. The interactive control information includes one or more target object categories contained in the drone image, as well as segmentation correction parameters corresponding to the target object categories. The method and drone of this embodiment utilize edge intelligent devices to detect and position targets in images. The edge intelligent device first processes the drone image to obtain candidate ROIs, and then utilizes the interactive control information and an image segmentation algorithm to further segment the candidate ROIs, thereby improving target detection accuracy. Furthermore, the model used in this embodiment utilizes lightweight processing to improve image processing efficiency. Therefore, the method of this embodiment has the advantages of high accuracy and real-time performance in target detection and positioning results. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary personnel in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0034] Figure 1 is a flowchart of the steps of the target detection and positioning method for urban scenes described in this embodiment;
[0035] Figure 2 This is a flowchart of real-time image transmission and video encoding and decoding of a drone in the method described in this embodiment;
[0036] Figure 3 This is a comparison chart of the TensorRT network optimization results in this embodiment;
[0037] Figure 4 This is a diagram showing the effect of implementing the interactive semantics and OpenCV segmentation algorithm in this embodiment;
[0038] Figure 5 Schematic diagram of the principle of image processing in this embodiment;
[0039] Figure 6 1 is a block diagram of the principle structure of the UAV described in this embodiment. DETAILED DESCRIPTION
[0040] The present invention provides a method for detecting and localizing an object in an urban setting, an unmanned aerial vehicle, and a storage medium. To clarify the objectives, technical solutions, and effects of the present invention, the present invention is further described below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.
[0041] In the embodiments and patent claims, unless otherwise specified, "a", "an" and "the" may refer to a single or plural number.
[0042] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features specified as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0043] In existing technologies, images and videos captured by drones are typically used for target detection and positioning in backend servers. However, backend servers often face challenges when performing semantic segmentation and recognition of images in complex scenarios:
[0044] 1. While semantic segmentation can classify pixels into broad categories such as “road” or “vehicle,” more detailed classification requires more complex models and larger datasets with diverse urban scenes.
[0045] 2. Urban road environments typically contain complex scenes, which often involve various objects approaching each other, occlusions, and different lighting conditions. When dealing with more complex scenes, existing deep learning models may find it difficult to accurately depict object boundaries, resulting in low target object recognition accuracy.
[0046] 3. The real-time deployment of semantic segmentation and target detection models for object recognition on urban roads requires high standards, and complex deep models are difficult to meet the strict latency requirements for real-time target recognition and positioning.
[0047] 4. Deep learning models must be robust to environmental conditions. Under adverse environmental conditions, the recognition performance of the object recognition system is low, and robustness needs to be improved through data enhancement and model adaptation.
[0048] Therefore, the existing technology uses a back-end server to process drone images, but there are still problems such as inaccurate target segmentation and low positioning accuracy.
[0049] In order to solve the above problems, this embodiment provides a method for target detection and positioning in urban scenes, a drone and a storage medium. By carrying an edge intelligent device on the drone, the edge intelligent device is used to perform image segmentation on the drone image, thereby realizing the recognition and positioning of the target object in the drone image. Furthermore, the edge intelligent device first performs target detection and instance segmentation on the drone image based on the image segmentation model to obtain candidate ROIs, and then optimizes the segmentation of the candidate ROIs according to the received interactive control information, thereby obtaining accurate recognition, detection and positioning results of the drone image. The present invention effectively solves the problems of typical target recognition accuracy and real-time computing of edge intelligent devices in existing complex urban scenes, and realizes real-time segmentation and positioning detection of typical targets such as urban roads, vehicles, buildings and vegetation based on the drone edge computing carrier.
[0050] The method and the drone of this embodiment are further described in detail below with reference to the accompanying drawings.
[0051] Combine Figure 1 As shown, this embodiment discloses a method for detecting and positioning targets in urban scenes, such as Figure 1 As shown, the method is applied to an edge intelligent device carried by a drone, wherein the drone is provided with an image acquisition device, and includes:
[0052] Step S1: Acquire a drone image of an urban scene from the image acquisition device.
[0053] The image acquisition device installed on the drone collects images or video streams in urban scenes. Based on the information transmission protocol of images between the edge intelligent device and the drone, the edge intelligent device obtains the images or video streams collected by the image acquisition device from the image acquisition device.
[0054] Specifically, the step of acquiring drone images in urban scenes includes:
[0055] Through the communication connection with the image acquisition device, the video image information is obtained from the image acquisition device, and the format of the video image information is converted into a recognizable data format using a coding algorithm to obtain the drone image. Figure 2 As shown in the figure, the stream management function uses IVideoStreamManager to manage stream channel settings and stream data output. The steps to enable the edge smart device to obtain the image or video stream from the image acquisition device include:
[0056] 1) Get available stream sources: Call getAvailableStreamSources to get available stream sources.
[0057] 2) Get available video channels: Call getAvailableVideoChannels to get available video channels.
[0058] 3) Bind the video source and open the stream channel:
[0059] Call startChannel in getAvailableVideoChannels to set the StreamSource obtained in step 1) to bind the stream source and stream channel, and start the current stream channel.
[0060] 4) If decoding is required for the connected device, a stream data listener can be added by calling addStreamDataListener to receive the stream data.
[0061] 5) If decoding by the connected device is not required, you can also use the IVideoDecoder provided by the device itself for decoding.
[0062] After the edge intelligent device obtains the video stream based on the above steps, it obtains the corresponding drone image based on the video stream. Specifically, during the video stream acquisition process, the drone's real-time image transmission interface and 4G communication module can ensure that the drone image enters the edge intelligent device in real time for calculation. At the same time, when the above video stream is transmitted, the drone video stream is converted into an OpenCV-encodable video stream and image through the 4G transmission network. In one implementation, the OpenCV video codec uses the IVideoStreamManager stream management class and the OpenCV video codec module to convert the camera data YUV format to RGB data.
[0063] Step S2: performing image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs.
[0064] In this step, the network model is used to perform instance segmentation and object detection on the drone imagery to obtain multiple candidate ROIs. ROI (region of interest) is a region of interest. In this step, the object detection model is used to perform instance segmentation and object detection on the drone imagery to obtain multiple candidate ROIs. Furthermore, to improve data processing efficiency, the image segmentation model is lightweighted in this step. This lightweighted object detection model is then used to perform instance segmentation and typical object detection in the drone imagery.
[0065] In this step, the edge intelligent device performs image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs, including the following steps:
[0066] Step S21: performing lightweight processing on the model accelerator and / or the target detection model to obtain a lightweight accelerator and a lightweight target detection model after lightweight processing;
[0067] Step S22: Use the lightweight target detection model to perform image instance segmentation and target object detection on the drone image, and use the lightweight accelerator to control the image instance segmentation and target object detection process to obtain multiple candidate ROIs.
[0068] In one implementation, the model accelerator is a TensorRT network, and the lightweight target detection model is a lightweight YOLOv8-seg model.
[0069] Specifically, the YOLOv8 network structure consists of three main parts: Backbone, Neck, and Head. The YOLOv8 backbone is based on the CSPDarkNet-53 network. The YOLOv8 neck uses PAN-FPN, an efficient and fast two-stream FPN (multi-scale object detection algorithm). The YOLOv8 head uses the Decoupled Head network structure of YOLOX. In this step, the YOLOv8 model is lightweighted to obtain the lightweight YOLOv8-SEG model.
[0070] Furthermore, in order to speed up the processing efficiency of the image segmentation model for drone images, the model accelerator is used in this step to accelerate the image detection process of the image segmentation model, thereby improving the image processing efficiency of the image segmentation model. Preferably, in this embodiment, the lightweight acceleration of TensorRT is used to improve the image processing efficiency. The lightweight processing of TensorRT is to merge and optimize the branch structure in the TensorRT network, combined with Figure 3 As shown, for example: a convolution layer, a bias layer and a re load layer. These three layers need to call the corresponding API of cuDNN three times, so these three layers can be merged together. TensorRT will merge some networks that can be merged, and the conv, BN, and Re lu layers of the current mainstream neural network are merged into one layer to form CBR.
[0071] In one implementation, the Gst-nvinfer inference plugin runs the yo lov8-seg model to implement the model's object detection functionality. The OSD plugin Gst-nvdsosd annotates images based on the model's output for subsequent display. For targets to be tracked, the tracking plugin Gst-nvtracker enables real-time tracking. The analysis plugin Gst-nvdsanalyt ics can be used to count objects that have crossed the boundary or are within the ROI region, as needed.
[0072] Step S3: Receive interactive control information, perform image segmentation correction on each candidate ROI according to the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image, and segmentation correction parameters corresponding to the target object categories.
[0073] In step S2 above, the target detection model is used to perform image instance segmentation and typical target detection on the drone image. However, due to the low accuracy of segmentation and detection, it cannot meet the requirements. Therefore, in this step, the received interactive control information is used to perform image segmentation correction on the candidate ROI obtained in the above step to obtain more accurate target detection and positioning results.
[0074] Specifically, the interactive control information received in this step is user input or automatically generated by AI, and includes the target object category being detected and the segmentation correction parameters corresponding to the target object category. Specifically, the target object category can be a large category, a small category, or a category containing typical objects, as needed. The segmentation correction parameters corresponding to the target object category are correction parameters corresponding to segmentation accuracy, and the accuracy of target object detection is controlled based on these segmentation correction parameters.
[0075] Furthermore, the step of receiving interactive control information, performing image segmentation correction on each candidate ROI according to the interactive control information, and obtaining target detection and positioning results includes:
[0076] Step S31: receiving interactive control information, and obtaining target object categories and segmentation correction parameters corresponding to each target object category contained in the interactive control information.
[0077] In this step, if the interactive control information is sent by the user, the method for receiving the interactive control information can be based on keyboard input, based on user input voice information, or based on the user forwarding it to the edge intelligent device through a mobile phone or other device. If the interactive control information is generated based on an AI device, it is necessary to use the target detection model to perform target detection on the drone image, identify the candidate ROI and input the result into the AI device. The AI device determines the interactive control information based on the result output by the target detection model and sends the interactive control information to the edge intelligent device.
[0078] Step S32: using an image segmentation algorithm based on the segmentation correction parameters, further segment the target object region in each candidate ROI that belongs to the target object category.
[0079] After the target object categories contained in the interactive control information and the segmentation correction parameters corresponding to the target object categories are determined, an image segmentation algorithm is used to perform further image segmentation on the candidate ROI based on the segmentation correction parameters.
[0080] Furthermore, the step of further segmenting the region where the target object belonging to the target object category is located in each candidate ROI using the image segmentation algorithm based on the segmentation correction parameter includes:
[0081] Step S321 : obtaining the region where the detection target object is located corresponding to the target object category in each candidate ROI, and the detection score corresponding to each detection target object.
[0082] When the target detection model recognizes the drone image and outputs the detection results, it also outputs the detection score corresponding to each target object. Therefore, after obtaining the target object category and corresponding area in each candidate ROI, the detection score corresponding to each detected target object can also be obtained.
[0083] Step S322: Determine the segmentation edge position of the region where the target object is located according to the segmentation correction parameter and the detection score, and perform further image segmentation on the candidate ROI based on the determined segmentation edge position.
[0084] The segmentation correction parameter in this embodiment is the score threshold corresponding to each target object. In the candidate ROI, the target detection result corresponding to each target object corresponds to a detection score, and each detection score represents the accuracy of the detection result. The larger the number of corresponding detection scores, the higher the accuracy of target object detection and positioning. In other words, if the number of corresponding detection scores is smaller, the accuracy of target object detection and positioning is lower.
[0085] After obtaining the detection score corresponding to each target object, it is judged according to the detection score whether it exceeds the score threshold corresponding to the corresponding category. If it is higher, no re-segmentation and optimization is required. If it is lower, the area corresponding to the target object is re-segmented and optimized.
[0086] Specifically, the step of confirming the segmentation edge position of the area where the target object is located according to the segmentation correction parameter and the detection score includes:
[0087] According to the segmentation correction parameters corresponding to the target object category and the detection score corresponding to the detected target object, the target objects to be segmented whose detection scores are lower than the segmentation correction parameters of the corresponding category are screened out; and the segmentation edge position of the area where the target objects to be segmented are located is determined according to the score correction parameters.
[0088] That is, first, the target objects that need to be further segmented and optimized are screened out based on the detection scores corresponding to each target object and its corresponding segmentation correction parameters. It can be imagined that the target objects that need to be further segmented and optimized can be on the candidate ROI or not. When the target object is on the candidate ROI, the area where it is located is further segmented and optimized. If it is not on the candidate ROI, it needs to be filled in to obtain the results of identification and positioning. After further segmentation optimization of the candidate ROI using this step, the effect achieved is as follows: Figure 4 shown.
[0089] Furthermore, in order to obtain more accurate detection and positioning results, after the step of further segmenting the target object region in each candidate ROI using the image segmentation algorithm based on the segmentation correction parameter, the method further includes:
[0090] The morphological algorithm is used to fill and segment the detection target image after image segmentation.
[0091] Morphological algorithms are image morphological operations that include image dilation, erosion, and opening and closing. The main uses of image dilation and erosion are to eliminate noise, segment independent image elements, connect adjacent elements in the image, find significant maxima or minima in the image, and calculate the image gradient. Dilation dilates the highlighted areas of an image, while erosion erodes the highlighted areas. In this step, based on the OpenCV image segmentation and morphological algorithms, the inaccurate edge segmentation problem in instance segmentation prediction is optimized, and the segmentation results of the trained categories are morphologically processed.
[0092] The present embodiment will be further described below in conjunction with specific application examples of the method of the present invention.
[0093] Flowchart of typical urban target recognition and positioning algorithm based on real-time edge computing of drone visible light images, combined with Figure 5 The detailed implementation steps are as follows:
[0094] H1, drone real-time image transmission, based on the OcuSync image transmission system and the edge smart device 4G module to read the YUV data from the drone camera.
[0095] H2, OpenCV video codec, use IVideoStreamManager stream management class and OpenCV video codec module to realize the conversion of YUV format to RGB data.
[0096] H3, video stream transmission and reading, uses the Nvstreammux plug-in to read the transcoded video stream data and push it to the yo lov8-seg inference module.
[0097] H4. Data normalization processing. Normalization technology helps solve the problems of gradient disappearance and gradient explosion, accelerates the convergence of the model, and improves the robustness and generalization ability of the model.
[0098] H5, channel number conversion, batch converts the original data into a tensor channel format consistent with the output, namely cx, cy, w, h, cls, which represent the bounding box information and category information respectively.
[0099] H6. Image data conversion: use the blobFromImage function to convert the image into neural network input data and preprocess the channel data.
[0100] H7. Load the engine model. After reading the yo lov8 model, you need to convert it into a dedicated engine model format that can be recognized by TensorRT.
[0101] Prior to this step, the best model was trained based on the Yo LOV8 instance segmentation model. A drone collected real-time street view images, including cars, trees, roads, buildings, grass, and people, captured from different heights within the same viewpoint. These images were then used with LabE LME to create a dataset for instance segmentation training. Training code was then developed to train this dataset on the Yo LOV8 x-seg model to obtain the best instance segmentation weight model. This training process was performed on the cloud server. After obtaining the best.pt model, it was copied locally for subsequent predictions.
[0102] H8. Model reasoning. In order to improve the unified adaptability and model decoupling of the edge-side algorithm framework, a plug-in approach is used in the overall algorithm framework to reproduce the functions of key modules, and the Nvinfer reasoning plug-in is used for model reasoning.
[0103] To facilitate subsequent morphological processing, in this step, change the "boxes" setting in the prediction settings in the default.yaml configuration file to "False." This ensures that only the post-identification boundary segmentation is performed in the information image after instance segmentation, without displaying the identification boxes. In the yo lov8 project, write a Python prediction script that calls the best model obtained in the above steps and performs preliminary predictions on the real-time image transmission JPG data. This generates an image file after instance segmentation, which serves as the candidate ROI for the next step.
[0104] H9. Data shows that in order to facilitate the display of subsequent processing results on the display terminal, the Nvdsosd plug-in is used to save the parameterized information of the model inference results.
[0105] H10, morphological processing, based on OpenCV image segmentation and morphological algorithms, optimizes the problem of inaccurate edge segmentation in instance segmentation prediction, and performs morphological processing on the segmentation results of trained categories.
[0106] In this step, interactive control information is received, and the interactive control information and the image segmentation algorithm in OpenCV are used to correct and optimize the area to be detected, and the morphological algorithm is used to fill and segment the area. The morphological processing of corrosion and expansion of the detected area can further correct and optimize the segmentation edges in the information picture obtained by the instance segmentation in the previous step. In the process of image segmentation optimization, when the foreground point is located in multiple masks, the background points can be used to filter out masks that are not related to the current task. By using a set of foreground / background points, multiple masks within the area of interest are selected. These masks will be merged into a single mask to completely mark the object of interest. In addition, OpenCV morphological operations can be used to improve the performance of mask merging, especially for large-scale target edge artifacts with better segmentation performance.
[0107] Using the tensorRT plug-in and model acceleration algorithm, the code and model are integrated into the edge intelligent device, which is then mounted on a drone to enable real-time image transmission and image processing tasks on the drone.
[0108] The method disclosed in the present invention is applied to visible light cameras and edge computing intelligent equipment devices carried by drones, and relates to a technical method for real-time image recognition and positioning of typical urban targets in the field of smart cities. First, image target detection and instance segmentation are implemented based on YOLOv8-seg, and the IOU intersection-over-union technology is used to implement the candidates for the pre-selected boxes. Then, based on the candidate box results, the semantic interaction method and the OpenCV image segmentation algorithm are used to implement the foreground and background segmentation, ROI template mask marking and morphological processing of the image. Finally, the recognition, detection and positioning of typical targets in the image are realized in real time on the edge intelligent device based on the lightweight model. The present invention effectively solves the problems of typical target recognition accuracy and real-time computing of edge devices in existing complex urban scenes, and realizes real-time segmentation and positioning detection of typical targets such as urban roads, vehicles, buildings, vegetation, etc. based on the drone edge computing carrier.
[0109] On the basis of the above method, the present invention also discloses a drone, such as Figure 6 As shown, it includes: an edge intelligent device 120 and an image acquisition device 110.
[0110] The image acquisition device 110 is used to capture drone images of urban scenes. The image acquisition device is a visible light camera.
[0111] The edge intelligent device 120 is used to obtain the drone image, perform image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs, and receive interactive control information, perform image segmentation correction on each candidate ROI according to the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image, and segmentation correction parameters corresponding to the target object categories.
[0112] The edge intelligent device obtains drone images from the image acquisition device and detects and identifies target objects in the drone images, thereby realizing instance segmentation and target object detection of the drone images.
[0113] Furthermore, the edge intelligent device includes: an information acquisition module.
[0114] The information acquisition module is used to obtain video image information from the image acquisition device through a communication connection with the image acquisition device, and use a coding and decoding algorithm to convert the format of the video image information into a recognizable data format to obtain a drone image.
[0115] Furthermore, the edge intelligent device also includes: a lightweight module and an image segmentation module.
[0116] The lightweight module is used to perform lightweight processing on the model accelerator and / or the target detection model to obtain a lightweight accelerator and a lightweight target detection model after lightweight processing;
[0117] The image segmentation module is used to perform image instance segmentation and target object detection on the drone image using a lightweight target detection model, and at the same time use the lightweight accelerator to control the process of image instance segmentation and target object detection to obtain multiple candidate ROIs.
[0118] In one embodiment, the model accelerator is a TensorRT network, and the lightweight target detection model is a lightweight YOLOv8-seg model.
[0119] Furthermore, the edge intelligent device also includes: an interactive information acquisition module and an image segmentation optimization module.
[0120] An interactive information acquisition module, configured to receive interactive control information and acquire target object categories and segmentation correction parameters corresponding to each target object category contained in the interactive control information;
[0121] The image segmentation optimization module is used to perform further image segmentation on the target object region belonging to the target object category in each candidate ROI based on the segmentation correction parameter using an image segmentation algorithm.
[0122] Specifically, the step of further segmenting the region where the target object belonging to the target object category is located in each candidate ROI using the image segmentation algorithm based on the segmentation correction parameter includes:
[0123] Obtain the detection target object area corresponding to the target object category in each candidate ROI, as well as the detection score corresponding to each detection target object;
[0124] The segmentation edge position of the area where the target object is located is determined according to the segmentation correction parameters and the detection score, and the candidate ROI is further segmented based on the determined segmentation edge position.
[0125] Furthermore, the step of confirming the segmentation edge position of the area where the target object is located according to the segmentation correction parameter and the detection score includes:
[0126] According to the segmentation correction parameter corresponding to the target object category and the detection score corresponding to the detected target object, the target objects to be segmented whose detection scores are lower than the segmentation correction parameter of the corresponding category are screened out;
[0127] The segmentation edge position of the region where the target object to be segmented is located is determined according to the score correction parameter.
[0128] Furthermore, the edge intelligent device also includes: a morphological algorithm processing module.
[0129] The morphological algorithm processing module is used to fill and segment the detection target image after image segmentation using a morphological algorithm.
[0130] The method and drone of this embodiment use the yo lov8-seg model to implement instance segmentation of the target in the captured image and obtain candidate ROIs. Based on the interactive semantic information of the acquired ROI, the execution of multimodal tasks is realized, and the recognition and detection tasks of any specified target object can be realized. The OpenCV image segmentation algorithm is adopted to optimize the accuracy of the segmentation edge and further improve the recognition and detection tasks of the candidate area. In this method, since the yo lov8-seg model can realize the target detection task and the target recognition task, but the recognition accuracy cannot meet the requirements, the candidate ROI is further subjected to edge morphological processing of the labeled category target to make the instance segmentation edge more accurate.
[0131] Furthermore, in order to realize the real-time image transmission and image processing tasks of edge intelligent devices on drones, tensorRT plug-ins and model acceleration algorithms are used to realize the plug-in and lightweight detection tasks, helping users to easily transplant the deep learning algorithms on the server and improve the efficiency of algorithm reasoning.
[0132] This embodiment also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the target detection and positioning method for urban scenes.
[0133] The method and drone disclosed in this embodiment use edge intelligent devices to achieve target detection and positioning of images. The edge intelligent device first processes the drone image to obtain candidate ROIs, and then uses interactive control information and image segmentation algorithms to achieve further image segmentation of the candidate ROIs, thereby improving the accuracy of target detection. The model used in this embodiment adopts lightweight processing to improve the efficiency of image processing. Therefore, the method of this embodiment has the advantages of high accuracy and real-time performance of target detection and positioning results.
[0134] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0135] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0136] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes additional implementations in which the order shown or discussed may not be followed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by a person skilled in the art to which the embodiments of the present application belong.
[0137] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can read and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting, or otherwise processing in a suitable manner as necessary, and then storing it in a computer memory.
[0138] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0139] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0140] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
Claims
1. A method for target detection and positioning in urban scenes, characterized in that: The method is applied to an edge intelligent device carried by a drone, wherein the drone is provided with an image acquisition device, and comprises: Acquire a drone image of an urban scene from the image acquisition device; Performing image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs; Receive interactive control information, perform image segmentation correction on each candidate ROI according to the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image and segmentation correction parameters corresponding to the target object categories; The steps of receiving interactive control information, performing image segmentation and correction on each candidate ROI according to the interactive control information, and obtaining target detection and positioning results include: receiving interactive control information, and obtaining target object categories and segmentation correction parameters contained in the interactive control information; Obtain the detection target object area corresponding to the target object category in each candidate ROI, as well as the detection score corresponding to each detection target object; Determine the segmentation edge position of the target object area based on the segmentation correction parameters and the detection score, and perform further image segmentation on the candidate ROI based on the determined segmentation edge position; The target detection result corresponding to each target object corresponds to a detection score, and each detection score represents the accuracy of the detection result. The larger the number of corresponding detection scores, the higher the accuracy of target object detection and positioning. After obtaining the detection score corresponding to each target object, it is judged according to the detection score whether it exceeds the score threshold corresponding to the corresponding category. If it is higher, there is no need to re-segment and optimize. If it is lower, the area corresponding to the target object is re-segmented and optimized.
2. The method for detecting and positioning objects in urban scenes according to claim 1, wherein: The step of acquiring the drone image in the urban scene includes: Through the communication connection with the image acquisition device, video image information is obtained from the image acquisition device, and the format of the video image information is converted into a recognizable data format using a coding and decoding algorithm to obtain the drone image.
3. The method for detecting and positioning objects in urban scenes according to claim 1, wherein: The edge intelligent device performs image instance segmentation and target object detection on the drone image to obtain multiple candidate ROIs, including the following steps: Performing lightweight processing on the model accelerator and / or the target detection model to obtain a lightweight accelerator and a lightweight target detection model after lightweight processing; A lightweight target detection model is used to perform image instance segmentation and target object detection on the drone image, and the lightweight accelerator is used to control the process of image instance segmentation and target object detection to obtain multiple candidate ROIs.
4. The method for detecting and positioning objects in urban scenes according to claim 3, wherein: The model accelerator is a TensorRT network, and the lightweight target detection model is a lightweight YOLOv8-seg model.
5. The method for detecting and positioning objects in urban scenes according to claim 1, wherein: The step of determining the segmentation edge position of the area where the target object is located according to the segmentation correction parameter and the detection score includes: According to the segmentation correction parameter corresponding to the target object category and the detection score corresponding to the detected target object, the target objects to be segmented whose detection scores are lower than the segmentation correction parameter of the corresponding category are screened out; The segmentation edge position of the region where the target object to be segmented is located is determined according to the score correction parameter.
6. The method for detecting and locating objects in urban scenes according to claim 5, characterized in that: After the step of further segmenting the target object region in each candidate ROI by using the image segmentation algorithm based on the segmentation correction parameter, the method further includes: The morphological algorithm is used to fill and segment the detection target image after image segmentation.
7. A drone, characterized in that: include: Equipped with edge intelligent devices and image acquisition equipment The image acquisition device is used to capture drone images in urban scenes; The edge intelligent device is configured to acquire a drone image of an urban scene from the image acquisition device, perform image instance segmentation and target object detection on the drone image to obtain a plurality of candidate ROIs, receive interactive control information, perform image segmentation correction on each candidate ROI based on the interactive control information, and obtain target detection and positioning results; the interactive control information includes one or more target object categories contained in the drone image and segmentation correction parameters corresponding to the target object categories; The steps of receiving interactive control information, performing image segmentation and correction on each candidate ROI according to the interactive control information, and obtaining target detection and positioning results include: receiving interactive control information, and obtaining target object categories and segmentation correction parameters contained in the interactive control information; Obtain the detection target object area corresponding to the target object category in each candidate ROI, as well as the detection score corresponding to each detection target object; Determine the segmentation edge position of the target object area based on the segmentation correction parameters and the detection score, and perform further image segmentation on the candidate ROI based on the determined segmentation edge position; The target detection result corresponding to each target object corresponds to a detection score, and each detection score represents the accuracy of the detection result. The larger the number of corresponding detection scores, the higher the accuracy of target object detection and positioning. After obtaining the detection score corresponding to each target object, it is judged according to the detection score whether it exceeds the score threshold corresponding to the corresponding category. If it is higher, there is no need to re-segment and optimize. If it is lower, the area corresponding to the target object is re-segmented and optimized.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the target detection and positioning method for urban scenes as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image segmentation method and device, electronic equipment and storage medium
CN114037710A
Unmanned aerial vehicle front-end defect identification method and system based on lightweight edge calculation
CN117274843A