Method, device and system for intelligently detecting target based on YOLO model

By integrating the YOLO model with an attention module and image segmentation technology on the server side, the problem of identifying small and edge targets in low-resolution images by UAV target detection systems has been solved, and the accurate extraction of target structure information in high-resolution images has been achieved, improving the efficiency and accuracy of UAV inspection tasks.

CN120877167APending Publication Date: 2025-10-31CHENGDU JOUAV AUTOMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510469083.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing UAV target detection systems struggle to identify small and edge targets in low-resolution video data and cannot accurately extract target structure information from high-resolution images, thus failing to meet the needs of refined inspection tasks.

Method used

The YOLO model, which incorporates spatial attention and/or channel attention modules, is integrated on the server side to detect images acquired by embedded devices. By combining image segmentation and detection box optimization techniques, it identifies small and edge targets in low-resolution images and extracts structural information from high-resolution images.

Benefits of technology

It improves the efficiency and accuracy of identifying small and edge targets in low-resolution images, balancing the efficiency and accuracy requirements of target detection, and is suitable for inspection scenarios with high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877167A_ABST
    Figure CN120877167A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and particularly discloses a method, a device and a system for intelligently detecting a target based on a YOLO model. The YOLO model integrating a backbone network and / or a multi-scale fusion network with an attention module is integrated on a server; the method is used for carrying out target detection operation on an image which identifies a new target from embedded equipment for the first time, not only can small targets and edge targets in a low-resolution image be accurately identified, but also structure information of targets in a high-resolution image can be accurately extracted, and the multi-target identification efficiency and accuracy are improved; the image data processing efficiency is improved, the requirements for the target detection efficiency and precision are balanced, and the method is particularly suitable for the inspection scene of a high-resolution image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method, device and system for intelligent target detection based on the YOLO model. Background Technology

[0002] With the rapid development of drone technology, the demand for its application in border monitoring, power line inspection, forest fire prevention, and other fields is constantly increasing. However, in these practical scenarios, the computing resources and transmission capabilities of drones limit the efficiency and accuracy of target detection and identification systems.

[0003] Currently, drones typically carry embedded devices and electro-optical pods to perform real-time target detection tasks. However, these embedded devices have limited computing power, making it difficult to run computationally complex target detection models, such as the DETR model based on the Transformer detection algorithm. Furthermore, to ensure real-time performance, the onboard unit usually chooses to run lightweight models, such as the YOLO model, for target detection on video data.

[0004] However, in practice, it has been found that existing YOLO models typically cannot identify small or edge targets in low-resolution video data, nor can they accurately extract structural information of targets in high-resolution images, making it difficult to meet the needs of refined inspection tasks. Therefore, there is an urgent need to propose a technical solution that can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images. Summary of the Invention

[0005] This invention provides a method, device, and system for intelligent target detection based on the YOLO model, which can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images.

[0006] To address the aforementioned technical problems, the first aspect of this invention discloses a method for intelligent target detection based on the YOLO model. The method is applied to a detection system, which includes a server. The method includes: The server obtains a first image sent by the embedded device with which it communicates, wherein the embedded device is integrated into the flight equipment, the first image is an image collected by the embedded device for the current scene during the flight of the flight equipment, and the first image contains the target first identified by the embedded device; The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image; The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. The server identifies information about at least one first target in the first image based on the detection results of the first image. All first targets include all targets identified for the first time, and the information of each first target includes the category of the first target.

[0007] As an optional implementation, in the first aspect of the present invention, before the server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image, the method further includes: The server performs an image segmentation operation on the acquired first image to obtain multiple block images; The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image, including: The server performs a detection operation on each block image based on the YOLO model to obtain the detection result of each block image; The server identifies information about at least one first target in the first image based on the detection results of the first image, including: The server identifies information about at least one first target in the first image based on the detection results of all the block images.

[0008] As an optional implementation, in the first aspect of the present invention, the detection result of each block image includes multiple detection boxes, the confidence level of each detection box, and the category; The server identifies information about at least one first target in the first image based on the detection results of all the block images, including: The server performs a stitching operation on the detection results of each block image to obtain the stitched first detection result; Based on the coordinates of each detection box in the first detection result, the server determines a group of detection boxes with an overlap greater than or equal to a preset overlap from all the detection boxes in the first detection result, and each group of detection boxes consists of at least two detection boxes; The server deletes all detection frames in each detection frame group except the one with the highest confidence, based on the confidence level of each detection frame in that detection frame group. After performing a detection box deletion operation on all the detection box groups, the server performs a stitching operation on the adjacent boxes of the first detection result to obtain a second detection result, and identifies information about at least one first target in the first image based on the second detection result.

[0009] As an optional implementation, in a first aspect of the present invention, the method includes: The server determines the number of blocks that the first image needs to be divided into, and obtains the image size of the first image and the size of all the first targets in the first image; The server determines the first step length in the horizontal direction and the second step length in the vertical direction based on the size of all the first targets in the first image. The server determines the block overlap during the first image segmentation process based on the first step length, the second step length, the image size of the first image, and the number of blocks. The server performs image segmentation on the received first image to obtain multiple block images, including: The server performs image segmentation on the received first image based on the segmentation overlap to obtain multiple block images.

[0010] As an optional implementation, in the first aspect of the present invention, the detection system further includes the embedded device; The method further includes: During the flight of the flight equipment, the embedded device performs an image acquisition operation on the current scene to obtain a video stream of the current scene; The embedded device performs object detection on the current frame image of the video stream of the current scene based on the YOLO model to obtain the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score of each detection box, and the category. When the current frame image is the first frame image of the video stream, the embedded device assigns a corresponding identifier to each target in the current frame image according to the detection result of the current frame image, and sends the current frame image to the server. The first image includes the first frame image of the video stream. When the current frame image is not the first frame image of the video stream, the embedded device obtains the detection result of the previous frame image, and performs a tracking operation on the target based on the detection result of the current frame image and the detection result of the previous frame image to obtain the target tracking result of the current frame image; The embedded device determines whether a new target appears in the current frame image based on the target tracking result of the current frame image. When it is determined that a new target appears in the current frame image, it performs an image acquisition operation at the target resolution to obtain a second image and sends the second image to the server. The first image includes the second image, wherein the target resolution is greater than the resolution of the video stream.

[0011] As an optional implementation, in the first aspect of the invention, the target performs a tracking operation to obtain the target tracking result of the current frame image, including... The embedded device analyzes the current information of each second target in the current frame image based on the detection result of the current frame image, and the current information of each second target includes the current position of the second target; The embedded device analyzes the historical position of each second target in the previous frame image based on the detection result of the previous frame image, and predicts the predicted position of each second target at the next moment based on the current position of each second target and the historical position of the second target. The embedded device performs a tracking operation on each second target based on the predicted position of the second target, and obtains the target tracking result of the current frame image; The embedded device determines whether a new target appears in the current frame image based on the target tracking result of the current frame image, including: When an identifier assignment operation has been performed on a target appearing in a historical frame image, the embedded device determines whether the target tracking result of the current frame image is used to indicate that there is a target in the current frame image that has not been assigned an identifier. If the result is yes, it determines that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier; or, The embedded device calculates the cross value between the current position of each second target and the predicted position of the second target based on the target tracking result of the current frame image, and determines whether the cross value corresponding to each second target is greater than or equal to a preset cross value. When it is determined that the cross value is greater than or equal to the preset cross value, it determines that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier.

[0012] A second aspect of this invention discloses a system for intelligent target detection based on the YOLO model, the system comprising a server, the server comprising: The first acquisition module is used to acquire a first image sent by an embedded device that is connected to it in communication, wherein the embedded device is integrated into the flight equipment, the first image is an image collected by the embedded device for the current scene during the flight of the flight equipment, and the first image contains a target that the embedded device first identifies; The first detection module is used to perform a detection operation on the first image based on the YOLO model to obtain the detection result of the first image; The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. The recognition module is used to recognize information of at least one first target in the first image based on the detection results of the first image, wherein all first targets include all targets identified for the first time, and the information of each first target includes the category of the first target.

[0013] As an optional implementation, in a second aspect of the present invention, the server further includes: The segmentation module is used to perform image segmentation on the acquired first image before the first detection module performs detection operation on the first image based on the YOLO model and obtains the detection result of the first image, thereby obtaining multiple block images; The specific method by which the first detection module performs detection operations on the first image based on the YOLO model to obtain the detection result of the first image includes: Based on the YOLO model, a detection operation is performed on each of the block images to obtain the detection result of each block image; The specific method by which the recognition module identifies information about at least one first target in the first image based on the detection results of the first image includes: Based on the detection results of all the block images, information on at least one first target in the first image is identified.

[0014] As an optional implementation, in a second aspect of the invention, the detection result of each block image includes multiple detection boxes, the confidence level of each detection box, and the category. The specific method by which the recognition module identifies information about at least one first target in the first image based on the detection results of all the block images includes: For the detection results of each of the image blocks, a stitching operation is performed to obtain the stitched first detection result; Based on the coordinates of each detection box in the first detection result, determine all detection box groups with an overlap greater than or equal to a preset overlap from all the detection boxes in the first detection result, wherein each detection box group consists of at least two detection boxes; Based on the confidence level of each detection box in each detection box group, all detection boxes in that detection box group except for the detection box with the highest confidence level are deleted. After performing a detection box deletion operation on all the detection box groups, a stitching operation is performed on the adjacent boxes of the first detection result to obtain a second detection result, and based on the second detection result, information of at least one first target in the first image is identified.

[0015] As an optional implementation, in a second aspect of the present invention, the server further includes: The determining module is used to determine the number of blocks that the first image needs to be divided into; The first acquisition module is further configured to acquire the image size of the first image and the size of all the first targets in the first image; The determining module is further configured to determine the first step length in the horizontal direction and the second step length in the vertical direction based on the size of all the first targets in the first image; The determining module is further configured to determine the block overlap during the first image segmentation process based on the first step length, the second step length, the image size of the first image, and the number of blocks. The specific method by which the segmentation module performs image segmentation on the received first image to obtain multiple block images includes: Based on the block overlap, the received first image is divided into multiple block images.

[0016] As an optional implementation, in a second aspect of the invention, the system further includes the embedded device, wherein the embedded device includes: The acquisition module is used to perform image acquisition operations on the current scene during the flight of the flight equipment to obtain a video stream of the current scene; The second detection module is used to perform target detection operation on the current frame image of the video stream of the current scene based on the YOLO model, and obtain the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score and category of each detection box. The setting module is used to assign a corresponding identifier to each target in the current frame image based on the detection result of the current frame image when the current frame image is the first frame image of the video stream; A sending module is used to send the current frame image to the server, wherein the first image includes the first frame image of the video stream; The second acquisition module is used to acquire the detection result of the previous frame image when the current frame image is a non-first frame image of the video stream; The tracking module is used to perform a tracking operation on the target based on the detection results of the current frame image and the detection results of the previous frame image, so as to obtain the target tracking result of the current frame image; The judgment module is used to determine whether a new target appears in the current frame image based on the target tracking result of the current frame image; The acquisition module is further configured to perform an image acquisition operation at a target resolution when it is determined that a new target appears in the current frame image, to obtain a second image, wherein the first image includes the second image, and the target resolution is greater than the resolution of the video stream; The sending module is also used to send the second image to the server.

[0017] As an optional implementation, in a second aspect of the invention, the tracking module performs a target tracking operation based on the detection results of the current frame image and the detection results of the previous frame image to obtain the target tracking result of the current frame image. The specific method for this is as follows: Based on the detection results of the current frame image, analyze the current information of each second target in the current frame image, where the current information of each second target includes the current position of the second target; Based on the detection results of the previous frame image, analyze the historical position of each second target in the previous frame image, and predict the predicted position of each second target at the next moment based on the current position and the historical position of each second target. Based on the predicted position of each second target, a tracking operation is performed on the second target to obtain the target tracking result of the current frame image; The specific method by which the judgment module determines whether a new target appears in the current frame image based on the target tracking result of the current frame image includes: When an identifier assignment operation has been performed on a target appearing in a historical frame image, it is determined whether the target tracking result of the current frame image is used to indicate that there is a target in the current frame image that has not been assigned an identifier. If the result is yes, it is determined that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier; or, Based on the target tracking results of the current frame image, calculate the cross value between the current position of each second target and the predicted position of the second target, and determine whether the cross value corresponding to each second target is greater than or equal to a preset cross value. When it is determined that the cross value is greater than or equal to the preset cross value, it is determined that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier.

[0018] A third aspect of the present invention discloses a server, the server comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps performed by the server in any of the methods for intelligent target detection based on the YOLO model disclosed in the first aspect of the present invention.

[0019] A fourth aspect of the present invention discloses an embedded device, the embedded device comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps performed by the embedded device in any of the methods for intelligent target detection based on the YOLO model disclosed in the first aspect of the present invention.

[0020] The fifth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps performed by the server in any of the methods for intelligent target detection based on the YOLO model disclosed in the first aspect of the present invention.

[0021] The sixth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps performed by the embedded device in any of the methods for intelligent target detection based on the YOLO model disclosed in the first aspect of the present invention.

[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: In this embodiment of the invention, the server acquires a first image sent by an embedded device connected to it. The embedded device is integrated into the flight equipment. The first image is an image captured by the embedded device for the current scene during the flight of the flight equipment, and the first image contains a target that the embedded device first identifies. The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image. Based on the detection result of the first image, the server identifies information of at least one first target in the first image. All first targets include all targets identified for the first time, and the information of each first target includes the category of the first target. The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. As can be seen, this invention integrates a YOLO model with a backbone network and / or multi-scale fusion network incorporated into the attention module on the server side. Using this model, target detection is performed on images where new targets are first identified from embedded devices. This not only accurately identifies small and edge targets in low-resolution images but also accurately extracts structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition. Furthermore, the collaborative processing of first performing target detection and capture through embedded devices and then performing refined recognition through the server side improves image data processing efficiency and balances the requirements of target detection efficiency and accuracy, making it particularly suitable for inspection scenarios involving high-resolution images. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating a method for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention. Figure 2 This is a flowchart illustrating another method for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a system for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of another system for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a server disclosed in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an embedded device disclosed in an embodiment of the present invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.

[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0028] This invention discloses a method, device, and system for intelligent target detection based on the YOLO model. By integrating a YOLO model with a backbone network and / or multi-scale fusion network into an attention module on a server, it performs target detection on images where new targets are first identified from an embedded device. This accurately identifies small and edge targets in low-resolution images and accurately extracts structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition. Furthermore, the collaborative processing of first performing target detection and capture on the embedded device and then performing refined recognition on the server improves image data processing efficiency, balancing the requirements of target detection efficiency and accuracy, making it particularly suitable for high-resolution image inspection scenarios. The following sections provide detailed explanations.

[0029] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention. Figure 1 The described method can be applied to any scenario requiring target detection, including but not limited to border patrol, power line inspection, and forest fire prevention. This scenario is equipped with a corresponding detection system, which includes a server and, more specifically, an embedded device that communicates with the server. This embedded device is integrated into a flight device that is performing a specific flight mission within the corresponding scenario. Figure 1 As shown, the method may include the following operations: 101. The server obtains the first image sent by the embedded device with which it communicates. The first image is an image collected by the embedded device for the current scene during the flight of the flight equipment, and the first image contains the target that the embedded device first recognizes.

[0030] In this embodiment of the invention, optionally, the embedded device includes a camera module, an optoelectronic pod, and an embedded computing unit. The camera module is integrated onto the optoelectronic pod, and the images acquired by the camera module are transmitted to the embedded computing unit. Further, after obtaining the first image, the first image is stored in a storage module and uploaded to the server via a transmission module. The optoelectronic pod supports zoom, focus, and real-time video stream capture at different resolutions such as 1080p. The camera module supports the acquisition of 4K, 6K, or higher resolution images.

[0031] 102. The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image. The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network based on the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network.

[0032] 103. Based on the detection results of the first image, the server identifies information about at least one first target in the first image. All first targets include all targets identified for the first time, and the information of each first target includes the category of the first target.

[0033] In this embodiment of the invention, the information for each first target further includes at least one of color, size, and position in the current scene.

[0034] It is evident that implementation Figure 1The described method integrates a YOLO model with a backbone network and / or a multi-scale fusion network incorporated into the attention module on the server. It is used to perform target detection on images where new targets are first identified from embedded devices. This method can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images, thus improving the efficiency and accuracy of multi-target recognition. Furthermore, the collaborative processing of first performing target detection and capture on the embedded device and then performing fine recognition on the server improves the efficiency of image data processing and balances the requirements of target detection efficiency and accuracy. It is particularly suitable for inspection scenarios with high-resolution images.

[0035] In an optional embodiment, before the server performs a detection operation on the first image based on the YOLO model and obtains the detection result of the first image, the method may further include the following steps: The server performs image segmentation on the first image obtained, resulting in multiple block images; The server-side performs detection operations on the first image based on the YOLO model, obtaining the detection results of the first image, including: The server uses the YOLO model to perform detection operations on each block of the image and obtains the detection results for each block of the image. The server identifies information about at least one first target in the first image based on the detection results of the first image, including: Based on the detection results of all block images, the server identifies information about at least one first target in the first image.

[0036] As can be seen, this optional embodiment improves the efficiency and accuracy of target detection and recognition by first dividing the image obtained from the embedded device into blocks, and then sequentially applying the above-mentioned YOLO model to the block images.

[0037] In this optional embodiment, as an optional implementation, the detection result of each block image includes multiple detection boxes, the confidence level of each detection box, and the category. The server identifies information about at least one first target in the first image based on the detection results of all block images, including: The server performs a stitching operation on the detection results of each image block to obtain the first stitched detection result; Based on the coordinates of each detection box in the first detection result, the server determines all groups of detection boxes with an overlap greater than or equal to a preset overlap from all the detection boxes in the first detection result. Each group of detection boxes consists of at least two detection boxes. The server deletes all detection boxes in each detection box group except the one with the highest confidence, based on the confidence level of each detection box in that group. After performing a detection box deletion operation on all detection box groups, the server performs a stitching operation on the adjacent boxes of the first detection result to obtain a second detection result, and identifies information about at least one first target in the first image based on the second detection result.

[0038] In this optional embodiment, the preset overlap degree can be determined based on the type of the current scene, that is, different scenes correspond to different preset overlap degrees.

[0039] In this optional embodiment, the detection box can be removed in any way, such as by non-maximum suppression.

[0040] As can be seen, this optional embodiment stitches together the target detection results of the block image by optimizing the stitching process to ensure the accuracy and integrity of the detection boxes. After stitching, the detection boxes are grouped according to their overlap. For each group of detection boxes, the detection box with the highest confidence is retained, and the other detection boxes are deleted. Then, adjacent detection boxes are stitched together, which improves the detection accuracy and data volume after stitching, thereby improving the efficiency and accuracy of target detection.

[0041] In another alternative embodiment, the method may further include the following steps: The server determines the number of blocks that the first image needs to be divided into, and obtains the image size of the first image and the size of all first targets in the first image; The server determines the first step length in the horizontal direction and the second step length in the vertical direction based on the size of all first targets in the first image. The server determines the block overlap during the first image segmentation process based on the first step length, the second step length, the image size of the first image, and the number of blocks. The server performs image segmentation on the received first image to obtain multiple block images, including: The server performs image segmentation on the received first image based on the segmentation overlap to obtain multiple block images.

[0042] In this optional embodiment, the number of blocks may be determined by at least one of the resolution of the first image, the image size, and the type of the current scene. The lower the resolution, the larger the size, and the higher the requirement for target detection accuracy, the more blocks are required.

[0043] Optionally, the formula for calculating the number of blocks is as follows:

[0044] In the formula, Block i,j Indicates the number of blocks in the first image; step xand step y These represent the horizontal and vertical step lengths, respectively, overlapping in width and height. x j•step y `i` and `j` represent the coordinates of the top-left corner of the current block image on the original image, respectively; `width` and `height` represent the size of the current block image, respectively. The formula for calculating the number of blocks is as follows: Starting from the first image `Image`, a rectangular region with width `width` and height `height` is extracted as the (i, j)th block image, beginning from the starting position (i•stepx, j•stepy). The overlap between adjacent block images is controlled by the horizontal and vertical step lengths and the block image size (width and height).

[0045] In this optional embodiment, the block overlap further includes horizontal overlap and vertical overlap. The formula for calculating the horizontal overlap is as follows:

[0046] The vertical overlap is calculated as follows:

[0047] In the formula, The horizontal and vertical overlap values ​​represent the horizontal and vertical overlap values, respectively, while W and H represent the width and height of the first image, respectively.

[0048] In this optional embodiment, both the vertical and horizontal overlap have corresponding overlap ranges, such as 10%-30%, to reduce the possibility of the target being cropped or missed during the cropping process. Optionally, the number of blocks required to divide the first image can be dynamically adjusted by the resolution of the first image, wherein the higher the resolution after segmentation, the smaller the block size.

[0049] As can be seen, the embodiments of the present invention determine the block overlap by combining multiple factors such as the number of blocks required for image segmentation, image size, target size in the image, and step length in the horizontal and vertical directions. This improves the accuracy and reliability of block overlap determination, increases the probability that targets in the image are assigned to the same image block, thereby improving the image segmentation accuracy and efficiency, and further improving the target detection efficiency and accuracy.

[0050] In this optional implementation, the server determines the number of blocks that the first image needs to be divided into, including: The server obtains the attribute parameters of the first image, wherein the attribute parameters of the first image include image resolution and / or image size; The server retrieves the attribute parameters of the integrated YOLO model. These parameters include the model type and the pruning ratio. The model type includes, but is not limited to, YOLO3, YOLO4, YOLO5, YOLO6, and YOLO8. The pruning ratio represents the proportion of pruned filters in the YOLO model to the original filters. It should be noted that other related descriptions of the pruning ratio are provided later and will not be repeated here.

[0051] The server determines the number of blocks needed to divide the first image based on the attribute parameters of the first image and the attribute parameters of the YOLO model. The larger the image size, the higher the image resolution, and the greater the pruning rate of the YOLO model, the more blocks are required.

[0052] As can be seen, this optional embodiment can also analyze the number of image blocks to be divided by the image resolution, size, type of model used, and pruning rate, thereby improving the accuracy of the number of image blocks and thus helping to further improve the accuracy of the division of the overlap between image blocks.

[0053] In this optional embodiment, the attribute parameters of the first image may further include the estimated size and complexity of each new sub-target in the image, such as texture complexity, shape complexity, etc.; the method may further include the following steps: The server obtains the recognition level of each new sub-target to be identified in the first image. The recognition level of each new sub-target includes the recognition type, recognition size, recognition color and / or recognition texture. The more content to be identified, the higher the required recognition level, and the more blocks are required. The server calculates the target count of all new sub-targets and the image ratio of the target count to the image size; The server updates the previously obtained block count based on the image ratio, the estimated size, complexity, and recognition accuracy of each new sub-target, resulting in an updated block count. This updated block count is used for overlap analysis. A smaller image ratio, smaller estimated size, higher complexity, and higher recognition accuracy result in a relatively larger number of blocks.

[0054] As can be seen, this optional embodiment can also correct the number of blocks obtained above based on the number, size and complexity of new sub-targets in the image, so as to further improve the accuracy and reliability of determining the number of blocks in the image, further improve the accuracy of block overlap analysis, further increase the possibility that new sub-targets are divided into a block image as much as possible, and further improve the target recognition efficiency and accuracy.

[0055] It should be noted that in all the aforementioned block divisions, it is necessary to ensure that the new sub-targets are kept in the same block image as much as possible.

[0056] In yet another alternative implementation, the method may further include the following steps: During flight, the embedded device performs image acquisition operations on the current scene to obtain a video stream of the current scene; The embedded device performs object detection on the current frame image of the video stream of the current scene based on the YOLO model, and obtains the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score of each detection box and the category. When the current frame image is the first frame image of the video stream, the embedded device assigns a corresponding identifier to each target in the current frame image based on the detection results of the current frame image, and sends the current frame image to the server. The first image includes the first frame image of the video stream. When the current frame image is not the first frame image of the video stream, the embedded device obtains the detection result of the previous frame image, and performs a tracking operation on the target based on the detection result of the current frame image and the detection result of the previous frame image to obtain the target tracking result of the current frame image; The embedded device determines whether a new target appears in the current frame image based on the target tracking result. When a new target appears in the current frame image, it performs an image acquisition operation at the target resolution to obtain a second image and sends the second image to the server. The first image includes the second image, wherein the target resolution is greater than the resolution of the video stream.

[0057] In this optional embodiment, optionally, when an image frame is acquired, the operation of this optional embodiment is performed on that image frame. The YOLO model here and the YOLO model integrated on the server can be the same YOLO model, or a more lightweight YOLO model. Regardless of the case, the target network of the YOLO model here incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network.

[0058] In this optional embodiment, the first frame image can be sent directly to the server, or it can be re-acquired at the target resolution and then sent to the server.

[0059] In this optional embodiment, if it is determined that no new target appears in the current frame image, the image continues to be acquired and the image is detected based on the YOLO model.

[0060] In this optional embodiment, when a new target is detected, the embedded device simultaneously sets a corresponding identifier for the new target so as to track the new target based on the corresponding identifier.

[0061] In this optional embodiment, the real-time detection frame rate used by the embedded device for image acquisition and detection may be ≥ a preset frame rate, such as 30 FPS.

[0062] As can be seen, this optional embodiment integrates the YOLO model, which incorporates the backbone network and / or multi-scale fusion network into the attention module, into the embedded device, thereby improving the computing power of the embedded device. Using this model to perform target detection on the acquired images, it can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition and reducing the rate of missed target detection. Furthermore, by first performing target detection and capture at a lower resolution on the embedded device, and then triggering higher-resolution image acquisition and sending to the server only when a new target is detected, it reduces the frequent backhaul of high-resolution images and the repeated backhaul of targets to the server, thus saving bandwidth and computing resources. Finally, the collaborative processing of the server for refined identification of images containing new targets improves image data processing efficiency and balances the requirements of target detection efficiency and accuracy, making it particularly suitable for high-resolution image inspection scenarios.

[0063] In this optional embodiment, as an optional implementation, the embedded device performs a target tracking operation based on the detection results of the current frame image and the detection results of the previous frame image to obtain the target tracking result of the current frame image, including: The embedded device analyzes the current information of each second target in the current frame image based on the detection results of the current frame image. The current information of each second target includes the current position of the second target. The embedded device analyzes the historical position of each of the second targets in the previous frame image based on the detection results of the previous frame image, and predicts the predicted position of each second target at the next moment based on the current position of each second target and the historical position of the second target. The embedded device performs a tracking operation on each second target based on the predicted position of that second target, and obtains the target tracking result of the current frame image.

[0064] In this optional embodiment, the position of each second target, such as its current position or historical position, can be obtained by analyzing the coordinates of the detection box of each second target. Further optionally, the predicted position of each second target at the next moment can be predicted using any method capable of position prediction, such as a Kalman filter; and the tracking operation for each second target can be performed using any method capable of target tracking, such as the Hungarian algorithm.

[0065] As can be seen, this optional embodiment predicts the target's position at the next moment by using the currently detected target's position and its historical position, which improves the accuracy and efficiency of target tracking. This is conducive to improving the accuracy and efficiency of discovering and distinguishing new targets, and further conducive to improving the accuracy and efficiency of multi-target identification.

[0066] In this optional embodiment, as an optional implementation, the embedded device determines whether a new target appears in the current frame image based on the target tracking result of the current frame image, including: When an identifier assignment operation has been performed on a target that appeared in a historical frame image, the embedded device determines whether the target tracking result of the current frame image is used to indicate that there is a target in the current frame image that has not been assigned an identifier. If the result is yes, it is determined that a new target has appeared in the current frame image, and the new target is a target that has not been assigned an identifier; or, The embedded device calculates the cross value between the current position of each second target and the predicted position of the second target based on the target tracking results of the current frame image, and determines whether the cross value corresponding to each second target is greater than or equal to a preset cross value. When it is determined that the cross value is greater than or equal to the preset cross value, it determines that a new target has appeared in the current frame image, and the new target is a target that has not been assigned an identifier.

[0067] In this optional embodiment, it is possible to determine that no new target has appeared in the current frame image when it is determined that all targets in the current frame image have been assigned corresponding identifiers, and / or when the cross value between the current position and the predicted position of each second target is small.

[0068] As can be seen, this optional embodiment improves the accuracy and efficiency of new target identification by identifying the target in the image or the intersection between the current position and its predicted position. This is beneficial to further improving the accuracy and efficiency of triggering high-resolution image acquisition, and thus further improving the efficiency and accuracy of accurately identifying new targets.

[0069] In yet another alternative embodiment, the YOLO model described above can be optimized in the following way: The attention module is integrated into the target network of the YOLO model that needs to be optimized. The target network of the YOLO model includes the backbone network of the YOLO model and / or the multi-scale fusion network of the YOLO model. The attention module includes a channel attention module and / or a spatial attention module. Based on the attention module, average pooling and max pooling operations matching the attention module are performed on each channel of the target feature map output from the target network to obtain average pooling and max pooling results matching the attention module. Based on the average pooling and max pooling results matched with the attention module and the target feature map corresponding to the target network, an attention feature map of the target feature map is generated; wherein, the attention feature map of the target feature map is used as the input feature map of the next layer network of the attention module.

[0070] In this optional embodiment, the attention module can be incorporated into the target network of the YOLO model, which can be understood as being incorporated into multiple network layers within the target network and / or into the last network layer of the target network.

[0071] In this optional embodiment, the next layer of the attention module can be either the last layer of the target network of the YOLO model or an intermediate layer of the target network.

[0072] As can be seen, this optional embodiment optimizes the YOLO model by integrating the attention module into the backbone network and / or multi-scale fusion network of the YOLO model. This results in a YOLO model that emphasizes important feature regions in the image, enabling the YOLO model to accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images. This improves the accuracy and efficiency of target recognition for various sizes in images, meeting the needs of refined inspection tasks. Furthermore, the optimized YOLO model not only has high accuracy but is also lightweight, allowing for flexible deployment on embedded devices and servers to meet target recognition needs in various scenarios.

[0073] In this optional embodiment, when the attention module includes a channel attention module, based on the attention module, average pooling and max pooling operations matching the attention module are performed on each channel of the target feature map output from the target network to obtain average pooling and max pooling results matching the attention module, including: Based on the channel attention module, average pooling and max pooling operations are performed on each channel of the target feature map output from the target network to obtain the channel average pooling result and channel max pooling result of each channel of the target feature map. Specifically, based on the average pooling and max pooling results matched with the attention module and the target feature map corresponding to the target network, an attention feature map of the target feature map is generated, including: For any channel of the target feature map, the target weight of the channel is determined based on the channel average pooling result and the channel max pooling result. Based on the target weights of each channel of the target feature map and the target feature map output by the target network, a channel attention feature map of the target feature map is generated. Specifically, for any channel of the target feature map, the target weight of the channel is determined based on the channel average pooling result and the channel max pooling result, including: For any channel of the target feature map, the channel average pooling result and the channel average pooling result for that channel are calculated. The max pooling result is summed to obtain the channel and pooling result for that channel. The channel and pooling result of the channel are input into a predetermined first fully connected layer, and a nonlinear transformation operation is performed on the channel and pooling result of the channel input into the first fully connected layer based on a predetermined first activation function to obtain the transformation result of the channel. The transformation result of this channel is input into a pre-determined second fully connected layer, and a linear transformation operation is performed on the transformation result of this channel input into the second fully connected layer based on a pre-determined second activation function to obtain the target weight of this channel.

[0074] In this optional embodiment, the target feature map F has a size of C×H×W, where C, H, and W are the number of channels, height, and width of the target feature map F, respectively.

[0075] For each channel of the target feature map F, the formulas for calculating the channel average pooling result and the channel max pooling result are as follows:

[0076] In the formula, G avg1 (F), G max1 F(c,i,j) represents the average pooling result and the max pooling result of any channel of the target feature map F, respectively. F(c,i,j) represents the coordinates of the i-th and j-th pixels of the target feature map F in that channel, i.e., its spatial location. Both the max pooling result and the average pooling result of that channel are C×1×1 vectors.

[0077] For any channel of the target feature map F, the transformation result of that channel and the calculation formula for the target weight are as follows:

[0078] In the formula, X0 is the transformation result of any channel of the target feature map F; M c (F) represents the target weight of this channel; W0 and W1 are the corresponding first fully connected layer and second fully connected layer, respectively; ReLU(), These represent the first and second activation functions, respectively. Optionally, ReLU() can also use the Leaky ReLU activation function, or the Mish or SiLU (Swish) activation function, to further improve the non-linear expressiveness and training stability of the YOLO model. Among these, the Mish activation function is smooth and continuous, has better gradient propagation properties, helps reduce gradient vanishing, and can better capture complex features, making it particularly suitable for complex background scenarios in UAV target detection.

[0079] In this optional embodiment, the output value of the target weight for each channel is in the range of [0,1]. The larger the output value, the more important the feature representation of the corresponding channel is in the target feature map F, that is, the greater its contribution to the task.

[0080] In this optional embodiment, after obtaining the target weights of each channel of the target feature map F, a vector of the target weights of the channels is generated, such as a vector of length C.

[0081] The formula for calculating the channel attention feature map of the target feature map is as follows:

[0082] In the formula, F out1 is the channel attention feature map of the target feature map; ⊙ represents element-wise multiplication, that is, the target weight of each channel is multiplied element-wise with the corresponding pixel coordinate (spatial position) in the target feature map F.

[0083] As can be seen, this optional embodiment can also extract important channel information by performing global average pooling and global max pooling on each channel of the feature map, and generate corresponding channel weights through a shared fully connected layer and activation function. That is, the feature map is weighted from the channel dimension. By assigning different weights to different channels, the importance of the feature map representation is enhanced. In this way, the YOLO model emphasizes the channels in the input feature map that make important contributions to the target task, while suppressing irrelevant or unimportant channels, thereby improving the expression of important features and improving the accuracy and reliability of the YOLO model in recognizing targets of various sizes.

[0084] This optional embodiment, as an optional implementation, when the attention module includes a spatial attention module, performs average pooling and max pooling operations matching the attention module on each channel of the target feature map output from the target network, based on the attention module, to obtain average pooling and max pooling results matching the attention module, including: Based on the spatial attention module, channel average pooling and channel max pooling operations are performed on the spatial location of each channel of the target feature map output from the target network to obtain the spatial average pooling result and the spatial max pooling result of the target feature map. Specifically, based on the average pooling and max pooling results matched with the attention module and the target feature map corresponding to the target network, an attention feature map of the target feature map is generated, including: Perform a concatenation operation on the spatial average pooling and spatial max pooling results of the target feature map to obtain the spatial concatenated feature result of the target feature map; Based on the splicing feature results of the target feature map and the target feature map output by the target network, a spatial attention feature map of the target feature map is generated. Specifically, based on the concatenated feature results of the target feature maps and the target feature map itself, a spatial attention feature map of the target feature map is generated, including: Based on the pre-determined first convolution kernel, a convolution operation is performed on the concatenated feature result of the target feature map to obtain the convolution result of the target feature map. Based on the pre-determined third activation function, a normalization operation is performed on the convolution result of the target feature map to obtain the normalized result corresponding to the target feature map. Based on the normalized result corresponding to the target feature map and the target feature map output by the target network, a spatial attention feature map of the target feature map is generated.

[0085] In this optional embodiment, the target feature map F has dimensions C×H×W, where C, H, and W are the number of channels, height, and width of the target feature map F, respectively. Channel-dimensional average pooling and max pooling are performed on all spatial locations of the target feature map F using the following formulas to obtain the spatial average pooling and spatial max pooling results of the target feature map F. The calculation formulas for channel-dimensional average pooling and channel-dimensional max pooling are as follows:

[0086] In the formula, G avg2 (F), G max2 (F) represents the spatial average pooling result and the spatial maximum pooling result of the target feature map F, both of which are 1×H×W feature maps.

[0087] Optionally, the formula for calculating the spatial stitching feature result of the target feature map is as follows:

[0088] In the formula, G(F) is the spatial splicing feature result of the target feature map, which is a 2×H×W tensor; Concat represents the splicing operation in the channel dimension.

[0089] Optionally, the formulas for calculating the normalized result of the target feature map and the spatial attention feature map of the target feature map are as follows:

[0090] In the formula, M s (F), F out2 These are the normalized result of the target feature map F and the spatial attention feature map of the target feature map F, respectively; Conv represents the convolution operation. It is the third activation function, such as Sigmoid, used to normalize the output to the range [0,1]; ⊙ represents element-wise multiplication, that is, multiplying the attention weight of each spatial location by the pixel value of the corresponding channel.

[0091] As can be seen, this optional embodiment can also compress the feature map from the channel dimension by using global average pooling and global max pooling to perform channel-dimensional average pooling and channel-dimensional max pooling on all spatial locations of each channel, thereby generating weights for each spatial location. This allows the YOLO model to focus more on regions that contribute to the task. This mechanism compresses the high-dimensional information of the feature map and converts it into two-dimensional features, thereby generating a spatial-level attention map to strengthen features at spatial locations. Combined with the aforementioned channel attention weighting, the YOLO model can more accurately focus on the target region and suppress irrelevant regions and unimportant features. In tasks such as multi-scale and small target detection, it can significantly improve the performance of the YOLO model and enhance the attention and expression of key features, ultimately improving the accuracy and robustness of the YOLO model, thereby improving the recognition accuracy of targets of various sizes and multiple targets, and improving the adaptability to various application scenarios.

[0092] In yet another optional embodiment, the method may further include the following steps: Obtain feature maps output from different target layers of the multi-scale fusion network of the YOLO model; For any two adjacent target layers, a second convolution kernel matching the attribute parameters of the feature maps of the two target layers is determined. Based on the second convolution kernel, a convolution operation matching the attribute parameters is performed on the feature map of the target layer matching the second convolution kernel to obtain the convolution feature map corresponding to the second convolution kernel. The attribute parameters of the convolution feature map corresponding to the second convolution kernel are the same as the attribute parameters of the feature map of the other target layer besides the target layer matching the second convolution kernel. The attribute parameters include the number of channels and / or spatial resolution. A feature fusion operation is performed on the convolutional feature map corresponding to the second convolutional kernel and the feature map of the other target layer of the two target layers, except the target layer that matches the second convolutional kernel, to obtain the fused first multi-scale fused feature map.

[0093] In this optional embodiment, when the attribute parameter includes the number of channels, if the number of channels between the two target layers is different, the number of channels in the target layer with more channels can be adjusted to match the number of channels in the target layer with fewer channels, or vice versa, to make the number of channels in the two target layers consistent. For example, if the lower-level feature map in the two target layers has 64 channels and the higher-level feature map has 256 channels, the number of channels in the lower-level feature map can be adjusted from 64 to 256 using a 1×1 convolutional kernel, thereby giving the two target layers the same number of channels and improving the accuracy of feature fusion.

[0094] When spatial resolution is included as an attribute parameter, if the spatial resolutions of two target layers are different, the spatial resolution of the target layer with higher spatial resolution can be downsampled to match that of the target layer with lower spatial resolution, or the spatial resolution of the target layer with lower spatial resolution can be upsampled to match that of the target layer with higher spatial resolution, thus making the spatial resolutions of the two target layers consistent. For example, if the spatial resolution of the lower-level feature map in two target layers is 128×128 and the spatial resolution of the higher-level feature map is 32×32, the spatial resolution of the lower-level feature map can be downsampled to 32×32 using a 3×3 convolution kernel or a 5×5 convolution kernel (e.g., using pooling or convolution operations). Alternatively, the spatial resolution of the higher-level feature map can be upsampled to 128×128 using a 3×3 convolution kernel with a stride or a 5×5 convolution kernel with a stride (e.g., using transposed convolution or bilinear interpolation), thus giving the two target layers the same spatial resolution and improving the accuracy of feature fusion.

[0095] Optionally, the fusion formula for the two target layer feature maps is as follows:

[0096] In the formula, F out3 F represents the first multi-scale fused feature map after fusion. low This represents the low-level feature map in the two target layers, containing details and spatial information; F high Represents the high-level feature maps in the two target layers, containing semantic and contextual information; Conv1 and Conv2 represent the channel transformation and downsampling / upsampling convolution operations on the features of the two target layer feature maps, respectively; ⊕ represents pixel-wise addition, used to fuse the features of the two target layers.

[0097] In this optional embodiment, the feature maps output by different target layers in the multi-scale fusion network may be derived from the aforementioned spatial attention feature maps, channel attention feature maps, or feature maps output by the backbone network.

[0098] As can be seen, this optional embodiment, by simultaneously utilizing both top-down and bottom-up fusion FPN methods, can not only transfer semantically rich high-level features to spatially high-resolution low-level features, but also gradually transfer detail-rich low-level features to semantically rich high-level features. This allows for closer interaction of feature information at different scales, achieving full fusion of multi-scale features, significantly improving the multi-scale detection performance of the YOLO model, and further enhancing the recognition accuracy of the YOLO model in small and multi-object scenarios, thus meeting the needs of refined object detection.

[0099] In yet another alternative implementation, the method may further include the following steps: Obtain feature maps output from different target layers of the multi-scale fusion network of the YOLO model; For any three adjacent target layers, a third convolutional kernel is determined based on the resolution of the feature maps of the higher-level target layers and the middle-level target layers. Based on this third convolutional kernel, a resolution convolution operation is performed on the feature maps of the higher-level target layers to obtain the transitive feature maps of the higher-level target layers. A fourth convolutional kernel is determined based on the resolution of the feature maps of the lower-level target layers and the middle-level target layers. Based on this fourth convolutional kernel, a resolution convolution operation is performed on the feature maps of the lower-level target layers to obtain the transitive feature maps of the lower-level target layers. The transitive feature maps of the lower-level target layers, the higher-level target layers, and the middle-level target layers have the same scale. The channel-dimensional concatenation operation is performed on the transfer feature maps of the target layer at the lower level, the target layer at the higher level, and the target layer at the middle level to obtain a channel concatenated feature map. Based on the pre-determined fifth convolution kernel, a convolution operation is performed on the channel concatenated feature map to obtain the second multi-scale fusion feature map.

[0100] In this optional embodiment, the second multi-scale fused feature map may be calculated using the following formula:

[0101] In the formula, F out4 F represents the second multi-scale fusion feature map; up This represents the transfer feature map of the target layer at a higher level; F down F represents the transfer feature map of the target layer at a lower level. horizontal This represents the feature map of the target layer in the middle layer; Concat represents the channel-dimensional concatenation operation; Conv4 represents the transitive feature map of the target layer in the higher layer, the transitive feature map of the target layer in the lower layer, and the feature map of the target layer in the middle layer. Figure 3Convolutional operations after fusing features.

[0102] As can be seen, this optional embodiment enhances the multi-directional smooth interaction of features at each layer by adding lateral connections in the up-down propagation path of FPN, directly transferring information across layers to features at the same scale, thereby improving the feature information sharing capability. Furthermore, through lateral connections and convolution operations, each layer of features contains richer contextual information, improving the YOLO model's ability to capture details, especially for the detection of small-sized targets. It also reduces detection omissions caused by differences in target size, improving the detection performance for small-sized targets and complex scenes. It is particularly suitable for complex scenes with small target areas and a large number of targets, such as drone scenarios.

[0103] In yet another optional embodiment, the method may further include the following steps: Determine multiple performance parameters of each filter in each of the multiple first network layers of the YOLO model, as well as scenario parameters for the application scenarios of the YOLO model; For any filter, based on the scenario parameters of the application scenario, assign a corresponding performance weight to each performance parameter of the filter, and calculate the importance of the filter based on each performance parameter and its performance weight; Based on the importance of all filters, filters with an importance less than or equal to a preset importance are removed from the YOLO model to obtain the pruned YOLO model.

[0104] In this optional embodiment, each filter's multiple performance parameters may include the loss change parameter of the YOLO model after pruning the filter, the L1 norm of the filter, and the scaling factor of the batch normalization of the filter. Optionally, different performance parameters may correspond to different performance weights for different application scenarios. Furthermore, the optimal combination of performance weights can be determined by adjusting the weights according to the specific application scenario through cross-validation or model performance evaluation experiments to ensure that the scores of different dimensions can reasonably reflect the importance of the filter. For example, for application scenarios more sensitive to feature distribution balance (such as object detection or segmentation), the performance weight of the scaling factor of the batch normalization is increased, and the priority of sensitivity analysis is controlled by adjusting the performance weight of the loss change parameter of the YOLO model after pruning the filter; if more attention is paid to the prediction accuracy of the pruned model, a higher performance weight is given to the loss change parameter of the YOLO model after pruning the filter; in image classification, the performance weight of the L1 norm of the filter weights can be appropriately increased to improve the sensitivity to weight strength. The sum of the performance weights corresponding to all performance parameters equals 1.

[0105]

[0106] In the formula, Importance(w i ) represents the importance of the i-th filter in the YOLO model; The L1 norm of the weight of the i-th filter represents the weight strength of filter i. The larger the L1 norm, the more significant the filter's contribution to the network. This represents the scaling factor of the i-th filter in the BatchNorm layer (batch normalization layer) of the YOLO model, reflecting the ability of the i-th filter to adjust the feature distribution. The larger the value, the more important the filter's impact on the performance of the YOLO model; α represents the change in the output error of the YOLO model after pruning the i-th filter, which is used to quantify the actual impact of the filter on the performance of the YOLO model. By calculating the difference in model output after pruning a certain filter, its importance can be assessed more accurately; α, β, and γ correspond to the performance weights of the performance parameters, respectively. Let represent the weight of the j-th row, k-th column, and c-th input channel of the i-th filter. Specifically, j and k represent the spatial dimensions of the convolution kernel (i.e., the height and width of the kernel), and c represents the index of the input channel. By summing the weights of all filters using a formula that iterates through all spatial locations and input channel weights, the total L1 norm of the filters is obtained.

[0107] In this optional embodiment, after obtaining the importance of each filter, they can be sorted, or not sorted, and filters with importance less than or equal to a preset importance can be deleted. Further, the preset importance can be determined based on the hardware performance of the device in which the YOLO model is integrated, where hardware performance includes computing power and / or memory / storage capacity. The stronger the computing power and memory / storage capacity, the lower the preset importance, i.e., the lower the pruning rate of the YOLO model. For example, high-performance hardware with high computing power can support a lower pruning rate (30%) while maintaining high accuracy; low-performance hardware with high computing power (such as embedded devices) requires a higher pruning rate (50%) to reduce computation; memory-constrained hardware requires an even higher pruning rate (50%) to reduce memory usage; hardware with low latency requirements (such as real-time systems) requires an even higher pruning rate to accelerate the inference process; high-performance processors are suitable for a 30% pruning rate, while edge devices require a 50% pruning rate.

[0108] As can be seen, this optional embodiment analyzes the importance of each filter by considering the loss change parameters of the YOLO model after pruning the corresponding filter, the L1 norm of the corresponding filter, the scaling factor of the batch normalization of the corresponding filter, the corresponding performance weight, and the sensitivity analysis combined with the YOLO model output. This improves the accuracy and reliability of determining the importance of each filter, filters (channels) with less impact on the final output of the YOLO model are screened and removed, thereby improving the pruning accuracy of the YOLO model, reducing the adverse effects on the performance of the YOLO model during pruning, and significantly reducing the number of parameters and computation of the YOLO model, reducing memory usage and inference latency. It maintains the prediction accuracy of the YOLO model while improving the inference speed of the YOLO model, making it particularly suitable for low-power computing environments, such as embedded devices and edge computing scenarios.

[0109] In this optional embodiment, determining the loss variation parameter corresponding to each filter among the multiple filters in each of the multiple first network layers of the YOLO model includes: Obtain the classification loss variation parameters, bounding box regression loss variation parameters, and / or confidence loss variation parameters of the filter; Based on the scenario parameters of the application scenario, set the corresponding loss change weights for the classification loss change parameters, bounding box regression loss change parameters, and / or confidence loss change parameters of the filter. Based on the classification loss variation parameters, bounding box regression loss variation parameters, and / or confidence loss variation parameters of the filter, and the corresponding loss variation weights, calculate the loss variation parameters of the YOLO model after pruning the filter.

[0110] In this optional embodiment, optionally, for any filter of the YOLO model, the formula for calculating the loss change parameter of the YOLO model after pruning the filter can be as follows:

[0111] In the formula, L pruned L original Let represent the loss parameters of the YOLO model after pruning the i-th filter and the original YOLO model, respectively.

[0112] In this optional embodiment, further optionally, for multi-task scenarios such as target detection, It can be further subdivided into:

[0113] In the formula, △L cls ΔL represents the classification loss variation parameter of the i-th filter; bbox The bounding box regression loss parameter represents the change in the bounding box of the i-th filter; △Ld This represents the confidence loss change parameter for the i-th filter; These represent the corresponding loss change weights, used to adjust the importance weights of different task losses on pruning.

[0114] In this optional embodiment, optionally, during the calculation At the same time, the contribution of each filter to the loss function can be obtained through gradient backpropagation, so as to more accurately evaluate the importance of the filter; and output sensitivity analysis ensures that filters that are sensitive to the performance of the model task are retained first, thereby effectively reducing the damage of pruning to the final model accuracy and optimizing inference efficiency.

[0115] As can be seen, this optional embodiment can also improve the accuracy of determining the loss change parameters of the filter by analyzing the classification loss change parameters, bounding box regression loss change parameters and / or confidence loss change parameters and their corresponding loss change weights. This is conducive to further improving the accuracy and reliability of determining the importance of the filter, and further reducing the weight of the YOLO model while ensuring its performance, which is conducive to more flexible deployment in lightweight embedded devices.

[0116] In yet another optional embodiment, the method may further include the following steps: From the pruned YOLO model, identify all second network layers from all first network layers that have undergone filter pruning; For any second network layer, determine the number of channels and the channel matrix of the second network layer, and update the number of channels and the channel matrix of the downstream network layer based on the number of channels and the channel matrix of the second network layer.

[0117] In this optional embodiment, the number of output channels after pruning is optionally determined. After removing unimportant filters, the number of output channels of the convolutional layer decreases. For example, if the original number of output channels is C, after pruning k channels, the number of output channels becomes C - k. At this time, the number of input channels of the downstream network needs to be updated, that is, the number of input channels of subsequent convolutional layers needs to be adjusted to the number of output channels of the pruned convolutional layer. For example, if the original number of input channels of the downstream layer was C, it is adjusted to C - k after pruning. The convolutional layer weights also need to be adjusted: after pruning, the shape of the weight matrix of the convolutional layer needs to be updated by removing the corresponding channels.

[0118] As can be seen, after pruning the YOLO model, this optional embodiment further adjusts the number of input channels and the channel matrix of the downstream network to ensure the structural consistency of the YOLO model, thereby ensuring the accuracy of the model.

[0119] In yet another optional embodiment, the method may further include the following steps: After updating the number of channels and the channel matrix of the downstream network layers of the second network layer, the feature map dataset is obtained, and the YOLO model is trained based on the feature dataset to obtain the trained YOLO model.

[0120] In this optional embodiment, the trained YOLO model is integrated into an embedded device and a server, with the embedded device being integrated into the flight equipment. The embedded device is used to acquire images and, together with the server, perform object detection on the acquired images based on the YOLO model.

[0121] In this optional embodiment, the feature map dataset may consist of feature maps from different scenes and may include targets of various sizes.

[0122] As can be seen, this optional embodiment further fine-tunes the YOLO model after pruning and structural adjustment, so that the model can further improve the balance between the accuracy, reliability and efficiency of the recognition of targets of different sizes and multiple targets while maintaining high accuracy. It is especially suitable for application scenarios with limited resources and high real-time requirements, such as flight mission scenarios.

[0123] Example 2 Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention. Figure 2 The described method can be applied to any scenario requiring target detection, including but not limited to border patrol, power line inspection, and forest fire prevention. This scenario is equipped with a corresponding detection system, which includes a server and, more specifically, an embedded device that communicates with the server. This embedded device is integrated into a flight device that is performing a specific flight mission within the corresponding scenario. Figure 2 As shown, the method may include the following operations: 201. During the flight of the flight equipment, the embedded device performs image acquisition operations on the current scene to obtain the video stream of the current scene.

[0124] 202. The embedded device performs object detection on the current frame image of the video stream of the current scene based on the YOLO model, and obtains the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score of each detection box and the category.

[0125] 203. When the current frame image is the first frame image of the video stream, the embedded device assigns a corresponding identifier to each target in the current frame image based on the detection results of the current frame image, and sends the current frame image to the server. The first image includes the first frame image of the video stream.

[0126] 204. When the current frame image is not the first frame image of the video stream, the embedded device obtains the detection result of the previous frame image, and performs a tracking operation on the target based on the detection result of the current frame image and the detection result of the previous frame image to obtain the target tracking result of the current frame image.

[0127] 205. Based on the target tracking results of the current frame image, the embedded device determines whether a new target appears in the current frame image. When it is determined that a new target appears in the current frame image, it performs an image acquisition operation at the target resolution to obtain a second image and sends the second image to the server. The first image includes the second image.

[0128] It is evident that implementation Figure 2 The described method integrates a YOLO model with a backbone network and / or multi-scale fusion network incorporated into the attention module into an embedded device, thereby improving the computing power of the embedded device. Using this model to perform target detection on acquired images, it can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition and reducing the rate of missed target detection. Furthermore, by first performing target detection and capture at a lower resolution on the embedded device, and then triggering higher-resolution image acquisition and sending to the server only when a new target is detected, the frequent backhaul of high-resolution images and repeated target backhauls to the server occur, thus saving bandwidth and computing resources. Finally, the server performs refined recognition of images containing new targets through collaborative processing, improving image data processing efficiency and balancing the requirements of target detection efficiency and accuracy, making it particularly suitable for high-resolution image inspection scenarios.

[0129] It should be noted that for other technical details regarding embedded devices, please refer to the specific description of the relevant content in Embodiment 1, which will not be repeated in this embodiment of the invention.

[0130] Example 3 Please see Figure 3 , Figure 3 This is a schematic diagram of a system for intelligent target detection based on the YOLO model, as disclosed in an embodiment of the present invention. This system can be applied to any scenario requiring target detection, including but not limited to border patrol, power line inspection, and forest fire prevention scenarios. Each scenario has a corresponding detection system. The system includes a server-side component. Figure 3 As shown, the server includes: The first acquisition module 301 is used to acquire a first image sent by the embedded device connected to it. The first image is an image collected by the embedded device for the current scene during the flight of the flight equipment, and the first image contains the target first identified by the embedded device.

[0131] In this embodiment of the invention, the embedded device is integrated into the flight equipment, wherein the flight equipment is performing a corresponding flight mission in the corresponding scenario.

[0132] The first detection module 302 is used to perform a detection operation on the first image based on the YOLO model to obtain the detection result of the first image. The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network based on the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. The recognition module 303 is used to recognize information of at least one first target in the first image based on the detection results of the first image. All first targets include all targets identified for the first time, and the information of each first target includes the category of the first target.

[0133] It is evident that implementation Figure 3 The described system for intelligent target detection based on the YOLO model integrates a YOLO model with an attention module incorporating the backbone network and / or multi-scale fusion network on the server. This system performs target detection on images where new targets are first identified from embedded devices. It accurately identifies small and edge targets in low-resolution images and accurately extracts structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition. Furthermore, the collaborative processing of first detecting and capturing targets on the embedded device and then refining the recognition on the server improves image data processing efficiency, balancing the requirements of target detection efficiency and accuracy. This makes it particularly suitable for high-resolution image inspection scenarios.

[0134] In an optional embodiment, Figure 4 This is a schematic diagram of the structure of a system for intelligent target detection based on the YOLO model disclosed in an embodiment of the present invention, as shown below. Figure 4 As shown, the server may also include: The segmentation module 304 is used to perform image segmentation on the first image before the first detection module 302 performs detection operation on the first image based on the YOLO model and obtains the detection result of the first image, thereby obtaining multiple block images; The first detection module 302, based on the YOLO model, performs detection operations on the first image to obtain the detection results of the first image. The specific methods include: Based on the YOLO model, a detection operation is performed on each block of the image to obtain the detection result for each block of the image; The specific method by which the recognition module 303 recognizes information about at least one first target in the first image based on the detection results of the first image includes: Based on the detection results of all block images, identify information about at least one first target in the first image.

[0135] It is evident that implementation Figure 4 The described device can first divide an image acquired from an embedded device into blocks, and then sequentially apply the YOLO model described above to the blocks of images for target detection and recognition, thereby improving the efficiency and accuracy of target detection and recognition.

[0136] In yet another optional embodiment, the detection result for each block image includes multiple detection boxes, the confidence level of each detection box, and the category; The specific method by which the recognition module 303 identifies information about at least one first target in the first image based on the detection results of all block images includes: For the detection results of each image block, a stitching operation is performed to obtain the first stitched detection result; Based on the coordinates of each detection box in the first detection result, determine all detection box groups with an overlap greater than or equal to a preset overlap from all detection boxes in the first detection result, with each detection box group consisting of at least two detection boxes; Based on the confidence level of each detection box in each detection box group, delete all detection boxes in that detection box group except for the detection box with the highest confidence level; After performing a detection box deletion operation on all detection box groups, a stitching operation is performed on the adjacent boxes of the first detection result to obtain a second detection result, and based on the second detection result, information of at least one first target in the first image is identified.

[0137] It is evident that implementation Figure 4 The described device can also stitch together the target detection results of the block image by optimizing the stitching process to ensure the accuracy and integrity of the detection boxes. After stitching, the detection boxes are grouped according to their overlap. For each group of detection boxes, the detection box with the highest confidence is retained, and the other detection boxes are deleted. Then, the adjacent detection boxes are stitched together, which improves the detection accuracy and data volume after stitching, thereby improving the efficiency and accuracy of target detection.

[0138] In yet another alternative embodiment, such as Figure 4As shown, the server may also include: The determining module 305 is used to determine the number of blocks that the first image needs to be divided into; The first acquisition module 301 is also used to acquire the image size of the first image and the size of all first targets in the first image; The determining module 305 is also used to determine the first step length in the horizontal direction and the second step length in the vertical direction based on the size of all the first targets in the first image; The determining module 305 is also used to determine the block overlap during the first image segmentation process based on the first step length, the second step length, the image size of the first image, and the number of blocks; The segmentation module 304 performs image segmentation on the received first image to obtain multiple image blocks in the following ways: Based on the block overlap, the received first image is divided into multiple block images.

[0139] It is evident that implementation Figure 4 The described device can also determine the block overlap by combining multiple factors such as the number of blocks required for image segmentation, image size, target size in the image, and step length in the horizontal and vertical directions. This improves the accuracy and reliability of block overlap determination, increases the probability that targets in the image are segmented into the same image block, thereby improving the image segmentation accuracy and efficiency, and further improving the target detection efficiency and accuracy.

[0140] In yet another alternative embodiment, such as Figure 4 As shown, the system may further include an embedded device that communicates with the server, wherein the embedded device includes: The acquisition module 306 is used to perform image acquisition operations on the current scene during the flight of the flight equipment to obtain a video stream of the current scene; The second detection module 307 is used to perform target detection operation on the current frame image of the video stream of the current scene based on the YOLO model, and obtain the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score and category of each detection box. The setting module 308 is used to assign a corresponding identifier to each target in the current frame image based on the detection results of the current frame image when the current frame image is the first frame image of the video stream. The sending module 309 is used to send the current frame image to the server, and the first image includes the first frame image of the video stream; The second acquisition module 310 is used to acquire the detection result of the previous frame image when the current frame image is a non-first frame image of the video stream; The tracking module 311 is used to perform a tracking operation on the target based on the detection results of the current frame image and the detection results of the previous frame image, so as to obtain the target tracking result of the current frame image; The judgment module 312 is used to determine whether a new target appears in the current frame image based on the target tracking result of the current frame image; The acquisition module 306 is also used to perform an image acquisition operation at the target resolution when it is determined that a new target appears in the current frame image, to obtain a second image, wherein the first image includes the second image, and the target resolution is greater than the resolution of the video stream; The sending module 309 is also used to send the second image to the server.

[0141] It is evident that implementation Figure 4 The described device can also integrate a YOLO model with an attention module, incorporating the backbone network and / or multi-scale fusion network into an embedded device, thereby improving the computing power of the embedded device. Using this device to perform target detection on acquired images, it can accurately identify small and edge targets in low-resolution images and accurately extract structural information of targets in high-resolution images, improving the efficiency and accuracy of multi-target recognition and reducing the rate of missed target detection. Furthermore, it first performs target detection and capture at a lower resolution using the embedded device, and only triggers higher-resolution image acquisition and transmission to the server when a new target is detected, reducing frequent backhauls of high-resolution images and repeated target transmissions to the server, thus saving bandwidth and computing resources. Finally, it performs collaborative processing on the server to perform refined recognition of images containing new targets, improving image data processing efficiency and balancing the requirements of target detection efficiency and accuracy, making it particularly suitable for high-resolution image inspection scenarios.

[0142] In this optional embodiment, optionally, the tracking module 311 performs a tracking operation on the target based on the detection results of the current frame image and the detection results of the previous frame image, and obtains the target tracking result of the current frame image in the following specific ways: Based on the detection results of the current frame image, analyze the current information of each second target in the current frame image. The current information of each second target includes the current position of the second target. Based on the detection results of the previous frame image, analyze the historical position of each of the second targets in the previous frame image, and predict the predicted position of each second target at the next moment based on the current position of each second target and the historical position of the second target. Based on the predicted position of each second target, a tracking operation is performed on the second target to obtain the target tracking result of the current frame image.

[0143] As can be seen, this optional embodiment predicts the target's position at the next moment by using the currently detected target's position and its historical position, which improves the accuracy and efficiency of target tracking. This is conducive to improving the accuracy and efficiency of discovering and distinguishing new targets, and further conducive to improving the accuracy and efficiency of multi-target identification.

[0144] In this optional embodiment, the specific method by which the determining module 312 determines whether a new target appears in the current frame image based on the target tracking result of the current frame image includes: When a tag assignment operation has been performed on a target that appeared in a historical frame image, it is determined whether the target tracking result of the current frame image is used to indicate that there is a target in the current frame image that has not been assigned a tag. If the result is yes, it is determined that a new target has appeared in the current frame image, and the new target is a target that has not been assigned a tag; or, Based on the target tracking results of the current frame image, calculate the cross value between the current position of each second target and the predicted position of the second target, and determine whether the cross value corresponding to each second target is greater than or equal to a preset cross value. When it is determined that it is greater than or equal to the preset cross value, it is determined that a new target has appeared in the current frame image, and the new target is a target that has not been assigned an identifier.

[0145] It is evident that implementation Figure 4 The described device can also identify new targets by identifying the target in the image or by the intersection between the current position and its predicted position, thereby improving the accuracy and efficiency of new target identification. This is conducive to further improving the accuracy and efficiency of triggering high-resolution image acquisition, and further conducive to further improving the efficiency and accuracy of accurate identification of new targets.

[0146] Example 4 Please see Figure 5 , Figure 5 This is a schematic diagram of a server structure disclosed in an embodiment of the present invention. This server can be applied to any scenario requiring target detection, including but not limited to border patrol, power line inspection, and forest fire prevention scenarios. Figure 5 As shown, the server may include: Memory 401 storing executable program code; Processor 402 coupled to memory 401; Furthermore, it may also include an input interface 403 and an output interface 404 coupled to the processor 402; The processor 402 calls the executable program code stored in the memory 401 to execute some or all of the steps performed by the server in the method for intelligent detection of targets based on the YOLO model disclosed in Embodiment 1 of the present invention.

[0147] Example 5 Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an embedded device disclosed in an embodiment of the present invention. This embedded device can be applied to any scenario requiring target detection, including but not limited to border patrol, power line inspection, and forest fire prevention scenarios. Figure 6 As shown, the embedded device may include: Memory 501 storing executable program code; Processor 502 coupled to memory 501; Furthermore, it may also include an input interface 503 and an output interface 504 coupled to the processor 502; The processor 502 calls the executable program code stored in the memory 501 to execute some or all of the steps performed by the embedded device in the method for intelligent detection of targets based on the YOLO model disclosed in Embodiment 1 of the present invention.

[0148] Example 6 This invention discloses a computer storage medium storing computer instructions. When these computer instructions are invoked, they are used to execute some or all of the steps performed by the server in the method for intelligent target detection based on the YOLO model disclosed in Embodiment 1 of this invention.

[0149] Example 7 This invention discloses a computer storage medium storing computer instructions. When these computer instructions are invoked, they are used to execute some or all of the steps performed by the embedded device in the method for intelligent target detection based on the YOLO model disclosed in Embodiment 1 of this invention.

[0150] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0151] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0152] Finally, it should be noted that the method, device, and system for intelligent target detection based on the YOLO model disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent target detection based on the YOLO model, characterized in that, The method is applied to a detection system, the detection system including a server, and the method includes: The server obtains a first image sent by the embedded device with which it communicates, wherein the embedded device is integrated into the flight equipment, the first image is an image collected by the embedded device for the current scene during the flight of the flight equipment, and the first image contains the target first identified by the embedded device; The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image; The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. The server identifies information about at least one first target in the first image based on the detection results of the first image. All first targets include all targets identified for the first time, and the information of each first target includes the category of the first target.

2. The method for intelligent target detection based on the YOLO model according to claim 1, characterized in that, Before the server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image, the method further includes: The server performs an image segmentation operation on the acquired first image to obtain multiple block images; The server performs a detection operation on the first image based on the YOLO model to obtain the detection result of the first image, including: The server performs a detection operation on each block image based on the YOLO model to obtain the detection result of each block image; The server identifies information about at least one first target in the first image based on the detection results of the first image, including: The server identifies information about at least one first target in the first image based on the detection results of all the block images.

3. The method for intelligent target detection based on the YOLO model according to claim 2, characterized in that, The detection result for each block image includes multiple detection boxes, the confidence level of each detection box, and the category. The server identifies information about at least one first target in the first image based on the detection results of all the block images, including: The server performs a stitching operation on the detection results of each block image to obtain the stitched first detection result; Based on the coordinates of each detection box in the first detection result, the server determines a group of detection boxes with an overlap greater than or equal to a preset overlap from all the detection boxes in the first detection result, and each group of detection boxes consists of at least two detection boxes; The server deletes all detection frames in each detection frame group except the one with the highest confidence, based on the confidence level of each detection frame in that detection frame group. After performing a detection box deletion operation on all the detection box groups, the server performs a stitching operation on the adjacent boxes of the first detection result to obtain a second detection result, and identifies information about at least one first target in the first image based on the second detection result.

4. The method for intelligent target detection based on the YOLO model according to claim 2 or 3, characterized in that, The method includes: The server determines the number of blocks that the first image needs to be divided into, and obtains the image size of the first image and the size of all the first targets in the first image; The server determines the first step length in the horizontal direction and the second step length in the vertical direction based on the size of all the first targets in the first image. The server determines the block overlap during the first image segmentation process based on the first step length, the second step length, the image size of the first image, and the number of blocks. The server performs image segmentation on the received first image to obtain multiple block images, including: The server performs image segmentation on the received first image based on the segmentation overlap to obtain multiple block images.

5. The method for intelligent target detection based on the YOLO model according to any one of claims 1-3, characterized in that, The detection system also includes the embedded device; The method further includes: During the flight of the flight equipment, the embedded device performs an image acquisition operation on the current scene to obtain a video stream of the current scene; The embedded device performs object detection on the current frame image of the video stream of the current scene based on the YOLO model to obtain the detection result of the current frame image. The detection result of the current frame image includes multiple detection boxes, the confidence score of each detection box, and the category. When the current frame image is the first frame image of the video stream, the embedded device assigns a corresponding identifier to each target in the current frame image according to the detection result of the current frame image, and sends the current frame image to the server. The first image includes the first frame image of the video stream. When the current frame image is not the first frame image of the video stream, the embedded device obtains the detection result of the previous frame image, and performs a tracking operation on the target based on the detection result of the current frame image and the detection result of the previous frame image to obtain the target tracking result of the current frame image; The embedded device determines whether a new target appears in the current frame image based on the target tracking result of the current frame image. When it is determined that a new target appears in the current frame image, it performs an image acquisition operation at the target resolution to obtain a second image and sends the second image to the server. The first image includes the second image, wherein the target resolution is greater than the resolution of the video stream.

6. The method for intelligent target detection based on the YOLO model according to claim 5, characterized in that, The embedded device performs a target tracking operation based on the detection results of the current frame image and the detection results of the previous frame image, obtaining the target tracking result of the current frame image, including: The embedded device analyzes the current information of each second target in the current frame image based on the detection result of the current frame image, and the current information of each second target includes the current position of the second target; The embedded device analyzes the historical position of each second target in the previous frame image based on the detection result of the previous frame image, and predicts the predicted position of each second target at the next moment based on the current position of each second target and the historical position of the second target. The embedded device performs a tracking operation on each second target based on the predicted position of the second target, and obtains the target tracking result of the current frame image.

7. The method for intelligent target detection based on the YOLO model according to claim 5 or 6, characterized in that, The embedded device determines whether a new target appears in the current frame image based on the target tracking result of the current frame image, including: When an identifier assignment operation has been performed on a target appearing in a historical frame image, the embedded device determines whether the target tracking result of the current frame image is used to indicate that there is a target in the current frame image that has not been assigned an identifier. If the result is yes, it determines that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier; or, The embedded device calculates the cross value between the current position of each second target and the predicted position of the second target based on the target tracking result of the current frame image, and determines whether the cross value corresponding to each second target is greater than or equal to a preset cross value. When it is determined that the cross value is greater than or equal to the preset cross value, it determines that a new target has appeared in the current frame image, and the new target is the target that has not been assigned an identifier.

8. A system for intelligent target detection based on the YOLO model, characterized in that, The system includes a server, which includes: The first acquisition module is used to acquire a first image sent by an embedded device that is connected to it in communication, wherein the embedded device is integrated into the flight equipment, and the first image is an image collected by the flight equipment in response to the current scene during flight. The first detection module is used to perform a detection operation on the first image based on the YOLO model to obtain the detection result of the first image; The target network of the YOLO model incorporates an attention module, which includes a spatial attention module and / or a channel attention module. The attention module is a module determined for each channel of the feature map output by the target network of the YOLO model. The target network of the YOLO model includes the backbone network of the YOLO model and / or a multi-scale fusion network. The recognition module is used to recognize information of at least one first target in the first image based on the detection results of the first image, wherein all first targets include all targets acquired for the first time, and the information of each first target includes the category of the first target.

9. A server-side component, characterized in that, The server includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method for intelligent target detection based on the YOLO model as described in any one of claims 1-7.

10. An embedded device, characterized in that, The embedded device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method for intelligent target detection based on the YOLO model as described in any one of claims 1-7.

Citation Information

Cited By

  • Elevator indicator light state determination method, electronic equipment and medium

    CN121542898A

  • An elevator indicator light state determination method, electronic device, and medium

    CN121542898B