Target detection method and system based on YOLOv9

By improving the YOLOv9 model, increasing the attention mechanism and depth separable convolution, optimizing the feature extraction and fusion process, the problem of high computing resource consumption and insensitive small object detection is solved, and the object detection effect with high accuracy and low resource consumption is achieved.

CN120070870AActive Publication Date: 2025-05-30DONGGUAN CITY COLLEGE

Patent Information

Application Number
CN202510277374.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-30
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The YOLOv9 model consumes high computing resources in object detection and is not sensitive enough to small object detection.

Method used

By improving the YOLOv9 model, the attention mechanism and depth separable convolution are increased, the feature extraction and fusion process is optimized, the computing resource requirements are reduced, and the sensitivity of small object detection is improved.

Benefits of technology

High-accurate object detection under the condition of fewer computing resources is achieved, the detection accuracy of small objects is improved, the number of model parameters is reduced, and the calculation efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070870A_ABST
    Figure CN120070870A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a target detection method and system based on YOLOv9. The YOLOv9 model is improved, particularly, convolution structure design optimization is carried out in the feature extraction process, and an attention module is added in the feature fusion process, so that computing resources in the model processing process can be reduced, the high requirement of an existing model for hardware configuration can be solved, and the sensitivity of small target recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and specifically provides an object detection method and system based on YOLOv9. Background Art

[0002] YOLOv9 is the latest version of the YOLO (You Only Look Once) series, which has achieved remarkable progress in the field of object detection. Refer to Figure 1 , which is the standard structure of the released YOLOv9 model. Programmable Gradient Information (PGI) and Generalized Efficient Layer Aggregation Network (GELAN) are introduced into the YOLOv9 model. These two technologies mark significant progress in the field of real-time object detection. PGI generates reliable gradients through auxiliary reversible branches, solving the problem of information loss in deep networks, while GELAN optimizes parameter utilization and computational efficiency, enabling YOLOv9 to adapt to various computing environments.

[0003] Although YOLOv9 has made remarkable progress in the field of object detection, there are still some drawbacks and limitations. As can be seen from Figure 1 , the network structure of YOLOv9 is relatively complex. As a large model, it requires more computing resources (such as GPU memory and processors) for inference, which may not be friendly to devices with limited resources. Moreover, for large models, they perform well in general object detection, but are not sensitive enough to small object detection, and the detection effect is not ideal. Summary of the Invention

[0004] In order to solve the technical problems described in the background art, the embodiments of this application provide an object detection method and system based on YOLOv9. By making structural changes to the existing YOLOv9 and adding multiple different modules, YOLOv9 can achieve high accuracy in detecting objects, especially small objects, with less computing resources.

[0005] In order to achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, a target detection method based on YOLOv9 is provided. The method includes: extracting features of an image to be detected through a feature extraction network to obtain multi-scale features of the image to be detected, and fusing the multi-scale features to obtain a feature map; the feature extraction network includes multiple feature extraction modules, and the multiple feature extraction modules are used to extract features of different scales of the image to be detected and output feature maps of different scales, and the feature maps of multiple different scales obtain more extensive context information through a downsampling module; each feature extraction module includes a 1×1 convolutional layer and two feature extraction sub-modules, and the two feature extraction sub-modules are used to receive two feature maps output by the 1×1 convolutional layer; passing the feature map through a feature fusion network for attention fusion, and obtaining a fusion feature through cross-stage connection; obtaining the predicted probability value corresponding to each category through a detection head for the fusion feature, selecting the maximum probability value among the predicted probability values as the prediction result, and displaying the bounding box coordinates in the image to be detected.

[0007] Further, the method further includes adjusting the size of the image to be detected, specifically including: adjusting the pixel size of the image to be detected to a target size through affine transformation, and adjusting the resolution of the adjusted image to be detected until it meets the resolution requirements of the feature extraction module.

[0008] Further, the feature extraction network further includes multiple convolutional layers, and the initial feature map after feature extraction through each convolutional layer enters the corresponding feature extraction sub-module.

[0009] Further, the feature extraction module receives the feature map processed by the convolutional layer and passes it through a 1×1 convolutional layer, which is used to convert the number of input channels from c1 to c3; the feature map output by the 1×1 convolutional layer is split into a first feature map and a second feature map, and the number of channels corresponding to each feature map is c3 / 2; the first feature map and the second feature map are respectively input into two feature extraction sub-modules.

[0010] Further, a 1×1 convolutional layer connected to the output of the feature extraction sub-module is further included in the feature extraction module.

[0011] Further, the feature extraction sub-module includes a first depthwise separable convolution and a second depthwise separable convolution. The first depthwise separable convolution is used to extract features, and the second depthwise separable convolution is used to merge the extracted features with the features extracted by the first depthwise separable convolution in the channel dimension; the merged features pass through a 1×1 convolutional layer connected to the output of the feature extraction sub-module, and the number of channels is converted from c3 + 2×c4 to the output channel number c2.

[0012] Furthermore, the feature fusion network includes a pooling module, an attention module, and a fusion module. The pooling module is used to divide the feature map into multiple grids of different sizes, perform max-pooling operations on each grid, and extract features of different scales. The attention module is configured with an attention mechanism to update the features of different scales based on attention weights to obtain updated features. The fusion module is used to concatenate the features of different scales and fuse them through a convolutional layer to obtain fused features.

[0013] Furthermore, the attention module is a multi-scale channel attention module.

[0014] In a second aspect, a target detection system based on YOLOv9 is provided. The system includes: an image acquisition end for acquiring a to-be-detected image within a target area; an image processing device for performing target detection on the to-be-detected image to obtain at least one classification result and display the bounding box coordinates in the to-be-detected image; and a result display end for displaying the to-be-detected image and the bounding box coordinates on the to-be-detected image.

[0015] Furthermore, the image processing device includes: a feature extraction module for extracting features of the to-be-detected image that has been resized through a feature extraction network to obtain multi-scale features of the to-be-detected image, and fusing the multi-scale features to obtain a feature map; a feature fusion module for performing attention fusion on the feature map through a feature fusion network and obtaining fused features through cross-stage connection; and a classification module for obtaining the predicted probability value corresponding to each category, selecting the maximum probability value among the predicted probability values as the prediction result, and displaying the bounding box coordinates in the to-be-detected image.

[0016] In the technical solution provided by the embodiments of the present application, by improving the YOLOv9 model, especially optimizing the convolutional structure design during the feature extraction process and adding an attention module during the feature fusion process, the computational resources in the model processing process can be reduced, and the high requirements for hardware configuration of the existing model can be solved, and the sensitivity for small target recognition can be improved. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] The methods, systems, and / or programs in the accompanying drawings will be further described according to exemplary embodiments. These exemplary embodiments will be described in detail with reference to the drawings. These exemplary embodiments are non-limiting exemplary embodiments, where example numbers represent similar mechanisms in various views of the drawings.

[0019] Figure 1 is a schematic structural diagram of the YOLOv9 model in the prior art.

[0020] Figure 2 is a schematic flowchart of an object detection method provided by an embodiment of the present application.

[0021] Figure 3 is a schematic structural diagram of a feature extraction module provided by an embodiment of the present application.

[0022] Figure 4 is a schematic structural diagram of a feature extraction sub-module provided by an embodiment of the present application.

[0023] Figure 5 is a schematic modular structural diagram of a feature fusion network provided by an embodiment of the present application.

[0024] Figure 6 is a schematic flowchart of an object detection method provided by another embodiment of the present application.

[0025] Figure 7 is a schematic structural diagram of an object detection system provided by an embodiment of the present application.

[0026] Figure 8 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed Embodiments

[0027] To better understand the above technical solutions, the technical solutions of the present application will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0028] In the following detailed description, many specific details are set forth by way of example in order to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present application may be practiced without these details. In other instances, well-known methods, procedures, systems, components, and / or circuits have been described at a relatively high level without detail in order to avoid unnecessarily obscuring aspects of the present application.

[0029] In this application, a flowchart is used to illustrate the execution process performed by the system according to the embodiments of this application. It should be clearly understood that the execution process of the flowchart may not be executed in sequence. On the contrary, these execution processes may be executed in reverse order or simultaneously. Additionally, at least one other execution process may be added to the flowchart. One or more execution processes may be deleted from the flowchart.

[0030] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0031] (1) Responsive to, which is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of the multiple executed operations.

[0032] (2) Based on, which is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of the multiple executed operations.

[0033] The embodiments of this application provide a target detection method, which is implemented through an artificial intelligence model. The model adopts the YOLO series of models, specifically the latest released YOLOv9 model. YOLOv9 not only performs excellently in target detection tasks but can also be applied to various computer vision tasks such as instance segmentation and panoramic segmentation. Its versatility makes it an ideal choice for various practical application scenarios, such as autonomous driving, video surveillance, and robot vision. Programmable Gradient Information (PGI) and Generalized Efficient Layer Aggregation Network (GELAN) are introduced into this model. Among them, PGI generates reliable gradients through auxiliary reversible branches, solving the problem of information loss in deep networks, while GELAN optimizes the parameter utilization rate and computational efficiency, enabling YOLOv9 to adapt to various computing environments.

[0034] However, the YOLOv9 model is currently the largest model with the most complex network structure and the most components in the YOLO family. The more complex the structure of the artificial intelligence model, the higher the required computational cost, especially the higher the requirement for the GPU. This leads to the need for a relatively high training environment and operating environment in order to enable the YOLOv9 model to achieve its optimal effect. Moreover, large models do not have an advantage in small target detection and are prone to missing and misdetecting small targets.

[0035] Therefore, in order to reduce the computational resource requirements of the YOLOv9 model and improve the detection accuracy of the model for small targets, the embodiments of the present application optimize the YOLOv9 model and provide an object detection method based on the optimized YOLOv9 model.

[0036] Specifically, in order to reduce the computational resource requirements, in this embodiment, on the basis of retaining the original architecture of the model, an attention mechanism is added to make the features better able to combine context information, enriching the expression of the features. On the basis of saving the original computational cost, it can ensure the integrity and richness of the model for feature extraction and feature fusion; moreover, on the basis of retaining the original architecture, the convolutional block is optimized, and depthwise separable convolutions are added on the basis of retaining some convolutions, improving the expression of features in the channels.

[0037] In this embodiment, refer to Figure 2 An object detection method using the optimized YOLOv9 model is provided, including the following steps:

[0038] Step S21. Extract features of the image to be detected through a feature extraction network to obtain multi-scale features of the image to be detected, and fuse the multi-scale features to obtain a feature map.

[0039] The optimized model provided in this embodiment mainly includes three structural sub-networks, namely a feature extraction network, a feature fusion network, and a detection head. The configuration methods of these three networks are the same in the YOLO series of models. The difference is that for different specific models, there are different modules in the above three sub-networks, and different components are configured in different modules.

[0040] Among them, the feature extraction network is mainly used to extract features of the image to be detected to obtain multi-scale features of the image to be detected. A plurality of feature extraction modules are provided in the feature extraction network, and convolutional layers are correspondingly provided for the feature extraction modules. Each convolutional layer is correspondingly provided for a feature extraction module. The feature image that has not entered the feature extraction module is called the initial feature map, and each initial feature map is obtained after feature extraction through the convolutional layer. Each feature extraction module is used to extract features of different scales of the image to be detected and thus output feature maps of different scales, and a plurality of feature maps of different scales are further obtained through downsampling to obtain more extensive context information.

[0041] In this embodiment, refer to Figure 3, each feature extraction module includes a 1*1 convolutional layer and two feature extraction sub-modules. The two feature extraction sub-modules are used to receive two feature maps output by the 1*1 convolutional layer. The feature maps processed by the convolutional layer first pass through the 1*1 convolutional layer, which is used to convert the number of input channels of the initial feature map from c1 to c3. And the features processed by the 1*1 convolutional layer are divided into a first feature map and a second feature map, and the number of channels corresponding to each feature map is c3 / 2.

[0042] The first feature map and the second feature map after segmentation are input into the corresponding feature extraction sub-modules. For the structure of the feature extraction sub-modules in this embodiment, please refer to Figure 4 . The feature extraction sub-module includes a first depthwise separable convolution and a second depthwise separable convolution. The first depthwise separable convolution is used to extract features from the received first feature map, and the second depthwise separable convolution is used to extract features from the received second feature map and merge the features extracted by the first depthwise separable convolution to obtain the merged features. The merged features pass through the 1*1 convolutional layer connected to the output of the feature extraction sub-module. This convolutional layer is used to convert the number of output channels from c3 + 2*c4 to the output channel number c2. That is, the features obtained through the first feature extraction sub-module and the second feature extraction sub-module respectively are concatenated and then the output channels are converted through the 1*1 convolutional layer.

[0043] It should be noted that the above feature extraction module is the Figure 1 feature extraction part in the RepNCSPELAN4 module in

[0044] Step S22. The feature maps are subjected to attention fusion through the feature fusion network and the fused features are obtained through cross-stage connection.

[0045] Regarding step S21, it is mainly used to extract multi-scale features. The extracted multi-scale features need to be fused to obtain the fused features, and the fused features are the features for final object detection.

[0046] In this embodiment, in order to reduce the model's demand for computing resources, a different method from the existing models is adopted when performing feature fusion. Refer to Figure 5 , in this embodiment, feature fusion is performed using a feature fusion network, that is, multiple features output after step S21 are fused through the feature fusion network; among them, the feature fusion network includes a pooling module, an attention module, and a fusion module.

[0047] Among them, the pooling module is used to divide the features output in step S21 into multiple networks of different sizes, and perform max pooling operations on each grid to segment the scale of the feature map and retain the main features, so as to realize the screening of features.

[0048] The features after screening are input backward to the attention module for weight update to obtain the updated features. Finally, the updated features pass through the fusion module to achieve connection and are fused through the convolutional layer to obtain the fused features.

[0049] Among them, in this embodiment, a multi-scale channel attention module is adopted for the attention module, and the pooling module includes a global average pooling sub-module and a global max pooling sub-module. Using different global pooling makes the extracted high-level features richer. Then, it first passes through a two-dimensional depthwise separable convolution with a kernel size of 1 to compress the channels to 1 / 8. After being weighted by the Relu activation function, it then passes through a two-dimensional depthwise separable convolution with a kernel size of 1 to restore the channels. Since the two-dimensional depthwise separable convolution can effectively reduce the number of parameters, improve the calculation efficiency, and at the same time maintain a strong feature expression ability, helping the network to better focus on details and local information, a two-dimensional depthwise separable convolution is adopted as the extractor of information between channels in this embodiment, focusing on the importance of features in different channels.

[0050] The weights output by the global average pooling branch and the weights output by the global max pooling branch are added together, and the Sigmoid function is used to re-weight the feature map. Finally, the re-weighted weights are multiplied by the original feature map through the broadcast mechanism and then output.

[0051] In this embodiment, the feature fusion network with the above structure can help the sparse representation ability of the model, making the model more robust and noise-resistant, and helping to improve the generalization performance of the model.

[0052] The multi-scale features extracted through step S21 need to be fused to obtain the fused features, and the fused features are for the final object detection

[0053] In step S23, the fused features pass through the detection head to obtain the predicted probability values corresponding to each category, and the maximum probability value in the predicted probability values is selected as the prediction result, and the bounding box coordinates are displayed in the image to be detected.

[0054] In this embodiment, the structure of the detection head can adopt the existing detection head of YOLOv9. Through the detection head, the corresponding predicted probability values can be output, and after screening, the bounding box coordinates are displayed in the image to be detected.

[0055] In summary, it can be seen that for the object detection method proposed in this embodiment, by improving the YOLOv9 model, optimizing the convolutional structure design during the feature extraction process, and adding an attention module during the feature fusion process, the computational resources in the model processing can be reduced, the high requirements for hardware configuration of the existing model can be solved, and the sensitivity for small object recognition can be improved.

[0056] To further prove the classification and detection performance of the object detection method in this embodiment for images, we also conducted experiments on the dataset. The experimental results show that compared with the existing YOLOv9 model architecture, the improved model has a 9.97% reduction in the number of parameters, a 6.29% increase in the inference speed, and a 0.62% and 0.91% increase in the mAP50 and mAP50-95 metrics respectively. These improvements prove that in a resource-constrained environment, through the experimentally designed attention mechanism and network structure, not only can the accuracy of the object detection model be improved, but also the number of model parameters can be significantly reduced and the computational efficiency can be increased.

[0057] Refer to Figure 6 , for the method provided in steps S21 - S23, a target detection system 60 is also provided. The system includes:

[0058] An image acquisition terminal 61 for acquiring the image to be detected within the target area.

[0059] An image processing device 62 for performing object detection on the image to be detected, obtaining at least one classification result, and displaying the bounding box coordinates in the image to be detected.

[0060] A result display terminal 63 for displaying the image to be detected and the bounding box coordinates on the image to be detected.

[0061] Refer to Figure 7 , another implementation is also provided. In this implementation, the image to be detected needs to be resized before the model processes it, so that the pixel size and resolution requirements of the image input to the feature extraction module meet the requirements of the model. Specifically, this method includes the following steps:

[0062] Step S71. Adjust the pixel size of the image to be detected to the target size through affine transformation, and adjust the resolution of the adjusted image to be detected until it meets the resolution requirements of the feature extraction module.

[0063] Step S72. Extract features of the image to be detected after size adjustment through the feature extraction network, obtain multi-scale features of the image to be detected, and fuse the multi-scale features to obtain a feature map.

[0064] Step S73. Perform attention fusion on the feature map through a feature fusion network, and obtain the fused feature through cross-stage connection.

[0065] Step S74. The fused feature obtains the predicted probability value corresponding to each category through the detection head, selects the maximum probability value among the predicted probability values as the prediction result, and displays the bounding box coordinates in the image to be detected.

[0066] Refer to Figure 8 , the above method can also be integrated into the provided terminal device 80. Due to the relatively large differences that may occur in the device due to configuration or performance, it may include one or more processors 801 and a memory 802. One or more application programs or data may be stored in the memory 802. Among them, the memory 802 can be short-term storage or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the terminal device. Further, the processor 801 can be set to communicate with the memory 802, and execute a series of computer-executable instructions in the memory 802 on the terminal device. The terminal device may also include one or more power supplies 803, one or more wired / wireless network interfaces 804, one or more input / output interfaces 805, one or more keyboards 806, etc.

[0067] In a specific embodiment, the terminal device includes a memory, and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the terminal device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions:

[0068] Extract the features of the image to be detected through a feature extraction network, obtain the multi-scale features of the image to be detected, and fuse the multi-scale features to obtain a feature map;

[0069] Perform attention fusion on the feature map through a feature fusion network, and obtain the fused feature through cross-stage connection;

[0070] The fused feature obtains the predicted probability value corresponding to each category through the detection head, selects the maximum probability value among the predicted probability values as the prediction result, and displays the bounding box coordinates in the image to be detected.

[0071] It is also used to perform the following computer-executable instructions:

[0072] Adjust the pixel size of the image to be detected to the target size through affine transformation, and adjust the resolution of the adjusted image to be detected until it meets the resolution requirements of the feature extraction module;

[0073] Extract features of the image to be detected after size adjustment through a feature extraction network to obtain multi-scale features of the image to be detected, and fuse the multi-scale features to obtain a feature map;

[0074] Perform attention fusion on the feature map through a feature fusion network, and obtain a fused feature through cross-stage connection;

[0075] The fused feature passes through a detection head to obtain the predicted probability value corresponding to each category, selects the maximum probability value among the predicted probability values as the prediction result, and displays the bounding box coordinates in the image to be detected.

[0076] The following is a specific introduction to each component of the processor:

[0077] Among them, in this embodiment, the processor is an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, for example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0078] Optionally, the processor can execute various functions by running or executing software programs stored in the memory and calling data stored in the memory, such as executing the above Figure 2 and / or Figure 7 shown method.

[0079] In a specific implementation, as an embodiment, the processor may include one or more microprocessors.

[0080] The memory is used to store the software program for executing the solution of the present application and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0081] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processing unit through the interface circuit of the processor. The embodiments of the present application do not make specific limitations on this.

[0082] It should be noted that the structure of the processor shown in this embodiment does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0083] In addition, the technical effects of the processor can refer to the technical effects of the method described in the above method embodiments and will not be elaborated here.

[0084] It should be understood that the processor in the embodiments of the present application may be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0085] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0086] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0087] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0088] It should be understood that in various embodiments of the present application, the order numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0089] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0090] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0091] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0092] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0093] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0094] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0095] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A target detection method based on YOLOv9, characterized in that: The method comprises: The image to be detected is subjected to feature extraction through a feature extraction network to obtain multi-scale features of the image to be detected, and the multi-scale features are fused to obtain a feature map; the feature extraction network includes a plurality of feature extraction modules, and the plurality of feature extraction modules are used to perform feature extraction of different scales on the image to be detected and output feature maps of different scales, and the plurality of feature maps of different scales are used to obtain more extensive context information through a downsampling module; each of the feature extraction modules includes a 1*1 convolution layer and two feature extraction submodules, and the two feature extraction submodules are used to receive two feature maps output by the 1*1 convolution layer; The feature map is passed through a feature fusion network for attention fusion, and a fusion feature is obtained through cross-stage connection; The fusion feature obtains the prediction probability value corresponding to each category through the detection head, selects the maximum probability value among the prediction probability values ​​as the prediction result, and displays the bounding box coordinates in the image to be detected.

2. The target detection method based on YOLOv9 according to claim 1, characterized in that The method also includes resizing the image to be detected, specifically including: adjusting the pixel size of the image to be detected to a target size through affine transformation, and adjusting the resolution of the adjusted image to be detected until it meets the resolution requirement of the feature extraction module.

3. The target detection method based on YOLOv9 according to claim 1, characterized in that The feature extraction network also includes multiple convolutional layers, and the initial feature map after feature extraction by each convolutional layer enters the corresponding feature extraction submodule.

4. The target detection method based on YOLOv9 according to claim 3, characterized in that The feature extraction module receives the feature map processed by the convolution layer and passes it through a 1*1 convolution layer to convert the number of input channels from c1 to c3; the features output by the 1*1 convolution layer are divided into a first feature map and a second feature map, and the number of channels corresponding to each feature map is c3 / 2; the first feature map and the second feature map are respectively input into two feature extraction sub-modules.

5. The target detection method based on YOLOv9 according to claim 4, characterized in that: The feature extraction module also includes a 1*1 convolution layer connected to the output of the feature extraction submodule.

6. The target detection method based on YOLOv9 according to claim 5, characterized in that: The feature extraction submodule includes a first depth-wise separable convolution and a second depth-wise separable convolution. The first depth-wise separable convolution is used to extract features, and the second depth-wise separable convolution is used to merge the extracted features with the features extracted by the first depth-wise separable convolution in the channel dimension; the merged features pass through a 1*1 convolutional layer connected to the output of the feature extraction submodule to convert the number of channels from c3+2*c4 to the output channel number c2.

7. The target detection method based on YOLOv9 according to claim 6, characterized in that: The feature fusion network includes a pooling module, an attention module and a fusion module; the pooling module is used to divide the feature map into multiple grids of different sizes, and perform a maximum pooling operation on each grid to extract features of different scales; the attention module is configured with an attention mechanism to update features of different scales based on attention weights to obtain updated features; the fusion module is used to connect features of different scales and fuse them through a convolutional layer to obtain fused features.

8. The target detection method based on YOLOv9 according to claim 7, characterized in that: The attention module is a multi-scale channel attention module.

9. A target detection system based on YOLOv9, characterized in that: The system comprises: The image acquisition end is used to obtain the image to be detected in the target area; An image processing device, used for performing target detection on the image to be detected, obtaining at least one classification result and displaying the coordinates of a bounding box in the image to be detected; The result display terminal is used to display the image to be detected and the coordinates of the boundary box on the image to be detected.

10. The target detection system based on YOLOv9 according to claim 9, characterized in that: The image processing device comprises: A feature extraction module is used to extract features from the resized image to be detected through a feature extraction network to obtain multi-scale features of the image to be detected, and to fuse the multi-scale features to obtain a feature map; A feature fusion module, used to perform attention fusion on the feature map through a feature fusion network, and obtain fused features through cross-stage connections; The classification module is used to obtain the predicted probability value corresponding to each category, select the maximum probability value among the predicted probability values ​​as the prediction result, and display the bounding box coordinates in the image to be detected.

Citation Information

Patent Citations

  • Underwater target detection method based on improved lightweight YOLOv9

    CN119068322A

  • Target identification method and device based on sonar image, equipment and storage medium

    CN119251631A

  • Method for detecting infrared ship target based on improved yolov7

    US20250078541A1

Cited By

  • Agricultural pest lightweight detection method and system based on dynamic channel attention

    CN120807889A

  • Lightweight detection method and system for agricultural pests based on dynamic channel attention

    CN120807889B