Target detection method and device, medium and product

By extracting and fusing features of different scales in image recognition, the error detection and missed detection problems of small-target object recognition are solved, and higher recognition accuracy is achieved.

CN120339580APending Publication Date: 2025-07-18广州算威科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419549.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing image recognition technology is difficult to effectively extract feature information of small target objects, and it is prone to false detection and missed detection.

Method used

By extracting the basic convolutional features of the input image, the first features of different scales are obtained, and the feature enhancement processing is performed to identify the target object by using the fusion features.

Benefits of technology

It improves the accuracy of feature information extraction of small target objects, reduces false detection and missed detection, and improves the accuracy of image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339580A_ABST
    Figure CN120339580A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device, a medium and a product. The target detection method comprises the following steps: extracting a basic convolution feature of an input image, and obtaining first features of different scales according to the basic convolution feature; performing feature enhancement processing on the first features to obtain second features, and fusing the second features of different scales to obtain fused features; according to the method and the device, the feature information of the small target object can be effectively extracted through multi-scale feature extraction, the influence of interference information on target recognition is reduced by using a feature enhancement and feature fusion mode, the occurrence of missing detection and false detection is reduced, and the accuracy of image recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology. Specifically, this application relates to a target detection method, device, medium, and product. Background Art

[0002] Image recognition refers to the technology of using a computer to process, analyze, and understand images to identify various different patterns of targets and objects, and it is a practical application of applying deep learning algorithms. At present, image recognition technology is generally divided into face recognition and commodity recognition. Face recognition is mainly used in security inspections, identity verification, and mobile payments; commodity recognition is mainly used in the process of commodity circulation, especially in unmanned retail fields such as unmanned shelves and intelligent retail cabinets.

[0003] In the field of image recognition, the recognition of small target objects in images (such as small target detection in remote sensing images) is often involved. However, since the small target objects occupy fewer pixels in the image, the target features are not obvious, making it difficult to extract effective feature information. Moreover, the background in the image usually contains a large amount of interfering information, such as buildings, vegetation, roads, etc., which are likely to be confused with small targets, resulting in false detection and missed detection. A method capable of effectively detecting small target objects is needed. Summary of the Invention

[0004] In view of the shortcomings of the existing methods, this application proposes a target detection method, device, medium, and product, which can solve the problem that it is difficult to effectively extract the feature information of small targets in the existing image recognition methods, and it is easy to have false detection and missed detection.

[0005] According to one aspect of the embodiments of this application, the embodiments of this application provide a target detection method, which includes:

[0006] Extract the basic convolutional features of the input image, and obtain first features of different scales according to the basic convolutional features;

[0007] Obtain second features corresponding to the first features, and fuse the second features of different scales to obtain a fused feature, where the second feature is obtained by performing feature enhancement processing on the first feature;

[0008] Use the fused feature to identify the target object in the input image.

[0009] In a possible implementation, the obtaining first features of different scales according to the basic convolutional features includes:

[0010] Use the basic convolutional features and the first processing module to extract the first features of the bottom layer scale of the input image;

[0011] According to the first feature of the underlying scale, the second processing module extracts the second feature of the middle scale, and uses the second feature of the middle scale and the third processing module to extract the third fusion feature.

[0012] In a possible implementation, the first processing module includes two CSP modules and one CBS module. The extracting the first feature of the underlying scale of the input image by using the basic convolutional feature and the first processing module includes:

[0013] Using two CSP modules to perform feature extraction on the basic convolutional feature to obtain a feature extraction result;

[0014] Inputting the feature extraction result into the CBS module, and obtaining the first feature of the underlying scale according to the result output by the CBS module.

[0015] In a possible implementation, the structure of the first processing module is the same as that of the second processing module, and the structure of the third processing module is different from that of the second processing module.

[0016] In a possible implementation, the obtaining the second feature corresponding to the first feature includes:

[0017] Using a feature enhancement module to perform feature enhancement processing on the first feature to obtain the second feature, and each feature enhancement module corresponds to processing the first feature of one scale;

[0018] The feature enhancement processing includes:

[0019] Using two branches of the feature enhancement module to perform feature extraction on the first feature to obtain two branch features;

[0020] Using the feature enhancement module to fuse the two branch features to obtain the second feature. In a possible implementation, the fusion feature includes a first fusion feature, a second fusion feature, and a third fusion feature. The fusing the second features of different scales to obtain a fusion feature includes:

[0021] Based on a feature fusion formula to obtain the second features of different scales, and the feature fusion formula is:

[0022]

[0023] In the formula, X ′ 1 is the first fusion feature, represents the multiplication operation, CSP represents the CSP module, X1 represents the second feature of the underlying scale, X2 represents the second feature of the middle scale, X3 represents the second feature of the high scale, X ′ 2 is the second fusion feature, X′ 3 represents the third fusion feature.

[0024] In a possible implementation, the using the fusion feature to identify a target object in the input image includes:

[0025] Inputting each fusion feature into a corresponding target head detection module, and obtaining information of the target object according to an output result of the target head detection module, where the target head detection module includes a CSP convolution module, a convolution module, and a two-dimensional convolution module connected in sequence, and the information includes a category of the target object and a position of a bounding box.

[0026] According to one aspect of an embodiment of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the method as described above.

[0027] According to one aspect of an embodiment of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the steps of the method as described above are implemented.

[0028] According to one aspect of an embodiment of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method as described above are implemented.

[0029] The beneficial technical effects brought by the technical solution provided by the embodiment of the present application include:

[0030] A target detection method provided by the present application has the beneficial effect that basic convolution features of an input image are extracted, and first features of different scales are obtained according to the basic convolution features; the first features are subjected to feature enhancement processing to obtain second features, and the second features of different scales are fused to obtain fusion features; the fusion features are used to identify a target object in the input image. The present application can effectively extract feature information of small target objects through multi-scale feature extraction, and reduce the influence of interference information on target recognition by means of feature enhancement and feature fusion, reduce the occurrence of missed detections and false detections, and improve the accuracy of image recognition.

[0031] Additional aspects and advantages of the present application will be given in part in the following description, which will become apparent from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0033] Figure 1 is a flowchart of the target detection method provided by the embodiment of the present application;

[0034] Figure 2 It is the overall block diagram of object detection provided by the embodiment of the present application;

[0035] Figure 3 It is the structural diagram of the feature enhancement module provided by the embodiment of the present application;

[0036] Figure 4 It is the structural diagram of the object detection head module provided by the embodiment of the present application;

[0037] Figure 5 It is the structural diagram of the CSP module provided by the embodiment of the present application;

[0038] Figure 6 It is the structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0039] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the implementation manners described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0040] Those skilled in the art of the present technology can understand that, unless specifically stated, the "the" and "this" used here may also include the plural form. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations, etc. supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here means at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0041] To make the purpose, technical solutions, and advantages of the present application clearer, the implementation manners of the present application will be further described in detail below with reference to the accompanying drawings.

[0042] The embodiment of the present application provides an object detection method, and the objects to which this method is applied can be mobile phones, laptop computers, servers, cloud platforms, cameras, and other objects that can be used for image recognition.

[0043] Optionally, the object to which the method of the present application is applied uses Mamba-YOLO to implement the object detection method of the present application. Among them, Mamba-YOLO may include a convolutional feature extraction module, a feature enhancement module, a feature fusion module, and a multi-head prediction module connected in sequence. Among them, the convolutional feature extraction module may include two CBS modules, a first processing module, a second processing module, and a third processing module. These two CBS modules are connected to the first processing module and transmit the extracted basic convolutional features to the first processing module. The multi-head prediction module includes three object head detection modules.

[0044] Optionally, the convolutional feature extraction module may form a convolutional neural network through two CBS modules, a first processing module, a second processing module, and a third processing module, use this convolutional neural network as a backbone network, and extract first features of different scales in the input image through the backbone network. It is also possible to form a lightweight convolutional neural network using this convolutional neural network. This lightweight convolutional neural network can be obtained by extracting important convolutional layers in the convolutional neural network.

[0045] As Figures 1-5 shown, the object detection method of the present application includes:

[0046] S101: Extract the basic convolutional features of the input image, and obtain first features of different scales according to the basic convolutional features.

[0047] Optionally, the input image may be a remote sensing image, or an image taken by a mobile phone, a camera, or other devices. The input image may include an object with the number of occupied pixels less than a preset number, and the method of the present application is used to identify the object.

[0048] Optionally, when extracting the basic convolutional features, two CBS (convolutional layer + batch normalization layer + SiLU activation function layer) modules may be used to extract features from the input image to obtain the basic convolutional features. Among them, the CBS module includes a convolutional layer, a batch normalization layer, and a SiLU (activation function) layer connected in sequence, and the feature extraction and linear transformation of the input image are realized through this CBS module.

[0049] Optionally, obtaining first features of different scales according to the basic convolutional features includes: using the basic convolutional features and the first processing module to extract the first features of the underlying scale of the input image; according to the first features of the underlying scale, the second processing module extracts the second features of the middle scale, and uses the second features of the middle scale and the third processing module to extract the third fusion features.

[0050] Optionally, the structures of the first processing module and the second processing module may be the same, and the structure of the third processing module may be different from that of the second processing module. The specific scale sizes of the bottom layer scale, the middle layer scale, and the high layer scale may be determined according to conditions such as the resolution of the input image, the size, the size of the target object in the input image, and the image recognition accuracy requirements.

[0051] Optionally, the first processing module includes two CSP (Cross Stage Partial Connections) modules and one CBS module. Using the basic convolutional features, the first processing module extracts the first features of the bottom layer scale of the input image, including: using the two CSP modules to perform feature extraction on the basic convolutional features to obtain a feature extraction result; inputting the feature extraction result into the CBS module, and obtaining the first features of the bottom layer scale according to the result output by the CBS module. Through the processing of the CSP module, the image recognition ability is improved and the computational amount is reduced.

[0052] In one embodiment, the structure of the second processing module is the same as that of the first processing module. The first features of the bottom layer scale are transmitted to the second processing module. Two CBS modules in the second processing module process the first features and transmit the processed first features to one CBS module. Through the processing of this CBS module, the first features of the middle layer scale are obtained.

[0053] Optionally, the CSP module includes a convolutional layer, a normalization layer, a LeakyReLU (Rectified Linear Unit function), a residual unit, and a splicing module. The CSP module in the second processing module divides the first features into two parts, uses the residual unit to process one part to increase the depth of the network and improve the target recognition ability, and uses the convolutional layer to perform convolutional processing on the other part. The processing results of the two parts are spliced through the splicing module, and the spliced result is processed using the normalization layer and LeakyReLU, and output to obtain the processing result.

[0054] In one embodiment, the third processing module may include three CSP modules and one SPPF (Spatial Pyramid Pooling - Fast) module (a feature pyramid module for pooling operations at different scales, splicing feature maps at different scales together to improve the detection ability for targets of different sizes). The first features of the middle layer scale are transmitted to the three CSP modules in the third processing module. After being processed by the three CSP modules, they are output to the SPPF module, and the features output by the SPPF are determined as the first features of the high layer scale. Among them, after obtaining the first features, the first features of each scale are output for subsequent processing.

[0055] S102: Obtain the second feature corresponding to the first feature, and fuse the second features of different scales to obtain a fused feature.

[0056] Optionally, obtaining the second feature corresponding to the first feature includes: performing feature enhancement processing on the first feature using a feature enhancement module to obtain the second feature, and each feature enhancement module corresponds to processing the first feature of one scale; the feature enhancement processing includes: extracting two branch features from the first feature using two branches of the feature enhancement module; and fusing the two branch features using the feature enhancement module to obtain the second feature.

[0057] Optionally, the number of feature enhancement modules can be three, which is the same as the number of first features extracted from the first image. When the number of first features is four or other numbers, the number of feature enhancement modules can also change following the number of first features.

[0058] Optionally, the feature enhancement module can be a Mamba feature enhancement module, which includes a token generator and two branches. One branch includes a Norm normalization layer, a one-dimensional convolutional layer, a LeakyReLU non-linear activation function (σ), a linear layer (Linear), a selective state space model (S6), and a Norm normalization layer connected in sequence. When processing the first feature, first generate multiple tokens (meaningful feature vectors) through the token generator, and the tokens are transmitted to the Norm normalization layer for normalization operation, then extract convolutional features through the one-dimensional convolutional layer, use the LeakyReLU non-linear activation function to enhance the non-linear representation ability of the convolutional features, perform a linear transformation through Linear, use the selective state space model to introduce a selective mechanism to dynamically adjust the attention to different parts of the input features, so as to more efficiently capture long-range dependencies, and use the Norm normalization layer directly connected to S6 to normalize the result output by S6. The other branch includes Linear and σ. When processing the first feature, Linear in this branch receives the tokens output by the token generator, processes the tokens, and transmits the processed result to σ, and fuses the result output by σ with the result output by the other branch to obtain the second feature.

[0059] Optionally, since the semantic information included in the second features of different scales is different, it is necessary to perform fusion processing on the semantic information of different scales. Among them, the fused feature includes a first fused feature, a second fused feature, and a third fused feature. Fusing the second features of different scales to obtain the fused feature includes: obtaining the second features of different scales based on the feature fusion formula, and the feature fusion formula is:

[0060]

[0061] In the formula, X ′ 1 is the first fusion feature, represents a multiplication operation, CSP represents the CSP module, X1 represents the second feature of the bottom layer scale, X2 represents the second feature of the middle layer scale, X3 represents the second feature of the high layer scale, X ′ 2 represents the second fusion feature, X ′ 3 represents the third fusion feature.

[0062] Optionally, the feature fusion module performs feature fusion based on the above feature fusion formula. Different-level second features are aggregated through this feature fusion formula to enhance the semantic representation of small targets. Specifically, as Figure 2 shown, the feature fusion module includes multiple CSP modules (every two CSP modules form a group) and a multiplication module that performs a multiplication operation. These modules form a double-layer pyramid network, and the effect of feature fusion is improved by means of channel reweighting. Among them, the pyramid network transmits semantic features through two paths: top-down and bottom-up, so as to fuse more information of different scales.

[0063] Optionally, after obtaining the fusion feature, the fusion feature is output to the multi-head detection module. The number of multi-head detection modules is the same as the number of types of fusion features, and different fusion features are output to the corresponding multi-head detection modules for processing.

[0064] S103: Identify the target object in the input image using the fusion feature.

[0065] Optionally, identifying the target object in the input image using the fusion feature includes: inputting each fusion feature into the corresponding target head detection module, and obtaining the information of the target object according to the output result of the target head detection module. The target head detection module includes a CSP convolution module, a convolution module, and a two-dimensional convolution module connected in sequence. The information includes the category of the target object and the position of the bounding box.

[0066] Optionally, the target head detection module includes a CSP convolution module, a convolution module, and a two-dimensional convolution module connected in sequence. The convolution module and the two-dimensional convolution module form a branch, and the output result of the CSP convolution module is transmitted to the two branches respectively. One branch calculates the bounding box loss, and the other branch calculates the classification loss. By synthesizing the output results of different target head detection modules, the recognition information of the target object is obtained.

[0067] Optionally, before the fusion feature is transmitted to the target head detection module, the spatial context awareness module can also be used to capture the global context information in the fusion feature, and the final feature map is output based on this global context awareness information. The target head detection module performs target detection on this final feature map to detect whether there is a target object, and after detecting the target object, outputs the position and category information of the target object.

[0068] In one embodiment, the method of this application is verified based on the USOD dataset, and the small object dataset USOD is constructed based on UNICORN2008. Using the visible light data of UNICORN2008, the dataset USOD is formed by filtering, segmentation, and manually adding annotations of small vehicle objects. USOD contains a total of 3000 images, among which there are 43378 vehicle instances. The ratio of the training set to the test set formed by USOD is 7:3. Among them, objects with a size less than 16×16 account for 96.3%, and objects with a size less than 32×32 account for 99.9%. In the experiment, the stochastic gradient descent optimizer is used, and the initial learning rate is set to 0.01, the momentum is 0.937, and the weight decay is 0.0005 to obtain the learning parameters. The batch size during training is set to 32. The normalized Wasserstein distance (NWD) loss is added as a supplement to the bounding box loss in the loss function of YOLOv5 of Mamba-YOLO. An adjustment weight is introduced for the CIOU loss and NWD loss in YOLOv5, and it is set to 0.5. The mean average precision (Map) is used as the standard evaluation metric, and it is divided into Map50 and Map50:95 according to different IOU. In addition, in order to measure the resource consumption degree of the model, the number of model parameters is also used as an evaluation metric. The comparative experimental results are shown in Table 1:

[0069] Table 1

[0070]

[0071]

[0072] The experimental results show that the solution of this application is superior to the existing solutions in terms of accuracy, recall, and various Map indicators. In addition, the number of parameters is also lower than several relatively advanced solutions, and higher detection accuracy can be achieved with fewer parameters, and the model of the present invention can be deployed under limited computing resources, which is widely applicable to actual scenarios.

[0073] The object detection method of the embodiment of this application extracts the basic convolutional features of the input image, obtains the first features of different scales according to the basic convolutional features; performs feature enhancement processing on the first features to obtain the second features, fuses the second features of different scales to obtain the fused features; uses the fused features to identify the target objects in the input image. This application can effectively extract the feature information of small target objects through multi-scale feature extraction, and use the methods of feature enhancement and feature fusion to reduce the impact of interference information on target recognition, reduce the occurrence of missed detections and false detections, and improve the accuracy of image recognition.

[0074] Based on the same inventive concept, the embodiment of this application provides an electronic device, such asFigure 6 As shown Figure 6 The electronic device 2000 shown in Figure 6 includes a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are communicatively connected, such as connected through a bus 2002.

[0075] The processor 2001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0076] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0077] The memory 2003 can be a ROM (Read-Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (random access memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read-Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0078] Optionally, the electronic device 2000 may further include a communication unit 2004. The communication unit 2004 can be used for receiving and sending signals. The communication unit 2004 can allow the electronic device 2000 to communicate with other devices wirelessly or wirelessly to exchange data. It should be noted that in practical applications, the communication unit 2004 is not limited to one.

[0079] Optionally, the electronic device 2000 may further include an input unit 2005. The input unit 2005 can be used for receiving input digital, character, image, and / or sound information, or generating key signal inputs related to the user settings and function controls of the electronic device 2000. The input unit 2005 can include, but is not limited to, one or more of a touch screen, a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), a trackball, a mouse, a joystick, a photographing device, a pickup, etc.

[0080] Optionally, the electronic device 2000 may further include an output unit 2006. The output unit 2006 can be used for outputting or presenting the information processed by the processor 2001. The output unit 2006 can include, but is not limited to, one or more of a display device, a speaker, a vibration device, etc.

[0081] Although the figure shows an electronic device 2000 with various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0082] Optionally, the memory 2003 is used to store a computer program for executing the solution of this application, and is controlled by the processor 2001 to execute. The processor 2001 is used to execute the computer program stored in the memory 2003 to implement the steps of any method provided in the embodiments of this application.

[0083] Based on the same inventive concept, the embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided in this application / implements the steps of various alternative embodiments of the method provided in this application.

[0084] Based on the same inventive concept, the embodiments of this application provide a computer program product, which includes a computer program. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided in this application / implements the steps of various alternative embodiments of the method provided in this application.

[0085] Those skilled in the art of this technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the related technologies that are the same as those disclosed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0086] In the description of this application, the directions or positional relationships indicated by the words "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are the exemplary directions or positional relationships based on the drawings, which are for the convenience of describing or simplifying the embodiments of this application, rather than indicating or implying that the device or component referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.

[0087] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0088] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0089] In the description of this specification, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0090] The above are only some implementation manners of the present application. It should be pointed out that for those of ordinary skill in the technical field, without departing from the technical concept of the present application's solution, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A target detection method, characterized in that, The method includes: Extracting the basic convolutional features of the input image, and obtaining first features of different scales according to the basic convolutional features; Obtaining second features corresponding to the first features, and fusing the second features of different scales to obtain a fused feature, where the second feature is obtained by performing feature enhancement processing on the first feature; Identifying the target object in the input image by using the fused feature.

2. The object detection method according to claim 1, wherein The obtaining of first features of different scales according to the basic convolutional features includes: Using the basic convolutional features and the first processing module to extract the first features of the bottom layer scale of the input image; According to the first features of the bottom layer scale and the second processing module, extracting the second features of the middle layer scale, and using the second features of the middle layer scale and the third processing module to extract the third fused feature.

3. The object detection method according to claim 2, characterized in that The first processing module includes two CSP modules and one CBS module. The using of the basic convolutional features and the first processing module to extract the first features of the bottom layer scale of the input image includes: Using two CSP modules to perform feature extraction on the basic convolutional features to obtain a feature extraction result; Inputting the feature extraction result into the CBS module, and obtaining the first features of the bottom layer scale according to the result output by the CBS module.

4. The object detection method according to claim 2 or 3, characterized in that, The structures of the first processing module and the second processing module are the same, and the structure of the third processing module is different from that of the second processing module.

5. The object detection method according to claim 1, wherein The obtaining of the second features corresponding to the first features includes: Using a feature enhancement module to perform feature enhancement processing on the first feature to obtain the second feature, and each feature enhancement module corresponds to processing the first feature of one scale; The feature enhancement processing includes: Using two branches of the feature enhancement module to perform feature extraction on the first feature to obtain two branch features; Using the feature enhancement module to fuse the two branch features to obtain the second feature.

6. The object detection method according to claim 1, characterized in that, The fused feature includes a first fused feature, a second fused feature, and a third fused feature. The fusing of the second features of different scales to obtain the fused feature includes: Obtaining the second features of different scales based on a feature fusion formula, and the feature fusion formula is: Wherein, X ′ 1 is the first fusion feature, represents a multiplication operation, CSP represents the CSP module, X1 represents the second feature at the underlying scale, X2 represents the second feature at the middle scale, X3 represents the second feature at the high scale, X ′ 2 represents the second fusion feature, X ′ 3 represents the third fusion feature.

7. The object detection method according to claim 1, wherein The using of the fused feature to identify the target object in the input image includes: Inputting each fused feature into a corresponding target head detection module, and obtaining information about the target object according to the output result of the target head detection module. The target head detection module includes a CSP convolutional module, a convolutional module, and a two-dimensional convolutional module connected in sequence. The information includes the category of the target object and the position of the bounding box.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.