Target detection method, electronic equipment, storage medium and program product
By adding a new target feature fusion layer to the YOLO network, the feature extraction capability is improved, the problem of insufficient detection accuracy of the YOLO network is solved, and higher target detection performance is achieved.
Patent Information
- Application Number
- CN202511081199.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-04
AI Technical Summary
The existing YOLO network has insufficient detection accuracy in target detection scenarios, making it difficult to meet the high requirements of users.
At least one target feature fusion layer is added to the YOLO network, through which the feature map is fused, the feature extraction capability is improved, and the target YOLO network is formed to train the target detection model.
More feature fusion is achieved through the newly added target feature fusion layer, improving the accuracy and performance of target detection.
Smart Images

Figure CN120580697A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a target detection method, electronic equipment, storage medium, and program product. Background Art
[0002] In related technologies, the YOLO network is often used for object detection. For example, in fiber-to-the-room (FTTR) scenarios, after a camera captures images or videos, the YOLO network can be used to identify various objects in the images or videos, such as people, pets, household items, and appliances.
[0003] As target detection scenarios become more complex, users have higher requirements for the accuracy of target detection. Therefore, it is necessary to improve the target detection performance of the YOLO network. Summary of the Invention
[0004] The present application provides a target detection method, electronic device, storage medium, and program product for at least solving the problem of how to improve the target detection performance of the YOLO network.
[0005] To solve the above technical problems, this application is implemented as follows: In a first aspect, a target detection method is provided, comprising: Acquire a first image to be subjected to target detection; Performing object detection on the first image according to a pre-trained object detection model to obtain a detection result; Among them, the target detection model is obtained based on the target YOLO network training, and the target YOLO network includes a backbone network layer, a neck network layer and a head network layer. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature map output by the backbone network layer, and the target feature fusion layer is used to fuse the feature map obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input of the head network layer.
[0006] In a second aspect, an electronic device is provided, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method as described in the first aspect.
[0007] According to a third aspect, a computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to the first aspect.
[0008] In a fourth aspect, a computer program product is provided, comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps in the method described in the first aspect.
[0009] In an embodiment of the present application, when performing target detection, a target detection model is trained using a target YOLO network, and at least one target feature fusion layer is added to the neck network of the target YOLO network. Each target feature fusion layer can fuse the feature map obtained before the target feature fusion layer, and the feature layer output by the last target feature fusion layer is used as the input of the head network layer in the YOLO network. In this way, by adding at least one target feature fusion layer to the YOLO network to perform feature fusion on the feature map, more feature fusion can be achieved. After the target detection model is obtained based on the target YOLO network training, when performing target detection on the first image according to the target detection model, more image information can be extracted based on the newly added target feature fusion layer, so that when performing target detection based on the extracted image information, the detection accuracy can be effectively improved, thereby improving the target detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0011] Figure 1 This is a flowchart of a target detection method according to an embodiment of the present application; Figure 2 It is a structural diagram of the YOLO network in related technology; Figure 3 This is a schematic diagram of the structure of the target YOLO network in one embodiment of the present application; Figure 4 This is a schematic diagram of the structure of the target feature fusion layer of an embodiment of the present application; Figure 5 2 is a schematic diagram of a YOLO network processing a first image in the related art; Figure 6 is a schematic diagram of a target YOLO network processing a first image according to an embodiment of the present application; Figure 7 This is a schematic structural diagram of an electronic device according to an embodiment of the present application; Figure 8It is a structural diagram of a target detection device according to an embodiment of the present application. DETAILED DESCRIPTION
[0012] In order to help those skilled in the art better understand the technical solutions of this application, the following will clearly and completely describe the technical solutions of this application in conjunction with the drawings of one or more embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0013] The terms "first," "second," and the like in this application and the claims are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that such terms are interchangeable where appropriate so that this application can be implemented in sequences other than those illustrated or described herein. In addition, the term "and / or" in this application and the claims refers to at least one of the connected objects, and the character " / " generally indicates that the connected objects are in an "or" relationship.
[0014] It should be noted that the target detection solution provided in the embodiments of this application can be applied to FTTR scenarios to identify various targets in images or videos captured by cameras, such as human figures, pets, household items, and electrical appliances. In addition, it can also be applied to other target detection scenarios, including but not limited to: (1) Intelligent security systems: such as video surveillance and intrusion detection, which can detect potential security threats more accurately.
[0015] (2) Autonomous driving technology: such as vehicle surrounding environment perception and obstacle recognition, which can achieve real-time and efficient target detection and improve the safety and reliability of autonomous driving.
[0016] (3) Industrial automation: such as quality control and anomaly detection on the production line, which can achieve more accurate quality control and anomaly detection and improve production efficiency.
[0017] (4) Medical imaging analysis: Auxiliary diagnostic tools can improve the early detection rate of diseases, such as detecting small lesions earlier and improving the early diagnosis rate of diseases.
[0018] (5) UAV navigation: Such as real-time target tracking and obstacle avoidance, which can achieve real-time and efficient target tracking and obstacle avoidance, and enhance the autonomous navigation capability of UAVs.
[0019] (6) Robot vision: such as autonomous navigation and object recognition, which can achieve more accurate object recognition and autonomous navigation, and improve the intelligence level of robots.
[0020] Optionally, in some implementations, the target detection solution provided in the embodiments of the present application may be combined with other technologies to further expand its application scope, including but not limited to: (1) Integration with edge computing: By deploying the target detection solution of this application on edge devices, low-latency and high-precision target detection can be achieved, which is suitable for various edge computing scenarios.
[0021] (2) Integration with cloud computing: By deploying the target detection solution of this application on the cloud, large-scale data processing and model training can be achieved, which is suitable for big data analysis and large-scale target detection tasks.
[0022] (3) Integration with the Internet of Things: By integrating the target detection solution of this application into the Internet of Things devices, intelligent perception and decision-making can be achieved, thereby improving the intelligence level of the Internet of Things system.
[0023] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0024] Figure 1 It is a flowchart of a target detection method according to an embodiment of the present application. Figure 1 The object detection method shown is described below.
[0025] Step S102: Acquire a first image to be subjected to target detection.
[0026] The first image may be obtained by image acquisition or video capture by an image acquisition device. The first image may or may not contain the target object. In the embodiment of the present application, target detection on the first image may be performed by detecting whether the first image contains the target object.
[0027] Step S104: performing target detection on the first image according to the pre-trained target detection model to obtain a detection result.
[0028] After acquiring the first image, the first image can be used as input to a pre-trained target detection model, and the target detection model can output a corresponding detection result.
[0029] In an embodiment of the present application, the target detection model is obtained based on the training of the target YOLO network. The target YOLO network is a network obtained by improving the structure of the YOLO network in the related art, and the improvement here includes adding at least one target feature fusion layer in the YOLO network. Specifically, the target YOLO network includes a backbone network layer (which can be expressed as a backbone network layer), a neck network layer (which can be expressed as a neck network layer) and a head network layer (which can be expressed as a head network layer). The backbone network layer is used to extract features from the input image, and the extracted feature map is used as the input of the neck network layer. The structure of the backbone network layer is the same as that of the backbone network layer in the YOLO network, and will not be described in detail here. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature maps output by the backbone network layer. Its network structure is the same as that of the neck network layer in the YOLO network, and will not be described in detail here. At least one target feature fusion layer is a newly added network structure. When there are multiple target feature fusion layers, the multiple target feature fusion layers are connected in series. Each target feature fusion layer is used to fuse the feature maps obtained by the previous target feature fusion layer. The feature map output by the last target feature fusion layer serves as the input of the head network layer. The head network layer is used to perform target detection based on the feature map output by the last target feature fusion layer and output the target detection result. The network structure of the head network layer is the same as that of the head network layer in the YOLO network and will not be described in detail here.
[0030] In order to understand the structure of the YOLO network and the structure of the target YOLO network, the following will be Figure 2 and Figure 3 The embodiment shown is described as an example.
[0031] Figure 2 It is a structural diagram of the YOLO network in related technology. Figure 2 The YOLO network shown includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer is used to extract features from the input image, generating multiple feature maps, which serve as input to the neck network layer. The neck network layer includes a feature fusion layer, which is used to fuse the multiple feature maps output by the backbone network layer to generate multiple fused feature maps, which serve as input to the head network layer. The head network layer processes these multiple fused feature maps to output object detection results.
[0032] Figure 3 Schematic diagram of the structure of the target YOLO network according to an embodiment of the present application. Figure 3The target YOLO network shown in the figure includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer is used to extract features from the input image and obtain multiple feature maps, which are used as inputs to the neck network layer. The neck network layer includes the original feature fusion layer (i.e. Figure 2 The network consists of two target feature fusion layers (shown here as an example, but one or more target feature fusion layers are also possible). The two target feature fusion layers are newly added. The original feature fusion layer is used to fuse the multiple feature maps output by the backbone network layer to obtain multiple fused feature maps. The first target feature fusion layer is used to fuse all or part of the feature maps output by the backbone network layer and the multiple feature maps output by the original feature fusion layer to obtain multiple fused feature maps. The second target feature fusion layer is used to fuse all or part of the feature maps output by the backbone network layer, the multiple feature maps output by the original feature fusion layer, and the multiple feature maps output by the first target feature fusion layer to obtain multiple fused feature maps. The multiple feature maps output by the second target feature fusion layer serve as input to the head network layer. The head network layer processes the multiple fused feature maps to obtain target detection results.
[0033] In the embodiment of the present application, the backbone network layer and the neck network layer of the target YOLO network each include one or more convolutional layers. For each target feature fusion layer in the neck network layer, the target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer, and the target feature fusion layer can be used to fuse the feature maps output by the convolutional layer before the target feature fusion layer.
[0034] In some embodiments, the target feature fusion layer is used to fuse the feature maps output by the convolution layer before the target feature fusion layer, specifically, it can be used to fuse the feature maps output by the target convolution layer. The target convolution layer is part or all of the convolution layers in the multiple convolution layers before the target feature fusion layer. The multiple convolution layers before the target feature fusion layer include the convolution layers in the backbone network layer and the convolution layers in each feature fusion layer between the backbone network layer and the target feature fusion layer (including the original feature fusion layer and other target feature fusion layers between the original feature fusion layer and the target feature fusion layer). That is to say, when the target feature fusion layer performs feature fusion, it can fuse the feature maps output by all the convolution layers before the target feature fusion layer, so as to maximize the target detection performance, or it can fuse the feature maps output by some of the convolution layers before the target feature fusion layer, so as to reduce the complexity while ensuring the performance as much as possible.
[0035] by Figure 3Taking the target YOLO network shown in the figure as an example, the backbone network layer, the original feature fusion layer and the two target feature fusion layers all include multiple convolutional layers ( Figure 3 Not shown), which can be represented as convolution layer A, convolution layer B, convolution layer C and convolution layer D respectively. For the first target feature fusion layer, the feature map output by convolution layer A and the feature map output by convolution layer B can be fused, or the feature map output by convolution layer B can be fused, or the feature map output by some convolution layers in convolution layer A and the feature map output by some convolution layers in convolution layer B can be fused. For the second target feature fusion layer, the feature map output by convolution layer A, the feature map output by convolution layer B and the feature map output by convolution layer C can be fused, or the feature map output by convolution layer A and the feature map output by convolution layer C can be fused, or the feature map output by some convolution layers in convolution layer B and the feature map output by convolution layer C can be fused. Examples are not given here one by one.
[0036] In some embodiments, the feature maps output by the target convolutional layer may include feature maps at multiple levels. Feature maps at different levels have different depths, resolutions, and semantic information. The target feature fusion layer is used to fuse the feature maps output by the target convolutional layer. The target feature fusion layer may be used to fuse feature maps of the same level output by the target convolutional layer.
[0037] For example, the target convolutional layer outputs feature maps of three different levels: a 20×20 feature map, a 40×40 feature map, and an 80×80 feature map. The target feature fusion layer fuses the feature maps of the same level output by the target convolutional layer, such as the 20×20 feature map, the 40×40 feature map, and the 80×80 feature map.
[0038] In order to achieve the purpose of fusing the feature maps output by the target convolution layer by the target feature fusion layer, in some embodiments, the target feature fusion layer may include a first path from bottom to top and a second path from top to bottom. The first path is used to transfer the feature information in the low-resolution feature map to the high-resolution feature map and fuse it with the feature map of the same level obtained before the first path. The second path is used to transfer the feature information in the high-resolution feature map to the low-resolution feature map and fuse it with the feature map of the same level obtained before the second path. When fusing feature maps of the same level, all feature maps of the same level may be fused, or some feature maps of all feature maps of the same level may be fused.
[0039] In some embodiments, the first path includes an upsampling layer, a multi-input connection layer, and a convolution layer from bottom to top. The upsampling layer is used to upsample the input low-resolution feature map to obtain a high-resolution feature map. The multi-input connection layer is used to splice the feature map sampled by the upsampling layer with the feature map of the same level obtained before the multi-input connection layer. The convolution layer is used to fuse the feature maps spliced by the multi-input connection layer. The number of upsampling layers, multi-input connection layers, and convolution layers can be one or more. In the case where there are multiple upsampling layers, multi-input connection layers, and convolution layers, one upsampling layer, one multi-input connection layer, and one convolution layer can constitute a subpath, and the first path can include multiple such subpaths.
[0040] In some embodiments, the second path includes, from top to bottom, a downsampling layer, a multi-input connection layer, and a convolutional layer. The downsampling layer is used to downsample the input high-resolution feature map to obtain a low-resolution feature map. The multi-input connection layer is used to splice the feature map sampled by the downsampling layer with the feature map of the same level obtained previously by the multi-input connection layer. The convolutional layer is used to fuse the feature maps spliced by the multi-input connection layer. The number of downsampling layers, multi-input connection layers, and convolutional layers can be one or more. When there are multiple downsampling layers, multi-input connection layers, and convolutional layers, one downsampling layer, one multi-input connection layer, and one convolutional layer can constitute a subpath, and the second path can include multiple such subpaths. Optionally, the number of subpaths included in the second path can be the same as the number of subpaths included in the first path. In this way, when performing feature fusion, the second path can fuse feature maps of different levels output by the first path separately.
[0041] To understand the structure of the target feature fusion layer, please refer to Figure 4 . Figure 4 Take the example of including a target feature fusion layer in the YOLO network. Figure 4 The target feature fusion layer in
[15] includes a first path and a second path. The first path consists of two subpaths from bottom to top, each of which includes an upsampling layer, a multi-input connection layer, and a convolutional layer. The second path consists of two subpaths from top to bottom, each of which includes a downsampling layer, a multi-input connection layer, and a convolutional layer.
[0042] For the first path, assuming the input image is a 20×20 feature map, the first path processes this 20×20 feature map as follows: The upsampling layer in the first subpath upsamples the 20×20 feature map to a 40×40 feature map, essentially upsampling the low-resolution feature map to a high-resolution feature map. This 40×40 feature map is then fed into the multi-input connection layer, which concatenates the 40×40 feature map with the previously obtained 40×40 feature map (i.e., all or part of the 40×40 feature map output by the convolutional layer preceding the first path). The concatenated feature maps are then fused through the convolutional layer to produce a fused 40×40 feature map. The fused 40×40 feature map is input into the second subpath of the first path. After upsampling in the upsampling layer, an 80×80 feature map is obtained. This 80×80 feature map is input into the multi-input connection layer, which concatenates this 80×80 feature map with the previously obtained 80×80 feature map (i.e., all or part of the 80×80 feature map output by the convolutional layer before the first path). The concatenated feature map is fused in the convolutional layer to obtain a fused 80×80 feature map. This fused 80×80 feature map serves as the input of the second path.
[0043] In the second path, processing of the fused 80×80 feature map involves: The downsampling layer in the first subpath downsamples the 80×80 feature map to obtain a 40×40 feature map, i.e., downsampling the high-resolution feature map to obtain a low-resolution feature map. This 40×40 feature map is then fed into a multi-input connection layer, which concatenates the 40×40 feature map with the previously obtained 40×40 feature map (i.e., all or part of the 40×40 feature map output by the convolutional layer preceding the second path). The concatenated feature maps are then fused through a convolutional layer to obtain a fused 40×40 feature map. The fused 40×40 feature map is input into the second subpath in the second path. After downsampling by the downsampling layer, a 20×20 feature map is obtained. The 20×20 feature map is input into the multi-input connection layer. The multi-input connection layer splices the 20×20 feature map with the previously obtained 20×20 feature map (that is, all or part of the 20×20 feature map output by the convolutional layer before the second path). The spliced feature maps are fused by the convolutional layer to obtain a fused 20×20 feature map.
[0044] The 80×80 feature map obtained by the first path, the 40×40 feature map output by the convolutional layer of the second path, and the 20×20 feature map are used as the input of the head network layer in the YOLO network.
[0045] It should be noted that in Figure 4In the second path shown, when the multi-entry connection layer fuses feature maps, in addition to fusing all or part of the feature maps output by the convolution layer before the target feature fusion layer (that is, before the first path), it can also fuse the feature maps output by the previous convolution layer in the target feature fusion layer, that is, it can also fuse the feature maps output by the convolution layer in the first path. For example Figure 4 In the second path shown, for the first multi-entry connection layer from top to bottom, when performing feature fusion, in addition to fusing all or part of the 40×40 feature maps output by the convolutional layer before the target feature fusion layer (that is, before the first path), the 40×40 feature maps output by the convolutional layer in the first path can also be fused.
[0046] As mentioned above, the target feature fusion layer can fuse the feature maps output by the target convolution layer (part or all of the convolution layers before the target feature fusion layer), and which convolution layers’ output feature maps are fused may affect the target detection effect. Therefore, before performing target detection on the first image according to the pre-trained target detection model, it is possible to first determine which convolution layers’ output feature maps are fused by each target feature fusion layer in the target detection model, that is, before performing target detection on the first image according to the pre-trained target detection model, it is possible to first determine the target convolution layer corresponding to each target feature fusion layer.
[0047] In some embodiments, before performing object detection on the first image according to the pre-trained object detection model, at least one of the following may be included: According to the preset configuration information, determine the target convolution layer corresponding to each target feature fusion layer; A target convolution layer corresponding to each target feature fusion layer is determined based on the detection information of the first image.
[0048] That is, when determining the target convolution layer corresponding to each target feature fusion layer, it can be determined by at least one of the above two methods.
[0049] In some embodiments, the target feature fusion layer may include multiple multi-input connection layers, each of which is used to stitch feature maps, such as Figure 4The target feature fusion layer shown includes four multi-entry connection layers, each of which can splice some or all of the feature maps output by the previous convolutional layer. Thus, when determining the target convolutional layer corresponding to each target feature fusion layer based on preset configuration information and / or detection information of the first image, the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer can be determined. That is, for each multi-entry connection layer in the target feature fusion layer, it is determined which convolutional layers' output feature maps are specifically spliced by the multi-entry connection layer.
[0050] The configuration information can be set by the user based on experience, actual needs, or the hardware resources of the device used for model inference. The configuration information can provide the target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer, that is, it can provide the specific convolution layer output feature maps of each multi-entry connection layer in each target feature fusion layer.
[0051] The detection information of the first image may include a detection task, for example, detecting whether a target object is contained in the first image. In determining the target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer based on the detection information of the first image, in some embodiments, the following may be included: For any multi-input connection layer in each target feature fusion layer, multiple convolutional layers before the multi-input connection layer are combined in different ways to obtain multiple convolutional layer sets, each convolutional layer set includes some or all of the multiple convolutional layers; For each convolutional layer set in the multiple convolutional layer sets, a target YOLO network is trained multiple times using a sample data set corresponding to the detection task of the first image, and a target convolutional layer set is determined from the multiple convolutional layer sets based on the training results of the multiple trainings, wherein, in each training process, the multi-entry connection layer in the target YOLO network is used to splice feature maps output by each convolutional layer in a convolutional layer set, the target convolutional layer set is the convolutional layer set used for the target training in the multiple trainings, and the target training is the training corresponding to the minimum loss function of the target YOLO network on the sample data set in the multiple trainings; The convolutional layer in the target convolutional layer set is determined as the target convolutional layer corresponding to the input feature of the multi-entry connection layer.
[0052] That is, when determining the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer based on the detection information of the first image, the target YOLO network can be trained on the sample dataset to determine which convolutional layer output feature maps can be fused by each multi-entry connection layer in each target feature fusion layer. To achieve better results, when training the YOLO network using the sample dataset, the sample dataset corresponding to the detection task of the first image can be used.
[0053] When combining multiple convolutional layers before the multi-input connection layer in different ways, you can use the following functions: .
[0054] in, x i Represents the feature map of the convolutional layer before the multi-entry connection layer or the output of the convolutional layer, L Represents the number of feature maps that can be fused by the multi-entry connection layer or the number of convolutional layers corresponding to the feature map. a i represent x i Whether the corresponding feature map is spliced by the multi-entry connection layer, a i 0 means no splicing, 1 means splicing. express The concatenation of the corresponding feature maps can be expressed as: .
[0055] Through Setting different values (0 or 1) can generate multiple convolutional layer sets. Each convolutional layer set includes some or all of the convolutional layers before the multi-entry connection layer. Different convolutional layer sets contain different convolutional layers. For example, for any two convolutional layer sets, the convolutional layers contained in the two convolutional layer sets can be completely different or partially different.
[0056] After obtaining multiple convolutional layer sets, the target YOLO network can be trained multiple times for each convolutional layer set using the sample dataset corresponding to the detection task. During each training process, the multi-entry connection layer in the target YOLO network will splice the feature maps output by each convolutional layer in a convolutional layer set. After the training is completed, the loss function of the target YOLO network on the sample dataset can be obtained for each training, thereby obtaining multiple loss functions (or error losses) corresponding to the multiple trainings. Afterwards, the sizes of these multiple loss functions can be compared, and the training corresponding to the smallest loss function can be determined as the target training, and the convolutional layer set used in the target training can be determined as the target convolutional layer set. The convolutional layers in this target convolutional layer set are the target convolutional layers corresponding to the input features of the multi-entry connection layer.
[0057] Through the above method, a target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer can be finally obtained.
[0058] In some embodiments, to ensure target detection performance, for any multi-input connection layer in each target feature fusion layer, the target convolution layer corresponding to the input features of the multi-input connection layer needs to include at least the convolution layer that is closest to the multi-input connection layer before the multi-input connection layer. For example, Figure 4 When performing feature fusion on the multi-entry connection layer in the first path shown in , it is necessary to fuse the feature maps output by at least one column of convolutional layers closest to the first path before the first path.
[0059] After determining the target convolution layer corresponding to the input feature of any multi-entry connection layer in each target feature fusion layer, when the target detection model is obtained according to the target YOLO network training, target detection can be performed on the first image according to the pre-trained target detection model to obtain a detection result.
[0060] In some embodiments, performing object detection on the first image according to a pre-trained object detection model to obtain a detection result may include: Sequentially extracting and fusing features of the first image through the backbone network layer and the original feature fusion layer in the neck network layer in the target detection model to obtain feature maps output by each convolution layer in the backbone network layer and the original feature fusion layer; Each target feature fusion layer in the neck network layer of the target detection model is sequentially passed through to fuse the feature maps output by the target convolution layer corresponding to the target feature fusion layer; The feature map output by the last target feature fusion layer is processed by the head network layer in the target detection model to obtain the detection result.
[0061] Specifically, the first image can be input into the target detection model, and then the various network components in the target detection model can process the first image. Before inputting the first image into the target detection model, the first image can be preprocessed to ensure that it meets the target detection model's input image requirements. Preprocessing the first image includes normalizing the first image. For example, if the standard input image size of the target detection model is 640×640, and the first image size is 1920×1080, the first image can be converted to a 640×640 image (to maintain the original image's aspect ratio) through resampling or other methods.
[0062] After the first image is input into the object detection model, the backbone network within the object detection model performs feature extraction on the first image. This feature extraction process extracts multiple levels of features, abstracting various objects in the image using feature maps at various levels, resulting in feature maps at different levels. The extracted shallow feature maps are close to the model input, possessing high resolution and rich detail (e.g., strong positional information), but weak semantic information. Deep feature maps are close to the model output, possessing low resolution and strong semantic information (capable of recognizing complex objects), but exhibiting significant loss of detail (e.g., weak positional information). After obtaining feature maps at different levels, the original feature fusion layer within the object detection model fuses the feature maps at the same level to produce a fused feature map.
[0063] Afterwards, each target feature fusion layer in the target detection model will fuse the feature maps output by the backbone network layer, the feature maps output by the original feature fusion layer, and the feature maps output by other target feature fusion layers before the target feature fusion layer. Specifically, since the target convolution layer corresponding to each target feature fusion layer has been determined based on the detection information of the first image and / or the preset configuration information in the aforementioned steps, the feature maps output by the target convolution layer can be fused when performing feature fusion.
[0064] After the feature maps are sequentially fused through all target feature fusion layers in the object detection model, the final target feature fusion layer outputs a fused feature map, which serves as input to the head network layer of the object detection model. The head network layer processes this fused feature map to obtain a detection result for the first image.
[0065] To facilitate understanding of the differences between the processing of the first image in the embodiment of the present application and the related art, please refer to Figure 5 and Figure 6 . Figure 5 Schematic diagram of the YOLO network processing the first image in the related art. Figure 62 is a schematic diagram of a target YOLO network processing a first image according to an embodiment of the present application.
[0066] Figure 5 In the YOLO network, the backbone network layer extracts features from the first image and obtains multiple feature maps at different levels. These feature maps are then fused by the neck network and directly output to the head network for processing. The target YOLO network of the present application embodiment adds a target feature fusion layer ( Figure 6 Taking the example of adding a new target feature fusion layer, the backbone network layer in the target YOLO network extracts features from the first image and obtains feature maps at multiple levels. These feature maps are then fused through the original feature fusion layer in the neck network (equivalent to Figure 5 After the feature fusion is performed on the neck network in the original feature fusion layer, it will also be fused through the target feature fusion layer. When the target feature fusion layer performs feature fusion, the feature maps output by the convolution layer in the backbone network layer and the original feature fusion layer can be fused ( Figure 6 The figure shows the fusion of all feature maps output by the convolutional layers in the backbone network layer and the original feature fusion layer. Alternatively, only some feature maps can be fused. During fusion, feature maps at the same level can be fused. The fused feature maps output by the target feature fusion layer are then fed into the head network for processing. This allows for more feature fusion by adding a target feature fusion layer to the YOLO network.
[0067] In some embodiments, the detection result for the first image may include the bounding box location, bounding box category, category probability, and confidence level. After obtaining the detection result, the detection result may be post-processed to improve detection accuracy. Post-processing of the detection result may include non-maximum suppression, such as removing redundant boxes and boxes with excessive overlap. The resulting final box may serve as the final detection result.
[0068] The target detection method provided in the embodiment of the present application is described below by taking target detection on images captured by a camera in an FTTR scenario as an example.
[0069] Step 1: Obtain the image to be detected captured by the camera.
[0070] The image to be detected is collected by a camera, and its size is generally 1920×1080.
[0071] Step 2: Preprocess the image to be detected.
[0072] Preprocessing here includes, but is not limited to, normalizing the image to be detected. For example, if the standard input image size for an object detection model is 640×640, and the image to be detected doesn't match this standard size, preprocessing can involve converting the image to the standard size. For example, resampling can be used to convert the image to 640×640.
[0073] Step 3: Input the preprocessed image to be detected into the target detection model, which is trained based on the target YOLO network.
[0074] When training an object detection model based on a target YOLO network, a corresponding sample dataset can be obtained based on the detection task, and the target YOLO network can be trained based on the sample dataset. For example, the detection task can be detecting whether an image contains a target object. In this embodiment, the detection task is related to FTTR scenarios, for example, detecting whether an image to be detected contains a human figure, pet, household item, or electrical appliance.
[0075] Compared to the YOLO network in related art, the target YOLO network adds at least one target feature fusion layer. Specifically, the target YOLO network includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer is used to extract features from the input image, and the extracted feature maps serve as the input to the neck network layer. The structure of the backbone network layer is the same as that of the backbone network layer in the YOLO network. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature maps output by the backbone network layer, and its network structure is the same as that of the neck network layer in the YOLO network. The at least one target feature fusion layer is a new network structure. If there are multiple target feature fusion layers, they are connected in series, with each target feature fusion layer fusing the feature maps obtained by the previous target feature fusion layer. The feature map output by the last target feature fusion layer serves as the input to the head network layer. The head network layer is used to detect objects based on the feature map output by the last target feature fusion layer and output the object detection results. The network structure of the head network layer is the same as that of the head network layer in the YOLO network and will not be described in detail here.
[0076] Before performing target detection based on the target detection model, you can first determine the target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer in the target detection model. The specific implementation method can be found in the above Figure 1 The relevant descriptions in the illustrated embodiments will not be described in detail here.
[0077] After that, the image to be detected can be input into the target detection model for processing. The target detection model's processing flow of the image to be detected can be found in Figure 6 The embodiment shown will not be described in detail here.
[0078] Step 4: Get the detection results output by the target detection model.
[0079] The detection results of the target detection model can include the bounding box position, bounding box category, category probability and confidence of the detected target object.
[0080] Step 5: Post-process the detection results of the target detection model to obtain the final target detection results.
[0081] To improve detection accuracy, the detection results of the object detection model can be post-processed. This post-processing can include non-maximum suppression, such as removing redundant frames and frames with excessive overlap. The resulting final frames serve as the final object detection results. Based on these object detection results, the target objects and their locations can be determined in the images captured by the camera in FTTR scenarios. Target objects can include people, pets, household items, and electrical appliances.
[0082] Since the target detection model used in performing target detection on the image to be detected is trained based on the target YOLO network, and the target YOLO network has added at least one target feature fusion layer compared to the YOLO network in the related technology, the at least one target feature fusion layer can realize more feature fusion. Therefore, when performing target detection on the image to be detected based on the target detection model, more image information can be extracted, so that the detection accuracy can be improved based on the extracted image information, thereby improving the target detection performance.
[0083] In an embodiment of the present application, when performing target detection, a target detection model is trained using a target YOLO network, and at least one target feature fusion layer is added to the neck network of the target YOLO network. Each target feature fusion layer can fuse the feature map obtained before the target feature fusion layer, and the feature layer output by the last target feature fusion layer is used as the input of the head network layer in the YOLO network. In this way, by adding at least one target feature fusion layer to the YOLO network to perform feature fusion on the feature map, more feature fusion can be achieved. After the target detection model is obtained based on the target YOLO network training, when performing target detection on the first image according to the target detection model, more image information can be extracted based on the newly added target feature fusion layer, so that when performing target detection based on the extracted image information, the detection accuracy can be effectively improved, thereby improving the target detection performance.
[0084] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0085] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 7 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0086] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0087] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0088] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a target detection device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Acquire a first image to be subjected to target detection; Performing object detection on the first image according to a pre-trained object detection model to obtain a detection result; Among them, the target detection model is obtained based on the target YOLO network training, and the target YOLO network includes a backbone network layer, a neck network layer and a head network layer. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature map output by the backbone network layer, and the target feature fusion layer is used to fuse the feature map obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input of the head network layer.
[0089] The above application Figure 7 The methods performed by the target detection device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0090] The electronic device may also perform Figure 1 method, and realize the target detection device in Figure 1 The functions of the illustrated embodiments will not be described in detail in this application.
[0091] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0092] The present application also proposes a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a portable electronic device including a plurality of application programs, enable the portable electronic device to execute Figure 1 The method of the embodiment shown is specifically used to perform the following operations: Acquire a first image to be subjected to target detection; Performing object detection on the first image according to a pre-trained object detection model to obtain a detection result; Among them, the target detection model is obtained based on the target YOLO network training, and the target YOLO network includes a backbone network layer, a neck network layer and a head network layer. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature map output by the backbone network layer, and the target feature fusion layer is used to fuse the feature map obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input of the head network layer.
[0093] Figure 8 This is a schematic diagram of the structure of a target detection device 80 according to an embodiment of the present application. Figure 8 In a software implementation, the target detection device 80 may include: an acquisition module 81 and a detection module 82, wherein: An acquisition module 81 acquires a first image to be subjected to target detection; A detection module 82 performs object detection on the first image according to a pre-trained object detection model to obtain a detection result; Among them, the target detection model is obtained based on the target YOLO network training, and the target YOLO network includes a backbone network layer, a neck network layer and a head network layer. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature map output by the backbone network layer, and the target feature fusion layer is used to fuse the feature map obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input of the head network layer.
[0094] The target detection device 80 provided in this application can also perform Figure 1 The method is implemented by the target detection device 80. Figure 1The functions of the illustrated embodiments will not be described in detail in this application.
[0095] The present application also proposes a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps in the above-mentioned target detection method embodiment.
[0096] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0097] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0098] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0099] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0100] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
Claims
1. A target detection method, comprising: Acquire a first image to be subjected to target detection; Performing object detection on the first image according to a pre-trained object detection model to obtain a detection result; Among them, the target detection model is obtained based on the target YOLO network training, and the target YOLO network includes a backbone network layer, a neck network layer and a head network layer. The neck network layer includes an original feature fusion layer and at least one target feature fusion layer. The original feature fusion layer is used to fuse the feature map output by the backbone network layer, and the target feature fusion layer is used to fuse the feature map obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input of the head network layer.
2. In the method as claimed in claim 1, the target feature fusion layer is used to fuse the feature maps output by the target convolutional layer, and the target convolutional layer is part or all of the convolutional layers in the multiple convolutional layers before the target feature fusion layer, and the multiple convolutional layers include the convolutional layers in the backbone network layer and the convolutional layers in each feature fusion layer between the backbone network layer and the target feature fusion layer.
3. In the method as claimed in claim 2, the feature map output by the target convolution layer includes feature maps of multiple different levels, and the feature maps of different levels have different depths, resolutions and semantic information. The target feature fusion layer is used to fuse feature maps of the same level.
4. According to the method described in any one of claims 1 to 3, the target feature fusion layer includes a first path from bottom to top and a second path from top to bottom, the first path is used to transfer the feature information in the low-resolution feature map to the high-resolution feature map and fuse it with the feature map of the same level obtained before the first path, and the second path is used to transfer the feature information in the high-resolution feature map to the low-resolution feature map and fuse it with the feature map of the same level obtained before the second path.
5. The method of claim 4, wherein the first path comprises an upsampling layer, a multi-input connection layer, and a convolutional layer, and the second path comprises a downsampling layer, a multi-input connection layer, and a convolutional layer; The upsampling layer is used to upsample the input low-resolution feature map to obtain a high-resolution feature map; The downsampling layer is used to downsample the input high-resolution feature map to obtain a low-resolution feature map; The multi-input connection layer is used to splice the sampled feature map with the feature map of the same level obtained before the multi-input connection layer; The convolutional layer is used to fuse the feature maps spliced from the multi-entry connection layers.
6. The method of claim 2, further comprising at least one of the following before performing object detection on the first image according to a pre-trained object detection model: Determining, based on the detection information of the first image, a target convolution layer corresponding to each of the target feature fusion layers; According to the preset configuration information, a target convolution layer corresponding to each of the target feature fusion layers is determined.
7. The method of claim 6, wherein determining a target convolution layer corresponding to each target feature fusion layer comprises: Determine a target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer.
8. The method of claim 7, wherein the detection information includes a detection task; determining, based on the detection information of the first image, a target convolution layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer, comprising: For any multi-input connection layer in each of the target feature fusion layers, combining multiple convolutional layers before the multi-input connection layer in different ways to obtain multiple convolutional layer sets, each of the convolutional layer sets including some or all of the multiple convolutional layers; For each of the multiple convolutional layer sets, the target YOLO network is trained multiple times using a sample data set corresponding to the detection task, and a target convolutional layer set is determined from the multiple convolutional layer sets based on the training results of the multiple trainings, wherein, during each training process, the multi-entry connection layer in the target YOLO network is used to splice feature maps output by each convolutional layer in one of the convolutional layer sets, the target convolutional layer set is the convolutional layer set used for the target training in the multiple trainings, and the target training is the training corresponding to the minimum loss function of the target YOLO network on the sample data set in the multiple trainings; A convolutional layer in the target convolutional layer set is determined as a target convolutional layer corresponding to the input feature of the multi-input connection layer.
9. The method according to claim 7 or 8, wherein, for any multi-input connection layer in each target feature fusion layer, the target convolution layer corresponding to the input features of the multi-input connection layer includes at least a convolution layer that is closest to the multi-input connection layer and precedes the multi-input connection layer.
10. The method according to claim 2, wherein performing object detection on the first image according to a pre-trained object detection model to obtain a detection result comprises: Sequentially extracting and fusing features of the first image through the backbone network layer and the original feature fusion layer in the neck network layer in the target detection model to obtain feature maps output by each convolution layer in the backbone network layer and the original feature fusion layer; Sequentially fusing the feature maps output by the target convolutional layer corresponding to each target feature fusion layer in the neck network layer in the target detection model; The detection result is obtained by processing the feature map output by the last target feature fusion layer through the head network layer in the target detection model.
11. An electronic device comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 10. 12 . A computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to claim 1 .
13. A computer program product, comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute part or all of the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Target detection model and method for fire fighting access occupancy target detection and application
CN114529825A
Unmanned aerial vehicle target detection model based on YOLOv5 network
CN114612835A
Traffic sign detection method based on improved YOLOv8n
CN118570767A
Target detection method and device in driving scene based on YOLOv8, storage medium and equipment
CN119206659A