Target detection method, electronic device, storage medium and program product

By adding a target feature fusion layer to the YOLO network, the problem of insufficient detection accuracy of the YOLO network is solved, and the performance of target detection is improved.

CN120580697BActive Publication Date: 2025-11-21ZTE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511081199.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-21
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

The existing YOLO network has insufficient detection accuracy in target detection scenarios, which cannot meet the high requirements of users.

Method used

Add at least one target feature fusion layer to the YOLO network to fuse feature maps and construct a target YOLO network for target detection.

Benefits of technology

By adding a target feature fusion layer, the accuracy of target detection was improved, and higher detection performance was achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580697B_ABST
    Figure CN120580697B_ABST
Patent Text Reader

Abstract

The application provides a target detection method, an electronic device, a storage medium and a program product. The method comprises: acquiring a first image to be subjected to target detection; and performing target detection on the first image according to a target detection model to obtain a detection result; the target detection model is obtained based on a target YOLO network, the target YOLO network comprises a backbone network layer, a neck network layer and a head network layer, the neck network layer comprises an original feature fusion layer and at least one target feature fusion layer, the original feature fusion layer is used for fusing feature maps output by the backbone network layer, the target feature fusion layer is used for fusing feature maps obtained before the target feature fusion layer, and a feature map output by a last target feature fusion layer is used as an input of the head network layer. By adding the target feature fusion layer to the YOLO network to fuse the feature maps, more feature fusion can be achieved, and the target detection performance can be improved when the target detection is performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a target detection method, electronic device, storage medium, and program product. Background Technology

[0002] In related technologies, YOLO networks are commonly used for target detection. For example, in a Fiber to the Room (FTTR) scenario, after capturing images or videos using cameras, YOLO networks can be used to identify various targets such as people, pets, furniture, and appliances in the images or videos.

[0003] As target detection scenarios become more complex, users are demanding higher accuracy from target detection, thus requiring improvements in the target detection performance of the YOLO network. Summary of the Invention

[0004] This application provides a target detection method, electronic device, storage medium, and program product, which at least addresses the problem of how to improve the target detection performance of the YOLO network.

[0005] To solve the above-mentioned technical problems, this application is implemented as follows:

[0006] Firstly, a target detection method is provided, including:

[0007] Acquire the first image to be detected;

[0008] The first image is subjected to object detection based on a pre-trained object detection model to obtain the detection results.

[0009] The target detection model is trained based on a target YOLO network, which includes a backbone network layer, a neck network layer, and a head network layer. The neck network layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer is used to fuse the feature maps output by the backbone network layer. The target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input to the head network layer.

[0010] Secondly, an electronic device is provided, comprising:

[0011] processor;

[0012] Memory used to store the processor's executable instructions;

[0013] The processor is configured to execute the instructions to implement the method as described in the first aspect.

[0014] Thirdly, a computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in the first aspect.

[0015] Fourthly, a computer program product is provided, the computer program product including a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of the method described in the first aspect.

[0016] In this embodiment, when performing target detection, a target YOLO network is used to train the target detection model. At least one target feature fusion layer is added to the neck network of the target YOLO network. Each target feature fusion layer can fuse the feature maps obtained before it. The feature map output by the last target feature fusion layer is used as the input to the head network layer of the YOLO network. Thus, by adding at least one target feature fusion layer to the YOLO network to fuse the feature maps, more feature fusion can be achieved. After training the target detection model based on the target YOLO network, when performing target detection on the first image according to the target detection model, more image information can be extracted based on the newly added target feature fusion layer. Therefore, when performing target detection based on the extracted image information, the detection accuracy can be effectively improved, thereby enhancing the target detection performance. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating an embodiment of the target detection method of this application;

[0019] Figure 2 This is a schematic diagram of the structure of the YOLO network in related technologies;

[0020] Figure 3 This is a schematic diagram of the structure of a target YOLO network according to an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of the target feature fusion layer in one embodiment of this application;

[0022] Figure 5 This is a schematic diagram of how the YOLO network processes the first image in related technologies;

[0023] Figure 6 This is a schematic diagram of a target YOLO network processing a first image according to an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of the structure of a target detection device according to an embodiment of this application. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings of one or more embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this application.

[0027] The terms "first," "second," etc., used in this application and the claims are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that this application can be implemented in orders other than those illustrated or described herein. Furthermore, in this application and the claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0028] It should be noted that the target detection scheme provided in this application embodiment can be applied to identify various targets in images or videos captured by a camera in FTTR scenarios, such as human figures, pets, furniture, and appliances. In addition, it can also be applied to other target detection scenarios, including but not limited to:

[0029] (1) Intelligent security systems: such as video surveillance and intrusion detection, can detect potential security threats more accurately.

[0030] (2) Autonomous driving technology: such as vehicle surrounding environment perception, obstacle recognition, etc., can realize real-time and efficient target detection, and improve the safety and reliability of autonomous driving.

[0031] (3) Industrial automation: such as quality control and anomaly detection on the production line, can achieve more precise quality control and anomaly detection, and improve production efficiency.

[0032] (4) Medical image analysis: such as auxiliary diagnostic tools to improve the early detection rate of diseases, can detect small lesions earlier and improve the early diagnosis rate of diseases.

[0033] (5) Drone navigation: such as real-time target tracking and obstacle avoidance, can achieve real-time and efficient target tracking and obstacle avoidance, and improve the autonomous navigation capability of drones.

[0034] (6) Robot vision: such as autonomous navigation and object recognition, can achieve more accurate object recognition and autonomous navigation, and improve the intelligence level of robots.

[0035] Optionally, in some implementations, the target detection scheme provided in this application can be combined with other technologies to further expand its application scope, including but not limited to:

[0036] (1) Combining with edge computing: By deploying the target detection scheme of this application on edge devices, low latency and high accuracy target detection can be achieved, which is suitable for various edge computing scenarios.

[0037] (2) Integration with cloud computing: By deploying the target detection scheme of this application in the cloud, large-scale data processing and model training can be achieved, which is suitable for big data analysis and large-scale target detection tasks.

[0038] (3) Integration with the Internet of Things: By integrating the target detection scheme of this application into Internet of Things devices, intelligent sensing and decision-making can be realized, thereby improving the intelligence level of the Internet of Things system.

[0039] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0040] Figure 1 This is a flowchart illustrating an embodiment of the target detection method of this application. Figure 1 The target detection method shown is described below.

[0041] Step S102: Obtain the first image to be detected.

[0042] The first image can be obtained by image acquisition or video capture using an image acquisition device. The first image may or may not contain the target object. In this embodiment, target detection is performed on the first image to determine whether the first image contains the target object.

[0043] Step S104: Perform target detection on the first image according to the pre-trained target detection model to obtain the detection result.

[0044] After obtaining the first image, it can be used as input to a pre-trained object detection model, which can then output the corresponding detection results.

[0045] In this embodiment, the object detection model is trained based on a target YOLO network. The target YOLO network is an improved version of the YOLO network in related technologies, with improvements including the addition of at least one target feature fusion layer. Specifically, the target YOLO network includes a backbone layer (which can be represented as a backbone network layer), a neck layer (which can be represented as a neck network layer), and a head layer (which can be represented as a head network layer). The backbone layer is used to extract features from the input image, and the extracted feature maps serve as the input to the neck layer. The structure of the backbone layer is the same as that of the backbone layer in the YOLO network, and will not be described in detail here. The neck layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer is used to fuse the feature maps output by the backbone layer, and its network structure is the same as that of the neck layer in the YOLO network, and will not be described in detail here. At least one target feature fusion layer is a newly added network structure. When there are multiple target feature fusion layers, they are connected in series. Each target feature fusion layer fuses the feature maps obtained before it. The feature map output by the last target feature fusion layer serves as the input to the head network layer. The head network layer performs target detection based on the feature map output by the last target feature fusion layer and outputs the target detection result. The network structure of the head network layer is the same as that of the head network layer in the YOLO network, and will not be described in detail here.

[0046] To facilitate understanding the structure of the YOLO network and the target YOLO network, the following will use... Figure 2 and Figure 3 The illustrated embodiment will be used as an example for explanation.

[0047] Figure 2 This is a schematic diagram of the structure of the YOLO network in related technologies. Figure 2 The YOLO network shown includes a backbone layer, a neck layer, and a head layer. The backbone layer extracts features from the input image, resulting in multiple feature maps, which serve as input to the neck layer. The neck layer includes a feature fusion layer that fuses the multiple feature maps output from the backbone layer, resulting in multiple fused feature maps, which serve as input to the head layer. The head layer processes these fused feature maps to output the object detection result.

[0048] Figure 3 This is a schematic diagram of the structure of a target YOLO network according to an embodiment of this application. Figure 3The target YOLO network shown includes a backbone layer, a neck layer, and a head layer. The backbone layer extracts features from the input image, resulting in multiple feature maps, which serve as input to the neck layer. The neck layer includes a raw feature fusion layer (i.e.,...). Figure 2 The system consists of a feature fusion layer (shown as an example) and two target feature fusion layers (two target feature fusion layers are used here as an example, but there can also be one or more target feature fusion layers). These two target feature fusion layers are newly added feature fusion layers. The original feature fusion layer is used to fuse multiple feature maps output by the backbone network layer to obtain multiple fused feature maps. The first target feature fusion layer is used to fuse all or part of the feature maps output by the backbone network layer and the original feature fusion layer to obtain multiple fused feature maps. The second target feature fusion layer is used to fuse all or part of the feature maps output by the backbone network layer, the original feature fusion layer, and the first target feature fusion layer to obtain multiple fused feature maps. The multiple feature maps output by the second target feature fusion layer serve as input to the head network layer. The head network layer processes the multiple fused feature maps to obtain the target detection result.

[0049] In this embodiment, both the backbone and neck layers of the target YOLO network include one or more convolutional layers. For each target feature fusion layer in the neck layer, the target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer, or it can be used to fuse the feature maps output by the convolutional layers preceding the target feature fusion layer.

[0050] In some implementations, the target feature fusion layer is used to fuse the feature maps output by the convolutional layers preceding it. Specifically, it may fuse the feature maps output by the target convolutional layer. The target convolutional layer is some or all of the multiple convolutional layers preceding the target feature fusion layer. These multiple convolutional layers include convolutional layers in the backbone network layer and convolutional layers in each feature fusion layer between the backbone network layer and the target feature fusion layer (including the original feature fusion layer and other target feature fusion layers between the original feature fusion layer and the target feature fusion layer). In other words, when performing feature fusion, the target feature fusion layer can fuse the feature maps output by all convolutional layers preceding it, thereby maximizing the target detection performance. Alternatively, it can fuse the feature maps output by only some of the convolutional layers preceding it, thereby reducing complexity while maintaining performance as much as possible.

[0051] by Figure 3Taking the target YOLO network shown as an example, the backbone network layer, the original feature fusion layer, and the two target feature fusion layers all include multiple convolutional layers. Figure 3 (Not shown), which can be represented as convolutional layer A, convolutional layer B, convolutional layer C, and convolutional layer D, respectively. For the first target feature fusion layer, the feature maps output by convolutional layer A and convolutional layer B can be fused, or the feature maps output by convolutional layer B can be fused, or a portion of the feature maps output by convolutional layers A and B can be fused. For the second target feature fusion layer, the feature maps output by convolutional layer A, convolutional layer B, and convolutional layer C can be fused, or the feature maps output by convolutional layer A and convolutional layer C can be fused, or a portion of the feature maps output by convolutional layers B and C can be fused. Examples of each approach will not be provided here.

[0052] In some implementations, the feature map output by the target convolutional layer may include multiple feature maps at different levels. These feature maps at different levels have different depths, resolutions, and semantic information. The target feature fusion layer is used to fuse the feature maps output by the target convolutional layer; specifically, it can be used to fuse feature maps of the same level output by the target convolutional layer.

[0053] For example, the feature map output by the target convolutional layer includes three different levels of feature maps: a 20×20 feature map, a 40×40 feature map, and an 80×80 feature map. The target feature fusion layer fuses the feature maps of the same level output by the target convolutional layer; this could involve fusing the 20×20 feature maps, the 40×40 feature maps, and the 80×80 feature maps.

[0054] To achieve the goal of fusing the feature maps output by the target convolutional layer, in some implementations, the target feature fusion layer may include a bottom-up first path and a top-down second path. The first path is used to pass feature information from the low-resolution feature map to the high-resolution feature map and then fuse it with the feature maps of the same level obtained before the first path. The second path is used to pass feature information from the high-resolution feature map to the low-resolution feature map and then fuse it with the feature maps of the same level obtained before the second path. When fusing feature maps of the same level, it may be possible to fuse all feature maps of the same level, or to fuse only a subset of feature maps from all feature maps of the same level.

[0055] In some implementations, the first path includes, from bottom to top, an upsampling layer, a multi-entry connection layer, and a convolutional layer. The upsampling layer upsamples the input low-resolution feature map to obtain a high-resolution feature map. The multi-entry connection layer concatenates the feature map obtained from the upsampling layer with feature maps of the same level obtained before the multi-entry connection layer. The convolutional layer fuses the concatenated feature maps obtained from the multi-entry connection layer. The number of upsampling layers, multi-entry connection layers, and convolutional layers can be one or more. When there are multiple upsampling layers, multi-entry connection layers, and convolutional layers, one upsampling layer, one multi-entry connection layer, and one convolutional layer can constitute a sub-path, and the first path can include multiple such sub-paths.

[0056] In some implementations, the second path includes, from top to bottom, a downsampling layer, a multi-entry connection layer, and a convolutional layer. The downsampling layer downsamples the input high-resolution feature map to obtain a low-resolution feature map. The multi-entry connection layer concatenates the feature map obtained from the downsampling layer with feature maps of the same level obtained before the multi-entry connection layer. The convolutional layer fuses the concatenated feature maps from the multi-entry connection layer. The number of downsampling layers, multi-entry connection layers, and convolutional layers can be one or more. When there are multiple downsampling layers, multi-entry connection layers, and convolutional layers, one downsampling layer, one multi-entry connection layer, and one convolutional layer can constitute a sub-path, and the second path can include multiple such sub-paths. Optionally, the number of sub-paths included in the second path can be the same as the number of sub-paths included in the first path, so that during feature fusion, the second path can fuse feature maps of different levels output by the first path separately.

[0057] To better understand the structure of the target feature fusion layer, please refer to [link / reference]. Figure 4 . Figure 4 Let's take a YOLO network that includes a target feature fusion layer as an example for illustration. Figure 4 The target feature fusion layer in the model includes a first path and a second path. The first path consists of two sub-paths from bottom to top, each of which includes an upsampling layer, a multi-entry connection layer, and a convolutional layer from bottom to top. The second path consists of two sub-paths from top to bottom, each of which includes a downsampling layer, a multi-entry connection layer, and a convolutional layer from top to bottom.

[0058] For the first path, assuming the input image is a 20×20 feature map, the processing of this 20×20 feature map in the first path includes: the upsampling layer in the first sub-path upsamples the 20×20 feature map to obtain a 40×40 feature map, that is, upsampling the low-resolution feature map to obtain a high-resolution feature map. After this 40×40 feature map is input into the multi-entry connection layer, the multi-entry connection layer concatenates this 40×40 feature map with the previously obtained 40×40 feature map (i.e., all or part of the 40×40 feature map output by the convolutional layer before the first path). The concatenated feature map is then fused by the convolutional layer to obtain the fused 40×40 feature map. The fused 40×40 feature map is input to the second sub-path of the first path. After upsampling by the upsampling layer, an 80×80 feature map is obtained. This 80×80 feature map is then input to the multi-entry connection layer. The multi-entry connection layer concatenates this 80×80 feature map with the previously obtained 80×80 feature map (i.e., all or part of the 80×80 feature map output by the convolutional layer before the first path). The concatenated feature map is then fused by the convolutional layer to obtain a fused 80×80 feature map. This fused 80×80 feature map is used as the input to the second path.

[0059] For the second path, the processing of the fused 80×80 feature map includes: the downsampling layer in the first sub-path downsampling the 80×80 feature map to obtain a 40×40 feature map, i.e., downsampling the high-resolution feature map to obtain a low-resolution feature map. This 40×40 feature map is then input into a multi-entry connection layer, which concatenates this 40×40 feature map with the previously obtained 40×40 feature map (i.e., all or part of the 40×40 feature map output by the convolutional layer before the second path). The concatenated feature map is then fused through a convolutional layer to obtain the fused 40×40 feature map. The fused 40×40 feature map is input into the second sub-path of the second path. After being downsampled by the downsampling layer, a 20×20 feature map is obtained. This 20×20 feature map is then input into the multi-entry connection layer. The multi-entry connection layer concatenates this 20×20 feature map with the previously obtained 20×20 feature map (i.e., all or part of the 20×20 feature map output by the convolutional layer before the second path). The concatenated feature map is then fused by the convolutional layer to obtain the fused 20×20 feature map.

[0060] The 80×80 feature map obtained from the first path, the 40×40 feature map and the 20×20 feature map output from the convolutional layer of the second path are used as inputs to the head network layers in the YOLO network.

[0061] It should be noted that, in Figure 4In the second path shown, when the multi-entry connection layer fuses feature maps, in addition to fusing all or part of the feature maps output by the convolutional layers before the target feature fusion layer (i.e., before the first path), it can also fuse the feature maps output by the convolutional layers preceding the target feature fusion layer, that is, it can also fuse the feature maps output by the convolutional layers in the first path. For example Figure 4 In the second path shown, for the first multi-entry connection layer from top to bottom, when performing feature fusion, in addition to fusing all or part of the 40×40 feature maps output by the convolutional layer before the target feature fusion layer (i.e. before the first path), the 40×40 feature maps output by the convolutional layer in the first path can also be fused.

[0062] As mentioned earlier, the target feature fusion layer can fuse the feature maps output by the target convolutional layer (some or all of the convolutional layers preceding the target feature fusion layer). However, the specific convolutional layers whose feature maps are fused may affect the target detection effect. Therefore, before performing target detection on the first image based on the pre-trained target detection model, it is possible to first determine which convolutional layers' feature maps are fused by each target feature fusion layer in the target detection model. That is, before performing target detection on the first image based on the pre-trained target detection model, it is possible to first determine the target convolutional layer corresponding to each target feature fusion layer.

[0063] In some implementations, prior to object detection of the first image based on a pre-trained object detection model, at least one of the following may be included:

[0064] Based on the preset configuration information, determine the target convolutional layer corresponding to each target feature fusion layer;

[0065] Based on the detection information of the first image, the target convolutional layer corresponding to each target feature fusion layer is determined.

[0066] In other words, when determining the target convolutional layer corresponding to each target feature fusion layer, it can be determined by at least one of the two methods mentioned above.

[0067] In some implementations, the target feature fusion layer may include multiple multi-entry connection layers, each used for concatenating feature maps, such as... Figure 4The target feature fusion layer shown includes four multi-entry connection layers, each of which can concatenate some or all of the feature maps output by the previous convolutional layers. Thus, when determining the target convolutional layer corresponding to each target feature fusion layer based on preset configuration information and / or detection information of the first image, it can be done by determining the target convolutional layer corresponding to the input features of each multi-entry connection layer within each target feature fusion layer. In other words, for each multi-entry connection layer in the target feature fusion layer, it determines which specific convolutional layer's output feature maps it concatenates.

[0068] The configuration information can be set by the user based on experience, actual needs, or the hardware resources of the device used for model inference. The configuration information can specify the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer, that is, specify which convolutional layer output feature maps are concatenated by each multi-entry connection layer in each target feature fusion layer.

[0069] The detection information of the first image may include a detection task, such as detecting whether the first image contains a certain target object. When determining the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer based on the detection information of the first image, in some implementations, this may include:

[0070] For any multi-entry connection layer in each target feature fusion layer, multiple convolutional layers preceding the multi-entry connection layer are combined in different ways to obtain multiple sets of convolutional layers. Each set of convolutional layers includes some or all of the convolutional layers in the multiple convolutional layers.

[0071] For each set of multiple convolutional layers, the target YOLO network is trained multiple times using a sample dataset corresponding to the detection task of the first image. The target convolutional layer set is determined from the multiple sets of convolutional layers based on the training results of the multiple training sessions. In each training process, the multi-entry connection layer in the target YOLO network is used to concatenate the feature maps output by each convolutional layer in a set of convolutional layers. The target convolutional layer set is the set of convolutional layers used for the target training in multiple training sessions. The target training is the training corresponding to the minimum loss function of the target YOLO network on the sample dataset in multiple training sessions.

[0072] The convolutional layers in the target convolutional layer set are determined as the target convolutional layers corresponding to the input features of the multi-entry connection layer.

[0073] In other words, when determining the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer based on the detection information of the first image, it is possible to determine which feature maps output by each convolutional layer in each target feature fusion layer can be fused by training the target YOLO network on a sample dataset. To achieve better results, when training the YOLO network using a sample dataset, a sample dataset corresponding to the detection task of the first image can be used.

[0074] When combining multiple convolutional layers preceding a multi-entry connection layer in different ways, the following function can be used:

[0075] .

[0076] in, x i The feature map representing the output of the convolutional layer or convolutional layer preceding the multi-entry connection layer. L This represents the number of feature maps that a multi-entry connection layer can fuse, or the number of convolutional layers corresponding to that feature map. a i represent x i Whether the corresponding feature map is spliced ​​by a multi-entry connection layer. a i A value of 0 indicates no concatenation, while a value of 1 indicates concatenation. express The concatenation of the corresponding feature maps can be represented as:

[0077] .

[0078] Through the By setting different values ​​(0 or 1), multiple sets of convolutional layers can be obtained. Each set of convolutional layers includes some or all of the convolutional layers in the multiple convolutional layers before the multi-entry connection layer. Different sets of convolutional layers contain different convolutional layers. For example, for any two sets of convolutional layers, the convolutional layers contained in the two sets can be completely different or partially different.

[0079] After obtaining multiple sets of convolutional layers, the target YOLO network can be trained multiple times using the sample dataset corresponding to the detection task for each set. During each training process, the multi-entry connection layer in the target YOLO network concatenates the feature maps output by each convolutional layer in the set. After training, the loss function of the target YOLO network on the sample dataset can be obtained for each training iteration, thus yielding multiple loss functions (or error losses) corresponding to the multiple training iterations. These multiple loss functions can then be compared, and the training corresponding to the smallest loss function is determined as the target training. The set of convolutional layers used in the target training is then defined as the target convolutional layer set. The convolutional layers in this target convolutional layer set are the target convolutional layers corresponding to the input features of the multi-entry connection layer.

[0080] Using the above method, we can finally obtain the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer.

[0081] In some implementations, to ensure target detection performance, for any multi-entry connection layer in each target feature fusion layer, the target convolutional layer corresponding to the input features of the multi-entry connection layer needs to include at least the convolutional layer closest to the multi-entry connection layer preceding it. For example, Figure 4 When performing feature fusion, the multi-entry connection layer in the first path shown in the figure needs to fuse at least the feature map output of the column of convolutional layers that are closest to the first path.

[0082] After determining the target convolutional layer corresponding to the input features of any multi-entry connection layer in each target feature fusion layer, and with the target detection model trained according to the target YOLO network, the first image can be targeted based on the pre-trained target detection model to obtain the detection result.

[0083] In some implementations, object detection is performed on the first image based on a pre-trained object detection model to obtain detection results, which may include:

[0084] The first image is sequentially processed through the backbone network layer and the original feature fusion layer in the neck network layer of the target detection model to extract and fuse features, thereby obtaining the feature maps output by each convolutional layer in the backbone network layer and the original feature fusion layer.

[0085] The feature maps output by the target convolutional layers corresponding to the target feature fusion layers are fused sequentially through each target feature fusion layer in the neck network layer of the target detection model.

[0086] The head network layer in the object detection model processes the feature map output by the last object feature fusion layer to obtain the detection result.

[0087] Specifically, the first image can be input into the object detection model, and then each network in the model can process the first image. Before inputting the first image into the object detection model, it can be preprocessed to ensure it meets the model's requirements for input images. Preprocessing of the first image includes normalization. For example, if the standard size of the input image for the object detection model is 640×640, while the first image is 1920×1080, it can be converted to a 640×640 image (maintaining the original image's aspect ratio) through resampling or other methods.

[0088] After the first image is input into the object detection model, the backbone network of the model performs feature extraction. During feature extraction, multi-level features can be extracted, abstracting various targets in the image into feature maps of different levels. Shallow feature maps are closer to the model input, have high resolution, and are rich in detail (e.g., strong positional information), but weaker semantic information. Deep feature maps are closer to the model output, have low resolution, strong semantic information (able to recognize complex objects), but suffer from significant loss of detail (e.g., weak positional information). After obtaining feature maps of different levels, the original feature fusion layer in the object detection model fuses feature maps of the same level to obtain a fused feature map.

[0089] Subsequently, each target feature fusion layer in the target detection model fuses the feature map output by the backbone network layer, the feature map output by the original feature fusion layer, and the feature maps output by other target feature fusion layers preceding that target feature fusion layer. Specifically, since the target convolutional layer corresponding to each target feature fusion layer has been determined in the aforementioned steps based on the detection information of the first image and / or the preset configuration information, the feature map output by the target convolutional layer can be fused during feature fusion.

[0090] After the feature maps are fused sequentially through all the target feature fusion layers in the target detection model, the last target feature fusion layer outputs a fused feature map, which can be used as input to the head network layer of the target detection model. The head network layer processes this fused feature map to obtain the detection result for the first image.

[0091] To facilitate understanding of the differences between the processing of the first image in this application embodiment and related technologies, please refer to... Figure 5 and Figure 6 . Figure 5 This is a schematic diagram of how the YOLO network processes the first image in related technologies. Figure 6This is a schematic diagram of the target YOLO network processing a first image according to an embodiment of this application.

[0092] Figure 5 In the YOLO network, after the backbone network layer extracts features from the first image and obtains multiple feature maps at different levels, these feature maps are fused by the neck network and then directly output to the head network for processing. However, the target YOLO network in this embodiment adds a target feature fusion layer (…). Figure 6 Taking the addition of a target feature fusion layer as an example, after the backbone network layer of the target YOLO network extracts features from the first image and obtains multiple feature maps at different levels, these feature maps are fused through the original feature fusion layer in the neck network (equivalent to...). Figure 5 After feature fusion in the backbone network (the neck network), it undergoes further fusion processing via the target feature fusion layer. During feature fusion in the target feature fusion layer, the feature maps output from the convolutional layers in the backbone network layer and the original feature fusion layer are fused. Figure 6 The diagram illustrates the fusion of all feature maps output from the convolutional layers in the backbone network layer and the original feature fusion layer (or only a subset of feature maps can be fused). During fusion, feature maps at the same level can be fused. The fused feature map output from the target feature fusion layer is then fed into the head network for processing. Thus, by adding a target feature fusion layer to the YOLO network, more feature fusions can be achieved.

[0093] In some implementations, the detection results for the first image may include bounding box locations, bounding box categories, category probabilities, and confidence levels. After obtaining the detection results, post-processing can be performed to improve detection accuracy. Post-processing may include non-maximum suppression, such as removing redundant boxes or boxes with excessive overlap. The resulting final bounding box can then be used as the final detection result.

[0094] The following uses the example of target detection in an FTTR scenario to illustrate the target detection method provided in this application.

[0095] Step 1: Acquire the image to be detected captured by the camera.

[0096] The image to be detected is captured by a camera, and its size is generally 1920×1080.

[0097] Step 2: Preprocess the image to be detected.

[0098] Preprocessing here includes, but is not limited to, normalizing the image to be detected. Taking a standard input image size of 640×640 for the object detection model as an example, since the size of the image to be detected differs from this standard size, the size of the image to be detected can be converted to the standard size during preprocessing. For example, the image to be detected can be converted to 640×640 through methods such as resampling.

[0099] Step 3: Input the preprocessed image to be detected into the target detection model, which is trained based on the target YOLO network.

[0100] When training an object detection model based on a target YOLO network, a corresponding sample dataset can be obtained according to the detection task, and the target YOLO network can be trained based on this sample dataset. The detection task could be, for example, detecting whether an image contains a specific target object. In this embodiment, the detection task is related to the FTTR scenario, such as detecting whether the image to be detected contains a human figure, pet, furniture, or appliances.

[0101] Compared to YOLO networks in related technologies, the target YOLO network adds at least one target feature fusion layer. Specifically, the target YOLO network includes a backbone layer, a neck layer, and a head layer. The backbone layer extracts features from the input image, and the extracted feature maps serve as input to the neck layer. The structure of the backbone layer is the same as that of the backbone layer in the YOLO network. The neck layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer fuses the feature maps output from the backbone layer, and its network structure is the same as that of the neck layer in the YOLO network. The at least one target feature fusion layer is a newly added network structure. When there are multiple target feature fusion layers, they are connected in series. Each target feature fusion layer fuses the feature maps obtained before it, and the feature map output from the last target feature fusion layer serves as input to the head layer. The head layer performs target detection based on the feature map output from the last target feature fusion layer and outputs the target detection result. The network structure of the head layer is the same as that of the head layer in the YOLO network, and will not be described in detail here.

[0102] Before performing object detection based on the object detection model, the target convolutional layer corresponding to the input features of each multi-entry connection layer in each target feature fusion layer of the object detection model can be determined first. For specific implementation methods, please refer to the above. Figure 1 The relevant descriptions in the illustrated embodiments will not be explained in detail here.

[0103] Next, the image to be detected can be input into the object detection model for processing. The processing flow of the object detection model for the image to be detected can be found in [link to relevant documentation]. Figure 6 The embodiments shown are not described in detail here.

[0104] Step 4: Obtain the detection results output by the object detection model.

[0105] The detection results of an object detection model can include the bounding box location, bounding box category, category probability, and confidence level of the detected object.

[0106] Step 5: Post-process the detection results of the target detection model to obtain the final target detection results.

[0107] To improve detection accuracy, post-processing can be applied to the detection results of the object detection model. Post-processing can include non-maximum suppression, such as removing redundant boxes and boxes with excessive overlap. The resulting bounding boxes can then be used as the final object detection result. Based on this result, it is possible to determine which target objects are present in the image captured by the camera in the FTTR scene and their locations. Target objects can be, for example, human figures, pets, furniture, and appliances.

[0108] Since the target detection model used in the image to be detected is trained based on the target YOLO network, and the target YOLO network adds at least one target feature fusion layer compared to the YOLO network in related technologies, the target feature fusion layer can achieve more feature fusion. Therefore, when performing target detection on the image to be detected based on the target detection model, more image information can be extracted, thereby improving the detection accuracy based on the extracted image information and thus improving the target detection performance.

[0109] In this embodiment, when performing target detection, a target YOLO network is used to train the target detection model. At least one target feature fusion layer is added to the neck network of the target YOLO network. Each target feature fusion layer can fuse the feature maps obtained before it. The feature map output by the last target feature fusion layer is used as the input to the head network layer of the YOLO network. Thus, by adding at least one target feature fusion layer to the YOLO network to fuse the feature maps, more feature fusion can be achieved. After training the target detection model based on the target YOLO network, when performing target detection on the first image according to the target detection model, more image information can be extracted based on the newly added target feature fusion layer. Therefore, when performing target detection based on the extracted image information, the detection accuracy can be effectively improved, thereby enhancing the target detection performance.

[0110] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0111] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 7 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0112] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0113] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0114] The processor reads the corresponding computer program from non-volatile memory into main memory and then executes it, forming a target detection device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:

[0115] Acquire the first image to be detected;

[0116] The first image is subjected to object detection based on a pre-trained object detection model to obtain the detection results.

[0117] The target detection model is trained based on a target YOLO network, which includes a backbone network layer, a neck network layer, and a head network layer. The neck network layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer is used to fuse the feature maps output by the backbone network layer. The target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input to the head network layer.

[0118] The above is as stated in this application. Figure 7 The target detection device disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0119] The electronic device can also perform Figure 1 The method, and realize the target detection device in Figure 1 The functions described in the illustrated embodiments will not be repeated here.

[0120] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0121] This application also discloses a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 The method of the illustrated embodiment is specifically used to perform the following operations:

[0122] Acquire the first image to be detected;

[0123] The first image is subjected to object detection based on a pre-trained object detection model to obtain the detection results.

[0124] The target detection model is trained based on a target YOLO network, which includes a backbone network layer, a neck network layer, and a head network layer. The neck network layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer is used to fuse the feature maps output by the backbone network layer. The target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input to the head network layer.

[0125] Figure 8 This is a schematic diagram of the structure of a target detection device 80 according to an embodiment of this application. Please refer to it. Figure 8 In one software implementation, the target detection device 80 may include: an acquisition module 81 and a detection module 82, wherein:

[0126] Module 81 acquires the first image to be detected;

[0127] Detection module 82 performs target detection on the first image according to a pre-trained target detection model to obtain detection results;

[0128] The target detection model is trained based on a target YOLO network, which includes a backbone network layer, a neck network layer, and a head network layer. The neck network layer includes a raw feature fusion layer and at least one target feature fusion layer. The raw feature fusion layer is used to fuse the feature maps output by the backbone network layer. The target feature fusion layer is used to fuse the feature maps obtained before the target feature fusion layer. The feature map output by the last target feature fusion layer is used as the input to the head network layer.

[0129] The target detection device 80 provided in this application can also perform... Figure 1 The method, and realize the target detection device 80 in Figure 1 The functions of the embodiments shown will not be described again in this application.

[0130] This application also proposes a computer program product comprising a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the above-described target detection method embodiments.

[0131] In summary, the above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0132] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0133] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0134] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0135] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

Claims

1. A target detection method, comprising: obtaining a first image to be subjected to target detection; performing target detection on the first image according to a pre-trained target detection model to obtain a detection result; wherein the target detection model is trained based on a target YOLO network, the target YOLO network comprising a backbone network layer, a neck network layer and a head network layer, the neck network layer comprising an original feature fusion layer and at least one target feature fusion layer, the original feature fusion layer being configured to fuse feature maps output by the backbone network layer, the target feature fusion layer being configured to fuse feature maps output by a target convolutional layer, the target convolutional layer being part or all of a plurality of convolutional layers before the target feature fusion layer, the plurality of convolutional layers comprising convolutional layers in the backbone network layer and convolutional layers in each feature fusion layer between the backbone network layer and the target feature fusion layer, and a feature map output by a last target feature fusion layer being input to the head network layer; before performing target detection on the first image according to the target detection model, the method further comprises: determining, according to detection information of the first image, a target convolutional layer corresponding to each target feature fusion layer; the detection information comprising a detection task, and the determination of the target convolutional layer corresponding to each target feature fusion layer comprising: for any multi-entry connection layer in each target feature fusion layer, combining a plurality of convolutional layers before the multi-entry connection layer in different ways to obtain a plurality of convolutional layer sets, each convolutional layer set comprising part or all of the plurality of convolutional layers; for each convolutional layer set in the plurality of convolutional layer sets, performing multiple training of the target YOLO network using a sample data set corresponding to the detection task, and determining a target convolutional layer set from the plurality of convolutional layer sets according to training results of the multiple training, wherein in each training process, the multi-entry connection layer in the target YOLO network is configured to splice feature maps output by convolutional layers in one convolutional layer set, the target convolutional layer set is a convolutional layer set used in a target training in the multiple training, and the target training is a training corresponding to a minimum loss function of the target YOLO network on the sample data set in the multiple training; determining convolutional layers in the target convolutional layer set as target convolutional layers corresponding to input features of the multi-entry connection layer.

2. The method of claim 1, wherein the feature maps output by the target convolutional layer comprise a plurality of feature maps of different levels, the feature maps of different levels having different depths, resolutions and semantic information, and the target feature fusion layer is configured to fuse feature maps of the same level. 3.The method of claim 1 or 2, wherein the target feature fusion layer comprises a first path from bottom to top and a second path from top to bottom, the first path is used to transmit feature information in a low resolution feature map to a high resolution feature map and fuse the high resolution feature map with a feature map of the same level obtained before the first path, and the second path is used to transmit feature information in a high resolution feature map to a low resolution feature map and fuse the low resolution feature map with a feature map of the same level obtained before the second path. 4.The method of claim 3, wherein the first path comprises an up-sampling layer, a multi-entry connection layer and a convolution layer, and the second path comprises a down-sampling layer, a multi-entry connection layer and a convolution layer; the up-sampling layer is used to up-sample an input low resolution feature map to obtain a high resolution feature map; the down-sampling layer is used to down-sample an input high resolution feature map to obtain a low resolution feature map; the multi-entry connection layer is used to splice a feature map obtained by sampling and a feature map of the same level obtained before the multi-entry connection layer; and the convolution layer is used to fuse the feature map spliced by the multi-entry connection layer. 5.The method of claim 1, before target detection is performed on the first image according to a pre-trained target detection model, the method further comprises: determining a target convolution layer corresponding to each target feature fusion layer according to pre-set configuration information. 6.The method of claim 1, for any multi-entry connection layer in each target feature fusion layer, a target convolution layer corresponding to an input feature of the multi-entry connection layer at least comprises a convolution layer closest to the multi-entry connection layer before the multi-entry connection layer. 7.The method of claim 1, wherein the target detection is performed on the first image according to a pre-trained target detection model to obtain a detection result, comprising: sequentially passing the first image through the backbone network layer and the original feature fusion layer in the neck network layer in the target detection model to perform feature extraction and fusion, and obtain feature maps output by each convolution layer in the backbone network layer and the original feature fusion layer; sequentially passing each target feature fusion layer in the neck network layer in the target detection model to fuse feature maps output by a target convolution layer corresponding to the target feature fusion layer; and passing a feature map output by a last target feature fusion layer through the head network layer in the target detection model to obtain the detection result. 8.An electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method of any one of claims 1 to 7. 9.A computer readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of any one of claims 1 to 7. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 10. A computer program product comprising a non-transitory computer readable storage medium having stored thereon a computer program, the computer program being operable to cause a computer to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target detection model and method for fire fighting access occupancy target detection and application

    CN114529825A