Target recognition network construction method and device based on feature fusion
By performing multiple feature extractions and feature fusions on sample data, the problem of poor detection performance of target detection algorithms in complex environments is solved, and the detection accuracy and recognition performance of small targets are improved.
Patent Information
- Application Number
- CN202511047568.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-28
AI Technical Summary
Existing target detection algorithms have poor detection performance in complex and highly adversarial environments, especially in multi-scale and small target recognition accuracy. Furthermore, one-stage detection algorithms are easily affected by background and lighting interference.
By performing multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging on the sample data, combined with dual cross-scale feature fusion, the fusion of deep and shallow features and cross-scale features is achieved, thereby improving detection performance.
It improves the model's ability to extract features from targets in complex backgrounds, enhances the detection accuracy and recognition performance of small targets, and solves the detection problem in complex environments.
Smart Images

Figure CN121033601A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a target recognition network construction method and device based on feature fusion. BACKGROUND
[0002] With the progress of visual sensors, artificial intelligence and other technologies, target detection technology based on machine vision and deep learning has been widely used in automatic driving, face recognition, aerospace and other fields.
[0003] The existing target detection algorithm mainly has two categories: two-stage target detection algorithm and one-stage target detection. The two-stage algorithm is mainly represented by the Region-based Convolutional Neural Network (R-CNN for short) series algorithm; although the detection accuracy is high, the detection speed is slow, the model occupies a large amount of memory, and the calculation amount is too large, which cannot meet the requirements of real-time detection of small targets. One-stage target detection algorithm mainly has You Only Look Once (YOLO for short) series, Single Shot Multibox Detector (SSD for short) and RetinaNet, etc.; the YOLO series has fast detection speed, can simultaneously predict the boundary frame and the category of the target, and completes all steps of target detection at one time.
[0004] However, in a complex and strong adversarial environment, the target to be detected is easily disturbed by factors such as background and light, which leads to the one-stage target detection algorithm being prone to recognition errors. And when the target has the characteristics of multi-scale and target size change, the detection difficulty of the one-stage target detection algorithm is increased, and the recognition accuracy for dense and small targets is not high.
[0005] Therefore, it is urgent to overcome the defects of the prior art in the technical field. SUMMARY
[0006] The technical problem to be solved by the present application is to provide a target recognition network construction method and device based on feature fusion, which aims to realize the fusion of deep and shallow layer features and cross-scale features by performing multiple feature extraction, sparse feature processing, connection processing and dynamic weighted average on sample data, and fusing each level feature based on double cross-scale feature fusion, thereby improving the detection performance and solving the problem of poor detection performance of the one-stage target detection algorithm in a complex and strong adversarial environment.
[0007] The present application adopts the following technical solutions: In a first aspect, the present application provides a target recognition network construction method based on feature fusion, comprising: collect a complex environment image, perform data augmentation and labeling on the complex environment image to generate sample data, and construct a data set based on the sample data; perform multiple feature extraction, sparse feature processing, connection processing, and dynamic weighted averaging on the sample data in the data set to obtain hierarchical features after each processing; fuse each hierarchical feature based on double cross-scale feature fusion to obtain a fusion feature map; use a head network to perform target detection output on each fusion feature map to obtain a sample detection result; process the sample detection result based on a loss function to obtain a training state; when the training state satisfies an end condition, obtain a target model to detect a complex environment image required for detection using the target model to obtain a target detection result.
[0008] Further, the double cross-scale feature fusion of each hierarchical feature to obtain a fusion feature map comprises: the last processed hierarchical feature is taken as a deep feature, and the hierarchical features other than the deep feature are taken as intermediate features; perform feature enhancement on the deep feature to obtain an enhanced feature; perform double cross-scale feature fusion on each intermediate feature and the enhanced feature respectively to obtain a corresponding fusion feature map.
[0009] Further, the double cross-scale feature fusion of each intermediate feature and the enhanced feature to obtain a corresponding fusion feature map comprises: divide the channels of each intermediate feature or the enhanced feature into multiple groups; standardize the features of each group of channels based on the standard deviation and mean value of the features of each group of channels and trainable parameters to obtain an intermediate feature; normalize the intermediate feature to obtain a normalized feature; use an activation function to process the product of the normalized feature and the intermediate feature to obtain a group weight value; process the features of each group of channels according to the group weight value to generate a fusion feature map.
[0010] Further, the processing of the features of each group of channels according to the group weight value to generate a fusion feature map comprises: multiply the features of each group of channels with the corresponding group weight value and the overall learnable weight value to obtain a weighted feature map; perform fusion operation on the weighted feature maps of each group of channels to generate a single feature map; map the single feature map to a preset range and perform convolution operation on the feature maps in the preset range to obtain a fusion feature map.
[0011] Furthermore, the process of performing multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging on the sample data in the dataset to obtain the hierarchical features after each processing includes: The sample data is subjected to a convolution operation to extract features and obtain the first feature; The first feature is subjected to sparse feature processing to obtain the second feature; The second feature is subjected to cross-stage feature fusion for connection processing to obtain the third feature; The third feature is dynamically weighted and averaged to obtain the hierarchical feature.
[0012] Furthermore, the step of performing sparse feature processing on the first feature to obtain the second feature includes: The first feature is divided into multiple grids according to the scaling factor; The multiple grids are rearranged along the channel dimension to generate sub-feature maps; By sequentially connecting all the sub-feature maps along the channel dimension, a new feature map is obtained; The new feature map is subjected to a non-stepping convolution operation to obtain the second feature.
[0013] Furthermore, the step of dynamically weighting the third feature to obtain hierarchical features includes: The third feature is subjected to global average pooling and processed using a multilayer perceptron and activation function to generate channel weights; The third feature is divided into G groups of semantic features; the three-dimensional features of the semantic features are mapped to the wide-dimensional and high-dimensional directions, and a single-kernel convolution operation is performed to obtain the features to be weighted; The feature to be weighted is multiplied by the channel weight, and a linear transformation is performed on the multiplied feature to obtain the first attention map; Perform multi-kernel convolution on the semantic features to obtain processed features; determine the product of the processed features and the channel weights as the second attention map; By aggregating the first attention map and the second attention map corresponding to the semantic features, hierarchical features are obtained.
[0014] Furthermore, the expression for the loss function is:
[0015] in, For outlier degree, and All are regulatory factors. The intersection-union ratio (IU) of the predicted bounding box and the ground truth bounding box. The width of the minimum bounding rectangle between the predicted bounding box and the ground truth bounding box. The height of the minimum bounding rectangle. The x-coordinate of the center point of the prediction box. The ordinate of the center point of the prediction box is... The x-coordinate of the center point of the real bounding box is given. The ordinate of the center point of the real bounding box is given by [reference]. This indicates that the moving average value is taken.
[0016] Secondly, the present invention also provides a target recognition network construction apparatus based on feature fusion, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor for performing the feature fusion-based target recognition network construction method described in the first aspect.
[0017] Thirdly, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions, which are executed by one or more processors to perform the target recognition network construction method based on feature fusion described in the first aspect.
[0018] Fourthly, a computer program product containing instructions is provided, which, when executed on a computer or processor, causes the computer or processor to perform the target recognition network construction method based on feature fusion as described in the first aspect.
[0019] Fifthly, the present invention also provides a target recognition network construction system based on feature fusion, including a target recognition network construction device based on feature fusion as described in the second aspect, and using the target recognition network construction method based on feature fusion as described in the first aspect to complete the interaction of the target recognition network construction device based on feature fusion as described in the second aspect.
[0020] Unlike existing technologies, the present invention has at least the following beneficial effects: This invention addresses the characteristics of targets in complex environment images, such as target occlusion, multi-scale, and target size variations. By performing multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging on sample data, features at different levels are extracted. Furthermore, based on dual cross-scale feature fusion, features at each level are fused to achieve the fusion of deep and shallow features and cross-scale features. This improves the model's feature extraction capability for targets in complex backgrounds, enhances the detection accuracy of small targets, and improves detection and recognition performance. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0022] Figure 1 This is a flowchart illustrating a target recognition network construction method based on feature fusion provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating a specific example of the network structure of a detection model provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating step 20 provided in an embodiment of the present invention; Figure 4 This is a flowchart illustrating step 202 provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating step 204 provided in an embodiment of the present invention; Figure 6 This is a flowchart illustrating step 30 provided in an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating a specific example of dual cross-scale feature fusion provided in an embodiment of the present invention; Figure 8 This is a flowchart illustrating step 303 provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the architecture of a target recognition network construction device based on feature fusion provided in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0024] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as openly inclusive, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples; that is, although they may be incorporated into embodiments or examples using the above terms for reasons such as order and position, it does not limit them to be incorporated in combination by a single embodiment or example.
[0025] In the description of this invention, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this disclosure.
[0026] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, for example, the description may use the prefix "A" or "B" to describe the same type of nouns as two independent entities. In this case, the corresponding features defined with "A" and "B" are used only to distinguish between similar entities and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features.
[0027] In describing some embodiments, the terms "coupled," "coupled," and "connected," and their derivative expressions, may be used. For example, the term "connected" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact with each other. Similarly, the term "coupled" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact. However, the terms "connected" or "coupled" may also refer to two or more components that do not have direct contact with each other but still cooperate or interact with each other, such as "optical coupling," "wireless connection," etc. The embodiments disclosed herein are not necessarily limited to the scope of this invention.
[0028] In the description of this invention, the expression “A and / or B” (where A and B are used to formally represent specific features) will be used. The corresponding expression includes the following three combinations: only A, only B, and a combination of A and B.
[0029] As used in this invention, “about,” “approximately,” or “approximately” includes the stated value and the average value within an acceptable range of deviation from a particular value, wherein the acceptable range of deviation is determined by a person skilled in the art taking into account the measurement under discussion and the error associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).
[0030] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0031] Example 1: On the one hand, target recognition and / or detection tasks in complex environments often deal with multi-target scenarios, meaning they need to handle situations where numerous targets appear simultaneously in complex environmental images. For example, in a busy traffic intersection scene, there are various vehicles (cars, trucks, motorcycles, etc.), pedestrians, traffic signs, and traffic lights; these targets vary in size, shape, and color, and mutual occlusion is quite common. It is necessary to accurately distinguish these different categories of targets and determine their positions and sizes; for example, to be able to simultaneously detect a speeding red sports car and a pedestrian wearing dark clothing crossing the street.
[0032] Complex environments can contain backgrounds that resemble the target's appearance, causing confusion. For example, in wildlife detection within a forest setting, the textures of tree branches and shrubs may resemble the patterns on an animal's fur. Furthermore, complex lighting conditions can affect detection; for instance, outdoor scenes may present shadows or direct sunlight. For instance, in a sunny desert environment, the strong light reflected from the sand may obscure the features of small targets (e.g., small reptiles), making them difficult to detect.
[0033] On the other hand, in highly adversarial environments, object detection systems may face attacks from adversarial examples. Attackers can mislead detection algorithms by subtly modifying complex environmental images. For example, by adding specific, carefully designed noise patterns to complex environmental images, targets that would otherwise be accurately detected (e.g., pedestrians) can be misclassified as other categories (e.g., trees) or completely undetectable. This noise may be almost imperceptible to human observers, but it can introduce significant errors for deep learning-based object detection models.
[0034] To solve the above problems, such as Figure 1 As shown, this embodiment of the invention provides a method for constructing a target recognition network based on feature fusion, including: Step 10: Collect complex environment images, perform data augmentation and annotation on the complex environment images to generate sample data, and construct a dataset based on the sample data.
[0035] First, complex environment images are collected. To expand the sample size, data augmentation is performed on the existing complex environment images, and targets in the augmented images are labeled to obtain sample data. A dataset for target recognition and / or target detection is constructed and divided into training and validation sets.
[0036] Step 20: Perform multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging on the sample data in the dataset to obtain the hierarchical features after each processing.
[0037] Then, a detection model is constructed based on a network structure with a multi-scale attention mechanism, and trained using the training set in the constructed dataset. Each processing step on the sample data includes: feature extraction, sparse feature processing, connection processing, and dynamic weighted averaging. The feature map after each processing step serves as the input feature map for feature extraction in the next processing step. The processes of feature extraction, sparse feature processing, connection processing, and dynamic weighted averaging will be explained separately below.
[0038] Step 30: Based on dual cross-scale feature fusion, each of the hierarchical features is fused to obtain a fused feature map; the head network is used to perform target detection on each of the fused feature maps to output the sample detection results.
[0039] The process based on dual cross-scale feature fusion will be explained below. Since the feature maps obtained after multiple processing steps in step 20 extract features from different levels of the sample data, cross-scale fusion of these features can better extract information about the target from the sample data.
[0040] Step 40: Process the sample detection results based on the loss function to obtain the training state; when the training state meets the termination condition, obtain the target model, and use the target model to detect the complex environment image to be detected, and obtain the target detection result.
[0041] The termination condition is selected by those skilled in the art based on the specific application scenario; in one embodiment, the training state can be the loss function value, and the termination condition can be that the loss function value converges to a certain value.
[0042] The target model is used for object recognition and / or object detection in images. Finally, the validation set is input into the trained target model to obtain the object detection results.
[0043] This invention addresses the characteristics of targets in complex environment images, such as target occlusion, multi-scale, and target size variations. By performing multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging on sample data, features at different levels are extracted. Furthermore, based on dual cross-scale feature fusion, features at each level are fused to achieve the fusion of deep and shallow features and cross-scale features. This improves the model's feature extraction capability for targets in complex backgrounds, enhances the detection accuracy of small targets, and improves detection and recognition performance.
[0044] The following is a detailed description of the target recognition network construction method based on feature fusion according to embodiments of the present invention: In one embodiment, step 10 includes: Construct an object recognition dataset and divide it into a training set and a validation set.
[0045] Specifically, due to the lack of publicly available datasets, a self-built target dataset is constructed, which mainly consists of images taken by drones. It can include multiple types of targets, and the specific type of target is selected by those skilled in the art based on the specific use case, without limitation here.
[0046] The collected images were augmented to enrich the military target image dataset. Data augmentation techniques such as image translation, random rotation, cropping, and Mosaic augmentation were used to expand the dataset, reduce overfitting, improve model robustness, and enhance generalization ability. The expanded dataset should contain more than 5000 images. The dataset was then labeled using annotation tools and divided into training and validation sets at a 9:1 ratio.
[0047] like Figure 2 The image shows a specific example of a network structure for the detection model according to an embodiment of the present invention; the following is based on... Figure 2 The network structure shown illustrates the training process for obtaining the target model: like Figure 3As shown, step 20 includes: Step 201: Perform a convolution operation on the sample data to extract features and obtain the first feature.
[0048] In one embodiment, the sample data can be convolved once to obtain a preprocessed feature map. Based on this, the feature map can be subjected to multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging to extract features at different levels.
[0049] The specific method of performing the convolution operation shall be selected by those skilled in the art based on the specific use case.
[0050] Step 202: Perform sparse feature processing on the first feature to obtain the second feature.
[0051] The following section will explain the specific processing flow for sparse feature processing.
[0052] Step 203: Perform cross-stage feature fusion on the second feature to perform connection processing and obtain the third feature.
[0053] In one embodiment, cross-stage feature fusion can be implemented using a convolution and concatenate with feature fusion (C2f) module.
[0054] Step 204: Perform a dynamic weighted average on the third feature to obtain the hierarchical feature.
[0055] The following section will explain the specific processing procedure for dynamic weighted averaging.
[0056] In one embodiment, the sample data is subjected to four rounds of feature extraction, sparse feature processing, connection processing, and dynamic weighted averaging to extract features at four different levels.
[0057] To illustrate the process of sparse feature processing, such as Figure 4 As shown, step 202 includes: Step 2021: Divide the first feature into multiple grids according to the scaling factor.
[0058] The size is The feature map X is cut into smaller grids according to the scaling factor.
[0059] For example, when the scaling factor is 2, the feature map X is divided into 4 parts of size 1. Sub-feature map Sub-feature map Sub-feature map Characteristic diagram .
[0060] Step 2022: Rearrange the multiple grids along the channel dimension to generate a sub-feature map.
[0061] These grids are rearranged along the channel dimension to create a sub-feature map.
[0062] Step 2023: Connect all the sub-feature maps sequentially along the channel dimension to obtain a new feature map.
[0063] For example, connecting sub-feature maps sequentially along the channel dimension. Sub-feature map Sub-feature map Characteristic diagram A new feature map is obtained. .
[0064] Step 2024: Perform a non-stepping convolution operation on the new feature map to obtain the second feature.
[0065] For example, after feature transformation, a non-stepping convolutional layer is applied to the feature map. Convert to feature map The step size of the non-stepping convolutional layer is... , This reduces the number of channels to a manageable size while maintaining spatial resolution.
[0066] Sparse feature processing enables models to reduce the size of spatial dimensions without losing information. Compared with traditional convolution operations, it retains more information within the channels and can effectively improve the feature extraction capability for small targets.
[0067] To illustrate the process of performing a dynamic weighted average, such as Figure 5 As shown, step 204 includes: Step 2041: Perform global average pooling on the third feature and process it using a multilayer perceptron and activation function to generate channel weights.
[0068] Step 2042: Divide the third feature into G groups of semantic features; map the three-dimensional features of the semantic features to the wide-dimensional and high-dimensional directions, and perform single-kernel convolution operation to obtain the features to be weighted.
[0069] The size is The feature tensor X (where H is the number of rows, W is the number of columns, and C is the number of channels) is divided into G groups of different semantic features for learning, which are represented as follows: .
[0070] Two-dimensional global pooling is performed using the following formula, which maps three-dimensional features to the wide and high-dimensional directions.
[0071]
[0072] in, It is the output associated with the c-th channel. This represents the value in the i-th row and j-th column of the c-th channel.
[0073] Step 2043: Multiply the feature to be weighted by the channel weight, and perform a linear transformation on the multiplied feature to obtain the first attention map.
[0074] The feature maps obtained above are subjected to single-kernel convolution operations, and then multiplied by matrix dot product operations to obtain the first spatial attention map with feature weights.
[0075] Step 2044: Perform multi-kernel convolution operation on the semantic features to obtain processed features; determine the product of the processed features and the channel weights as the second attention map.
[0076] The grouped feature tensors are processed using a 3×3 large kernel convolution, and then multiplied by matrix dot product to obtain the second spatial attention map.
[0077] Step 2045: Aggregate the first attention map and the second attention map corresponding to the semantic features to obtain hierarchical features.
[0078] The output of the two-dimensional global average pooling is linearly transformed using the Softmax function of Gaussian mapping, and the two spatial attention weights in each group are aggregated to calculate the output feature map of each group; in an optional embodiment, the feature map is finally fitted using the Sigmoid function.
[0079] This invention embodiment introduces a parallel Convolutional kernels expand the feature space of cross-channel interaction, accelerate response speed, and facilitate the capture of feature information at different scales, thereby enhancing the model's focus on the target area and improving the model's feature extraction capability and efficiency for military targets in complex battlefield environments.
[0080] To illustrate the process of fusing hierarchical features, such as Figure 6 As shown, in step 30, the process of fusing each hierarchical feature based on dual cross-scale features to obtain a fused feature map includes: Step 301: Take the hierarchical features after the last processing as deep features, and take all hierarchical features other than the deep features as intermediate features.
[0081] like Figure 2As shown, the hierarchical features after the last processing are the features output by the deepest dynamic weighted average module in the backbone network, while the intermediate features are the features output by the dynamic weighted average modules of the remaining layers in the backbone network.
[0082] For ease of description, Figure 2 The term "fusion" is used to refer to "dual cross-scale feature fusion".
[0083] Step 302: Perform feature enhancement on the deep features to obtain enhanced features.
[0084] In one embodiment, feature enhancement can be implemented using Spatial Pyramid Pooling-Fast (SPPF).
[0085] Step 303: Perform dual cross-scale feature fusion on each of the intermediate features and the enhanced features to obtain the corresponding fused feature map.
[0086] like Figure 2 As shown, in one embodiment, dual cross-scale feature fusion can combine cross-stage feature fusion and upsampling operation; wherein, upsampling operation can be performed on the enhanced features, and then dual cross-scale feature fusion and cross-stage feature fusion can be performed sequentially; wherein, the dual cross-scale feature fusion in this stage refers to: fusing the penultimate intermediate feature and the upsampled feature.
[0087] Then, the obtained features are subjected to two more upsampling operations, dual cross-scale feature fusion, and cross-stage feature fusion to obtain shallow features. In the first repetition, dual cross-scale feature fusion refers to fusing the third-to-last intermediate feature with the output feature of the previous upsampling module; in the second repetition, dual cross-scale feature fusion refers to fusing the fourth-to-last intermediate feature with the output feature of the previous upsampling module.
[0088] The extracted shallow features are output through the head network (i.e., Figure 2 The head network outputs features of size 160×160. This shallow feature is then subjected to three repetitions of "convolution operation, dual cross-scale feature fusion, and cross-stage feature fusion"; the features obtained in the first iteration are used as the mid-layer features and output through the head network (i.e., Figure 2 The head network outputs features of size 80×80; the features obtained the second time are used as deep features and output through the head network (i.e., Figure 2 The head network outputs features of size 40×40; the features obtained in the third step are used as the final features and output through the head network (i.e., Figure 2The head network outputs features of size 20×20. Specifically, in the first repetition, dual cross-scale feature fusion refers to fusing the shallow features (160×160) processed by the first convolutional module with features of size 80×80; in the second repetition, dual cross-scale feature fusion refers to fusing the output features of the second convolutional module with features of size 40×40; and in the third repetition, dual cross-scale feature fusion refers to fusing the enhanced features with the output features of the third convolutional module.
[0089] To illustrate the process of performing dual cross-scale feature fusion, such as Figure 7 and Figure 8 As shown, step 303 includes: Step 3031: Divide the channels of each of the intermediate features or the enhancement features into multiple groups.
[0090] First, the number of channels in each input feature map is separated into n groups.
[0091] This invention addresses the problem of poor performance of batch normalization (BN) when the number of channels is small by using group normalization.
[0092] Step 3032: Based on the standard deviation and mean of the features of each group of channels, as well as the trainable parameters, standardize the features of the channels to obtain intermediate features.
[0093] Then, the weight values of each group of feature maps are obtained through grouping standardization and normalization operations.
[0094] Step 3033: Normalize the intermediate features to obtain normalized features; use an activation function to process the product of the normalized features and the intermediate features to obtain the group weight value.
[0095] The activation function is determined by those skilled in the art based on the specific application scenario; in one embodiment, the activation function can be the Sigmoid activation function; the Sigmoid activation function maps the feature value of the product of the normalized feature and the intermediate feature to the range (0,1). The formula is shown below:
[0096] in, The input feature map (i.e., each intermediate feature or enhancement feature). This is a grouping standardization operation in an embodiment of the present invention. For normalization function, The standard deviation of the characteristics of each channel group is represented by the standard deviation of ... This represents the average value of the characteristics of each channel group. and These are trainable parameters, all of which are network parameters; It is a small positive number added to balance division, and its specific value is determined by those skilled in the art based on the specific use case.
[0097] It's important to note that each feature map group is not a sub-feature map containing multiple channels, but rather the channels of the feature map are divided into several groups. For example, assuming a feature map has 16 channels, i.e., n=4, these 16 channels are divided into 4 groups, each containing 4 channels. This grouping is done along the channel dimension, dividing the channels into several subsets. After grouping, the channels in each group are continuous channels in the original feature map, rather than splitting the original feature map into multiple independent sub-feature maps. Each feature map group still retains the height and width dimensions of the original feature map; only the number of channels is divided into different groups. For each feature map group, the mean and variance of all elements within that group are calculated. For example, if a feature map group has a size of H×W×C (where C represents the number of channels in each group), then the mean and variance of this data group in the spatial dimensions of height H and width W and the channel dimension are calculated. The calculated mean and variance are then used to standardize the feature map group. Through standardization, the data distribution of each feature map group is adjusted to a state with zero mean and unit variance.
[0098] Step 3034: Process the features of each group channel according to the group weight values to generate a fused feature map.
[0099] In one embodiment, step 3034 includes: The features of each channel are multiplied by the corresponding group weight value and the overall learnable weight to obtain a weighted feature map; the weighted feature maps of each channel are fused to generate a single feature map; the single feature map is mapped to a preset range, and the feature maps within the preset range are convolved to obtain a fused feature map.
[0100] The activation function is determined by those skilled in the art based on the specific application scenario; in one embodiment, the activation function can be a ReLU activation function; the ReLU activation function maps the feature values to a suitable range. Finally, a weighted feature fusion output feature map is obtained through 1x1 convolution, and the size of the output feature map is the same as the input size of each feature fusion. The formula is:
[0101] in, The characteristics of the i-th group of channels For the first The group weight values corresponding to the features of the group channels The overall learnable weights are network parameters. It is a relatively small positive number, which shall be determined by those skilled in the art based on the specific use case; The weight values are obtained by weighting the groups.
[0102] in, , Indicates the first Normalized weight values of the group channels, Represents the original first Group learnable weights, and Each represents the current channel group number.
[0103] The dual cross-scale feature fusion method of this invention solves the problem of target scale variation by fusing deep and shallow features and cross-scale features, thereby improving the model recognition accuracy.
[0104] In one embodiment, the expression for the loss function is:
[0105] in, The outlier is the outlier score; a higher outlier score indicates lower quality of the sample data. and All are regulatory factors; among them, regulatory factors and regulatory factors The specific value is selected by those skilled in the art based on the specific application scenario, and is not limited here. In object detection, the detection model often identifies the detected targets in the sample data in the form of bounding boxes; during the training phase, the detection model outputs information such as the coordinates of the predicted bounding boxes of each detected target during the iteration process; and the sample data of the dataset corresponds to the coordinates of the ground truth bounding boxes. The specific method for calculating the intersection-union ratio (IU) between the predicted bounding box and the ground truth bounding box shall be selected by those skilled in the art based on the specific application scenario. The width of the minimum bounding rectangle between the predicted bounding box and the ground truth bounding box. The height of the minimum bounding rectangle. The x-coordinate of the center point of the prediction box. The ordinate of the center point of the prediction box is... The x-coordinate of the center point of the real bounding box is given. The ordinate of the center point of the real bounding box is given. This indicates taking the moving average; for example, Indicates taking The moving average value is used; the specific method for obtaining the moving average value shall be selected by those skilled in the art based on the specific application scenario, and is not limited here.
[0106] In existing technologies, traditional concatenation operations simply stitch together semantically strong high-level feature maps with detailed low-level feature maps, ignoring the differences in contribution of different feature maps to a specific target. However, the dual cross-scale feature fusion method of this invention performs weighted fusion of different features: assigning higher weights to low-level feature maps and enhancing key details of small targets (e.g., edge textures), thus improving the recognition performance of small targets. To better fuse deep and shallow features and cross-scale features, this invention employs a dual cross-scale feature fusion approach. By grouping feature channels, it can identify and suppress background-related channels (e.g., repetitive textures and illumination noise), preserving the integrity of details and solving the problem of 90% detail loss in traditional methods; even under natural environmental interference, the target recognition accuracy is still guaranteed.
[0107] This invention addresses the real-time requirements of target recognition by considering algorithm convergence efficiency, achieving a balance between target recognition effectiveness and real-time performance. It mitigates the adverse effects of low-quality data on gradients from small targets, accelerates network convergence, and solves the problem of low recognition accuracy when the model encounters significant differences in target size, shape, and other features.
[0108] In one alternative embodiment, during the training phase, the target model is trained using the Stochastic Gradient Descent (SGD) optimization algorithm, the momentum parameter is set to 0, the learning rate is set using a linear learning rate warm-up strategy, and the number of iterations is initialized.
[0109] After obtaining the target model, the validation set is input into the target model to obtain the target detection results; in an optional embodiment, the performance of the target model is evaluated by target recognition performance evaluation metrics such as (accuracy, recall, class average precision, etc.).
[0110] Example 2: like Figure 9 The diagram shown is an architectural schematic of a target recognition network construction device based on feature fusion according to an embodiment of the present invention. This target recognition network construction device based on feature fusion includes one or more processors 21 and a memory 22. Figure 9 Take a processor 21 as an example.
[0111] Processor 21 and memory 22 can be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.
[0112] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the target recognition network construction method based on feature fusion in this embodiment. The processor 21 executes the target recognition network construction method based on feature fusion by running the non-volatile software program and instructions stored in the memory 22.
[0113] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may optionally include memory remotely located relative to processor 21, which can be connected to processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0114] The program instructions / modules are stored in the memory 22. When executed by one or more processors 21, they execute the target recognition network construction method based on feature fusion in the above embodiments, for example, executing each step of the target recognition network construction method based on feature fusion in the embodiments of the present invention described above.
[0115] This invention also provides a non-volatile computer storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 9 A processor 21 may enable one or more processors to execute the feature fusion-based target recognition network construction method in the specific embodiments of the present invention, for example, to execute the various steps of the feature fusion-based target recognition network construction method described above in the embodiments of the present invention; it may also implement Figure 9 The various modules and units described above; or the target recognition network construction method based on feature fusion in the specific embodiments of the present invention, for example, executing the various steps of the target recognition network construction method based on feature fusion in the embodiments of the present invention described above; can also achieve Figure 9 The aforementioned modules and units.
[0116] It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.
[0117] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0118] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a target recognition network based on feature fusion, characterized in that, include: Collect images of complex environments, perform data augmentation and annotation on the images to generate sample data, and construct a dataset based on the sample data; The sample data in the dataset undergoes multiple feature extractions, sparse feature processing, connection processing, and dynamic weighted averaging to obtain hierarchical features after each processing step. Each of the hierarchical features is fused based on dual cross-scale features to obtain a fused feature map; a head network is used to perform target detection on each of the fused feature maps to output the sample detection results; The sample detection results are processed based on the loss function to obtain the training state; When the training state meets the termination condition, the target model is obtained, which is then used to detect the complex environment image to be detected, and the target detection result is obtained.
2. The target recognition network construction method based on feature fusion according to claim 1, characterized in that, The method includes: The hierarchical features after the last processing are taken as deep features, and all hierarchical features other than the deep features are taken as intermediate features. The deep features are enhanced to obtain enhanced features; For each of the intermediate features and the enhanced features, a dual cross-scale feature fusion is performed to obtain the corresponding fused feature map.
3. The target recognition network construction method based on feature fusion according to claim 2, characterized in that, The method includes: The channels of each of the intermediate features or the enhanced features are divided into multiple groups; Based on the standard deviation and mean of the features of each group of channels, as well as the trainable parameters, the features of the channels are standardized to obtain intermediate features; The intermediate features are normalized to obtain normalized features; the product of the normalized features and the intermediate features is processed using an activation function to obtain the group weight value. The features of each group of channels are processed according to the group weight values to generate a fused feature map.
4. The target recognition network construction method based on feature fusion according to claim 3, characterized in that, The method includes: Multiply the features of each channel group by the corresponding group weight value and the overall learnable weight to obtain the weighted feature map; The weighted feature maps of each group of channels are fused to generate a single feature map; The single feature map is mapped to a preset range, and a convolution operation is performed on the feature map within the preset range to obtain a fused feature map.
5. The target recognition network construction method based on feature fusion according to claim 1, characterized in that, The method includes: The sample data is subjected to a convolution operation to extract features and obtain the first feature; The first feature is subjected to sparse feature processing to obtain the second feature; The second feature is subjected to cross-stage feature fusion for connection processing to obtain the third feature; The third feature is dynamically weighted and averaged to obtain the hierarchical feature.
6. The target recognition network construction method based on feature fusion according to claim 5, characterized in that, The method includes: The first feature is divided into multiple grids according to the scaling factor; The multiple grids are rearranged along the channel dimension to generate sub-feature maps; By sequentially connecting all the sub-feature maps along the channel dimension, a new feature map is obtained; The new feature map is subjected to a non-stepping convolution operation to obtain the second feature.
7. The target recognition network construction method based on feature fusion according to claim 6, characterized in that, The method includes: The third feature is subjected to global average pooling and processed using a multilayer perceptron and activation function to generate channel weights; The third feature is divided into G groups of semantic features; the three-dimensional features of the semantic features are mapped to the wide-dimensional and high-dimensional directions, and a single-kernel convolution operation is performed to obtain the features to be weighted; The feature to be weighted is multiplied by the channel weight, and a linear transformation is performed on the multiplied feature to obtain the first attention map; Perform multi-kernel convolution on the semantic features to obtain processed features; determine the product of the processed features and the channel weights as the second attention map; By aggregating the first attention map and the second attention map corresponding to the semantic features, hierarchical features are obtained.
8. The method for constructing a target recognition network based on feature fusion according to any one of claims 1-7, characterized in that, The expression for the loss function is: in, For outlier degree, and All are regulatory factors. The intersection-union ratio (IU) of the predicted bounding box and the ground truth bounding box. The width of the minimum bounding rectangle between the predicted bounding box and the ground truth bounding box. The height of the minimum bounding rectangle. The x-coordinate of the center point of the prediction box. The ordinate of the center point of the prediction box is... The x-coordinate of the center point of the real bounding box is given. The ordinate of the center point of the real bounding box is given by [reference]. This indicates that the moving average value is taken.
9. A target recognition network construction device based on feature fusion, characterized in that, The target recognition network construction device based on feature fusion includes at least one processor and a memory, which are connected via a data bus. The memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to implement the target recognition network construction method based on feature fusion as described in any one of claims 1-8.
10. A non-volatile computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which are executed by one or more processors to perform the target recognition network construction method based on feature fusion as described in any one of claims 1-8.