Adaptive Feature Fusion Method and System in Convolutional Neural Network

Through the adaptive feature fusion method, lightweight convolution branching and activation/normalization operations are used to solve the problem of conflicting training targets of adjacent scale feature extraction layers in the single-stage detection framework, and the detection accuracy and training convergence are improved.

CN114092760BActive Publication Date: 2025-06-10CRSC COMM & INFORMATION GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111310425.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-06-10
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

In the object detection, the existing single-stage detection framework forces the target samples to the feature extraction layer of a specific scale, resulting in conflicts in the training targets of the feature extraction layer of adjacent scales, affecting the detection accuracy and training convergence.

Method used

Adaptive feature fusion method is adopted to obtain the weight coefficients of different scale features through lightweight convolution branches, and nonlinear activation and linear normalization are performed to realize adaptive weighted fusion and splicing of different scale features.

Benefits of technology

The adaptability and convergence of convolutional neural networks to different training goals is improved, the overall accuracy of deep learning algorithms is improved, and manpower, material resources and time costs are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092760B_ABST
    Figure CN114092760B_ABST
Patent Text Reader

Abstract

The present invention relates to an adaptive feature fusion method and system in a convolutional neural network, which includes: obtaining weight coefficients of features of each scale of the current feature fusion layer; activating and normalizing the weight coefficients of the features of each scale of the current feature fusion layer; performing weighted fusion on the features of each scale in the current feature fusion layer, and splicing the results after weighted fusion to obtain an adaptive feature fusion result, thereby completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy. While the present invention improves the adaptability and convergence of feature fusion for different training objectives and the overall accuracy of deep learning algorithms, it can effectively save labor, material and time costs. The present invention can be widely applied in the fields of artificial intelligence technologies such as object detection, tracking, and semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an adaptive feature fusion method and system in a convolutional neural network for deep learning. Background Art

[0002] In recent years, with the application of deep learning based on convolutional neural networks (abbreviated as CNN), significant progress has been made in the research of image classification, object detection, semantic segmentation, etc. in the field of computer vision. Compared with algorithms based on handcrafted features, convolutional neural networks can very robustly learn expressive features, so they are widely used in the object detection process to extract object features.

[0003] Solutions for object detection have gradually converged under two mainstream frameworks: one is the two-stage detection frameworks represented by R-CNN, Fast-RCNN, Faster-RCNN, and R-FCN, and the other is the one-stage detection frameworks represented by YOLO, SSD, Retina-Net, etc. In particular, due to the huge speed advantage of the one-stage detection framework, it has been more widely applied in industry. In the one-stage detection framework, in order to improve the detection accuracy, SSD creatively tried to extract features of different scales in parallel from multiple different convolutional layers from low to high to deal with the detection of targets of different sizes respectively, thus achieving significantly better results than the previous one-stage framework that only extracts features from a single convolutional layer at the end, represented by YOLO. However, due to the limitations of its own network structure and working principle, it is difficult for the one-stage detection framework to utilize cascaded multiple detection opportunities like Faster RCNN, cut out the region of interest (RoI) where the target may exist in the convolutional map through preliminary detection, and then normalize the convolutional map of this region to a specified size as the initial input for subsequent detection, so as to precisely match the receptive field of the detector with the scale of the target features in a step-by-step progressive manner. Eventually, its adaptability to targets of different sizes is weaker than that of the two-stage detection framework. Recently, with the method of using a feature pyramid based on structures such as FPN to extract feature information of different scales in multiple layers, adjusting the convolutional maps corresponding to the feature information of different scales to the same size by using downsampling and top-down pathways, and finally fusing the deep semantic information and shallow location information, which is represented by Retina-Net, gradually becoming a common configuration of the one-stage detection framework, the adaptability of the one-stage detection framework to the target scale has been significantly improved. In the feature pyramid structure, typical feature fusion methods include the element-wise mode of adding the convolutional maps of features of each scale from different convolutional layers in the network pixel by pixel, or the concat mode of splicing the convolutional maps of features of each scale from different convolutional layers in the network, etc.

[0004] However, the single-stage detection framework with a feature pyramid is not absolutely perfect. Connecting the detectors for detecting targets of different sizes to the feature extraction convolutional layers at the corresponding feature scales (hereinafter referred to as the feature extraction layers at the corresponding scales) respectively will bring a new defect: during the training process, since the target samples are forced to be corresponded to the detectors behind the feature extraction layers at the corresponding scales according to the size of the target calibration box, although the adjacent feature extraction layers at the adjacent scales can also extract some target features at the target position in the original image, this type of algorithm forces the detectors behind these adjacent feature extraction layers to judge that the features of the target and its category do not exist. Eventually, it has a negative impact on both the judgment of the target category and the regression of the position by the detector, not only affecting the detection accuracy, but also this approach of giving contradictory training objectives between adjacent feature extraction layers will further lead to the difficulty of the algorithm training to converge. To solve the above problems, some researchers have proposed a data-driven feature fusion strategy, namely Adaptive Spatial Feature Fusion (ASFF), which can flexibly change the weights of features at different scales when fusing features of each layer. The ASFF strategy adaptively learns and adjusts the weights of features at different scales from each feature extraction layer at each position on the convolutional map through a series of learnable parameters, effectively alleviating the contradiction between the training objectives of the detectors corresponding to adjacent feature extraction layers during the backpropagation of errors. Classic algorithms such as YOLOv3 using this method have achieved a better trade-off between speed and accuracy than the standard version on the MS COCO dataset.

[0005] In the Adaptive Spatial Feature Fusion (ASFF) method, for each pixel (i, j) on the convolutional map of the feature fusion layer l, the weights of features at different scales from each feature extraction layer (this method uses a total of 3 scales of features) at this pixel position are adaptively learned. Suppose represents the value of the feature vector from the nth layer to the lth layer in the deep learning network at the pixel (i, j) position of the convolutional map of the lth layer, then the value of the fused result at the pixel (i, j) position in the convolutional map of the feature fusion layer l can be calculated by the following formula:

[0006]

[0007] where, respectively represent the normalized weight coefficients when features at different scales from 3 feature extraction layers are adaptively fused at the pixel (i, j) position of the convolutional map of the feature fusion layer l.

[0008] This method obtains three convolutional maps by adding an additional convolutional branch behind each of the three feature extraction layers, respectively, and takes the values at the pixel (i, j) position in the three convolutional maps as the weight coefficients of the features at three scales, then uses the softmax formula for normalization, normalizes the values of each weight coefficient itself to the interval [0, 1], and normalizes the sum of the values of each weight coefficient to 1, finally obtaining the normalized weight coefficients for feature fusion of the features at three scales

[0009]

[0010]

[0011]

[0012] Through the above operations, features from different scales can be adaptively fused at any feature fusion layer l, and the fused result can be used as the initial input for the subsequent detector to further improve the detection accuracy.

[0013] After 2020, with the development of feature fusion technology, more and more new generation object detection algorithms represented by YOLOV4 have gradually proved that the concat mode of vector splicing of features at various scales from different feature extraction layers in the network, because it more completely retains the detailed information of features at various scales, so the effect of feature fusion based on the vector splicing mode on improving the accuracy of subsequent object detection and other tasks is significantly better than the traditional mode of element-wise addition of features at various scales in units of pixels. However, the existing adaptive spatial feature fusion (ASFF) method can only be applied to the traditional feature fusion mode based on element-wise addition of pixels, and cannot be used for the above-mentioned mode based on vector splicing (concat). Currently, for the above-mentioned feature fusion mode based on vector splicing (concat), there is no feature fusion method that can adaptively adjust the weights of features at different scales. As a result, it is artificially forced that features at various scales are spliced with equal weights and participate in feature fusion. This approach cannot avoid contradictions between the training objectives of the detectors corresponding to adjacent feature extraction layers, which is obviously not conducive to improving the adaptability and convergence of convolutional neural networks with feature fusion structures to different training objectives, and ultimately affects the overall accuracy of object detection.

[0014] In summary, in the convolutional neural network of deep learning algorithms aimed at object detection, tracking, semantic segmentation, etc., for the features of various scales from different feature extraction layers, to develop an adaptive feature fusion method based on the vector concatenation (concat) feature fusion mode that can flexibly adjust the weights of different-scale features through self-learning, thereby avoiding contradictions in the training objectives of each feature extraction layer and maintaining the computational complexity of the algorithm at a relatively low level has become an urgent problem to be solved. Summary of the Invention

[0015] In the feature fusion mode based on vector concatenation (concat) in the convolutional neural network, where the features of each scale are artificially forced to participate in feature fusion with equal weights, it is impossible to avoid contradictions between the training objectives of the detectors corresponding to adjacent feature extraction layers, which is not conducive to improving the adaptability and convergence of the convolutional neural network with feature fusion to different training objectives, and ultimately affects the overall accuracy of deep learning. The purpose of the present invention is to provide an adaptive feature fusion method and system in the convolutional neural network, which can effectively save human, material and time costs while improving the adaptability and convergence of feature fusion to different training objectives and the overall accuracy of the deep learning algorithm.

[0016] To achieve the above object, the present invention adopts the following technical solutions: An adaptive feature fusion method in a convolutional neural network, which includes: obtaining the weight coefficients of the features of each scale of the current feature fusion layer; activating and normalizing the weight coefficients of the features of each scale of the current feature fusion layer; performing weighted fusion on the features of each scale in the current feature fusion layer, and concatenating the results of the weighted fusion to obtain an adaptive feature fusion result, completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy.

[0017] Further, the obtaining of the weight coefficients of the features of each scale of the current feature fusion layer includes:

[0018] At the current feature fusion layer, fusing the features of different scales from different feature extraction layers, and scaling the convolution maps corresponding to the features of all scales to the same size through downsampling or upsampling operations;

[0019] Sending the convolution maps of the features of different scales from each feature extraction layer to a lightweight convolution branch respectively;

[0020] Taking the values of the results of different convolution branches at any pixel position as the weight coefficients of the features of each scale at the pixel position of the convolution image of the current feature fusion layer.

[0021] Further, the activation and normalization of the weight coefficients of the features at each scale of the current feature fusion layer includes:

[0022] Non-linearly activate the weight coefficients of the features at each scale at any pixel position on the convolution map of the current feature fusion layer;

[0023] Linearly normalize the weight coefficients of the non-linearly activated features to obtain the normalized weight coefficients of the features at each scale;

[0024] Obtain the normalized weight coefficients at all pixel positions on the convolution map of the current feature fusion layer.

[0025] Further, the weighted fusion of the features at each scale in the current feature fusion layer includes:

[0026] At any pixel position on the convolution map of the current feature fusion layer, respectively perform weighted fusion of the features at each scale with the features at other scales;

[0027] Obtain the results of the weighted fusion of the features at each scale with the features at other scales at all pixel positions on the convolution map of the current feature fusion layer.

[0028] Further, the weighted fusion of the features at each scale with the features at other scales at any pixel position on the convolution map of the current feature fusion layer includes:

[0029] If the normalized weight coefficient α of the feature at the m-th scale at the pixel (i, j) position on the convolution map of the feature fusion layer l l,m,ij is greater than or equal to the mean 1 / M of the normalized weight coefficients of all M different scales of features, then the result of the weighted fusion of the feature at this scale with the features at other scales at the (i, j) position is equal to itself.

[0030] Further, the weighted fusion of the features at each scale with the features at other scales at any pixel position on the convolution map of the current feature fusion layer includes:

[0031] If the normalized weight coefficient α of the feature at the m-th scale at the pixel (i, j) position on the convolution map of the feature fusion layer l l,m,ij is less than the mean 1 / M of the normalized weight coefficients of all M different scales of features, then the feature x at this scale at the (i, j) position l,m,ij after the weighted fusion with the features at other scales is:

[0032]

[0033] where, It represents the weighted mean of the features with normalized weight coefficients greater than 1 / M among the features of other scales at the position of the convolutional image pixel (i, j) in the feature fusion layer l.

[0034] Furthermore, the weighted mean is calculated as follows:

[0035]

[0036] where Max[*, *] represents the value of taking the larger one between the two in the brackets.

[0037] Furthermore, if the normalized weight coefficient α of the feature of the m-th scale at the position of the convolutional image pixel (i, j) in the feature fusion layer l l,m,ij is less than the mean 1 / M of the normalized weight coefficients of all M different-scale features, and M equals 2, then the result of weighted fusion of the feature x of this scale at the position (i, j) l,m,ij with the features of other scales is:

[0038]

[0039] where x l,n≠m,ij represents the one that is different from the m-th among the 2-scale features at the position of the convolutional image pixel (i, j) in the feature fusion layer l.

[0040] Furthermore, the splicing of the results of weighted fusion includes:

[0041] Splicing the results of weighted fusion of the features of each scale from M feature extraction layers on the convolutional map of the feature fusion layer l with the features of other scales in the order of 1, …, M on a preset dimension to obtain the adaptive fusion result Y l of the features of each scale:

[0042] Y l =(X l,1 , X l,2 , …, X l,M )

[0043] where X l,1 , X l,2 , …, X l,M are respectively vector matrices composed of the numerical values of the results of weighted fusion of the features of each scale at the feature fusion layer l with the features of other scales at all pixel positions in their own convolutional maps. Concatenating the respective vector matrices in the concat mode on a preset dimension forms the new vector matrix Y lIt is the adaptive feature fusion result of the features at each scale of the feature fusion layer l.

[0044] An adaptive feature fusion system in a convolutional neural network, which includes: a weight coefficient acquisition module, a weight coefficient activation and normalization module, and a feature weighted fusion and splicing module;

[0045] The weight coefficient acquisition module is used to acquire the weight coefficients of the features at each scale of the current feature fusion layer;

[0046] The weight coefficient activation and normalization module is used to activate and normalize the weight coefficients of the features at each scale of the current feature fusion layer;

[0047] The feature weighted fusion and splicing module performs weighted fusion on the features at each scale in the current feature fusion layer, and splices the results after weighted fusion to obtain an adaptive feature fusion result, completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy.

[0048] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0049] 1. Different from the previous adaptive spatial feature fusion method represented by ASFF in the traditional feature fusion mode based on element-wise addition of pixels, the present invention relies on lightweight convolutional branches and a simple calculation process to achieve adaptive feature fusion in a more advanced feature fusion mode based on vector concat, thereby improving the adaptability and convergence of the convolutional neural network to different training objectives, as well as the overall accuracy of the deep learning algorithm.

[0050] 2. Through non-linear activation and linear normalization operations, the present invention ensures that the values of the weight coefficients of the features at each scale are between 0 and 1 and the sum is equal to 1. In particular, by using the saturation region of the non-linear activation function, it avoids the violent oscillation caused by the rapid further amplification of the gap between those weight coefficients with larger values during training. Then, by using linear normalization to reduce the computational amount, the stability and efficiency of weight coefficient calculation are improved.

[0051] 3. The present invention integrates the loss of generating the weight coefficients of the features at each scale with the entire convolutional neural network through its lightweight convolutional branches and participates in end-to-end training, without the need for additional manual operations such as complex sample calibration or parameter adjustment according to intermediate results during the training process.

[0052] 4. The present invention can be conveniently nested in the convolutional neural network of algorithms with feature fusion structures such as object detection, tracking, and semantic segmentation. The improvement of the accuracy of related algorithms does not come at the cost of significantly sacrificing the running speed. The lightweight convolutional branch structure and simple and efficient feature weighted fusion calculation ensure that the running speed of the algorithm after adding the present invention is close to that of the corresponding original algorithm.

[0053] In summary, the present invention can be widely used in the fields of artificial intelligence technologies such as object detection, tracking, and semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a schematic diagram of the overall process of the fusion method in an embodiment of the present invention;

[0055] Figure 2 is a schematic diagram of the principle of weighted fusion and splicing of features at each scale in the feature fusion layer in an embodiment of the present invention (taking the scenario of 3 feature extraction layers and 3 feature fusion layers as an example). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.

[0057] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0058] The adaptive feature fusion method and system in the convolutional neural network provided by the present invention are used for feature fusion relying on adaptive weights in the convolutional neural network for deep learning. Its main steps include: obtaining the weight coefficients of the features of each scale in the current feature fusion layer; activating and normalizing the weight coefficients of the features of each scale in the current feature fusion layer; performing weighted fusion on the features of each scale in the current feature fusion layer; splicing the results after weighted fusion of the features of each scale in the current feature fusion layer; and obtaining the adaptive feature fusion results of all feature fusion layers in the convolutional neural network. The present invention uses the results of lightweight convolutional branches, combines operations such as activation and normalization, generates normalized weight coefficients for the features of different scales from each feature extraction layer, and then uses the normalized weight coefficients to perform adaptive weighted fusion and splicing on the features of each scale, thereby solving the problem of adaptive feature fusion in the vector splicing mode. While improving the adaptability and convergence of the convolutional neural network to different training objectives and the overall accuracy of the deep learning algorithm, it can effectively save human, material and time costs.

[0059] In an embodiment of the present invention, an adaptive feature fusion method in a convolutional neural network is provided. In this embodiment, taking the application of this method to a terminal as an example, it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, a lightweight convolutional branch is used to generate weight coefficients for the features of different scales from each feature extraction layer at the feature fusion layer of the convolutional neural network, and through non-linear activation and normalization operations, ensure that the values of each weight coefficient are between 0 and 1 and the sum is equal to 1. Then, the above weight coefficients are used to perform weighted fusion on the features of each scale with the features of other scales respectively. Finally, the results after weighted fusion of the features of each scale are spliced to obtain the adaptive feature fusion results. The above method solves the problem of adaptive feature fusion in the feature fusion operation based on the vector splicing (concat) mode in the convolutional neural network, improves the accuracy of feature fusion without significantly increasing the computational complexity, and thereby enhances the overall performance of the deep learning algorithm.

[0060] Specifically, as Figure 1 shown, in this embodiment, the method includes the following steps:

[0061] Step 1: Obtain the weight coefficients of the features of each scale in the current feature fusion layer;

[0062] In this embodiment, the obtaining of the weight coefficient includes the following steps:

[0063] Step 1.1: At the current feature fusion layer, fuse the features of different scales from different feature extraction layers, and scale the convolution maps corresponding to the features of all scales to the same size through downsampling or upsampling operations;

[0064] Step 1.2: Send the convolution maps of the features of different scales from each feature extraction layer to a lightweight convolution branch respectively;

[0065] Step 1.3: Take the values at any pixel position of the results of different convolution branches as the weight coefficients of the features of each scale at the pixel position of the convolution image of the current feature fusion layer.

[0066] Specifically: Assume that at the current feature fusion layer l in the convolutional neural network, it is necessary to fuse the features of M scales from M different feature extraction layers. Scale (resize) the convolution maps corresponding to the features of all scales to the same size through downsampling or upsampling operations, and then send the convolution maps of the features of different scales from each feature extraction layer to a lightweight convolution branch with a convolution kernel size of 1*1. Take the values at any pixel (i, j) position of the results of the above M convolution branches as the weight coefficients λ of the features of each scale at the pixel (i, j) position of the convolution image of the feature fusion layer l l,1,ij , λ l,2, ij, …, λ l,M,ij .

[0067] Through the weight coefficients calculated from its results, the above lightweight convolution branch participates in the backpropagation of the reverse error in the training process of the deep learning basic network in a way that affects the weighted fusion result of the features of each feature extraction layer by the feature fusion layer. The entire training process is an end-to-end mode and does not require additional manual intervention (such as additional labeled samples or specified hyperparameters, etc.).

[0068] Step 2: Activate and normalize the weight coefficients of the features of each scale of the current feature fusion layer;

[0069] Step 2.1: Perform non-linear activation on the weight coefficients of the features of each scale at any pixel position on the convolution map of the current feature fusion layer:

[0070] First, to avoid the gap between those larger values in the weight coefficients of features at various scales from being further amplified too quickly, which may lead to unstable oscillations during the training process, and to ensure that the weight coefficients of features at each scale are greater than 0, the non-linear activation function Sigmoid is used to non-linearly activate the weight coefficients of features at each scale, so that the change of the activated weight coefficients shows a relatively rapid linear trend near the center point of the value range; while in the region far from the center point, it shows a relatively slow non-linear saturation trend. Taking the weight coefficient λ of the feature at the m-th (m ∈ [1, …, M]) scale at the position (i, j) of the convolutional image pixels in the feature fusion layer l as an example, its activated weight coefficient is obtained using the following formula l,m,ij For example, the following formula is used to obtain its activated weight coefficient

[0071]

[0072] Repeat the above process to obtain the activated weight coefficients of features at each scale at the position (i, j) of the convolutional image pixels in the feature fusion layer l

[0073] Step 2.2: Linearly normalize the weight coefficients of the non-linearly activated features to obtain the normalized weight coefficients of features at each scale:

[0074] Since in Step 2.1, the violent oscillations caused by the rapid further amplification of the gap between larger weight coefficients during training have been avoided through the saturation region of the non-linear activation function, and the values of the non-linearly activated weight coefficients of features at each scale are ensured to be greater than zero, this step directly uses linear normalization to ensure that the sum of the weight coefficients of features from different scales at the position (i, j) of the convolutional image pixels is equal to 1. The reasons for not using non-linear normalization functions represented by SoftMax such as ASFF in the present invention include reducing the computational amount and avoiding negative impacts on the saturation region of the non-linear activation function. Taking the activated weight coefficient of the feature at the m-th (m ∈ [1, …, M]) scale at the position (i, j) of the convolutional image pixels in the feature fusion layer l as an example, its normalized weight coefficient α is obtained using the following formula For example, the following formula is used to obtain its normalized weight coefficient α l,m,ij :

[0075]

[0076] Since the values of the activated weight coefficients of features at each scale are greater than zero, there is no situation where the denominator of the above formula is equal to 0

[0077] Repeat the above process to obtain the normalized weight coefficients α of features at each scale at the position (i, j) of the convolutional image pixels in the feature fusion layer l l,1,ij, α l,2, ij, …, α l,M,ij 。

[0078] Step 2.3: Obtain the normalized weight coefficients at all pixel positions on the convolutional map of the current feature fusion layer;

[0079] Repeat the operations in Step 2.1 and Step 2.2 at each pixel position on the convolutional map of the feature fusion layer l until the normalized weight coefficients at all pixel positions are obtained.

[0080] Step 3: Perform weighted fusion on the features of each scale at the current feature fusion layer (as Figure 2 shown), and splice the results after weighted fusion to obtain the adaptive feature fusion result, completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy.

[0081] Among them, the weighted fusion includes the following steps:

[0082] Step 3.1.1: At any pixel position on the convolutional map of the current feature fusion layer, perform weighted fusion of the features of each scale with the features of other scales respectively:

[0083] This step uses the normalized weight coefficients α l,1,ij , α l,2,ij , …, α l,M,ij of the M-scale features at the pixel (i, j) position of the convolutional map of the feature fusion layer l to perform weighted fusion on the features of each scale x l,1,ij , xl ,2,ij , …, x l,M,ij . For the value x l,m,ij of the above-mentioned arbitrary m-th (m ∈ [1, …, M]) scale feature at the pixel (i, j) position of the convolutional map of the feature fusion layer l, the specific method of weighted fusion is as follows:

[0084] If the normalized weight coefficient α l,m,ij of the m-th scale feature at the pixel (i, j) position of the convolutional map of the feature fusion layer l is greater than or equal to the mean value 1 / M of the normalized weight coefficients of all M-scale features, then the result of weighted fusion of the feature of this scale with the features of other scales at the (i, j) position is equal to itself.

[0085]

[0086] Conversely, if the normalized weight coefficient α l,m,ij of the m-th scale feature at the pixel (i, j) position of the convolutional map of the feature fusion layer l is less than the mean value 1 / M of the normalized weight coefficients of all M-scale features, then the feature x of this scale at the (i, j) positionl,m,ij The result after weighted fusion with features of other scales can be calculated using the following formula.

[0087]

[0088] Wherein, represents the weighted mean of features with normalized weight coefficients greater than 1 / M among the features of other scales at the position of the convolutional image pixel (i, j) in the feature fusion layer l, and it can be calculated using the following formula.

[0089]

[0090] Wherein, Max[*, *] represents the larger value between the two in the brackets. Since the prerequisite for the execution of this formula is that the non-linear activation normalized weight coefficient α l,m,ij of the features of the m-th scale is less than 1 / M, so at least one of the non-linear activation normalized weight coefficients α l,m,ij of all scales of features is greater than 1 / M, that is, there is no situation where the denominator of the above formula is equal to 0.

[0091] Particularly, when the number M of features of different scales is equal to 2, Formula 6 can be further simplified to the following form:

[0092]

[0093] Wherein, x l,n≠m,ij represents the one that is inconsistent with the m-th one among the 2-scale features at the position of the convolutional image pixel (i, j) in the feature fusion layer l. At this time, it is not necessary to use Formula 7 to further calculate the weighted mean of features with normalized weight coefficients greater than 1 / M.

[0094] At the position of the convolutional image pixel (i, j) in the feature fusion layer l, the above operation is repeated for the features of each scale until the result of weighted fusion of the features of each scale with the features of other scales is obtained.

[0095] Step 3.1.2: Obtain the result of weighted fusion of the features of each scale with the features of other scales at all pixel positions on the convolutional map of the current feature fusion layer:

[0096] The operation in Step 3.1.1 is repeated at each pixel position of the convolutional map of the feature fusion layer l until the result of weighted fusion of the features of each scale with the features of other scales at all pixel positions is obtained.

[0097] In this embodiment, the results of weighted fusion of the features of each scale in the current feature fusion layer are concatenated (as Figure 2 shown), including:

[0098] At the convolutional map of the feature fusion layer l, the weighted fusion results of the features of each scale from M feature extraction layers with the features of other scales are concatenated in the preset dimension in the order of 1, …, M to obtain the adaptive fusion result Y of the features of each scale. l .

[0099] Y l = (X l,1 , X l,2 , …, X l,M ) (9)

[0100] where X l,1 , X l,2 , …, X l,M are respectively vector matrices composed of the numerical values of the weighted fusion results of the features of each scale at the feature fusion layer l with the features of other scales at all pixel positions in their own convolutional maps. The above-mentioned vector matrices are concatenated in the concat mode in the preset dimension to form a new vector matrix Y l which is the adaptive feature fusion result of the features of each scale of the feature fusion layer l.

[0101] Repeat steps 1 to 3 for each feature fusion layer in the convolutional neural network that requires feature fusion until the adaptive feature fusion results at all feature fusion layers are obtained.

[0102] In an embodiment of the present invention, an adaptive feature fusion system in a convolutional neural network is provided, which includes: a weight coefficient acquisition module, a weight coefficient activation and normalization module, and a feature weighted fusion and concatenation module;

[0103] The weight coefficient acquisition module is used to acquire the weight coefficients of the features of each scale of the current feature fusion layer;

[0104] The weight coefficient activation and normalization module is used to activate and normalize the weight coefficients of the features of each scale of the current feature fusion layer;

[0105] The feature weighted fusion and concatenation module performs weighted fusion on the features of each scale at the current feature fusion layer and concatenates the weighted fusion results to obtain an adaptive feature fusion result, completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy.

[0106] The system provided in this embodiment is used to execute the above-mentioned method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive feature fusion method in a convolutional neural network, used for object detection, tracking, and semantic segmentation, Characterized in that, It includes: Obtain the weight coefficients of the features of each scale in the current feature fusion layer; Activate and normalize the weight coefficients of the features of each scale in the current feature fusion layer; Perform weighted fusion on the features of each scale in the current feature fusion layer, and splice the results of the weighted fusion to obtain the adaptive feature fusion result, complete the adaptive feature fusion in the convolutional neural network, and improve the detection accuracy; The performing weighted fusion on the features of each scale in the current feature fusion layer includes: At any pixel position on the convolution map of the current feature fusion layer, perform weighted fusion on the features of each scale with the features of other scales respectively; Obtain the results of weighted fusion of the features of each scale with the features of other scales at all pixel positions on the convolution map of the current feature fusion layer; The performing weighted fusion on the features of each scale with the features of other scales at any pixel position on the convolution map of the current feature fusion layer includes: If the normalized weight coefficient α of the feature at the m-th scale at the convolutional image pixel (i, j) of the feature fusion layer l l,m,ij is greater than or equal to the mean 1 / M of the normalized weight coefficients of the features of all M different scales, then the result of weighted fusion of the feature of this scale and the features of other scales at the (i, j) position is equal to itself.

2. The method according to claim 1, Characterized in that, The obtaining the weight coefficients of the features of each scale in the current feature fusion layer includes: At the current feature fusion layer, fuse the features of different scales from different feature extraction layers, and scale the convolution maps corresponding to the features of all scales to the same size through downsampling or upsampling operations; Send the convolution maps of the features of different scales from each feature extraction layer to a lightweight convolution branch respectively; Take the values at any pixel position of the results of different convolution branches as the weight coefficients of the features of each scale at the pixel position of the convolution map of the current feature fusion layer.

3. The method according to claim 1, Characterized in that: The activating and normalizing the weight coefficients of the features of each scale in the current feature fusion layer includes: Perform non-linear activation on the weight coefficients of the features of each scale at any pixel position on the convolution map of the current feature fusion layer; Perform linear normalization on the weight coefficients of the features after non-linear activation to obtain the normalized weight coefficients of the features of each scale; Obtain the normalized weight coefficients at all pixel positions on the convolution map of the current feature fusion layer.

4. The method according to claim 1, Characterized in that: The performing weighted fusion on the features of each scale with the features of other scales at any pixel position on the convolution map of the current feature fusion layer includes: If the normalized weight coefficient α of the feature at the m-th scale at the convolutional image pixel (i, j) of the feature fusion layer l l,m,ij is less than the mean 1 / M of the normalized weight coefficients of the features of all M different scales, then the feature x of this scale at the (i, j) position l,m,ij after weighted fusion with the features of other scales is: Among them, represents the weighted mean of the features with normalized weight coefficients greater than 1 / M among the features of other scales at the position of the convolutional image pixel (i, j) in the feature fusion layer l.

5. The method according to claim 4, Characterized in that: The weighted mean is calculated as follows: Among them, Max[*, *] represents taking the larger value between the two in the brackets.

6. The method according to claim 4, Characterized in that: If the normalized weight coefficient α of the feature of the m-th scale at the convolutional image pixel (u, j) position of the feature fusion layer l l,m,ij is less than the mean value 1 / M of the normalized weight coefficients of the features of all M different scales, and when M is equal to 2, then the feature x of this scale at the (i, j) position l,m,ij after weighted fusion with the features of other scales is: where x l,n≠m,ij represents the one out of the two-scale features at the convolutional image pixel (i, j) position in the feature fusion layer l that is inconsistent with the m-th one.

7. The method according to claim 1, Characterized in that: The splicing the results of the weighted fusion includes: The weighted fusion results of the features of each scale from the M feature extraction layers on the convolutional map of the feature fusion layer l and the features of other scales are concatenated in the preset dimension in the order of 1,..., M to obtain the adaptive fusion result Y of the features of each scale l : Y l = (X l,1 , X l,2 , …, X l,M ) Among them, X l,1 , X l,2 , …, X l,M are respectively vector matrices composed of the numerical values of the results of weighted fusion of the features at each scale at all pixel positions in their own convolution maps with the features of other scales at the feature fusion layer l. The new vector matrix Y l formed by concatenating the respective vector matrices in the concat mode in a preset dimension is the adaptive feature fusion result of the features at each scale of the feature fusion layer l.

8. An adaptive feature fusion system in a convolutional neural network, used for object detection, tracking, and semantic segmentation, Characterized in that, It includes: A weight coefficient acquisition module, a weight coefficient activation and normalization module, and a feature weighted fusion and splicing module; The weight coefficient acquisition module is used to acquire the weight coefficients of the features at each scale of the current feature fusion layer; The weight coefficient activation and normalization module is used to activate and normalize the weight coefficients of the features at each scale of the current feature fusion layer; The feature weighted fusion and splicing module performs weighted fusion on the features at each scale in the current feature fusion layer, and splices the results of the weighted fusion to obtain an adaptive feature fusion result, completing the adaptive feature fusion in the convolutional neural network and improving the detection accuracy; The performing weighted fusion on the features at each scale in the current feature fusion layer includes: At any pixel position on the convolution map of the current feature fusion layer, the features at each scale are respectively weighted and fused with the features at other scales; Obtain the results of the weighted fusion of the features at each scale with the features at other scales at all pixel positions on the convolution map of the current feature extraction fusion layer; The respectively performing weighted fusion on the features at each scale with the features at other scales at any pixel position on the convolution map of the current feature fusion layer includes: If the normalized weight coefficient α of the feature at the m-th scale at the convolutional image pixel (i, j) of the feature fusion layer l l,m,ij is greater than or equal to the mean 1 / M of the normalized weight coefficients of the features of all M different scales, then the result of the weighted fusion of the feature of this scale and the features of other scales at the (i, j) position is equal to itself.

Citation Information

Patent Citations

  • Pedestrian re-identification method and device, computer equipment and storage medium

    CN112183295A