An Object Detection Method Based on Histogram of Oriented Gradients and Improved Capsule Network

The integration of direction gradient histograms with an improved CapsNet structure addresses the limitations of traditional CNNs and CapsNets by enriching input information and reducing redundancy, resulting in enhanced object detection accuracy and robustness.

CN114677558BActive Publication Date: 2025-07-15HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210234245.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-07-15
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

In the prior art, traditional convolutional neural networks have insufficient attention to the position and direction information of image elements in object detection, the pooling layer information is lost, and capsule networks have problems such as redundancy and high algorithm complexity when detecting the same type of object.

Method used

The directional gradient histogram is used to combine with the improved capsule network, comprehensive features are extracted through the parallel convolution network, and the digital capsule layer retains position and direction information, the parallel convolution layer is fused with the HOG feature, de-redundant main capsule network, and deconvolution image reconstruction network to realize image reconstruction.

Benefits of technology

It improves the accuracy and robustness of object detection, reduces network redundancy, optimizes training time, and enhances the ability to extract and detect image information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677558B_ABST
    Figure CN114677558B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method based on Histogram of Oriented Gradients (HOG) and improved capsule network. By fusing the HOG of the target image and the convolutional feature map in parallel to combine the edge contour features of the image and the field of view features of the convolutional kernel, and then using this as the input of the improved capsule network. The improved capsule network extracts comprehensive features using a parallel convolutional network, forms feature vectors through a redundancy-reducing capsule network, and realizes image reconstruction using a deconvolutional image reconstruction network to train the network model. Finally, a convolutional layer of 3*3*256 and two parallel 1*1 convolutional kernels are used to extract the center point feature map of the detection box and the scale feature map of the detection box, form the corresponding target box and output the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly to a target detection method based on histogram of oriented gradients and improved capsule network. Background Art

[0002] Traditional convolutional neural networks extract target features through convolutional operations and achieve network learning through backpropagation to achieve the purpose of target detection, and have achieved remarkable results in target detection tasks. However, traditional convolutional neural networks have problems such as insufficient attention to information such as the relative position and direction of elements in the image and information loss in the pooling layer, and factors such as obstacle occlusion and bad weather will have a serious impact on detecting and recognizing targets. The capsule neural network proposes the concept of capsules to replace neurons in some convolutional neural networks, and expands the scalars of traditional neural networks into vectors, effectively overcoming the deficiencies of traditional neural networks. Therefore, applying the capsule neural network to target detection has a higher overall recognition rate, and the capsule neural network has higher accuracy and robustness to meet the target detection tasks under different influencing factors.

[0003] However, there is only one capsule of a given type at a given position in the traditional capsule network. Therefore, if two capsules in a capsule network are too close to each other, two objects of the same type cannot be detected. Chen Lichao proposed in the article "Capsule Network with Convolutional Features of Histogram of Oriented Gradients for Vehicle Type Classification under Traffic Surveillance" to preprocess the original image using the histogram of oriented gradients, which solves the problem of not being able to detect two objects of the same type to a certain extent. However, due to the lack of improvement in the structure of the capsule network itself, there are problems such as insufficient extraction of original image information, redundancy in the capsule structure, and high algorithm complexity. Summary of the Invention

[0004] The purpose of the present invention is to solve the technical problems of insufficient extraction of original image information, redundancy in the capsule structure, and high algorithm complexity in the prior art, and to provide a target detection method based on histogram of oriented gradients and improved capsule network.

[0005] A target detection method based on histogram of oriented gradients and improved capsule network includes the following steps:

[0006] 1) Obtain the original target image, mark the target position using a marking tool, and then randomly select different images as the training set;

[0007] 2) Parallelly fuse the histogram of oriented gradients (HOG) of the original image and the convolutional feature map to combine the edge contour features of the image and the visual field features of the convolutional kernel, and then use this as the input of the improved capsule network;

[0008] 3) The improved capsule network uses a parallel convolutional network to extract comprehensive features, forms a feature vector through a redundancy-removing capsule network, and realizes image reconstruction using a deconvolution image reconstruction network;

[0009] 4) Use a convolutional layer of 3*3*256 and two parallel 1*1 convolutional kernels to extract the center point feature map of the detection box and the scale feature map of the detection box, form the corresponding target box and output the detection result.

[0010] Preferably, the specific steps of the histogram of oriented gradients (HOG) and convolutional feature map parallel fusion algorithm in step 2) are as follows:

[0011] 2-1) Normalization; first divide the target image into 4 cells, divide each cell into 9 blocks, and perform Gamma normalization processing, and at the same time perform parameter optimization on the Gamma correction value; the normalization formula is as follows:

[0012]

[0013]

[0014] In this formula, τ represents the feature vector of the block, ε takes a small constant value, represents the i-th block in the j-th cell of the target image, and f represents the target image after normalizing all block feature vectors.

[0015] 2-2) Select the detection window from the normalized image. Select a detection window with the same aspect ratio as the image and not exceeding half of the image size.

[0016] 2-3) Select blocks from the window. Select rectangular blocks with equal length and width according to the detection window.

[0017] 2-4) Divide cell units within the block. Use a square cell with a size of 8*8 pixels within the rectangular block as the smallest unit for feature extraction within the block to divide the block.

[0018] 2-5) Perform directional projection within the cell. Divide 9 directions within the cell, and extract directional information every 20° as an angle range. The directional information is obtained by performing a convolution operation on the grayscale image I and the gradient template U in the x horizontal direction and the y vertical direction. The mathematical formula is as follows:

[0019] G x (x,y) = H(x + 1,y) - H(x - 1,y)

[0020] G y (x,y) = H(x,y + 1) - H(x,y - 1)

[0021]

[0022]

[0023] In this formula, H(x, y) represents the gray value at the corresponding coordinates, G x represents the horizontal direction gradient value, G y represents the vertical direction gradient value, G represents the gradient amplitude, and α represents the gradient direction.

[0024] 2 - 6) Normalize within the cell. Count the actual number of direction angles in each direction angle range within the cell to obtain a direction histogram, and select the direction angle with the most concentrated angle direction as the direction of the cell.

[0025] 2 - 7) Construct HOG features within the block. Count the actual number of direction angles in each cell's direction angle range within the block to obtain a direction histogram, and select the direction angle with the most concentrated angle direction as the direction of the block.

[0026] 2 - 8) If it has not reached the last block, return to step 2 - 3).

[0027] 2 - 9) If it has not reached the last window, return to step 2 - 2), otherwise obtain the direction gradient histogram.

[0028] 2 - 10) Single - row convolution. Input the original image, and use a single - row convolutional layer to extract features from the original image to obtain a convolutional feature map.

[0029] 2 - 11) Feature splicing and parallel fusion. Connect the dimensions of the convolutional feature map and the direction gradient histogram, and connect the two feature maps in the third dimension to obtain an image of 28 * 28 * 1.

[0030] The pooling layer structure adopted in traditional convolutional neural networks does not consider the relative spatial relationship, resulting in the loss of some valuable information in this layer. To solve this problem, a digital capsule layer is used in the capsule network to perform the function of the pooling layer, and vector neurons are proposed. The vector neurons store information such as direction and position in vector form and continuously transmit this information in the network, enabling the capsule network to be sensitive to changes in the position and direction information of elements in the image.

[0031] The improved capsule network based on the histogram of oriented gradients proposed by the present invention performs convolution layer and HOG feature extraction in parallel. Its image pre - processing part uses HOC - C features. While retaining the advantages of the original capsule network, the introduction of the histogram of oriented gradients strengthens the extraction of edge feature information of the target detection image. By the gradient direction and amplitude, the distinguishability between two overly close capsule networks is increased, and to a certain extent, the problem of being unable to detect two objects of the same type is solved.

[0032] Preferably, the specific steps for improving the capsule network algorithm in step 3) are as follows:

[0033] 3-1) Use a parallel convolutional network to extract comprehensive features.

[0034] 3-2) Use a redundant capsule network to generate feature vectors.

[0035] 3-3) Use a deconvolutional image reconstruction network to restore the original image and evaluate the network loss.

[0036] Preferably, the specific steps for the parallel convolutional network algorithm in step 3-1) are as follows:

[0037] 3-1-1) First, use a parallel convolutional neural network as the feature extraction network. Take the image after parallel fusion as the input, with the image size being 28*28*1. The parallel convolutional neural network uses 4 convolutional kernels in the convolutional layer, with the convolutional kernel sizes being 3, 5, 7, and 9 respectively, the number of convolutional kernels being selected as 32, and the stride being 2.

[0038] 3-1-2) Boundary padding. Adjust the padding size to perform boundary padding on the original matrix.

[0039] 3-1-3) Feature extraction. The nonlinear function in the feature extraction layer uses the PReLU function, and its mathematical formula is as follows:

[0040] PReLU(x) = max(0, x) + α * min(0, x)

[0041] In this formula, α is the learning rate.

[0042] 3-1-4) Feature tensor connection. Connect the feature tensors in the third dimension to obtain a feature tensor of 14*14*128.

[0043] Preferably, the specific steps for the redundant capsule network algorithm in step 3-2) are as follows:

[0044] 3-2-1) Input redundancy removal. Use the output of the parallel convolutional network as the input to the redundant main capsule network, and use a 1*1 convolutional kernel to remove redundant capsules, so that the 16*16 feature map after feature extraction is transformed into a 14*14 feature image, and the number of capsules is reduced to 196.

[0045] 3-2-2) Input capsule scalar u i .

[0046] 3-2-3) The vector obtained by multiplying the input capsule vector by the transformation matrix The mathematical formula is as follows

[0047]

[0048] In this formula, W ij is the transformation matrix.

[0049] 3-2-4) Weighted sum of the vector and the coupling coefficient c ij is performed to obtain the weighted sum s j , and the mathematical formula is as follows:

[0050]

[0051] 3-2-5) Use a non-linear function to compress s j and perform forward propagation. The mathematical formula is as follows:

[0052]

[0053] In this formula, s j represents the weighted sum, and v j represents the non-linear compression function.

[0054] 3-2-6) Update the coupling coefficient c ij using the softmax equation. The mathematical formula is as follows:

[0055]

[0056] In this formula represents the coupling coefficient after dynamic routing update, The initial value of is set to 0. v j represents the non-linear compression function, and the vector is the result of multiplying the capsule vector by the transformation matrix.

[0057] 3-2-7) If the number of updates of v j = the number of capsules k, then the finally obtained is the finally output v j , that is, the feature vector representing the j-th category. Otherwise, return to step 3-2-2).

[0058] Preferably, the specific steps of the deconvolution image reconstruction network algorithm in step 3-3) are as follows:

[0059] 3-3-1) Image input. The size of the input image is 14*14, the feature input is 5*5, the convolution kernel sizes are 3, 5, 7, 9 respectively, the number of convolution kernels is selected to be 32, the stride is 2, and the output image is 28*28 by adjusting the padding. The formula for the size of the input and the output image after deconvolution is as follows:

[0060]

[0061] In this formula represents the size of the output image, s represents the stride, represents the size of the input image, k represents the size of the convolutional kernel, and p represents the padding, that is, the padding size.

[0062] 3-3-2) Encapsulation information splitting. For the 6272 neurons in the capsule, they are transformed into a tensor of 14*14*32 through a fully connected layer, and the tensor is combined with the connection in the third dimension in the parallel convolutional network so that the tensor of 14*14*32 becomes a tensor of 14*14*160.

[0063] 3-3-3) Redundancy removal deconvolution. Use a 1*1 convolutional kernel to remove redundancy from the tensor of 14*14*160, and finally obtain a tensor of 14*14*32, and perform deconvolution operation with the corresponding convolutional kernel in the parallel convolutional layer to generate the final reconstructed image.

[0064] Preferably, the specific steps of the step 4) of extracting the center feature map of the detection box and the scale feature map of the detection box, forming the corresponding target box and outputting the detection result are as follows:

[0065] 4-1) Image input. Use the reconstructed image as the input of a convolutional layer with a size of 3*3 and an output channel of 256.

[0066] 4-2) Extract the center feature map. Use a 1*1 convolutional kernel to extract the center feature map of the detection box, mark the coordinates of the object center point as positive values, and mark the coordinates of non-object center points as negative values.

[0067] 4-3) Extract the scale feature map. Use parallel 1*1 convolutional kernels to extract the scale feature map of the detection box. The scale in the scale feature map is the length and width of the detection box.

[0068] 4-4) Output the image. Fuse the center feature map and the scale feature map to complete object target detection.

[0069] The beneficial effects of the present invention are as follows:

[0070] Aiming at the problem that the extraction of original image information in the prior art is not rich enough, the present invention improves the capsule network structure itself, adds a parallel convolutional network layer in front of the main capsule layer. After the parallel convolutional network extracts convolutional features, it directly fuses two kinds of image feature information, namely the convolutional feature map of the original image and the histogram of oriented gradients, in the way of dimension vector connection rather than the way of the network, so that the input information of the network is richer and closer to the original image information.

[0071] In view of the problems of redundancy in the existing capsule structure and high algorithm complexity, the present invention performs a redundancy removal operation on the main capsule. The redundant main capsule network simplifies the network structure, reduces the parameters compared with the original main capsule network, optimizes the time required for network training, and improves the algorithm efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a schematic diagram of the overall structure of the object detection model in the present invention;

[0073] Figure 2 It is a flow chart of the histogram of oriented gradients (HOG) and convolutional feature map parallel model in the present invention;

[0074] Figure 3 It is an improved capsule network architecture diagram in the present invention;

[0075] Figure 4 It is a schematic diagram of the parallel convolutional network in the present invention;

[0076] Figure 5 It is a schematic diagram of the redundant capsule network in the present invention;

[0077] Figure 6 It is a schematic diagram of the deconvolution image reconstruction network in the present invention;

[0078] Figure 7 It is the original image;

[0079] Figure 8 It is a schematic diagram of the HOG histogram processed image;

[0080] Figure 9 It is a schematic diagram of the detection result of the pedestrian image using the present method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0082] As Figure 1 shown, an object detection method based on the histogram of oriented gradients and the improved capsule network mainly includes the following steps:

[0083] 1) Obtain the original image of the object, as Figure 7 shown.

[0084] 2) As Figure 2 , the histogram of oriented gradients of the original image and the convolutional feature map are fused in parallel to combine the edge contour features of the image and the field of view features of the convolutional kernel, and then this is used as the input of the improved capsule network;

[0085] 2-1) Normalization. First, divide the target image into 4 cells, then divide each cell into 9 blocks, and perform Gamma normalization. At the same time, optimize the parameters of the Gamma correction value. The normalization formula is as follows:

[0086]

[0087]

[0088] In this formula, τ represents the feature vector of the block, ε takes a small constant value, represents the i-th block in the j-th cell of the target image, and f represents the target image after normalizing all block feature vectors.

[0089] 2-2), Select the detection window from the normalized image. Select a detection window with an aspect ratio equal to that of the image and not exceeding half of the image size.

[0090] 2-3), Select blocks from the window. Select rectangular blocks with equal length and width according to the detection window.

[0091] 2-4), Divide cell units within the block. Use a square cell with a size of 8*8 pixels within the rectangular block as the minimum unit for feature extraction within the block to divide the block.

[0092] 2-5), Perform direction projection within the cell. Divide 9 directions within the cell, and extract direction information every 20° as an angle range. The direction information is obtained by performing a convolution operation on the grayscale image I and the gradient template U in the x horizontal direction and the y vertical direction. The mathematical formula is as follows:

[0093] G x (x,y) = H(x + 1,y) - H(x - 1,y)

[0094] G y (x,y) = H(x,y + 1) - H(x,y - 1)

[0095]

[0096]

[0097] In this formula, H(x,y) represents the grayscale value at the corresponding coordinates, G x represents the horizontal direction gradient value, G y represents the vertical direction gradient value, G represents the gradient amplitude, and α represents the gradient direction.

[0098] 2-6) Perform normalization within the cell. Count the actual number of direction angles in each direction angle range within the cell to obtain a direction histogram, as Figure 8As shown in the figure, select the direction angle with the most concentrated angle direction as the direction of the cell.

[0099] 2 - 7) Construct HOG features within the block. Count the actual number of direction angles within the range of the direction angles of each cell within the block to obtain a direction histogram, and select the direction angle with the most concentrated angle direction as the direction of the block.

[0100] 2 - 8) If it has not reached the last block, return to step 3).

[0101] 2 - 9) If it has not reached the last window, return to step 2), otherwise obtain the histogram of oriented gradients.

[0102] 2 - 10) Single - row convolution. Input the original image, and use a single - row convolutional layer to extract features from the original image to obtain a convolutional feature map.

[0103] 2 - 11) Feature concatenation and parallel fusion. Concatenate the dimensions of the convolutional feature map and the histogram of oriented gradients, and connect the two feature maps in the third dimension to obtain an image of 28 * 28 * 1.

[0104] 3) As Figure 3 , 4 , the improved capsule network uses a parallel convolutional network to extract comprehensive features, forms feature vectors through a redundant - removing capsule network, and realizes image reconstruction using a de - convolutional image reconstruction network;

[0105] 3 - 1) Use a parallel convolutional network to extract comprehensive features;

[0106] 3 - 1 - 1) First, use a parallel convolutional neural network as the feature extraction network. Take the image after parallel fusion as the input, and the image size is 28 * 28 * 1. The parallel convolutional neural network uses 4 convolutional kernels in the convolutional layer, and the sizes of the convolutional kernels are 3, 5, 7, and 9 respectively. The number of convolutional kernels is selected as 32, and the stride is 2.

[0107] 3 - 1 - 2) Boundary padding. Adjust the padding size to perform boundary padding on the original matrix.

[0108] 3 - 1 - 3) Feature extraction. The non - linear function of the feature extraction layer uses the PReLU function, and its mathematical formula is as follows:

[0109] PReLU(x) = max(0, x)+α * min(0, x)

[0110] In this formula, α is the learning rate.

[0111] 3 - 1 - 4) Feature tensor connection. Connect the feature tensors in the third dimension to obtain a feature tensor of 14 * 14 * 128.

[0112] 3-2) As shown in Figure 5 , a feature vector is generated using a redundant capsule network;

[0113] 3-2-1) Input redundancy removal. The output of a parallel convolutional network is used as the input to the redundant main capsule network. A 1*1 convolutional kernel is used to remove redundant capsules, converting the 16*16 feature map after feature extraction into a 14*14 feature image and reducing the number of capsules to 196.

[0114] 3-2-2) Input capsule scalar u i .

[0115] 3-2-3) The vector obtained by multiplying the input capsule vector by the transformation matrix The mathematical formula is as follows:

[0116]

[0117] In this formula, W ij is the transformation matrix.

[0118] 3-2-4) The vector is weighted and summed with the coupling coefficient c ij to obtain the weighted sum s j , and the mathematical formula is as follows:

[0119]

[0120] 3-2-5) The nonlinear function is used to compress s j and perform forward propagation. The mathematical formula is as follows:

[0121]

[0122] In this formula, s j represents the weighted sum, and v j represents the nonlinear compression function.

[0123] 3-2-6) The coupling coefficient c ij is updated using the softmax equation. The mathematical formula is as follows:

[0124]

[0125] In this formula, represents the coupling coefficient after dynamic routing update, and is set to 0 initially. v j represents the nonlinear compression function, and the vector is the result of multiplying the capsule vector by the transformation matrix.

[0126] 3-2-7) If the update count of v j = the number of capsules k, then the finally obtained This is the finally output v j , which is the feature vector representing the j-th category. Otherwise, return to step 3-2-2).

[0127] 3-3) As Figure 6 , use the deconvolution image reconstruction network to restore the original image.

[0128] 3-3-1) Image input. The size of the input image is 14*14, the feature input is 5*5, the convolutional kernel sizes are 3, 5, 7, 9 respectively, the number of convolutional kernels is selected as 32, the stride is 2, and by adjusting the padding, the output image is 28*28. The formula for the size of the input and the output image after deconvolution is as follows:

[0129]

[0130] In this formula represents the size of the output image, s represents the stride, represents the size of the input image, k represents the convolutional kernel size, and p represents the padding, that is, the padding size.

[0131] 3-3-2) Encapsulation information splitting. For the 6272 neurons in the capsule, convert them into a tensor of 14*14*32 through a fully connected layer, and combine this tensor with the connection in the third dimension in the parallel convolutional network to make the tensor of 14*14*32 become a tensor of 14*14*160.

[0132] 3-3-3) Redundancy removal deconvolution. Use a 1*1 convolutional kernel to perform redundancy removal on the tensor of 14*14*160, and finally obtain a tensor of 14*14*32, and perform deconvolution operation with the corresponding convolutional kernel in the parallel convolutional layer to generate the final reconstructed image.

[0133] 4) Use a convolutional layer of 3*3*256 and two parallel 1*1 convolutional kernels to extract the center point feature map of the detection box and the scale feature map of the detection box, form the corresponding target box and output the detection result.

[0134] 4-1) Image input. Use the reconstructed image as the input of a convolutional layer with a size of 3*3 and an output channel of 256.

[0135] 4-2) Center feature map. Use a 1*1 convolutional kernel to extract the center point feature map of the detection box, mark the coordinates of the object center point as positive, and mark the coordinates of the non-object center point as negative.

[0136] 4-3) Scale feature map. Use parallel 1*1 convolutional kernels to extract the scale feature map of the detection box. The scale in the scale feature map is the length and width of the detection box.

[0137] 4-4) Output image. The object detection is completed by fusing the central feature map and the scale feature map, as Figure 9 shown.

Claims

1. A target detection method based on histogram of oriented gradients and improved capsule network, characterized in that, It includes the following steps: 1) Obtain the target original image, use an annotation tool to annotate the target position, and then randomly select different images as the training set; 2) Combine the histogram of oriented gradients of the original image and the convolutional feature map in parallel to integrate the edge contour features of the image and the visual field features of the convolutional kernel, and then use this as the input to the improved capsule network; 3) The improved capsule network uses a parallel convolutional network to extract comprehensive features, forms feature vectors through a redundant capsule network removal, and uses a deconvolution image reconstruction network to achieve image reconstruction; The specific steps of the improved capsule network algorithm in step 3) are as follows: 3-1) Use a parallel convolutional network to extract comprehensive features; 3-2) Use a redundant capsule network removal to generate feature vectors; 3-3) Use a deconvolution image reconstruction network to restore the original image and evaluate the network loss; The specific steps of the parallel convolutional network algorithm in step 3-1) are as follows: 3-1-1) First, use a parallel convolutional neural network as the feature extraction network, take the image after parallel fusion as the input, the image size is 28*28*1, the parallel convolutional neural network uses 4 convolutional kernels in the convolutional layer, the convolutional kernel sizes are 3, 5, 7, and 9 respectively, the number of convolutional kernels is selected as 32, and the stride is 2; 3-1-2) Boundary padding, adjust the padding size to perform boundary padding on the original matrix; 3-1-3) Feature extraction, the non-linear function of the feature extraction layer uses the PReLU function, and its mathematical formula is as follows: PReLU(x) = max(0, x) + α * min(0, x) In this formula, α is the learning rate; 3-1-4) Feature tensor connection, connect the feature tensors in the third dimension to obtain a feature tensor of 14*14*128; The specific steps of the redundant capsule network removal algorithm in step 3-2) are as follows: 3-2-1) Input redundancy removal, use the output of the parallel convolutional network as the input to the redundant main capsule network removal, use a 1*1 convolutional kernel to remove redundant capsules so that the 16*16 feature map after feature extraction is transformed into a 14*14 feature image, and reduce the number of capsules to 196; 3-2-2) Input capsule scalar u i ; The vector obtained by multiplying the input capsule vector by the transformation matrix The mathematical formula is as follows In this formula, W ij is the transformation matrix; (3-2-4) Weight the vector with the coupling coefficient c ij and perform weighted summation to obtain the weighted sum s j , and the mathematical formula is as follows: 3-2-5) Compress s using a non-linear function and perform forward propagation. The mathematical formula is as follows: j and perform forward propagation. The mathematical formula is as follows: In this formula, s j represents the weighted sum, and v j represents the non-linear compression function; 3-2-6) Update the coupling coefficient c using the softmax equation ij , and the mathematical formula is as follows: In this formula represents the coupling coefficient after dynamic routing update, the initial value of which is set to 0, v j represents a non-linear compression function, and the vector is obtained by multiplying the capsule vector by the transformation matrix; 3-2-7) If the number of updates to v j = the number of capsules k, then the finally obtained is the finally output v j , that is, the feature vector representing the j-th category. Otherwise, return to step 3-2-2); The specific steps of the deconvolution image reconstruction network algorithm in step 3-3) are as follows: 3-3-1) Image input, the input image size is 14*14, the feature input is 5*5, the convolutional kernel sizes are 3, 5, 7, and 9 respectively, the number of convolutional kernels is selected as 32, the stride is 2, and by adjusting the padding, the output image is 28*28. The formula for the input and the output image size after deconvolution is as follows: In this formula represents the output image size, s represents the stride, represents the input image size, k represents the convolutional kernel size, and p represents the padding, that is, the padding size; 3-3-2) Encapsulation information splitting, for the 6272 neurons in the capsule, convert them into a tensor of 14*14*32 through a fully connected layer, and combine this tensor with the connection in the parallel convolutional network in the third dimension so that the 14*14*32 tensor becomes a 14*14*160 tensor; 3-3-3) Redundant deconvolution removal, perform redundancy removal on the 14*14*160 tensor using a 1*1 convolutional kernel, and finally obtain a 14*14*32 tensor, perform deconvolution operation with the corresponding convolutional kernel in the parallel convolutional layer to generate the final reconstructed image; 4) Use a 3×3×256 convolutional layer and two parallel 1×1 convolutional kernels to extract the center feature map of the detection box and the scale feature map of the detection box, form the corresponding target box, and output the detection result.

2. The object detection method based on histogram of oriented gradients and improved capsule network according to claim 1, wherein: The specific steps of the algorithm for fusing the histogram of oriented gradients and the convolutional feature map in parallel in step 2) are as follows: 2-1) Normalization: First, divide the target image into 4 cells, divide each cell into 9 blocks, and perform Gamma normalization processing. At the same time, optimize the parameters of the Gamma correction value. The normalization formula is as follows: In this formula, τ represents the feature vector of the block, and ε takes a relatively small constant value. represents the i-th block in the j-th cell of the target image, and f represents the target image after normalizing all block feature vectors. 2-2) Select the detection window from the normalized image, and select the detection window whose aspect ratio is equal to that of the image and does not exceed half of the image size. 2-3) Select the blocks from the window, and select rectangular blocks with equal length and width according to the detection window. 2-4) Divide cell units within the block. Use a square cell with a size of 8×8 pixels within the rectangular block as the minimum unit for feature extraction within the block to divide the block. 2-5) Perform direction projection within the cell. Divide 9 directions within the cell, and extract direction information every 20° as an angle range. The direction information is obtained by performing a convolution operation on the grayscale image I and the gradient template U in the x horizontal direction and the y vertical direction. The mathematical formula is as follows: G x (x,y) = H(x + 1,y) - H(x - 1,y) G y (x,y) = H(x,y + 1) - H(x,y - 1) In this formula, H(x, y) represents the gray value at the corresponding coordinates, and G x represents the horizontal direction gradient value, and G y represents the vertical direction gradient value, G represents the gradient amplitude, and α represents the gradient direction; 2-6) Perform normalization within the cell. Count the actual number of direction angles in each direction angle range within the cell to obtain the direction histogram, and select the direction angle with the most concentrated angle direction as the direction of the cell. 2-7) Construct the HOG feature within the block. Count the actual number of direction angles in each direction angle range of each cell within the block to obtain the direction histogram, and select the direction angle with the most concentrated angle direction as the direction of the block. 2-8) If it has not reached the last block, return to step 2-3). 2-9) If it has not reached the last window, return to step 2-2), otherwise obtain the histogram of oriented gradients. 2-10) Single-row convolution: Input the original image, and use a single-row convolutional layer to extract features from the original image to obtain the convolutional feature map. 2-11) Feature splicing and parallel fusion: Connect the dimensions of the convolutional feature map and the histogram of oriented gradients, and connect the two feature maps in the third dimension to obtain an image of 28×28×1.

3. The object detection method based on histogram of oriented gradients and improved capsule network according to claim 1, wherein: The specific steps of step 4) for extracting the center feature map of the detection box and the scale feature map of the detection box, forming the corresponding target box, and outputting the detection result are as follows: 4-1) Image input: Use the reconstructed image as the input of a convolutional layer with a size of 3×3 and an output channel of 256. 4-2) Extract the center feature map: Use a 1×1 convolutional kernel to extract the center feature map of the detection box. Mark the coordinates of the object center point as positive values, and mark the coordinates of non-object center points as negative values. 4-3) Extract the scale feature map: Use parallel 1×1 convolutional kernels to extract the scale feature map of the detection box. The scale in the scale feature map is the length and width of the detection box. 4-4) Output image: Fuse the center feature map and the scale feature map to complete object target detection.

Citation Information

Patent Citations

  • A method for classifying and recognizing capsule network image based of improved reconstruct network

    CN108985316A

  • Traffic sign recognition method based on capsule neural network

    CN111428556A