Image processing method and device, electronic equipment, storage medium and vehicle
By designing a residual network based on an inverter operator in the image recognition algorithm, combining the advantages of convolutional layer and inverter layer, the problem of missing short-distance visual relationship in the prior art is solved, and more efficient image feature extraction and recognition accuracy is achieved.
Patent Information
- Application Number
- CN202311672977.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art captures long-distance visual relationships only through the inverted layer in image recognition algorithms, resulting in the loss of the short-distance visual relationship capabilities captured by the convolutional layer, thereby limiting the accuracy of image recognition.
A residual network based on the inverter operator is designed, which includes a convolutional layer and an inverter layer. The short-distance visual relationship capability is retained through the convolutional layer, and the long-distance visual feature relationship is captured through the inverter layer.
By combining the advantages of convolutional layer and inverse layer, more abundant and accurate image features are extracted, which avoids the loss of short-distance visual relationship ability caused by single capture of long-distance visual relationships and improves the accuracy of image recognition.
Smart Images

Figure CN120125865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image related technologies, and particularly to an image processing method, apparatus, electronic device, storage medium, and vehicle. Background Art
[0002] In current intelligent vehicles, image recognition algorithms play important roles in various directions such as vehicle automation testing, assisted driving, and intelligent space. The accuracy of image recognition directly affects the accuracy of the robotic arm in vehicle automation testing, the judgment of traffic conditions in assisted driving, and the in-vehicle human-machine touchless interaction experience in intelligent space. Therefore, it is particularly important to improve the accuracy of image recognition algorithms.
[0003] Most existing solutions for optimizing image recognition algorithms are based on fine-tuning the original model or optimizing images when making datasets. However, in the image feature extraction part of the algorithm structure, the convolution of the traditional Convolutional Neural Network (CNN) is still used. The spatial invariance of CNN convolution limits the ability of the convolution kernel to capture long-distance visual relationships, and the improvement of algorithm accuracy is very limited.
[0004] Therefore, the prior art proposes to optimize the image feature extraction part in the image recognition algorithm. To improve the ability to capture long-distance visual relationships, an involution operator is introduced, and the convolutional layer is replaced with an involution layer based on the involution operator.
[0005] However, the existing image recognition algorithms only consider the ability to capture long-distance visual relationships. Therefore, by directly replacing all convolutional layers with involution layers, the ability of the convolutional layer to capture short-distance visual relationships is lost. Summary of the Invention
[0006] Based on this, it is necessary to provide an image processing method, apparatus, electronic device, storage medium, and vehicle to solve the technical problem that the existing technology using the involution operator loses the ability to capture short-distance visual relationships.
[0007] The present invention provides an image processing method, including:
[0008] Obtain a feature image;
[0009] Input the feature image into a residual network based on an involution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch and a second branch. The first branch includes at least a convolutional layer, and the second branch includes at least an involution layer based on involution operator operations.
[0010] Further, the residual network includes a plurality of serially connected residual blocks, and each of the residual blocks includes the first branch, the second branch, or the first branch and the second branch whose outputs are added together;
[0011] The step of inputting the feature image into the residual network based on the convolution operator to obtain the feature vector of the feature image includes:
[0012] Inputting the feature image into the first residual block, and inputting the output data of the previous residual block into the next residual block, and taking the output data of the last residual block as the feature vector of the feature image.
[0013] Further, the residual network further includes a direct connection branch, and the output of the direct connection branch is added to the output of the first branch and / or the output of the second branch;
[0014] The step of inputting the feature image into the residual network based on the convolution operator to obtain the feature vector of the feature image includes:
[0015] Inputting the feature image into the first branch, the second branch and / or the direct connection branch, and adding the output data of the first branch, the output data of the second branch and the output data of the direct connection branch as the feature vector of the feature image.
[0016] Even further, the residual network includes a plurality of serially connected residual blocks, and each of the residual blocks includes:
[0017] The first branch; or
[0018] The second branch; or
[0019] The direct connection branch; or
[0020] The first branch and the direct connection branch whose outputs are added together; or
[0021] The second branch and the direct connection branch whose outputs are added together; or
[0022] The first branch, the second branch, and the direct connection branch whose outputs are added together;
[0023] The step of inputting the feature image into the first branch, the second branch and / or the direct connection branch, and adding the output data of the first branch, the output data of the second branch and the output data of the direct connection branch as the feature vector of the feature image includes:
[0024] Input the feature image into the first branch, the second branch, and / or the direct connection branch of the first residual block. After adding the output data of the first branch, the output data of the second branch, and / or the output data of the direct connection branch of the previous residual block, input the result into the next residual block to obtain the output data of the last residual block as the feature vector of the feature image.
[0025] Furthermore, the residual network further includes an image dimensionality reduction layer, and the output data of the image dimensionality reduction layer is the input of the first residual block among the multiple cascaded residual blocks;
[0026] Inputting the feature image into the first residual block includes:
[0027] Input the feature image into the image dimensionality reduction layer for dimensionality reduction operation to obtain a dimensionality-reduced feature image;
[0028] Input the dimensionality-reduced feature image into the first residual block.
[0029] Still further, it further includes:
[0030] Perform image recognition and / or image classification based on the feature vector.
[0031] The present invention provides an image processing apparatus, including:
[0032] An image acquisition module for acquiring a feature image;
[0033] A feature vector calculation module for inputting the feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch and a second branch. The first branch includes at least a convolutional layer, and the second branch includes at least a convolution layer based on convolution operator operations.
[0034] The present invention provides an electronic device, including:
[0035] At least one processor; and,
[0036] A memory communicatively connected to at least one of the processors; wherein,
[0037] The memory stores instructions executable by at least one of the processors. The instructions are executed by at least one of the processors so that at least one of the processors can execute the image processing method as described above.
[0038] The present invention provides a storage medium that stores computer instructions. When a computer executes the computer instructions, it is used to execute all steps of the image processing method as described above.
[0039] The present invention provides a vehicle, including the image processing device as described above, or the electronic device as described above.
[0040] The present invention inputs a feature image into a residual network based on an involution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch and a second branch. The first branch includes at least a convolutional layer, and the second branch includes at least an involution layer based on an involution operator operation. In the residual network of the present invention, the first branch including the convolutional layer and the second branch including the involution layer not only retain the short-distance visual relationship ability through the convolutional layer, but also capture the long-distance visual feature relationship through the involution layer, avoiding the problem of the lack of short-distance visual relationship ability caused by only the long-distance visual feature relationship, and can extract richer and more accurate image features and output a feature vector more fitting the image. Description of the Drawings
[0041] Figure 1 is a flowchart of an image processing method according to an embodiment of the present invention;
[0042] Figure 2 is a flowchart of an image processing method according to another embodiment of the present invention;
[0043] Figure 3 is a schematic diagram of a residual network according to an embodiment of the present invention;
[0044] Figure 4 is a residual block including only the first branch according to an embodiment of the present invention;
[0045] Figure 5 is a residual block including only the second branch according to an embodiment of the present invention;
[0046] Figure 6 is a residual block including the first branch and the second branch according to an embodiment of the present invention;
[0047] Figure 7 is a residual block including the first branch and a direct connection branch according to an embodiment of the present invention;
[0048] Figure 8 is a residual block including the second branch and a direct connection branch according to an embodiment of the present invention;
[0049] Figure 9 is a residual block including the first branch, the second branch and a direct connection branch according to an embodiment of the present invention;
[0050] Figure 10 is a schematic diagram of an involution operator operation;
[0051] Figure 11 is a flowchart of an image recognition method based on involution convolution according to the best embodiment of the present invention;
[0052] Figure 12 Schematic diagram of a residual block including a first branch and a second branch according to an example of the present invention;
[0053] Figure 13 Schematic diagram of a residual block including a first branch, a second branch, and a direct connection branch according to another example of the present invention;
[0054] Figure 14 Schematic diagram of an image processing device according to an embodiment of the present invention;
[0055] Figure 15 Schematic diagram of the hardware structure of an electronic device according to the present invention. Detailed implementation manners
[0056] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings. The same components are denoted by the same reference numerals. It should be noted that the terms "front", "rear", "left", "right", "upper", and "lower" used in the following description refer to the directions in the drawings, and the terms "inner" and "outer" refer to the directions towards or away from the geometric center of a specific component, respectively.
[0057] As Figure 1 shown is a flowchart of an image processing method according to an embodiment of the present invention, including:
[0058] Step S101, obtaining a feature image;
[0059] Step S102, inputting the feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch 11 and a second branch 12. The first branch 11 includes at least a convolutional layer 111, and the second branch 12 includes at least a convolution layer 121 based on convolution operator operations.
[0060] Specifically, the present invention can be applied to an electronic device with processing capabilities, such as an electronic control unit (ECU) of a vehicle or an extended domain control unit (XCU).
[0061] Specifically, first, the electronic device executes step S101 to obtain a feature image. For example, multiple image sequence frames can be continuously obtained, and each image sequence frame serves as a feature image.
[0062] Then, the electronic device executes step S102, inputs the feature image into a residual network based on a convolution operator, and obtains a feature vector of the feature image.
[0063] Among them, the residual network includes at least a first branch 11 and a second branch 12. The first branch 11 includes at least a convolutional layer 111, and the second branch 12 includes at least an involution layer 121.
[0064] In some embodiments, the first branch 11 sequentially includes a first dimensionality reduction layer 110, the convolutional layer 111, and a first dimensionality increase layer 112.
[0065] Specifically, the first branch 11 sequentially includes a first dimensionality reduction layer 110, a convolutional layer 111, and a first dimensionality increase layer 112.
[0066] Among them, the first dimensionality reduction layer 110 is used to reduce the dimension of the input feature image. As an example, the first dimensionality reduction layer 110 performs a 1×1 CNN convolution operation.
[0067] After being processed by the first dimensionality reduction layer 110, the convolutional layer 111 performs a CNN convolution operation to extract image features. The convolutional layer 111 uses an existing CNN convolution, for example, a 3×3 CNN convolution kernel for convolution. Among them, 3×3 is the resolution, that is, the number of horizontal pixel points multiplied by the number of vertical pixel points.
[0068] Finally, the first dimensionality increase layer 112 restores the dimension of the image features. As an example, the first dimensionality increase layer 112 performs a 1×1 CNN convolution operation.
[0069] By performing CNN convolution through the convolutional layer, the ability to retain short-distance visual relationships can be preserved.
[0070] In some embodiments, the second branch 12 sequentially includes a second dimensionality reduction layer 120, the involution layer 121, and a second dimensionality increase layer 122.
[0071] Specifically, the second branch 12 sequentially includes a second dimensionality reduction layer 120, an involution layer 121, and a second dimensionality increase layer 122.
[0072] Among them, the second dimensionality reduction layer 120 is used to reduce the dimension of the input feature image. As an example, the second dimensionality reduction layer 120 performs a 1×1 CNN convolution operation.
[0073] After being processed by the second dimensionality reduction layer 120, the involution layer 121 performs an involution convolution operation to extract image features through an involution operator.
[0074] Such as Figure 10The figure shows a schematic diagram of the operation of the Involution operator. The Involution operator can perform a new type of convolution operation, leveraging the internal relationships of the input data to flexibly learn the spatial specific relationships between different positions. The advantage of this spatial specificity enables Involution convolution to better handle the information interaction between different positions, thereby enhancing the expressiveness of the model. Involution convolution can also be referred to as inverse feature convolution or folding convolution. The generation formula for the Involution convolution kernel is:
[0075] H i,j = φ(X Ψi,j ) (1)
[0076] where Ψ i,j is an index set in the neighborhood of coordinates (i, j), φ is a linear transformation function, H i,j is the corresponding Involution convolution kernel, and X Ψi,j represents an image patch on the feature image that contains Ψ i,j .
[0077] Figure 10 The left half in shows the generation process of the Involution convolution kernel, which changes from 1×1×C to 1×1×K 2 through φ, and then is reshaped into an Involution kernel of K×K×1. The right half is the process of generating a new feature image after obtaining the Involution kernel. First, the K×K×1 Involution kernel is multiplied pixel by pixel with the pixels 43 in the K×K neighborhood of the current pixel 42 in the original feature image, and then this operation is repeated on C channels. Finally, a three-dimensional matrix 44 of K×K×C is obtained. The width and height two dimensions in the three-dimensional matrix 44 are summed, keeping the channel dimension, to obtain a vector 45 of the new feature image 1×1×C, and the values of C channels at 1 pixel position can be generated at once.
[0078] Through Involution convolution, the feature relationships of long-distance vision can be captured, and it is better to reconstruct the three-dimensional human motion sequence from a long sequence of videos.
[0079] Finally, the second upsampling layer 122 restores the dimension of the image features. As an example, the second upsampling layer 122 performs a 1×1 CNN convolution operation.
[0080] The present invention inputs a feature image into a residual network based on an involution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch and a second branch. The first branch includes at least a convolutional layer, and the second branch includes at least an involution layer based on involution operator operations. In the residual network of the present invention, the first branch including the convolutional layer and the second branch including the involution layer not only retain the short-distance visual relationship ability through the convolutional layer, but also capture the long-distance visual feature relationship through the involution layer, avoiding the problem of the lack of short-distance visual relationship ability caused by only the long-distance visual feature relationship, enabling more abundant and accurate image features to be extracted and a feature vector more fitting the image to be output.
[0081] In one embodiment, the residual network includes a plurality of residual blocks 1 connected in series, and each residual block 1 includes the first branch 11, the second branch 12, or the first branch 11 and the second branch 12 whose outputs are added together;
[0082] The step of inputting the feature image into the residual network based on the involution operator to obtain the feature vector of the feature image includes:
[0083] Input the feature image into the first residual block 1, and input the output data of the previous residual block 1 into the next residual block 1, and obtain the output data of the last residual block 1 as the feature vector of the feature image.
[0084] Specifically, the residual network includes a plurality of residual blocks (Residual blocks) 1 connected in series. As Figure 3 shown, a plurality of residual blocks 1 are connected in series in sequence. The output of the previous residual block 1 serves as the input of the next residual block 1. When inputting the feature image into the residual network, input the feature image into the first residual block 1, and input the output data of the previous residual block 1 into the next residual block 1, and obtain the output data of the last residual block 1 as the feature vector of the feature image.
[0085] Each residual block 1 includes the first branch 11, the second branch 12, or the first branch 11 and the second branch 12 whose outputs are added together, that is:
[0086] As Figure 4 shown, the residual block 1 may only include the first branch 11; or
[0087] As Figure 5 shown, the residual block 1 may only include the second branch 12; or
[0088] As Figure 6 shown, the residual block 1 includes the first branch 11 and the second branch 12. The data input to the residual block 1 is respectively processed by the first branch 11 and the second branch 12, and the outputs of the first branch 11 and the second branch are added together as the output of the residual block 1.
[0089] However, the entire residual network includes at least one first branch 11 and one second branch 12. That is, the residual network:
[0090] includes at least one residual block 1 including the first branch 11 and one residual block 1 including the second branch 12; or
[0091] includes at least one residual block 1 including the first branch 11 and the second branch 12.
[0092] Preferably, the residual network includes four residual blocks 1.
[0093] In some embodiments, the residual network includes a plurality of serially connected residual blocks 1, and each residual block 1 includes a first branch 11 and a second branch 12 whose outputs are added together.
[0094] In this embodiment, the first branch and / or the second branch is set in each residual block, and the short-distance visual relationship ability and / or the ability to capture the feature relationship of long-distance vision are retained in the same residual block.
[0095] As Figure 2 shown is a flowchart of a working process of an image processing method in another embodiment of the present invention, including:
[0096] Step S201, obtaining a feature image.
[0097] Step S202, inputting the feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch 11, a second branch 12, and a direct connection branch 13. The first branch 11 includes at least a convolutional layer 111, the second branch 12 includes at least a convolution layer 121 based on convolution operator operations, and the output of the direct connection branch is added to the output of the first branch 11 and / or the output of the second branch 12;
[0098] The step of inputting the feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image includes:
[0099] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13, and adding the output data of the first branch 11, the output data of the second branch 12, and the output data of the direct connection branch 13 as the feature vector of the feature image.
[0100] In one of the embodiments, the residual network includes a plurality of serially connected residual blocks 1, and each of the residual blocks 1 includes: the first branch 11, the second branch 12, and / or the direct connection branch 13;
[0101] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13, and adding the output data of the first branch 11, the output data of the second branch 12, and the output data of the direct connection branch 13 as the feature vector of the feature image includes:
[0102] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13 of the first residual block 1. After adding the output data of the first branch 11, the output data of the second branch 12, and / or the output data of the direct connection branch 13 of the previous residual block 1, input it into the next residual block 1, and obtain the output data of the last residual block 1 as the feature vector of the feature image.
[0103] In one embodiment, the residual network further includes an image dimensionality reduction layer, and the output data of the image dimensionality reduction layer is the input of the first residual block 1 in the plurality of cascaded residual blocks;
[0104] Inputting the feature image into the first residual block 1 includes:
[0105] Inputting the feature image into the image dimensionality reduction layer for dimensionality reduction operation to obtain a dimensionality-reduced feature image;
[0106] Inputting the dimensionality-reduced feature image into the first residual block 1.
[0107] Step S203, performing image recognition and / or image classification based on the feature vector.
[0108] Specifically, first execute step S201 to obtain a feature image.
[0109] In some embodiments, before obtaining the feature image, first perform image preprocessing on the original image to obtain a to-be-processed image. The image preprocessing adopts existing digital image processing methods. The purpose is to reduce the influence of interference and noise factors on the original image and program the original image so that it is suitable for the computer to extract image features. During the image processing, image enhancement is used to highlight the main structure of the image, reduce the noise in the image, and change parameters such as the brightness, color distribution, and contrast of the original image. Image enhancement improves the clarity and quality of the image, makes the contours of the objects in the image clearer, and the details more obvious, which is more conducive to image feature extraction.
[0110] Then, execute step S202 to input the feature image into the residual network based on the convolution operator to obtain the feature vector of the feature image.
[0111] Among them, the residual network at least includes a first branch 11, a second branch 12, and a direct connection branch 13. The first branch 11 at least includes a convolutional layer 111. The second branch 12 at least includes a convolution layer 121 based on a convolution operator operation. The output of the direct connection branch is added to the output of the first branch 11 and / or the output of the second branch 12.
[0112] In this embodiment, the residual network may further include a direct connection branch 13. The direct connection branch 13 outputs after adding the input of the residual network to the output of the first branch 11, or outputs after adding the input of the residual network to the output of the second branch 12, or outputs after adding the input of the residual network to the outputs of the first branch 11 and the second branch 12.
[0113] In this embodiment, through the direct connection branch, the input of the residual network is added to the output, so as to retain the original data of the image and further enrich the image features.
[0114] In one embodiment, the residual network includes a plurality of residual blocks 1 connected in series, and each of the residual blocks 1 includes: the first branch 11, the second branch 12, and / or the direct connection branch 13;
[0115] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13, and adding the output data of the first branch 11, the output data of the second branch 12, and the output data of the direct connection branch 13 as the feature vector of the feature image includes:
[0116] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13 of the first residual block 1. After adding the output data of the first branch 11, the output data of the second branch 12, and / or the output data of the direct connection branch 13 of the previous residual block 1, inputting it into the next residual block 1, and obtaining the output data of the last residual block 1 as the feature vector of the feature image.
[0117] Specifically, when the residual network includes a plurality of residual blocks 1 connected in series, and each residual block 1 includes a first branch 11, a second branch 12, and / or a direct connection branch 13, that is:
[0118] As Figure 4 shown, the residual block 1 may only include the first branch 11; or
[0119] As Figure 5 shown, the residual block 1 may only include the second branch 12; or
[0120] As Figure 6 shown, the residual block 1 includes the first branch 11 and the second branch 12; or
[0121] As shown Figure 7 in the figure, the residual block 1 includes a first branch 11 and a direct connection branch 13; or
[0122] As shown Figure 8 in the figure, the residual block 1 includes a second branch 12 and a direct connection branch 13; or
[0123] As shown Figure 9 in the figure, the residual block 1 includes a first branch 11, a second branch 12 and a direct connection branch 13.
[0124] When the residual block 1 includes the first branch 11 and the second branch 12, the data input to the residual block 1 is respectively processed by the first branch 11 and the second branch 12, and the outputs of the first branch 11 and the second branch are added as the output of the residual block 1.
[0125] When the residual block 1 includes the first branch 11 and the direct connection branch 13, the data input to the residual block 1 is processed by the first branch 11, and the output of the first branch 11 and the input of the residual block 1 are added as the output of the residual block 1.
[0126] When the residual block 1 includes the second branch 12 and the direct connection branch 13, the data input to the residual block 1 is processed by the second branch 12, and the output of the second branch 12 and the input of the residual block 1 are added as the output of the residual block 1.
[0127] Preferably, the residual network includes four residual blocks 1.
[0128] In some embodiments, the residual network includes a plurality of residual blocks 1 connected in series, and each residual block 1 includes a first branch 11, a second branch 12 and a direct connection branch 13. The data input to the residual block 1 is respectively processed by the first branch 11 and the second branch 12, and the sum of the outputs of the first branch 11 and the second branch is added to the input of the residual block 1 as the output of the residual block 1.
[0129] In this embodiment, a first branch, a second branch and / or a direct connection branch are set in the residual block, and the short-distance visual relationship ability, the ability to capture long-distance visual feature relationships through convolutional layers, or the original data of the image is retained through the direct connection branch in the same residual block.
[0130] In one of the embodiments, the residual network further includes an image downsampling layer, and the output data of the image downsampling layer is the input of the first residual block 1 in the plurality of residual blocks connected in series;
[0131] Inputting the feature image into the first said residual block 1 includes:
[0132] Input the feature image into the image dimensionality reduction layer for dimensionality reduction operation to obtain a dimensionality-reduced feature image;
[0133] Input the dimensionality-reduced feature image into the first residual block 1.
[0134] Specifically, for the feature image input to the residual network, first perform a dimensionality reduction operation through the image dimensionality reduction layer, and then output the feature image after the dimensionality reduction operation to the residual block 1. The dimensionality reduction operation includes a convolution operation and a pooling operation.
[0135] For example, for an input feature image with a size of 224×224, first perform a convolution operation with a convolution kernel of 7×7 and a stride of 2, and then perform a max pooling operation with a convolution kernel of 3×3 and a stride of 2 for downsampling to obtain a feature image with a resolution of 56x56.
[0136] This embodiment reduces the dimension of the feature image through a dimensionality reduction operation to reduce the computational complexity.
[0137] In some embodiments, the residual network further includes: a pooling layer connected to the output of the last residual block 1.
[0138] Specifically, the feature image processed by the last residual block 1 is subjected to a pooling operation through the pooling layer. The pooling layer is preferably a global average pooling layer.
[0139] The pooling layer divides the feature image into several blocks, and the size of each block is preferably 2x2. For each block, calculate the average value of all elements therein, and use this average value as the representative value of the block. Use all the representative values to form a one-dimensional vector as the output of the global average pooling layer. This vector is the feature vector of the entire feature image, which reflects the feature information of this frame of image.
[0140] This embodiment reduces information redundancy and prevents overfitting by adding a pooling layer.
[0141] After the above steps, more abundant and accurate image features can be extracted to output a feature vector, and then step S203 is executed to perform image recognition and / or image classification based on the feature vector.
[0142] Specifically, use the classification part of the existing image recognition algorithm to classify the objects in the image based on the feature vector output by the residual network, and finally output the object category and confidence level. The image recognition algorithm can be the Yolov5 (YouOnly Look Once) algorithm.
[0143] In this embodiment, a convolution operator is introduced in the image feature extraction step of the image recognition algorithm to capture the feature relationships of long-distance vision. At the same time, the ability of short-distance vision relationships is retained through the convolutional layer, and the original data of the image is retained through the direct connection branch, thereby enhancing the image feature extraction ability and improving the accuracy of image recognition or image classification. The image recognition algorithm of this embodiment can improve the accuracy of image recognition in related fields such as assisted driving, thereby improving practicability and safety. In the current era of intelligent vehicles, improving the accuracy of the image recognition algorithm also improves the market competitiveness to a certain extent.
[0144] As Figure 11 shown in the working flowchart of an image recognition method based on involution convolution according to the best embodiment of the present invention, which includes:
[0145] Step S1101, perform image preprocessing on the input image.
[0146] Step S1102, perform image feature extraction on the preprocessed image to obtain a feature vector.
[0147] Step S1103, classify the feature vector through image classification to obtain the corresponding category.
[0148] Specifically, step S1101 is image preprocessing. Digital image processing is used for image preprocessing. The purpose is to reduce the influence of interference and noise factors on the original image and program the original image so that it is suitable for computer to perform image feature extraction. During the image processing, image enhancement is used to highlight the main structure of the image, reduce the noise in the image, and change parameters such as the brightness, color distribution, and contrast of the original image. Image enhancement improves the clarity and quality of the image, makes the contours of the objects in the image clearer, and the details more obvious, which is more conducive to image feature extraction.
[0149] Steps S1102 and S1103 are the image recognition part. In this embodiment, a self-attention mechanism is introduced in the image feature extraction of step S1102, and a convolution operator is added to the residual network, such as a 50-layer Residual Network 50 (ResNet50), to obtain the feature vector of the preprocessed image.
[0150] Specifically, in the image feature extraction step, for the preprocessed image sequence frames I1, I2... In, in order to obtain the feature vector Fi of each image, a ResNet50 deep learning network is used, and a self-attention mechanism Involution operator is introduced. The specific steps are as follows:
[0151] (1) Image dimensionality reduction. For the input image sequence frames I1, I2…In with a size of 224×224, first perform a convolution operation with a convolution kernel of 7×7 and a stride of 2, and then perform a max-pooling operation with a convolution kernel of 3×3 and a stride of 2 for downsampling to obtain a feature image with a resolution of 56x56.
[0152] (2) Image feature extraction. For the obtained feature image with a resolution of 56x56, use 4 cascaded residual blocks 1 for feature extraction. Inside each residual block 1, there are multiple residual blocks 1 with a Bottleneck structure to reduce the feature channel dimension and computational complexity. To improve the effectiveness of the extracted feature vectors, in this embodiment, a self-attention mechanism convolution operator is introduced into the residual block 1, and its module structure is as Figure 12 or Figure 13 shown.
[0153] For the residual block 1 as Figure 12 shown, it includes a first branch 11 and a second branch 12, where:
[0154] In the first branch 11:
[0155] For the input feature image, first perform a 1×1 CNN convolution operation by the first dimensionality reduction layer 110 to reduce the dimension of the input feature image.
[0156] The second layer of the first branch 11 is a convolution layer 111, and the convolution layer 111 performs a 3×3 CNN convolution operation to extract image features from the feature image output by the first dimensionality reduction layer 110.
[0157] The third layer of the first branch 11 is the first dimensionality increase layer 112, which performs a 1×1 CNN convolution operation to increase the dimension of the feature image output by the convolution layer 111 through a 1×1 CNN convolution operation. After convolution by the first dimensionality increase layer 112, the output feature image obtains higher-level feature information.
[0158] In the second branch 12:
[0159] For the input feature image, first perform a 1×1 CNN convolution operation by the second dimensionality reduction layer 120 to reduce the dimension of the input feature image.
[0160] The second layer of the second branch 12 is an involution layer 121, and the involution layer 121 performs an involution convolution operation. For the feature image output by the first dimensionality reduction layer 110 operation, an n×n involution convolution operation is performed using an involution convolution kernel to extract image features, where n×n is the resolution of the feature image input to the involution layer 121.
[0161] The third layer of the second branch 12 is the second dimensionality increasing layer 122, which performs 1×1 CNN convolution operation. For the feature image output by the convolution operation of the involution layer 121, 1×1 CNN convolution operation is performed for dimensionality increase. After convolution by the second dimensionality increasing layer 122, the output feature image obtains higher-level feature information.
[0162] Then, the outputs of the first dimensionality increasing layer 112 and the second dimensionality increasing layer 122 are added together and used as the output of the residual block 1.
[0163] For Figure 13 the shown residual block 1, it includes a first branch 11, a second branch 12, and a direct connection branch 13, where:
[0164] In the first branch 11:
[0165] For the input feature image, first, the first dimensionality reduction layer 110 performs 1×1 CNN convolution operation to reduce the dimension of the input feature image.
[0166] The second layer of the first branch 11 is the convolution layer 111, and the convolution layer 111 performs 3×3 CNN convolution operation. For the feature image output by the operation of the first dimensionality reduction layer 110, CNN convolution is performed to extract image features.
[0167] The third layer of the first branch 11 is the first dimensionality increasing layer 112, which performs 1×1 CNN convolution operation. For the feature image output by the convolution operation of the convolution layer 111, 1×1 CNN convolution operation is performed for dimensionality increase. After convolution by the first dimensionality increasing layer 112, the output feature image obtains higher-level feature information.
[0168] In the second branch 12:
[0169] For the input feature image, first, the second dimensionality reduction layer 120 performs 1×1 CNN convolution operation to reduce the dimension of the input feature image.
[0170] The second layer of the second branch 12 is the involution layer 121, and the involution layer 121 performs involution convolution operation. For the feature image output by the operation of the first dimensionality reduction layer 110, an n×n involution convolution operation is performed using an involution convolution kernel to extract image features, where n×n is the resolution of the feature image input to the involution layer 121.
[0171] The third layer of the second branch 12 is the second dimensionality increasing layer 122, which performs 1×1 CNN convolution operation. For the feature image output by the involution convolution operation of the involution layer 121, 1×1 CNN convolution operation is performed for dimensionality increase. After convolution by the second dimensionality increasing layer 122, the output feature image obtains higher-level feature information.
[0172] In the direct connection branch 13, the input feature image is output through an identity mapping, that is, the input feature image is output without being changed.
[0173] Then, after the outputs of the first upsampling layer 112, the second upsampling layer 122, and the direct connection branch 13 are added together, they are used as the output of the residual block 1.
[0174] Specifically, Figure 12 and Figure 13 the convolutional layer 121 of both Figure 12 uses the involution operator as an invertible convolution. At the same time, the convolutional layer 111 is retained. Among them, the convolutional layer 111 uses a 3×3 CNN small convolution kernel. The spatial span limitation of the small convolution kernel deprives the convolution kernel of the ability to adapt to different visual patterns at different spatial positions, and there will be computational redundancy between the convolution kernels corresponding to multiple channels. Therefore, through the convolutional layer 121, involution convolution is used to capture the feature relationships of long-distance vision, and better reconstruct the three-dimensional human motion sequence from the long sequence of videos. And if more abundant image features are needed, then as
[0175] (3) Output of the feature vector. The feature image processed by the residual block 1 is pooled using the global average pooling layer. This layer divides the feature image into several blocks, each block has a size of 2x2. For each block, the average value of all elements in it is calculated, and this average value is used as the representative value of the block. All the representative values are used to form a one-dimensional vector, which is used as the output of the global average pooling layer. This vector is the feature vector of the entire feature image, which reflects the feature information of this frame of image.
[0176] Through the above steps, more abundant and accurate image features can be extracted to output the feature vector. Then, according to the classification part of the image recognition algorithm, the objects in the image are classified, and finally the object category and confidence are output.
[0177] When conducting experimental verification, a publicly available dataset can be selected for testing. The verification result is judged according to the accuracy of the image recognition algorithm on the test set.
[0178] Based on the same inventive concept, as Figure 14 shown in the figure is a schematic diagram of an image processing device according to an embodiment of the present invention, including:
[0179] An image acquisition module 1401, configured to acquire a feature image;
[0180] A feature vector calculation module 1402 is configured to input the feature image into a residual network based on an involution operator to obtain a feature vector of the feature image. The residual network at least includes a first branch 11 and a second branch 12. The first branch 11 at least includes a convolutional layer 111, and the second branch 12 at least includes an involution layer 121 that operates based on an involution operator.
[0181] In the present invention, a feature image is input into a residual network based on an involution operator to obtain a feature vector of the feature image. The residual network at least includes a first branch and a second branch. The first branch at least includes a convolutional layer, and the second branch at least includes an involution layer that operates based on an involution operator. In the residual network of the present invention, the first branch including a convolutional layer and the second branch including an involution layer not only retain the short-distance visual relationship ability through the convolutional layer, but also capture the long-distance visual feature relationship through the involution layer, avoiding the problem of the lack of short-distance visual relationship ability caused by only the long-distance visual feature relationship, and can extract richer and more accurate image features and output a feature vector that better fits the image.
[0182] In one embodiment, the residual network includes a plurality of cascaded residual blocks 1, and each of the residual blocks 1 includes the first branch 11, the second branch 12, or the first branch 11 and the second branch 12 whose outputs are added together;
[0183] The step of inputting the feature image into a residual network based on an involution operator to obtain a feature vector of the feature image includes:
[0184] Input the feature image into the first residual block 1, and input the output data of the previous residual block 1 into the next residual block 1, and obtain the output data of the last residual block 1 as the feature vector of the feature image.
[0185] In one embodiment, the residual network further includes a direct connection branch 13, and the output of the direct connection branch is added to the output of the first branch 11 and / or the output of the second branch 12;
[0186] The step of inputting the feature image into a residual network based on an involution operator to obtain a feature vector of the feature image includes:
[0187] Input the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13, and add the output data of the first branch 11, the output data of the second branch 12, and the output data of the direct connection branch 13 as the feature vector of the feature image.
[0188] In one embodiment, the residual network includes a plurality of serially connected residual blocks 1, and each of the residual blocks 1 includes: the first branch 11, the second branch 12, and / or the direct connection branch 13;
[0189] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13, and adding the output data of the first branch 11, the output data of the second branch 12, and the output data of the direct connection branch 13 as the feature vector of the feature image, includes:
[0190] Inputting the feature image into the first branch 11, the second branch 12, and / or the direct connection branch 13 of the first residual block 1, adding the output data of the first branch 11, the output data of the second branch 12, and / or the output data of the direct connection branch 13 of the previous residual block 1, and inputting the result into the next residual block 1, and taking the output data of the last residual block 1 as the feature vector of the feature image.
[0191] In one embodiment, the residual network further includes an image dimensionality reduction layer, and the output data of the image dimensionality reduction layer is the input of the first residual block 1 in the plurality of serially connected residual blocks;
[0192] Inputting the feature image into the first residual block 1 includes:
[0193] Inputting the feature image into the image dimensionality reduction layer for dimensionality reduction operation to obtain a dimensionality-reduced feature image;
[0194] Inputting the dimensionality-reduced feature image into the first residual block 1.
[0195] In one embodiment, it further includes an identification and classification module for:
[0196] Performing image recognition and / or image classification based on the feature vector.
[0197] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0198] As Figure 15 shown is a schematic hardware structure diagram of an electronic device according to the present invention, including:
[0199] At least one processor 1501; and,
[0200] A memory 1502 communicatively connected to at least one of the processors 1501; wherein,
[0201] The memory 1502 stores instructions executable by at least one of the processors. The instructions are executed by at least one of the processors to enable at least one of the processors to execute the image processing method as described above.
[0202] Figure 15 Taking one processor 1501 as an example.
[0203] The electronic device may further include: an input device 1503 and a display device 1504.
[0204] The processor 1501, the memory 1502, the input device 1503, and the display device 1504 may be connected by a bus or other means. In the figure, connection by a bus is taken as an example.
[0205] As a non-volatile computer-readable storage medium, the memory 1502 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the image processing method in the embodiments of the present application. For example, Figure 1 、 Figure 2 The method flow shown. By running the non-volatile software programs, instructions, and modules stored in the memory 1502, the processor 1501 executes various functional applications and data processing, that is, implements the image processing method in the above embodiments.
[0206] The memory 1502 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the image processing method, etc. In addition, the memory 1502 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1502 may optionally include a memory remotely provided relative to the processor 1501, and these remote memories can be connected to the device executing the image processing method through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0207] The input device 1503 can receive input user clicks and generate signal inputs related to user settings and function controls of the image processing method. The display device 1504 may include a display screen and other display devices.
[0208] When the one or more modules are stored in the memory 1502 and run by the one or more processors 1501, the image processing method in any of the above method embodiments is executed.
[0209] The present invention inputs a feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image. The residual network includes at least a first branch and a second branch. The first branch includes at least a convolutional layer, and the second branch includes at least a convolution layer based on the operation of the convolution operator. In the residual network of the present invention, the first branch including the convolutional layer and the second branch including the convolution layer not only retain the short-distance visual relationship ability through the convolutional layer, but also capture the long-distance visual feature relationship through the convolution layer, avoiding the problem of the lack of short-distance visual relationship ability caused by only the long-distance visual feature relationship, and can extract richer and more accurate image features and output a feature vector more suitable for the image.
[0210] An embodiment of the present invention provides a storage medium that stores computer instructions. When a computer executes the computer instructions, it is used to execute all steps of the image processing method described above.
[0211] In the context of the present disclosure, the storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The storage medium may be a machine-readable signal medium or a machine-readable storage medium. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0212] An embodiment of the present invention provides a vehicle, including the image processing device described above or the electronic device described above. It can be understood that the vehicle may also include: a processor, a memory, and a computer program. Among them, the computer program is stored in the memory and is configured to be executed by the processor to implement the image processing method provided by the embodiments of the present disclosure. Among them, the processor and the memory have been described in the embodiments shown Figure 15 and will not be elaborated here.
[0213] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. An image processing method, characterized in that, comprising: obtaining a feature image; inputting the feature image into a residual network based on a convolution operator to obtain a feature vector of the feature image, the residual network at least including a first branch and a second branch, the first branch at least including a convolutional layer, and the second branch at least including a convolution layer based on a convolution operator operation.
2. The image processing method according to claim 1, characterized in that, the residual network includes a plurality of serially connected residual blocks, and each of the residual blocks includes the first branch, the second branch, or the first branch and the second branch whose outputs are added; The step of inputting the feature image into the residual network based on the convolution operator to obtain the feature vector of the feature image includes: inputting the feature image into the first of the residual blocks, and inputting the output data of the previous residual block into the next residual block to obtain the output data of the last residual block as the feature vector of the feature image.
3. The image processing method according to claim 1, characterized in that, the residual network further includes a direct connection branch, and the output of the direct connection branch is added to the output of the first branch and / or the output of the second branch; The step of inputting the feature image into the residual network based on the convolution operator to obtain the feature vector of the feature image includes: inputting the feature image into the first branch, the second branch, and / or the direct connection branch, and adding the output data of the first branch, the output data of the second branch, and the output data of the direct connection branch as the feature vector of the feature image.
4. The image processing method according to claim 3, characterized in that, the residual network includes a plurality of serially connected residual blocks, and each of the residual blocks includes: the first branch, the second branch, and / or the direct connection branch; The step of inputting the feature image into the first branch, the second branch, and / or the direct connection branch, and adding the output data of the first branch, the output data of the second branch, and the output data of the direct connection branch as the feature vector of the feature image includes: inputting the feature image into the first branch, the second branch, and / or the direct connection branch of the first of the residual blocks, adding the output data of the first branch, the output data of the second branch, and / or the output data of the direct connection branch of the previous residual block, and inputting the result into the next residual block to obtain the output data of the last residual block as the feature vector of the feature image.
5. The image processing method according to claim 2, characterized in that, the residual network further includes an image dimensionality reduction layer, and the output data of the image dimensionality reduction layer is the input of the first of the plurality of serially connected residual blocks; The step of inputting the feature image into the first of the residual blocks includes: inputting the feature image into the image dimensionality reduction layer for dimensionality reduction operation to obtain a dimensionality-reduced feature image; inputting the dimensionality-reduced feature image into the first of the residual blocks.
6. The image processing method according to any one of claims 1 to 5, characterized in that, further comprising: Perform image recognition and / or image classification based on the feature vector.
7. An image processing apparatus, characterized in that it includes: an image acquisition module for acquiring a feature image; a feature vector calculation module for inputting the feature image into a residual network based on an involution operator to obtain a feature vector of the feature image, the residual network at least includes a first branch and a second branch, the first branch at least includes a convolutional layer, and the second branch at least includes an involution layer based on involution operator operations.
8. An electronic device, characterized in that it includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method according to any one of claims 1 to 6.
9. A storage medium, characterized in that the storage medium stores computer instructions, and when a computer executes the computer instructions, it is used to execute all steps of the image processing method according to any one of claims 1 to 6.
10. A vehicle, characterized in that it includes the image processing apparatus according to claim 7, or the electronic device according to claim 8.