Image processing method, model, computer device and storage medium
By combining a multi-receptive-field convolutional module and a prediction module, the problem of insufficient feature extraction in complex scenes of image processing models is solved, thereby improving the accuracy of image classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2026-03-27
AI Technical Summary
Existing image processing models are insufficient in their ability to extract image features in complex scenes, leading to inaccurate classification.
A multi-receptive-field convolution module is adopted, which uses dilated convolution kernels with different dilation rates to perform dilated convolution on image features and fuses the output features of multiple dilated convolution kernels. The prediction module of the image processing model is then used for prediction.
It improves the accuracy of image prediction results, can extract feature information from multi-scale receptive fields, and enhances the richness of image features.
Smart Images

Figure CN115841582B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image processing method, an image processing model, a computer device and a computer readable storage medium. BACKGROUND
[0002] With the development of computer technology and the wide application of computer vision principles, image processing technology has also developed. Image processing technology has important application value in intelligent security fields such as target detection, target recognition, image classification, image editing, etc.
[0003] Taking image classification as an example, the accuracy of image classification can directly affect the performance of downstream tasks of subsequent other visual processing. In the prior art, an image processing model classifies by extracting features of an image. Due to the complexity of the shooting scene of the image, the quality of the image, etc., the feature extraction capability of the image processing model for the image is not high, resulting in inaccurate classification. SUMMARY
[0004] The technical problem solved by the present application is to provide an image processing method, model, computer device and storage medium, which can extract more rich feature information of an image, thereby improving the accuracy of the prediction result.
[0005] To solve the above problems, the first aspect of the present application provides an image processing method, which comprises: acquiring first image features of a target image; processing the first image features by using a multi-receptive field convolution module of an image processing model to obtain second image features, wherein the multi-receptive field convolution module comprises a plurality of atrous convolution kernels with different atrous rates, the atrous convolution kernel is used for performing atrous convolution on the input features to obtain atrous convolution features corresponding to the receptive field and the atrous rate, and the second image features are obtained by fusing the atrous convolution features output by the plurality of atrous convolution kernels; and predicting the second image features by using a prediction module of the image processing model to obtain a prediction result of the target image.
[0006] To solve the above problems, the second aspect of the present application provides an image processing model, which comprises a feature extraction module, a multi-receptive field convolution module and a prediction module, wherein the feature extraction module is used for acquiring first image features of a target image; the multi-receptive field convolution module is used for processing the first image features to obtain second image features, wherein the multi-receptive field convolution module comprises a plurality of atrous convolution kernels with different atrous rates, the atrous convolution kernel is used for performing atrous convolution on the input features to obtain atrous convolution features corresponding to the receptive field and the atrous rate, and the second image features are obtained by fusing the atrous convolution features output by the plurality of atrous convolution kernels; and the prediction module is used for predicting the second image features to obtain a prediction result of the target image.
[0007] To solve the above problems, the third aspect of the present application provides a computer device, which comprises a memory and a processor coupled with each other, the memory stores program data, and the processor is configured to execute the program data to implement any step of the image processing method.
[0008] To solve the above problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program data capable of being executed by a processor, and the program data is used to implement any step of the image processing method.
[0009] The above scheme, by acquiring the first image feature of the target image; the multi-receptive field convolution module of the image processing model is used to process the first image feature, and the second image feature is obtained. Since the multi-receptive field convolution module includes a plurality of hollow convolution kernels with different hollow rates, the hollow convolution kernel is used to perform hollow convolution on the input feature to obtain the hollow convolution feature corresponding to the receptive field and the hollow rate, the feature information of the receptive field with different scales can be extracted, and the second image feature is obtained by fusing the hollow convolution features output by the plurality of hollow convolution kernels. The characteristic information of the multi-scale receptive field of the target image can be fused, and more rich feature information of the image can be extracted, so that the prediction module of the image processing model is used to predict the second image feature, and the prediction result of the target image is obtained. The accuracy of the prediction result can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor. Among them:
[0011] Figure 1 is a structural schematic diagram of an embodiment of the image processing model of the present application;
[0012] Figure 2 is a flowchart of an embodiment of the image processing method of the present application;
[0013] Figure 3 is a structural schematic diagram of an embodiment of the feature extraction module of the present application;
[0014] Figure 4 is a flowchart of an embodiment of step S12 in the present application; Figure 1
[0015] Figure 5 is a flowchart of an embodiment of step S21 in the present application; Figure 4
[0016] Figure 6 is a structural schematic diagram of an embodiment of the receptive field convolutional network of the present application;
[0017] Figure 7 is a structural schematic diagram of an embodiment of the receptive field convolutional network of the present application; Figure 1 is a flowchart of another embodiment of step S12 in the method of the present application;
[0018] Figure 8 is a structural schematic diagram of an embodiment of the multi-receptive field convolutional module of the present application;
[0019] Figure 9 is a structural schematic diagram of an embodiment of the multi-receptive field convolutional module of the present application; Figure 1 is a flowchart of still another embodiment of step S12 in the method of the present application;
[0020] Figure 10 is a structural schematic diagram of another embodiment of the multi-receptive field convolutional module of the present application;
[0021] Figure 11 is a network structure schematic diagram of an embodiment of the image processing model of the present application;
[0022] Figure 12 is a structural schematic diagram of an embodiment of the computer device of the present application;
[0023] Figure 13 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0025] The terms “first” and “second” in the present application are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with “first” and “second” can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of “multiple” is at least two, such as two, three, etc., unless otherwise specifically limited. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0026] Reference to an "embodiment" in this application means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that the embodiments described herein are merely examples from a multitude of embodiments that are in substantial compliance with the principles of the application.
[0027] The term "and / or" herein is merely used to describe an associated relationship between associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" herein means two or more than two. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0028] The application provides the following embodiments, which will be specifically described below.
[0029] Please refer to Figure 1 , Figure 1 is a structural schematic diagram of an embodiment of an image processing model of the application.
[0030] The image processing model 100 can include a feature extraction module 101, a multi-receptive field convolution module 102, and a prediction module 103, wherein the feature extraction module 101, the multi-receptive field convolution module 102, and the prediction module 103 are sequentially connected.
[0031] The feature extraction module 101 is configured to obtain a first image feature of a target image.
[0032] The first image feature can be input into the multi-receptive field convolution module 102, and the multi-receptive field convolution module 102 is configured to process the first image feature to obtain a second image feature, wherein the multi-receptive field convolution module includes a plurality of dilated convolution kernels with different dilated rates, the dilated convolution kernel is configured to perform dilated convolution on the input feature to obtain a dilated convolution feature corresponding to the receptive field and the dilated rate, and the second image feature is obtained by fusing the dilated convolution features output by the plurality of dilated convolution kernels.
[0033] The second image feature can be input into the prediction module 103, and the prediction module 103 is configured to predict the second image feature to obtain a prediction result of the target image.
[0034] The specific processing process of the image processing model of the application on the target image can refer to the implementation process of the following embodiments.
[0035] Referring to Figure 2 , Figure 2 is a flowchart of an embodiment of an image processing method of the present application. The method can include the following steps:
[0036] S11: Obtain first image features of a target image.
[0037] In some embodiments, the image processing method of the present embodiment can be performed using an image processing model, which can be an image classification model to classify images.
[0038] In some embodiments, the feature extraction module of the image processing model can be used to extract features of the target image to obtain the first image features. For example, the feature extraction module can use a convolution kernel to perform convolution processing on the target image to obtain the first image features. For example, the feature extraction module can use a feature extraction algorithm to extract features of the target image to obtain the first image features. For example, the feature extraction algorithm can be LBP algorithm (Local Binary Patterns), SIFT operator (Scale-invariant feature transform), HOG feature extraction algorithm (Histogram of Oriented Gradient), etc.
[0039] In some embodiments, the feature extraction module includes a plurality of convolution layers connected in sequence, wherein the step length of the convolution kernel of at least one convolution layer in the plurality of convolution layers is an integer greater than 1, and the plurality of convolution layers are used to perform convolution processing on the target image in sequence, so that the first image features of the target image can be extracted more completely.
[0040] In some embodiments, the step length of the convolution kernel of one convolution layer in the plurality of convolution layers is an integer greater than 1, for example, the step length of the convolution kernel of the first convolution layer is an integer greater than 1, such as 2, so that the target image can be down-sampled.
[0041] Referring to Figure 3For example, the feature extraction module includes three convolutional layers connected in sequence, which can be denoted as a first convolutional layer, a second convolutional layer, and a third convolutional layer, respectively. Each convolutional layer includes a convolution kernel. In the first convolutional layer (Conv 3x3, s=2), the size of the convolution kernel can be set to 3x3, and the step size s of the convolution kernel can be 2. Through the first convolutional layer, the downsampling task of the target image can be implemented. The size of the convolution kernel included in the second convolutional layer (Conv 3x3, s=1) and the third convolutional layer (Conv 3x3, s=1) can be set to 3x3, and the step size of the convolution kernel can be set to 1. The target image can be input into the first convolutional layer, the output of the first convolutional layer can be used as the input of the second convolutional layer, the output of the second convolutional layer can be used as the input of the third convolutional layer, and the output of the third convolutional layer can be used as the first image feature.
[0042] S12: processing the first image feature by using a multi-receptive field convolution module of the image processing model to obtain a second image feature, wherein the multi-receptive field convolution module includes a plurality of dilated convolution kernels with different dilated rates, the dilated convolution kernel is used for dilated convolution on the input feature to obtain a dilated convolution feature corresponding to the receptive field and the dilated rate, and the second image feature is obtained by fusing the dilated convolution features output by the plurality of dilated convolution kernels.
[0043] The multi-receptive field convolution module of the image processing model includes a plurality of dilated convolution kernels with different dilated rates, and the dilated convolution kernel is used for dilated convolution on the input feature to obtain a dilated convolution feature corresponding to the receptive field and the dilated rate, wherein the dilated convolution can make the output more dense, and without increasing the calculation amount, the field of view of the convolution kernel (the size of the convolution kernel becomes larger) is expanded.
[0044] The input feature can be the first image feature. After the input feature is dilated convolved by the plurality of dilated convolution kernels with different dilated rates included in the multi-receptive field convolution module to obtain a dilated convolution feature corresponding to the receptive field and the dilated rate, the second image feature can be obtained by fusing the dilated convolution features output by the plurality of dilated convolution kernels.
[0045] In some embodiments, the image processing model can include a plurality of multi-receptive field convolution modules, and the first image feature is processed by using the plurality of multi-receptive field convolution modules in sequence to obtain the second image feature.
[0046] S13: predicting the second image feature by using a prediction module of the image processing model to obtain a prediction result of the target image.
[0047] The prediction module of the image processing model is used to predict the second image feature, wherein the prediction module comprises a global average pooling layer and a full connection layer, the global average pooling layer and the full connection layer are sequentially connected, the input feature of the global average pooling layer is the second image feature, the output feature of the global average pooling layer is used as the input feature of the full connection layer, and the output feature of the full connection layer is used as the prediction result of the target image.
[0048] In some embodiments, the image processing model is an image classification model, and the prediction result of the target image is a classification result, for example, the image can be classified in target detection, target recognition, target tracking, image deduplication, etc.
[0049] In this embodiment, the first image feature of the target image is obtained, and the multi-receptive field convolution module of the image processing model is used to process the first image feature to obtain the second image feature. Since the multi-receptive field convolution module comprises a plurality of dilated convolution kernels with different dilated rates, the dilated convolution kernel is used to perform dilated convolution on the input feature to obtain dilated convolution features corresponding to the receptive field and the dilated rate, the feature information of different scale receptive fields can be extracted, and the second image feature is obtained by fusing the dilated convolution features output by the plurality of dilated convolution kernels, the characteristic information of the multi-scale receptive field of the target image can be fused, and more rich feature information of the image can be extracted, so that the prediction module of the image processing model is used to predict the second image feature to obtain the prediction result of the target image, and the accuracy of the prediction result can be improved.
[0050] In some embodiments, referring to Figure 4 Step S12 of the above embodiments can be further extended. The multi-receptive field convolution module of the image processing model is used to process the first image feature to obtain the second image feature, and the embodiment can comprise the following steps:
[0051] S21: for each receptive field convolution network, the network input feature is dilated convolved by the receptive field convolution network to obtain the dilated convolution features corresponding to each dilated convolution kernel in the receptive field convolution network, the dilated convolution features corresponding to each dilated convolution kernel are fused to obtain a first fused feature, and the network output feature is obtained based on the first fused feature; wherein the network input feature of the first receptive field convolution network is the first image feature, and the network output feature of the last receptive field convolution network is directly or after processing as the second image feature.
[0052] In some embodiments, the multi-receptive field convolution module comprises at least one receptive field convolution network connected in sequence, each receptive field convolution network comprising a plurality of parallelly connected hole convolution kernels. For each receptive field convolution network, the network input features are subjected to hole convolution by the receptive field convolution network to obtain hole convolution features corresponding to each hole convolution kernel in the receptive field convolution network, and the hole convolution features corresponding to each hole convolution kernel are fused to obtain first fused features, and the network output features are obtained based on the first fused features.
[0053] In some embodiments, the network input features of the first receptive field convolution network are the first image features, the network output features of the first receptive field convolution network are taken as the network input features of the second receptive field convolution network, and the network output features of the last receptive field convolution network can be directly or after processing taken as the second image features.
[0054] In some embodiments, the number of hole convolution kernels contained in each receptive field convolution network is different, and multiple different scale feature information of the target image can be extracted.
[0055] In some embodiments, the number of hole convolution kernels contained in each receptive field convolution network is in the same order as the connection order of each receptive field convolution network. That is, the number of hole convolution kernels contained in the receptive field convolution network decreases in turn according to the connection order, and the hole rate of the hole convolution kernel decreases in turn. In this way, as the network deepens, the network output features extracted contain a large enough receptive field, and rich feature information of the target image is extracted. By reducing the number of hole convolution kernels, the amount of parameters added to the network can be minimized.
[0056] In some embodiments, the multiple receptive field convolution networks can comprise networks of the same structure. For example, the multiple receptive field convolution networks are first receptive field convolution networks, and the network output features of the last first receptive field convolution network can be directly or after processing taken as the second image features. For example, the multiple receptive field convolution networks are second receptive field convolution networks, and the network output features of the last second receptive field convolution network can be taken after processing as the second image features.
[0057] In some embodiments, one of the receptive field convolution networks is a first receptive field convolution network, and at least one receptive field convolution network other than the first receptive field convolution network is a second receptive field convolution network. For example, the multi-receptive field convolution module comprises a first receptive field convolution module and multiple second receptive field convolution modules, wherein the network output features of the first receptive field convolution module are taken as the network input features of the first second receptive field convolution module, and the network output features of the last second receptive field convolution network are directly or after processing taken as the second image features.
[0058] In some embodiments, referring to Figure 5 The step S21 of the above embodiments can be further extended. The receptive field convolution network is used to perform a dilated convolution on the network input feature to obtain a dilated convolution feature corresponding to each dilated convolution kernel of the receptive field convolution network, and the dilated convolution features corresponding to each dilated convolution kernel are fused to obtain a first fused feature, and the network output feature is obtained based on the first fused feature. The present embodiment can include the following steps:
[0059] S211: The dimension reduction convolution layer is used to perform dimension reduction processing on the network input feature to obtain a dimension reduction input feature.
[0060] In some embodiments, referring to Figure 6 , the receptive field convolution network includes a dimension reduction convolution layer (Conv 1x1, s=1), a dilated convolution layer (ConvHD, s=2 or s=1) and a first dimension increase convolution layer (Conv 1x1, s=1) connected in sequence, and the dilated convolution layer of the receptive field convolution network is composed of a plurality of dilated convolution kernels connected in parallel in the receptive field convolution network.
[0061] In some embodiments, when the receptive field convolution network is a first receptive field convolution network, the step length of the dilated convolution layer (ConvHD, s=2) is 2, and when the receptive field convolution network is a second receptive field convolution network, the step length of the dilated convolution layer (ConvHD, s=1) is 1.
[0062] In some embodiments, the dimension reduction convolution layer (Conv 1x1, s=1) of the receptive field convolution network includes a convolution kernel with a size of 1x1 and a step length of 1, which is used to perform dimension reduction processing on the network input feature to obtain a dimension reduction input feature.
[0063] In some embodiments, the network input feature of the first receptive field convolution network is a first image feature.
[0064] S212: Each dilated convolution kernel in the dilated convolution layer is used to perform a dilated convolution on the dimension reduction input feature to obtain a dilated convolution feature corresponding to each dilated convolution kernel, and the dilated convolution features corresponding to each dilated convolution kernel are fused to obtain a first fused feature.
[0065] In some embodiments, the dilated convolution layer of the receptive field convolution network is composed of a plurality of dilated convolution kernels connected in parallel in the receptive field convolution network, and the plurality of dilated convolution kernels have different dilation rates. For example, the dilated convolution layer (ConvHD, s=1) is composed of four dilated convolution kernels with a size of 3x3 {(Conv 3x3, s=1, d=1), (Conv 3x3, s=1, d=2), (Conv 3x3, s=1, d=3), (Conv 3x3, s=1, d=4)}, and the dilation rates are 1, 2, 3 and 4, respectively.
[0066] The dimension-reduced input features are respectively subjected to the hole convolution by using each hole convolution kernel in the hole convolution layer, to obtain the hole convolution features corresponding to each hole convolution kernel. Since the hole rates of the hole convolution kernels are different, the receptive fields of different scales can be extracted by using the hole convolution on the dimension-reduced input features by the multiple hole convolution kernels in parallel. The hole convolution kernel with a larger hole rate extracts the receptive field of the hole convolution feature with a larger size.
[0067] Then, the hole convolution features corresponding to each hole convolution kernel are fused to obtain the first fused feature. Specifically, the hole convolution features corresponding to each hole convolution kernel are subjected to the weighting processing to obtain the first fused feature. For example, different weights can be set for the hole convolution features corresponding to the hole convolution kernels with different hole rates, or the weights of the hole convolution features corresponding to each hole convolution kernel are the same, and the first fused feature is obtained by combining the weights in the form of weight accumulation.
[0068] In this step, in order to reduce the parameter quantity and the floating point calculation quantity of the model, different group convolution with different group sizes is used in each of the four parallel hole convolution kernels. In the receptive field convolution network, the input of each hole convolution kernel is the dimension-reduced input feature. When the 4 parallel 3x3 hole convolution kernels perform the hole convolution operation, the complete feature information can be obtained, and the loss of information is avoided. The feature information (hole convolution feature) output by the 4 parallel 3x3 hole convolution kernels is combined in the form of weight accumulation and used as the input of the first dimension-increasing convolution layer. Such a combination manner can ensure that the original fine-grained feature information is not lost, and at the same time, the first fused feature output contains information of different receptive field sizes, so that the multi-scale feature extraction capability of the model is finally realized.
[0069] S213: The first fused feature is subjected to the dimension-increasing processing by using the first dimension-increasing convolution layer, to obtain the network output feature.
[0070] The first dimension-increasing convolution layer (Conv 1x1, s=1) of the receptive field convolution network contains a convolution kernel with a size of 1x1, and the step length is 1, which is used for the dimension-increasing processing of the first fused feature to obtain the network output feature.
[0071] Based on the network output feature, the second image feature can be obtained.
[0072] In this embodiment, the hole convolution layer (ConvHD, s=2) uses the hole convolution kernel with a step length of 2 to perform the hole convolution operation, which takes into account the downsampling function of the network model, and at the same time, avoids the problem that the feature information is lost seriously due to the dimension-reducing operation in the downsampling process (dimension-reducing processing) by using the dimension-reducing convolution layer (Conv 1x1, s=1).
[0073] In some embodiments, please refer to Figure 7 The step S12 of the above embodiment can be further extended. One of the multi-receptive field convolution modules is a first receptive field convolution network. The multi-receptive field convolution module of the image processing model is used to process the first image feature to obtain a second image feature. The embodiment can include the following steps:
[0074] S31: The dimension reduction convolution layer is used to perform dimension reduction processing on the network input feature to obtain a dimension reduction input feature.
[0075] In some embodiments, one of the receptive field convolution networks is a first receptive field convolution network, and at least one receptive field convolution network of the first receptive field convolution network is a second receptive field convolution network. The network input feature of the first receptive field convolution network is a first image feature.
[0076] S32: Each of the hole convolution kernels in the hole convolution layer is used to perform hole convolution on the dimension reduction input feature to obtain a hole convolution feature corresponding to each of the hole convolution kernels, and the hole convolution features corresponding to each of the hole convolution kernels are fused to obtain a first fusion feature.
[0077] S33: The first dimension increase convolution layer is used to perform dimension increase processing on the first fusion feature to obtain a network output feature.
[0078] The first receptive field convolution network includes a dimension reduction convolution layer (Conv 1x1, s=1), a hole convolution layer (Conv HD, s=2 or s=1) and a first dimension increase convolution layer (Conv 1x1, s=1) connected in sequence. The hole convolution layer of the receptive field convolution network is composed of a plurality of hole convolution kernels (for example, (Conv 3x3, s=1, d=1), (Conv 3x3, s=1, d=2), (Conv 3x3, s=1, d=3), (Conv 3x3, s=1, d=4)) connected in parallel in the receptive field convolution network, and the hole rates are 1, 2, 3 and 4 respectively.
[0079] In some embodiments, the first receptive field convolution network is the first receptive field convolution network;
[0080] The first receptive field network can be used to perform the above steps S31 to S33. The specific implementation process of steps S31 to S33 can refer to the specific implementation process of steps S211 to S213 described above, and the present application will not be repeated here.
[0081] In addition, the embodiment can further include the following steps:
[0082] S34: The down-sampling module is used to perform down-sampling processing on the network input feature of the first receptive field convolution network to obtain a target down-sampling feature.
[0083] In some embodiments, the one receptive field convolutional network is a first receptive field convolutional network, the first receptive field convolutional network is a first receptive field convolutional network. At least one receptive field convolutional network other than the first receptive field convolutional network is a second receptive field convolutional network. The network input feature of the first receptive field convolutional network is a first image feature.
[0084] Please refer to Figure 8 , the multi-receptive field convolutional module further comprises a down-sampling module, the down-sampling module is connected in parallel with the first receptive field convolutional network, and the step length of the hole convolution kernel contained in the first receptive field convolutional network is the same as the down-sampling multiple of the down-sampling module.
[0085] In some embodiments, the down-sampling module comprises an average pooling layer (AvgPooling 3x3, s=2) and a second dimension increasing convolutional layer (Conv 1x1, s=1) connected in sequence.
[0086] Among them, the average pooling layer uses a convolution kernel size of 3x3, and sets its step length to 2, and the step length of 2 can represent the down-sampling multiple of the down-sampling module is 2. Through this design, the down-sampling task of the first image feature can be completed.
[0087] The step length of the hole convolution kernel (ConvHD, s=2) contained in the first receptive field convolutional network is the same as the step length of the average pooling layer (AvgPooling 3x3, s=2). The step length of the average pooling layer can be set according to the required down-sampling multiple, which is not limited in the present application.
[0088] The down-sampling module is used to down-sample the first image feature to obtain a target down-sampling feature. Specifically, the connection order of the down-sampling module is: the input of the average pooling layer is the output of the feature extraction module, that is, the first image feature, and the output of the average pooling layer is the input of the second dimension increasing convolutional layer. The average pooling layer (AvgPooling 3x3, s=2) can be used to perform an average pooling operation on the network input feature (the first image feature) of the first receptive field convolutional network to obtain an initial down-sampling feature, and the dimension of the initial down-sampling feature is smaller than that of the network input feature of the first receptive field convolutional network. The second dimension increasing convolutional layer (Conv 1x1, s=1) is a convolution kernel with a step length of 1 and a size of 1x1, and the second dimension increasing convolutional layer (Conv 1x1, s=1) is used to perform dimension increasing processing on the initial down-sampling feature to obtain the target down-sampling feature.
[0089] S35: The target sampling feature is fused with the network output feature of the first receptive field convolutional network to obtain a second fusion feature, and the second fusion feature is used as the network input feature of the next receptive field convolutional network of the first receptive field convolutional network.
[0090] The target sampling feature output by the downsampling module is fused with the network output feature of the first receptive field convolutional network. The fusion manner can be weighted summation, or vector addition, etc. The second fusion feature is obtained through fusion.
[0091] In the first receptive field convolutional network, multiple dilated convolution kernels with a step of 2 are used for dilated convolution operation in this embodiment. The downsampling function of the model can be considered, and the problem of serious loss of feature information caused by the need for dimension reduction operation in the downsampling process of the dimension increasing convolution layer can be avoided.
[0092] In addition, since the downsampling module and the first receptive field convolutional network are connected in parallel, and the first image feature is input into the downsampling module and the first receptive field convolutional network for processing respectively, and the target sampling feature is fused with the network output feature of the first receptive field convolutional network to obtain the second fusion feature, the downsampling task can be completed without loss of feature information, and the calculation amount of the modules is not increased too much.
[0093] In some embodiments, the second fusion feature is used as the network input feature of the next receptive field convolutional network of the first receptive field convolutional network. The next receptive field convolutional network can be the second receptive field convolutional network. The second fusion feature is processed based on the subsequent receptive field convolutional network, and the second image feature can be obtained. For example, the network output feature of the last receptive field convolutional network is directly or processed as the second image feature.
[0094] In some embodiments, referring to Figure 9 The step S12 of the above embodiment can be further extended. One of the receptive field convolutional networks of the multi-receptive field convolutional module is the first receptive field convolutional network, and at least one of the receptive field convolutional networks of the first receptive field convolutional network is the second receptive field convolutional network. The first image feature is processed by the multi-receptive field convolutional module of the image processing model to obtain the second image feature. The embodiment can include the following steps:
[0095] S41: For each second receptive field convolutional network, the network output feature of the second receptive field convolutional network is processed by the attention module corresponding to the second receptive field convolutional network to obtain a second output feature, wherein the second output feature is used as the network input feature of the next second receptive field convolutional network of the second receptive field convolutional network. The second output feature of the last second receptive field convolutional network is used as the second image feature.
[0096] In some embodiments, one of the receptive field convolution networks of the multi-receptive field convolution module is a first receptive field convolution network, at least one of the receptive field convolution networks other than the first receptive field convolution network is a second receptive field convolution network, and the second receptive field convolution network includes all of the receptive field convolution networks other than the first receptive field convolution network.
[0097] Referring to Figure 10 The multi-receptive field module further includes at least one attention module (ECANetBlock) corresponding to each second receptive field convolution network. The second receptive field convolution network has the same structure as the first receptive field convolution network, for example, the second receptive field convolution network includes sequentially connected dimension reduction convolution layers (Conv 1x1, s=1), a dilated convolution layer (ConvHD, s=2 or s=1), and a first dimension increasing convolution layer (Conv 1x1, s=1), and the dilated convolution layer of the receptive field convolution network is composed of a plurality of dilated convolution kernels (for example, (Conv 3x3, s=1, d=1), (Conv 3x3, s=1, d=2), (Conv 3x3, s=1, d=3), (Conv 3x3, s=1, d=4)) connected in parallel in the receptive field convolution network, where the dilated rates are 1, 2, 3, and 4, respectively.
[0098] There are multiple second receptive field convolution networks, the number of dilated convolution kernels included in each second receptive field convolution network is different, and the number of dilated convolution kernels included in the second receptive field convolution network decreases in turn according to the connection order of the second receptive field convolution network.
[0099] For the first second receptive field convolution network, the network output feature of the second receptive field convolution network is input into the attention module, so that the network output feature of the second receptive field convolution network is processed by the attention module corresponding to the second receptive field convolution network to obtain a second output feature, wherein the second output feature is used as the network input feature of the next second receptive field convolution network of the second receptive field convolution network, until the second output feature of the last second receptive field convolution network is used as the second image feature.
[0100] In this embodiment, the input of the attention module is set as the output of the receptive field convolution network, which can make the network assign higher weight values to the parts of the feature information that are more important without increasing the number of parameters, so that the model pays more attention to the main information part of the image.
[0101] For the above-mentioned embodiments, the following provides an image processing method as an example to illustrate the above-mentioned embodiments.
[0102] Referring to Figure 11The image processing model can include a feature extraction module, a multi-receptive field convolution module, and a prediction module. The image processing model is an image classification model, and the prediction result is a classification result.
[0103] Specifically, the feature extraction module includes three convolution layers connected in sequence, which are (Conv 3x3, s=2), (Conv 3x3, s=1), and (Conv 3x3, s=1) in turn. The first image feature of the target image can be obtained, and the down-sampling task processing of the target image can be realized.
[0104] The multi-receptive field convolution module includes a plurality of receptive field convolution networks, a down-sampling module, and an attention module. The first receptive field convolution network is a first receptive field convolution network, and the remaining other receptive field convolution networks are second receptive field convolution networks. The down-sampling module is connected in parallel with the first receptive field convolution network, and the step length of the hole convolution kernel included in the first receptive field convolution network is the same as the down-sampling multiple of the down-sampling module. Each second receptive field convolution network can be one-to-one corresponding to at least one attention module.
[0105] The first receptive field convolution network and / or the first second receptive field convolution network include a dimension reduction convolution layer (Conv 1x1, s=1), a hole convolution layer (ConvHD, s=2 or s=1), and a first dimension increasing convolution layer (Conv 1x1, s=1) connected in sequence. The hole convolution layer of the receptive field convolution network is composed of a plurality of hole convolution kernels (such as (Conv 3x3, s=1, d=1), (Conv 3x3, s=1, d=2), (Conv 3x3, s=1, d=3), (Conv 3x3, s=1, d=4)) connected in parallel in the receptive field convolution network.
[0106] The first receptive field convolution network processes the first image feature to obtain a first fusion feature.
[0107] The down-sampling module includes an average pooling layer (AvgPooling 3x3, s=2) and a second dimension increasing convolution layer (Conv 1x1, s=1), and performs down-sampling processing on the first image feature to obtain a target down-sampling feature.
[0108] The target down-sampling feature and the first fusion feature output by the first receptive field convolution network are fused to obtain a second fusion feature.
[0109] The second receptive field convolution network processes the second fusion feature to obtain a second output feature, and the second output feature of the last second receptive field convolution network is directly taken as or processed as a second image feature.
[0110] The image processing model can include a first receptive field convolutional network and a down-sampling module connected in parallel, and N second receptive field convolutional networks and an attention module connected in sequence. The above-mentioned first receptive field convolutional network and down-sampling module connected in parallel, and N second receptive field convolutional networks and an attention module connected in sequence can be M, N and M are integers greater than 1. According to the specific application scenario, the values of N and M can be set at each stage, and the number of N and M contained in each stage can be different.
[0111] For example, a receptive field convolutional network including 4 stages, in the first stage, the first receptive field convolutional network and the down-sampling module are connected in parallel, and the N second receptive field convolutional networks and the attention module are connected in sequence, the first receptive field convolutional network and / or the second receptive field convolutional network includes 4 hole convolution kernels connected in parallel {for example, (Conv 3×3, s=1, d=1), (Conv 3×3, s=1, d=2), (Conv 3×3, s=1, d=3), (Conv 3×3, s=1, d=4)}.
[0112] In the second stage, the first receptive field convolutional network and / or the second receptive field convolutional network includes 3 hole convolution kernels connected in parallel {for example, (Conv 3×3, s=1, d=1), (Conv 3×3, s=1, d=2), (Conv 3×3, s=1, d=3)}.
[0113] In the third stage and the fourth stage, the number of hole convolution kernels connected in parallel in the first receptive field convolutional network and / or the second receptive field convolutional network decreases in turn, for example, the third stage includes 2 hole convolution kernels connected in parallel {for example, (Conv 3×3, s=1, d=1), (Conv 3×3, s=1, d=2)}. The fourth stage includes 1 hole convolution kernel (Conv 3×3, s=1, d=1).
[0114] In some embodiments, only the down-sampling operation in the second step is used before the start of each stage, and the network in the remaining stages does not perform the down-sampling operation. The above-mentioned four stages can be repeated multiple times respectively. With the deepening of the network depth, 4 kinds of hole rate convolution operations are used in the first stage, 3 kinds of hole rate convolution operations are used in the second stage, 2 kinds of hole rate convolution operations are used in the third stage, and 1 kind of hole rate convolution operation is used in the fourth stage, and the maximum hole rate convolution operation is used in turn. As the network depth deepens, the extracted feature information is very rich and the receptive field is large enough, and using multi-scale convolution only increases the parameter quantity of the network, and the improvement of the network performance is not very large. Through the above-mentioned manner, the parameter quantity and the floating point calculation quantity of the network model are greatly reduced, and the performance index of the image processing model is improved.
[0115] As an example, the image processing model of the present application is an image classification model, and the prediction result is a classification result. The same target image is classified by using the above-mentioned image classification model and other image classification models in the prior art, respectively, to obtain the classification results of each model. Comparative analysis is performed on the above-mentioned image classification model and the image classification models in the prior art (ResNet, pre-act.ResNet, iResNet, NL-ResNet, SE-ResNet, ECA-ResNet, ResNeXt, PyConvResNet), and the comparative analysis results are shown in Table 1 as follows:
[0116]
[0117] Table 1 Comparative table of the image classification model of the present application and the image classification model in the prior art
[0118] Among them, Top 1 represents the accuracy rate of Top 1, Top 5 represents the accuracy rate of Top 5, and floating point calculation amount (FLOPs, Floating Point Operations) represents the number of floating point operations, which can be used to measure the complexity of the above-mentioned image classification model.
[0119] As can be seen from the above, the parameter amount and the floating point calculation amount of the image classification model of the present application are 14.81M and 3.27G, respectively, which greatly reduces the parameter amount and the floating point calculation amount compared with the image classification model in the prior art, and also can improve the accuracy rate of image classification to a certain extent.
[0120] For the above-mentioned embodiments, the present application provides a computer device, please refer to Figure 12 , Figure 12 is a structural schematic diagram of an embodiment of the computer device of the present application. The computer device 50 comprises a memory 51 and a processor 52, wherein the memory 51 and the processor 52 are coupled to each other, the memory 51 stores program data, and the processor 52 is used to execute the program data to realize the steps in any embodiment of the above-mentioned image processing method.
[0121] In this embodiment, the processor 52 can also be called CPU (Central Processing Unit, central processing unit). The processor 52 can be an integrated circuit chip with signal processing capability. The processor 52 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 52 can also be any conventional processor or the like.
[0122] The specific implementation of this embodiment can refer to the implementation process of the above-mentioned embodiments, which will not be repeated here.
[0123] For the method of the above-mentioned embodiments, it can be realized in the form of a computer program, and thus the present application proposes a computer readable storage medium, please refer to Figure 13 , Figure 13 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 60 stores program data 61 capable of being executed by a processor, and the program data 61 can be executed by the processor to realize the steps of any embodiment of the above-mentioned image processing method.
[0124] The specific implementation of this embodiment can refer to the implementation process of the above-mentioned embodiments, which will not be repeated here.
[0125] The computer readable storage medium 60 of the present embodiment can be a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc. which can store program data 61, or it can also be a server storing the program data 61. The server can send the stored program data 61 to other devices for running, or it can also run the stored program data 61.
[0126] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be realized by other ways. For example, the above-mentioned device implementation is only schematic, for example, the division of the module or unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0127] The unit described as a separate component can be or can not be physically separated, and the component displayed as a unit can be or can not be a physical unit, that is, it can be located in one place, or it can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to realize the purpose of the present embodiment scheme.
[0128] In addition, each of the functional units in the embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0129] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium, which is a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for making an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor execute all or part of the steps of the methods of the embodiments of the present application.
[0130] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be realized by a general computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a computer readable storage medium and executed by a computing device, or they can be made into each integrated circuit module, or multiple modules or steps among them can be made into a single integrated circuit module to realize. Therefore, the present application is not limited to any specific hardware and software combination.
[0131] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. An image processing method, characterized by, The method comprises: obtaining a first image feature of a target image; processing the first image feature by using a multi-receptive field convolution module of an image processing model to obtain a second image feature, wherein the multi-receptive field convolution module comprises a plurality of hole convolution kernels with different hole rates, the hole convolution kernels are used for performing hole convolution on input features to obtain hole convolution features with receptive fields corresponding to the hole rates, and the second image feature is obtained by fusing hole convolution features output by the plurality of hole convolution kernels; predicting the second image feature by using a prediction module of the image processing model to obtain a prediction result of the target image; wherein the multi-receptive field convolution module comprises at least one receptive field convolution network and a down-sampling module connected in sequence, each of the receptive field convolution networks comprises a plurality of hole convolution kernels connected in parallel, one of the receptive field convolution networks is a first receptive field convolution network, the down-sampling module is connected in parallel with the first receptive field convolution network, the step length of the hole convolution kernels contained in the first receptive field convolution network is the same as the down-sampling multiple of the down-sampling module; and the processing of the first image feature by using the multi-receptive field convolution module of the image processing model to obtain the second image feature comprises: performing down-sampling processing on network input features of the first receptive field convolution network by using the down-sampling module to obtain target down-sampling features; fusing the target down-sampling features and network output features of the first receptive field convolution network to obtain second fusion features, the second fusion features serving as network input features of a next receptive field convolution network of the first receptive field convolution network; wherein network input features of the first receptive field convolution network are the first image features, and network output features of the last receptive field convolution network are directly or after processing as the second image features.
2. The method of claim 1, wherein, The processing of the first image feature by using the multi-receptive field convolution module of the image processing model to obtain the second image feature further comprises: for each of the receptive field convolution networks, performing hole convolution on network input features by using the receptive field convolution network to obtain hole convolution features corresponding to the hole convolution kernels in the receptive field convolution network, fusing the hole convolution features corresponding to the hole convolution kernels to obtain first fusion features, and obtaining network output features based on the first fusion features.
3. The method of claim 2, wherein, The receptive field convolution networks are a plurality, and the number of the hole convolution kernels contained in each of the receptive field convolution networks is different.
4. The method of claim 3, wherein, The high-low order of the number of the hole convolution kernels contained in each of the receptive field convolution networks is the same as the connection order of each of the receptive field convolution networks.
5. The method of claim 2, wherein, The receptive field convolution network comprises a dimension reduction convolution layer, a hole convolution layer and a first dimension increase convolution layer connected in sequence, and the hole convolution layer of the receptive field convolution network is composed of the plurality of hole convolution kernels connected in parallel in the receptive field convolution network; The network input feature is subjected to the dilated convolution of the receptive field convolution network, to obtain the dilated convolution feature corresponding to each dilated convolution kernel in the receptive field convolution network, and the dilated convolution features corresponding to each dilated convolution kernel are fused to obtain a first fused feature, and the network output feature is obtained based on the first fused feature, comprising: The network input feature is subjected to dimension reduction processing by the dimension reduction convolution layer to obtain a dimension reduction input feature; Each dilated convolution kernel in the dilated convolution layer is used to subject the dimension reduction input feature to dilated convolution, to obtain the dilated convolution feature corresponding to each dilated convolution kernel, and the dilated convolution features corresponding to each dilated convolution kernel are fused to obtain the first fused feature; The first fused feature is subjected to dimension increasing processing by the first dimension increasing convolution layer to obtain the network output feature.
6. The method according to claim 2 or 5, characterized in that, The dilated convolution features corresponding to each dilated convolution kernel are subjected to weighting processing to obtain the first fused feature. The first receptive field convolution network is the first receptive field convolution network; 7. The method of claim 1, wherein, And / or, the downsampling module comprises a mean pooling layer and a second dimension increasing convolution layer connected in sequence, and the network input feature of the first receptive field convolution network is subjected to downsampling processing by the downsampling module to obtain a target downsampling feature, comprising: The network input feature of the first receptive field convolution network is subjected to a mean pooling operation by the mean pooling layer to obtain an initial downsampling feature, and the dimension of the initial downsampling feature is smaller than that of the network input feature of the first receptive field convolution network; The initial downsampling feature is subjected to dimension increasing processing by the second dimension increasing convolution layer to obtain the target downsampling feature. At least one of the receptive field convolution networks other than the first receptive field convolution network is a second receptive field convolution network, and the multi-receptive field convolution module further comprises at least one attention module corresponding to each second receptive field convolution network one by one; 8. The method of claim 1, wherein, The first image feature is processed by the multi-receptive field convolution module of the image processing model to obtain a second image feature, further comprising: For each second receptive field convolution network, the network output feature of the second receptive field convolution network is subjected to attention processing by the attention module corresponding to the second receptive field convolution network to obtain a second output feature, wherein the second output feature serves as the network input feature of the next second receptive field convolution network of the second receptive field convolution network, and the second output feature of the last second receptive field convolution network serves as the second image feature. The second receptive field convolution network comprises all the receptive field convolution networks other than the first receptive field convolution network.
9. The method of claim 8, wherein, The first image feature of the target image is obtained, comprising:
10. The method of claim 1, wherein, The feature extraction module of the image processing model is used to extract features from the target image to obtain the first image feature. 11. The method of claim 10, wherein, The feature extraction module comprises a plurality of convolution layers connected in sequence, wherein a step length of a convolution kernel of at least one of the plurality of convolution layers is an integer greater than 1.
12. The method of claim 1, wherein, The image processing model is an image classification model, and the prediction result is a classification result.
13. An image processing model, characterized in that, The method comprises: a feature extraction module configured to obtain a first image feature of a target image; a multi-receptive field convolution module configured to process the first image feature to obtain a second image feature, wherein the multi-receptive field convolution module comprises a plurality of hole convolution kernels with different hole rates, the hole convolution kernel is configured to perform hole convolution on an input feature to obtain a hole convolution feature corresponding to a receptive field and the hole rate, and the second image feature is obtained by fusing hole convolution features output by the plurality of hole convolution kernels; a prediction module configured to predict the second image feature to obtain a prediction result of the target image. The multi-receptive field convolution module comprises at least one receptive field convolution network and a down-sampling module connected in sequence, each of the receptive field convolution networks comprises a plurality of hole convolution kernels connected in parallel, one of the receptive field convolution networks is a first receptive field convolution network, the down-sampling module is connected in parallel with the first receptive field convolution network, a step length of a hole convolution kernel included in the first receptive field convolution network is the same as a down-sampling multiple of the down-sampling module, and the multi-receptive field convolution module processes the first image feature to obtain a second image feature, comprising: performing down-sampling processing on a network input feature of the first receptive field convolution network by using the down-sampling module to obtain a target down-sampling feature; fusing the target down-sampling feature and a network output feature of the first receptive field convolution network to obtain a second fusion feature, and the second fusion feature is used as a network input feature of a next receptive field convolution network of the first receptive field convolution network.
14. A computer device, comprising: The memory and the processor are coupled to each other, the memory stores program data, and the processor is configured to execute the program data to implement the steps of the method of any one of claims 1 to 12.
15. A computer readable storage medium characterized by: The memory stores program data capable of being executed by the processor, and the program data is used to implement the steps of the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Dense multi-scale target detection system and method
CN112529098A