Image processing method and device based on artificial intelligence, electronic equipment, computer readable storage medium and computer program product
By performing multi-scale feature extraction and scale adjustment on the image in image processing, the problems of poor capture capabilities and environmental sensitivity in the prior art are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411824212.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to capture complex features such as textures and defects when detecting objects in image processing, and is very sensitive to environmental conditions such as lights and backgrounds, resulting in unstable detection results.
By acquiring images, the feature extraction process of channel dimensions and position dimensions is performed on different scales, multi-scale image features are obtained, and these features are scaled to enhance the richness of the features.
It improves the accuracy and robustness of object detection, enhances the network's detection capabilities in complex environments, and can more effectively handle image detection tasks in multi-scale and complex environments.
Smart Images

Figure CN119942171A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to an image processing method, device, electronic device, computer-readable storage medium and computer program product based on artificial intelligence. Background Art
[0002] The object detection task in image processing methods is an important research direction in the field of artificial intelligence. Its purpose is to accurately identify and locate the target object of interest in an image or video. This technology is of great significance for improving the accuracy of image understanding and accurately identifying the target object in the image.
[0003] In the related technology, detection based on traditional methods only considers limited features of the target object in the image, such as color and shape. This limited method has poor capture capabilities for more complex features, such as texture and defects, and is very sensitive to environmental conditions, such as lighting and background. Changes in environmental factors may affect the detection results and reduce the robustness of the method. Summary of the invention
[0004] In order to solve the technical problems existing in the related technologies, the embodiments of the present application provide an image processing method, device, electronic device, computer-readable storage medium and computer program product based on artificial intelligence.
[0005] To achieve the above purpose, the technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides an artificial intelligence-based image processing method, the method comprising:
[0007] acquiring a first image;
[0008] Based on the channel dimension and the position dimension, perform feature extraction processing on the first image at N scales to obtain first image features corresponding to the N scales one by one, where N is a positive integer;
[0009] Performing a scale adjustment process on the M first image features to obtain M second image features, where M is a positive integer less than or equal to N;
[0010] Based on the N first image features and the M second image features, an object detection result of the first image is obtained.
[0011] The embodiment of the present application provides an image processing device based on artificial intelligence, comprising:
[0012] A data acquisition module, used for acquiring a first image;
[0013] A first image feature acquisition module, configured to perform feature extraction processing of N scales on the first image based on a channel dimension and a position dimension, to obtain first image features corresponding to the N scales one by one, where N is a positive integer;
[0014] A second image feature acquisition module, configured to perform scale adjustment processing on the M first image features to obtain M second image features, where M is a positive integer less than or equal to N;
[0015] The target detection module is used to obtain the target detection result of the first image based on the N first image features and the M second image features.
[0016] An embodiment of the present application provides an electronic device, including:
[0017] A memory for storing computer executable instructions;
[0018] The processor is used to implement the artificial intelligence-based image processing method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0019] An embodiment of the present application also provides a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions, when executed by a processor, implement the artificial intelligence-based image processing method provided in the embodiment of the present application.
[0020] An embodiment of the present application also provides a computer program product, including computer executable instructions, characterized in that when the computer executable instructions are executed by a processor, the artificial intelligence-based image processing method provided in the embodiment of the present application is implemented.
[0021] The artificial intelligence-based image processing method, device, electronic device, computer-readable storage medium and computer program product provided in the embodiments of the present application perform feature extraction processing on the first image at N scales based on channel dimensions and position dimensions to obtain first image features corresponding to the N scales one by one. These image features of different scales include channel information and position information, which helps to subsequently improve the accuracy of detection results. The M first image features are scaled to obtain M second image features, thereby making the scales of the features richer. Based on the N first image features and the M second image features, the target detection result of the first image is obtained. Target detection is performed based on these image features of rich scales, thereby enhancing the detection capability of the network in complex environments and improving the accuracy of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1-Figure 4A flowchart of an image processing method based on artificial intelligence according to an embodiment of the present application;
[0023] Figure 5 This is a flow chart of a method for detecting apples before thinning according to an embodiment of the present application;
[0024] Figure 6 A schematic diagram of the YOLOv5 network structure of an embodiment of the present application;
[0025] Figure 7 A schematic diagram of an apple in a complex environment according to an embodiment of the present application;
[0026] Figure 8 This is a schematic diagram of small target data enhancement in an embodiment of the present application;
[0027] Fig. 9 This is a schematic diagram of detection results of different detection distances according to an embodiment of the present application;
[0028] Fig.10 This is a schematic diagram of detection results under different lighting conditions according to an embodiment of the present application;
[0029] Fig.11 This is a schematic diagram of the detection result of an environment without occlusion according to an embodiment of the present application;
[0030] Fig.12 A schematic diagram of the structure of an image processing device based on artificial intelligence according to an embodiment of the present application;
[0031] Fig.13 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0034] The object detection task in image processing methods is an important research direction in the field of artificial intelligence. Its purpose is to accurately identify and locate the target object of interest in an image or video. This technology is of great significance for improving the accuracy of image understanding and accurately identifying the target object in the image. The solution in related technologies is to use traditional methods and deep learning methods to achieve it.
[0035] The first is to use traditional methods. Image detection technology based on traditional methods mainly uses computer vision and image processing technology to analyze and process images. This method usually requires analyzing the shape, texture and other features of the target object, and through a series of preprocessing, feature extraction and regression operations, to achieve image detection and recognition.
[0036] The second is the image detection method based on deep learning, which can automatically extract useful features from the image, and then use the extracted features to obtain the category and location information of the target object. It has the characteristics of faster detection speed, higher accuracy and stronger robustness. Image detection based on deep learning methods can be divided into two categories: two-stage and single-stage.
[0037] Image detection based on traditional methods only considers limited features of the target object, such as color and shape. This limited method has poor capture capabilities for more complex features, such as texture, skin defects, etc., and is very sensitive to environmental conditions, such as lighting and background. Changes in environmental factors may affect the detection results and reduce the robustness of the method.
[0038] Deep learning-based methods often cannot achieve a balance between detection time and detection accuracy in a specific detection environment, and when applying the trained model to different environments, they may face the challenge of transfer learning. At this time, the model may need to be readjusted to adapt to the new application environment.
[0039] Based on this, the embodiment of the present application proposes an image processing method based on artificial intelligence. In various embodiments of the present application, by acquiring an image, performing feature extraction processing on the image at different scales, and obtaining the features of the channel dimension and position dimension corresponding to the different scales, the features of the channel dimension and position dimension corresponding to the different scales are scale-adjusted to obtain the image features corresponding to the adjusted different scales, and the image features corresponding to the different scales and the adjusted image features corresponding to the different scales are used for target detection, which can realize image detection at multiple scales and improve the accuracy of detection. The features of the channel dimension and position dimension at different scales are effectively extracted, and the contextual information of a series of scales is obtained, which enhances the detection ability of the network in complex environments.
[0040] The present application embodiment provides an image processing method based on artificial intelligence. Figure 1 , Figure 1 is a flowchart of an image processing method based on artificial intelligence provided by an embodiment of the present application, which will be combined with Figure 1 Steps 101 to 104 are shown for explanation.
[0041] In step 101, a first image is acquired;
[0042] In step 102, based on the channel dimension and the position dimension, perform feature extraction processing on the first image at N scales to obtain first image features corresponding one-to-one to the N scales, where N is a positive integer.
[0043] In some embodiments, referring to Figure 2 , Figure 1 step 102 shown in can be implemented by the following steps 1021 to 1022, which will be described below in conjunction with Figure 2 for illustration.
[0044] In step 1021, perform feature extraction processing on the input of the n-th feature extraction network through the n-th feature extraction network among N cascaded feature extraction networks.
[0045] In step 1022, transmit the n-th feature extraction result output by the n-th feature extraction network to the (n + 1)-th feature extraction network to continue feature extraction processing to obtain the (n + 1)-th feature extraction result corresponding to the (n + 1)-th feature extraction network.
[0046] As an example, the feature extraction results respectively output by N cascaded feature extraction networks are the first image features corresponding one-to-one to the N scales. n is an integer variable starting from 1 and increasing. The value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th feature extraction network is the first image. When n takes the value of 2 ≤ n < N, the input of the n-th feature extraction network is the (n - 1)-th feature extraction result output by the (n - 1)-th feature extraction network. When n meets the configuration condition, the feature extraction processing is feature extraction processing based on the channel dimension and the position dimension.
[0047] As an example, the feature extraction processing of some feature extraction networks is feature extraction processing based on the channel dimension and the position dimension, or the feature extraction processing of all feature extraction networks is feature extraction processing based on the channel dimension and the position dimension. Specifically, when n meets the configuration condition, for example, the configuration condition here can be that n is an even number, or the configuration condition here can be the specific value of n. The feature extraction processing is feature extraction processing based on the channel dimension and the position dimension.
[0048] As an example, take the value of N as 3 and the configuration condition that n is 3 as an example: when the value of N is 3, there are three feature extraction networks representing the current cascade, such as A→B→C, and the configuration condition n is 3, which means that the C feature extraction network is subjected to feature extraction processing based on the channel dimension and the position dimension, specifically: the input of the A feature extraction network is the first image, and feature extraction is performed on the first image to obtain the feature extraction result of the A feature extraction network; the feature extraction result of the A feature extraction network is input into the B feature extraction network for feature extraction, and the feature extraction result of the B feature extraction network is obtained; the feature extraction result of the B feature extraction network is input into the C feature extraction network for feature extraction based on the channel dimension and the position dimension, and the feature extraction result of the C feature extraction network based on the channel dimension and the position dimension is obtained.
[0049] In some embodiments, see Figure 3 , when n meets the configuration conditions, Figure 2 Step 1021 shown in FIG. 1021 can be implemented by steps 1021A to 1021C as follows. Figure 3 Provide explanation.
[0050] In step 1021A, feature extraction processing based on the channel dimension is performed on the input of the nth feature extraction network to obtain the channel dimension features corresponding to the nth feature extraction network.
[0051] In some embodiments, step 1021A can be implemented by the following technical solutions: performing maximum pooling processing and average pooling processing on the input of the nth feature extraction network, respectively, to obtain a maximum pooling result and an average pooling result; performing multi-layer perception processing on the maximum pooling result and the average pooling result, respectively, to obtain a first multi-layer perception result corresponding to the maximum pooling result and a second multi-layer perception result corresponding to the average pooling result; performing fusion processing on the first multi-layer perception result and the second multi-layer perception result to obtain a first fusion result, and based on the first fusion result and the input of the nth feature extraction network, determining the channel dimension features corresponding to the nth feature extraction network.
[0052] As an example, after obtaining the original input features of the feature extraction network, the original input features are mapped to F∈R C×H×W F∈R C×H×W , and as the input of the feature extraction network, the module generates the channel attention map M in turn. c ∈R C×1×1 , the calculation formula for channel dimension feature extraction is shown in the following formula (1).
[0053]
[0054] Among them, F is the input of the feature extraction network; M c is the channel attention map, i.e., the first fusion result, which can be obtained by the following formula (2); F' is the channel dimension feature, which is used in the subsequent deep feature extraction module. represents element-by-element multiplication. During the multiplication process, the channel extraction module will be broadcasted along the spatial dimension. The broadcast content is the features extracted by the channel extraction module, and the obtained weights are applied to the channels of the original feature map, emphasizing important channels and suppressing unimportant channels, so that the convolutional neural network can focus more on important information when processing features.
[0055] As an example, spatial information is extracted from the input of the feature extraction network, and two spatial context information descriptors F are generated using average pooling (GAP) and maximum pooling (GMP). ap , F mp , these two descriptors represent the average pooling features and maximum pooling features of the input of the feature extraction network respectively. ap , F mp The input shared network generates channel attention mapping to M c ∈R C×1×1 The shared network includes a multilayer perceptron and a hidden layer. To reduce overhead, the hidden layer size is set to R C / r×1×1 , where R represents the receptive field. In a convolutional neural network, the receptive field refers to the area of the input image that a certain point on the input of the feature extraction network can see, that is, the point on the input of the feature extraction network is calculated by the receptive field size area in the input image. C represents the number of neurons; r is the dimension reduction coefficient r used in the hidden layer to reduce the calculation parameters. The calculation method of the channel attention map is shown in the following formula (2).
[0056]
[0057] Among them, F is the input of the feature extraction network, S represents the Sigmiod activation function, and W 0 ∈R C / r×C , W 1 ∈R C×C / r , AvgPool represents average pooling, and MaxPool represents maximum pooling. MLP() represents a multi-layer perceptron model, which can perform multi-layer perceptual processing. MLP(AvgPool(F)) is the second multi-layer perceptual result, and MLP(MaxPool(F)) is the first multi-layer perceptual result. C (F) is the first fusion result.
[0058] In step 1021B, feature extraction processing based on the position dimension is performed on the input of the nth feature extraction network to obtain the position dimension feature corresponding to the nth feature extraction network.
[0059] In some embodiments, step 1021B can be implemented by the following technical solutions: performing average pooling processing on the first image features in the horizontal direction to obtain horizontal feature information; performing average pooling processing on the first image features in the vertical direction to obtain vertical feature information; fusing the horizontal feature information and the vertical feature information to obtain a second fusion result; expanding the second fusion result along the spatial dimension to obtain a horizontal weight in the horizontal direction and a vertical weight in the vertical direction; based on the horizontal weight, the vertical weight and the input of the nth feature extraction network, determining the position dimension feature corresponding to the nth feature extraction network.
[0060] As an example, in the process of deep feature extraction, the channel feature operation uses average pooling and maximum pooling, but the extracted global information is compressed into two spatial information, namely F ap , F mp , it is difficult to effectively retain the position information that plays a key role in the spatial structural features of the image during processing. In order to solve this shortcoming of the channel feature extraction module, the position feature module averages the input of the given feature extraction network along the X-axis direction to obtain horizontal feature information, and averages the input along the Y-axis direction to obtain vertical feature information, and encodes each channel in the horizontal coordinate direction and the vertical coordinate direction along (H, 1) and (1, W) to obtain the position information of the feature map. Among them, H and W are the height and width determined by the input size of the feature extraction network. In the process of obtaining the position dimension feature, the output of the cth channel at the height h, that is, the horizontal feature information, is shown in the following formula (3). Similarly, the output of the cth channel with a width of w, that is, the vertical feature information, is shown in the following formula (4).
[0061]
[0062] Among them, x c (h,i) represents the feature information at the cth channel and at the height i of h, expressed as a matrix. c (j,w) represents the feature information at the cth channel and width j, represented by a matrix. is the horizontal feature information, is the vertical feature information. and Input into the shared 1×1 convolution transformation function F1, the result is the second fusion result. The calculation formula is shown in the following formula (5), which can save accurate position information. σ is the nonlinear Sigmiod activation function, f∈RC / r×(H+W) represents the intermediate feature map that encodes the position information in the horizontal and vertical directions. r is used to represent the ratio of the size of the generated feature block, dividing f into two independent tensors f along the spatial dimension h ∈R C / r×H and f w ∈r C / r×w , using two 1×1 convolution transformations F h and F w f h 、f w Transformed into a tensor with the same number of channels, we get g h and g w , the specific calculation method is shown in the following formula (6) to formula (7). Then, the output g h Expand and use it as the horizontal weight in the horizontal direction, and output g w Expand and use it as the vertical weight in the vertical direction. Finally, the calculation of the position dimension feature is shown in the following formula (8). The deep feature extraction module obtains the target feature map of interest through the channel feature extraction module, and then combines it with the position perception information extracted from the position feature, which can effectively improve the feature extraction capability and thus improve the detection effect of the network in complex environments.
[0063]
[0064] g h =σ(F h (f h )) (6)
[0065] g w =σ(F w (f w )) (7)
[0066]
[0067] Where f is the second fusion result, δ is the nonlinear activation function, x c (i,j) is the feature of channel c at (i,j), g h The expansion of channel c is the horizontal weight in the horizontal direction, g w The expansion of channel c is the vertical weight in the vertical direction; y c (i,j) is the position dimension feature.
[0068] In step 1021C, the nth channel dimension feature and the nth position dimension feature are combined into the nth feature extraction result.
[0069] Through steps 1021A to 1021C, the position dimension features and channel dimension features of the input of the feature extraction network can be obtained. By obtaining and combining the features of these two dimensions, the feature extraction capability can be effectively improved, thereby improving the detection effect in complex environments.
[0070] Through step 1021 to step 1022, features of the first image can be extracted based on the channel dimension and the position dimension at N scales, and context information of a series of scales, namely, channel dimension features and position dimension features, can be obtained, thereby enhancing detection capabilities in complex environments.
[0071] Continue to see Figure 1 In step 103, the M first image features are scaled to obtain M second image features. It should be noted that M is a positive integer less than or equal to N.
[0072] In some embodiments, step 103 can be implemented by the following technical solution: performing feature map expansion processing on the M first image features of the scale to obtain M third image features; fusing the M third image features with the M first image features to obtain the M second image features, wherein the image features involved in the fusion operation have the same initial scale or have different initial scales.
[0073] As an example, it is necessary to resize M of the N first images so that the first images are enlarged, which is helpful for extracting feature information of different scales. For example, for an image with an input resolution of 640×640 pixels, the final scales of the feature maps after feature extraction are 80×80, 40×40, and 20×20, respectively. In order to better detect images of different scales, the feature maps are upsampled after the feature extraction network of the original network, which can further expand the feature maps and help extract feature information of images of different scales. At the same time, the feature maps of size 160×160 extracted by the network are concat-fused with the second-layer feature maps in the Backbone feature extraction network, that is, two or more tensors are connected together in a certain dimension to generate a larger tensor. After such improvements, the model will eventually generate feature maps of four scales of 8×8, 16×16, 32×32, and 64×64.
[0074] In step 104, an object detection result of the first image is obtained based on the N first image features and the M second image features.
[0075] In some embodiments, see Figure 4 The artificial intelligence-based image processing method provided in the embodiment of the present application can also execute the following steps 105 to 110.
[0076] In step 105, a first image sample and a true detection label of the first image sample are obtained.
[0077] In some embodiments, step 105 may be implemented by the following technical solutions: obtaining a second image sample; performing data enhancement processing based on the second image sample to obtain an enhanced image sample set; and obtaining the first image sample from the enhanced image sample set.
[0078] As an example, since some of the targets to be detected are small, it is speculated that the reason why small targets are missed is that the overlap between the real target and the predicted box is far below the expected IOU threshold of 0.45. The small target data enhancement process can be simply understood as increasing the number of samples so that the number of small-sized targets in the same image gradually increases. The corresponding number of candidate boxes (anchors) is 0. If a small target is copied, a new small target will appear in the image, and the corresponding number of anchors will increase to 3. By making multiple copies, it helps to greatly increase the probability of detecting this small target, because the corresponding number of anchors also increases. This provides more small target samples during model training and improves the detection accuracy of small targets.
[0079] In step 106, based on the channel dimension and the position dimension, feature extraction processing is performed on the first image sample at N scales to obtain first image sample features of the first image sample at the N scales. It should be noted that N is a positive integer.
[0080] In step 107, the scale adjustment process is performed on the M first image sample features of the scale to obtain M second image sample features.
[0081] In step 108, based on the N first image sample features and the M second image sample features, an object detection result of the first image sample is obtained.
[0082] In some embodiments, the implementation process of steps 106 to 108 is the same as Figure 1 Steps 102 to 104 are shown to be the same and will not be described again here.
[0083] Continue to see Figure 4 In step 109, a loss function is determined based on the target detection result of the first image sample and the true detection label of the first image sample.
[0084] As an example, in the actual image sample collection process, the area occupied by the target is small, and most of them are interference factors unrelated to the target. For an input image sample, due to the complex environment, thousands of pre-selected boxes may be generated, but only a very small part of them actually contain the target, which brings about the problem of imbalanced category distribution. Therefore, in order to solve the imbalance of positive and negative sample areas during the detection process, focal loss is used to calculate the classification loss. The specific calculation method is shown in the following formula (9).
[0085] L=-α t (1-p t ) γ ln(p t ) (9)
[0086] Among them, p t Represents the probability of predicting the target, when p t When the value of tends to 1, it indicates that the sample is easy to distinguish. At this time, the modulation factor (1-p t ) γ The value of tends to 0, indicating that its contribution to the loss is small, so it can be calculated by α t To suppress the imbalance in the number of positive and negative samples, the balance factor is used to adjust the impact between positive and negative samples to prevent excessive contribution of negative samples to the loss. t =α; for negative samples α t =1-α. Usually, the value of α is between [0,1], indicating the weight ratio of positive and negative samples. γ is a parameter with a value range of [0,5]. It can be used to control the imbalance of the number of simple or difficult samples. γ is the focus factor, which is used to adjust the weight of easy and difficult samples. By introducing the focus factor γ, FocalLoss adjusts the model's attention to easy and difficult samples. When γ increases, the loss contribution of easy-to-classify samples will be further reduced, while the loss contribution of difficult-to-classify samples will increase relatively.
[0087] In step 110, the initialized image processing model is updated based on the loss function.
[0088] In some embodiments, the image processing model obtained in step 110 is an image processing model used to execute the artificial intelligence-based image processing method provided in the embodiments of the present application.
[0089] As an example, to update the model according to the loss function, the gradient descent method can be used to calculate the gradient of the loss function to the model parameters and update the parameters in the opposite direction of the gradient. At the same time, the optimizer such as Adam can be used to adjust the learning rate, and regularization terms can be added to prevent overfitting, thereby gradually improving the model performance. There is no limitation on the method of updating the model according to the loss function.
[0090] The artificial intelligence-based image processing method, device, electronic device, computer-readable storage medium, and computer program product provided in the embodiments of the present application acquire an image, extract features of the image at different scales, obtain image features corresponding to channel dimensions and position dimensions at different scales, then adjust the scales of the image features corresponding to the different scales to obtain the image features corresponding to the adjusted different scales, and use the image features corresponding to the different scales and the image features corresponding to the adjusted different scales for target detection, so as to detect images at multiple scales and improve the accuracy of detection. By extracting features of channel dimensions and position dimensions at different scales, contextual information of a series of scales can be obtained, thereby enhancing the detection capability of the network in complex environments.
[0091] The present application is described below in conjunction with application examples.
[0092] The growth process of apples can be divided into the young fruit stage, thinning stage, growth stage, maturity stage and picking stage. Apples at different growth stages show different characteristics such as color and size. Implementing apple detection before thinning can not only provide a basis for yield estimation, intelligent thinning, fertilization, etc., reduce labor costs, but also enhance the automation and automation level of orchard production management. There are many methods for apple detection in related technologies. The more mainstream solutions are to use traditional methods and deep learning methods to achieve it.
[0093] The first is to use traditional methods. Fruit detection technology based on traditional methods mainly uses computer vision and image processing technology to analyze and process fruit images. This method usually needs to analyze the shape, texture and other features of the fruit, and through a series of preprocessing, feature extraction and regression operations, to achieve the detection and recognition of apples.
[0094] The second is the fruit detection method based on deep learning, which can automatically extract useful features from the image and then use the extracted features to obtain the category and location information of the fruit. It has the characteristics of faster detection speed, higher accuracy and stronger robustness. Fruit detection based on deep learning methods can be divided into two categories: dual-stage and single-stage.
[0095] Apple detection based on traditional methods only considers limited features of apples, such as color and shape. This limited method has poor capture capabilities for more complex features, such as texture, skin defects, etc., and is very sensitive to environmental conditions, such as lighting and background. Changes in environmental factors may affect the detection results and reduce the robustness of the method.
[0096] Deep learning-based methods often cannot achieve a balance between detection time and detection accuracy in a specific detection environment, and applying trained models in different orchards or environments may encounter difficulties in transfer learning, and the model may need to be re-adapted to the new environment.
[0097] The embodiment of the present application proposes a method for detecting apples before thinning. The single-stage target detection method can achieve a balance in detection accuracy while ensuring the detection speed, and can effectively solve the problem that the detection distance and angle in the natural environment will cause the size of the apple to change, and the problem that the complex detection environment brings to the detection. With the support of the terminal system of the embodiment of the present application, it is possible to quickly detect and analyze apples before thinning, which will further promote the application of deep learning technology in the agricultural field and provide technical support for the intelligent and mechanized management and production of apple orchards.
[0098] Compared with the prior art, the embodiments of the present application include at least the following improvements:
[0099] Since most of the public fruit datasets are fruits in the ripening period, there is less apple data available before fruit thinning. Therefore, 1,668 original images were collected and produced using the POSCAL VOC2007 dataset format. The entire dataset contains 13,792 images and approximately 700,000 labels.
[0100] Aiming at the complex environmental problem of apple detection in complex environments, a multi-scale network model for apple detection before thinning based on attention guidance is proposed.
[0101] See also Figure 5 , Figure 5 The content shown is a process of a method for detecting apples before thinning, including an input terminal 201 , a feature extraction terminal 202 , a fusion terminal 203 , and a detection output terminal 204 .
[0102] The embodiment of the present application is aimed at the apple fruit thinning field in the fruit tree planting field, and proposes a single-stage pre-thinning apple detection technology. The whole process is divided into: self-made data set, network model design, network model training, and performance evaluation.
[0103] Step 1: Make your own dataset:
[0104] The image data used was collected from an orchard, and 862 images were selected from 1668 apple images. LabelImg software was used for image annotation. The minimum bounding rectangle was used to ensure that there was only one apple in each rectangular box and that there were as few background elements as possible. A small number of training images may cause overfitting or non-convergence of deep learning methods, and using data enhancement to increase the number of training images can be used to overcome this problem. In the process of data set preparation, MATLAB software and Photoshop image processing tools were used to enhance the data set. The specific enhancement methods included multi-angle rotation, horizontal mirroring, vertical mirroring, brightness change, and blur processing. After enhancement by the above methods, the amount of data in the training data set increased by 15 times, from 862 images to 13792 images. Then, LabelImg software was used to label the images, and the POSCAL VOC2007 data set format was used to generate an ".xml" file. Each image contained about 50 labels, and the entire data set contained about 700,000 labels.
[0105] Step 2 Network model design:
[0106] The original YOLOv5 has four different network structures: YOLOv5s, YOLOv5m, YOLOv5l and YOLOv5x. These four network structures differ in network depth and width. In order to save resources and make the network model lightweight, the YOLOv5s structure is selected. Figure 6 , the network structure mainly consists of four parts: input end 301 (Input), backbone end 302 (Backbone), neck end 303 (Neck) and head end 304 (Head). In the process of target detection, YOLOv5 first preprocesses the image with a resolution of 608×608 pixels, inputs it into the network model, and performs preliminary feature extraction on the image. Then, the extracted features are fused, and after further feature extraction at the neck end, three feature maps of different scales can be obtained, further improving the diversity and robustness of the features. Finally, the generated feature map is sent to different detection layers, the target image generates the corresponding prediction box, and non-maximum suppression NMS (Non-Maximum Suppression) is processed to suppress those prediction boxes with low confidence, and finally the target is detected. Figure 6In 305 shown, there are various components of the entire YOLOv5. The Focus module is a convolutional neural network layer for feature extraction, which is used to compress and combine the information in the input feature map to extract a higher-level feature representation. It is used as the first convolution layer in the network to downsample the input feature map to reduce the amount of calculation and the amount of parameters. The convolution layer is usually composed of three modules: convolution, batch normalization, and activation function. The spatial pyramid pooling structure extracts multi-scale context information by performing pooling operations on the feature map at different scales. This can enhance the model's detection ability for targets of different sizes and improve the accuracy of detection. The C3 module consists of three convolutional layers and multiple bottleneck modules, and the number of these bottleneck modules is determined by the configuration file. The Bottleneck module is one of the basic components in the residual network (ResNetResNet), which is used to reduce the computational complexity of the model and improve the feature extraction capability.
[0107] The apple detection scene before thinning is usually complicated. Due to factors such as the shape, position and light of the fruit, there may be occlusion. The changing external environment will also weaken the features of the apple surface, causing some features to be insignificant, resulting in reduced detection accuracy or missed detection. Traditional manually designed detection methods can only capture limited features of apples and can only detect apples in specific environments. The environment for collecting apple images is very complex. See Figure 7 . Figure 7 The contents shown in (a)-(f) show apple images collected in different environments.
[0108] Based on the above analysis, a multi-scale apple detection method before thinning based on attention guidance is proposed. Firstly, the target detection layer is improved on the basis of the YOLOv5 network to perform multi-scale detection; secondly, the channel feature and position feature information are combined to generate a deep feature module DFEM (Depth Feature Extraction Module) and embedded in the Backbone feature extraction network, that is, Figure 6 The seventh layer in the backbone 302 makes the network pay more attention to the characteristic information of the target of interest; then, the joint loss function of Focal loss and EIOU loss is used to solve the imbalance problem of positive and negative samples in the detection process. Finally, an experimental comparative analysis is carried out on the self-made apple dataset to verify the effectiveness of the proposed method.
[0109] (1) Improved target detection layer: In the original YOLOv5 network, there are three target detection layers at the head end, which correspond to three sets of initialized anchor values. The anchor value is a predefined rectangular box used to sample regions of different positions, scales, and aspect ratios in the image as candidate regions for the target detection model. The input image has a resolution of 640×640 pixels, and the final scales of the feature maps after feature extraction are 80×80, 40×40, and 20×20. In order to better detect apples of different scales, the feature maps are upsampled after the feature extraction network of the original network, which can expand the feature maps and help extract the feature information of apples of different scales. At the same time, the feature maps of size 160×160 extracted by the network are concat fused with the second layer feature maps in the Backbone feature extraction network, that is, two or more tensors are connected together in a certain dimension to generate a larger tensor. After this improvement, the model will eventually generate four feature maps with scales of 8×8, 16×16, 32×32, and 64×64, which will be sent to different detection layers for multi-scale detection, which will help improve the detection effect of apples.
[0110] (2) Depth Feature Extraction Module (DFEM): For image processing tasks, traditional object detection models usually take the entire image as input and use convolutional neural networks to extract features. When processing image areas that are beyond the receptive field of the convolution operation, the model may not be able to obtain global contextual information. For complex scenes with a large amount of background information, the model may allocate too much attention to the background and ignore the key features of the target object. In order to accurately realize apple detection before thinning, the deep feature extraction module (DFEM) is integrated into the YOLOv5 network model. The DFEM module is divided into channel feature extraction and position feature information extraction. This module can effectively realize feature extraction, obtain contextual information of a series of scales, and enhance the network's detection ability in complex environments.
[0111] After the X-th layer feature map is input, the channel feature extraction module first maps the feature to F∈R C×H×W F∈R C×H×W , and as input features, the module generates channel attention maps M in turn. c ∈R C×1×1 The calculation formula for the entire channel feature extraction is shown in the following formula (10).
[0112]
[0113] Among them, F is the feature; M cis the channel attention map, which can be obtained by the following formula (11); F' is the result of channel feature extraction, which is used in the subsequent deep feature extraction module. represents element-by-element multiplication. During the multiplication process, the channel extraction module will be broadcasted along the spatial dimension. The broadcast content is the features extracted by the channel extraction module, and the obtained weights are applied to the channels of the original feature map, emphasizing important channels and suppressing unimportant channels, so that the convolutional neural network can focus more on important information when processing features.
[0114] The deep feature extraction module uses the relationship between channels to generate channel attention features. Since each channel can obtain the characteristics of a specific object, the generated channel attention map plays a key role in the final apple detection effect. This can effectively suppress complex information outside the target, which is beneficial to the detection effect of apples. In order to effectively calculate the channel attention, the module extracts spatial information from the feature map and generates two spatial context information descriptors F using average pooling (GAP) and maximum pooling (GMP). ap , F mp , these two descriptors represent the average pooling features and the maximum pooling features of the feature map respectively.
[0115] The two generated spatial information feature maps are input into the shared network to generate channel attention maps to M c ∈R C×1×1 , that is, F ap , F mp Enter the shared network and the result is M c ∈R C×1×1 The shared network includes a multilayer perceptron and a hidden layer. To reduce overhead, the hidden layer size is set to R C / r×1×1 , where R represents the receptive field. In a convolutional neural network, the receptive field refers to the area of the input image that a point on the feature map can see, that is, the point on the feature map is calculated by the receptive field size area in the input image. C represents the number of neurons; r is the dimension reduction coefficient r used in the hidden layer to reduce the calculation parameters. The calculation method of the channel attention map is shown in the following formula (11).
[0116]
[0117] Among them, S represents the Sigmiod activation function W 0 ∈R C / r×C , W 1 ∈R C×C / r , AvgPool and MaxPool represent average pooling and maximum pooling respectively. MLP() represents the multi-layer perceptron model.
[0118] Position information is also very important for target detection tasks. In the process of deep feature extraction, channel feature operations use average pooling and maximum pooling, but the extracted global information is compressed into two spatial information, namely F ap , F mp , it is difficult to effectively retain the position information that plays a key role in the spatial structural features of the image during processing. In order to solve this shortcoming of the channel feature extraction module, the position feature module uses XAP, i.e. average pooling in the X-axis direction, and YAP, i.e. average pooling in the Y-axis direction, to encode each channel along (H, 1) and (1, W) in the horizontal coordinate direction and vertical coordinate direction respectively to obtain the position information of the feature map. Among them, H and W are the height and width determined by the size of the input image.
[0119] The output of the cth channel at height h in the process of acquiring position information is shown in the following formula (12).
[0120] Similarly, the output of the cth channel with width w is given by the following formula (13).
[0121]
[0122] Among them, x c (h,i) represents the feature information at the cth channel and at the height i of h, expressed as a matrix. c (j,w) represents the feature information at the cth channel and width j, represented by a matrix.
[0123] XAP and YAP in position feature extraction aggregate features along the horizontal and vertical spatial directions respectively, thereby helping the network to locate the target of interest more accurately.
[0124] The generated perceptual feature maps in two directions are used to generate the position information aggregation map. First, and Input into the shared 1×1 convolution transformation function F1, the calculation formula is shown in the following formula (14), which can save accurate position information, σ is the nonlinear Sigmiod activation function, f∈R C / r×(H+W) represents the intermediate feature map that encodes the position information in the horizontal and vertical directions. r is used to represent the ratio of the size of the generated feature block, dividing f into two independent tensors f along the spatial dimension h ∈R C / r×H and f w ∈r C / r×w , using two 1×1 convolution transformations F h and F w f h 、f w Transformed into a tensor with the same number of channels, we get gh and g w For specific calculation methods, please refer to the following formulas (15) and (16). Then, the output g h and g w Expand and use them as attention weights respectively. Finally, the output of the position feature module is shown in the following formula (17). The deep feature extraction module obtains the target feature map of interest through the channel feature extraction module, and then combines it with the position perception information extracted by the position feature, which can effectively improve the feature extraction capability and thus improve the apple detection effect of the network in complex environments.
[0125]
[0126] g h =σ(F h (f h )) (15)
[0127] g w =σ(F w (f w )) (16)
[0128]
[0129] Among them, δ is a nonlinear activation function, x c (i,j) is the feature of channel c at (i,j), g h In the expansion of channel c, g w The expansion in channel c.
[0130] (3) Joint loss function: In an actual orchard, the area occupied by apple targets is relatively small, and most of the orchard is branches, leaves, and complex backgrounds. For an input apple image, due to the complex environment, thousands of pre-selected boxes may be generated, but only a small number of them actually contain the target, which brings about the problem of imbalanced category distribution. Therefore, in order to solve the imbalance of positive and negative sample areas during the detection process, focal loss is used to calculate the classification loss. The specific calculation method is shown in the following formula (18).
[0131] L=-α t (1-p t ) γ ln(p t ) (18)
[0132] Among them, p t Represents the probability of predicting the apple target, when p t When the value of tends to 1, it indicates that the sample is easy to distinguish. At this time, the modulation factor (1-pt ) γ The value of tends to 0, indicating that its contribution to the loss is small, so it can be calculated by α t To suppress the imbalance in the number of positive and negative samples, the balance factor is used to adjust the impact between positive and negative samples to prevent excessive contribution of negative samples to the loss. t =α; for negative samples α t =1-α. Usually, the value of α is between [0,1], indicating the weight ratio of positive and negative samples. γ is a parameter with a value range of [0,5]. It can be used to control the imbalance of the number of simple or difficult samples. γ is the focal factor, which is used to adjust the weight of easy and difficult samples. By introducing the focal factor γ, Focal Loss adjusts the model's attention to easy and difficult samples. When γ increases, the loss contribution of easy-to-classify samples will be further reduced, while the loss contribution of difficult-to-classify samples will increase relatively.
[0133] Step 3 Network model training:
[0134] The apples before thinning are relatively small. It is speculated that the small object is missed because the overlap between the real object and the predicted box is much lower than the expected IOU threshold of 0.45. The small object data enhancement process can be seen in Figure 8 , which can be simply understood as increasing the number of samples so that the number of small-sized apples in the same image gradually increases, that is, the process of copying 401 and generating 402 to 404. The corresponding number of candidate boxes (anchors) is 0. If a small target is copied, a new small target will appear in the image, and the corresponding number of anchors will increase to 3. By making three copies, it helps to greatly increase the probability of detecting this small target, because the corresponding number of anchors also increases. This provides more small target samples during model training and improves the detection accuracy of small targets.
[0135] In the initialization stage of the experiment, in order to reduce the training cost and time of the model, the initial weights of the feature extraction network of the improved YOLO_Net still use the pre-trained yolo5s.pt. During the training process, 8274 (60%) images were randomly selected from 13792 images as the training set, 2759 (20%) as the validation set, and the remaining 2759 (20%) as the test set. The number of model training times was set to 300; the training batch size was set to 8, and setting it too large may exceed the size of the video memory; the size of the input image was 640640 at the beginning; the learning rate was set to 0.01; the confidence threshold was set to 0.25; the intersection-over-union ratio threshold was used in the process of non-maximum suppression iterative screening, and was set to 0.45 based on experience.
[0136] See also Fig. 9,It can be seen that the improved YOLOv5 can recognize objects that the unimproved YOLOv5 cannot recognize at different distances.
[0137] Step 4 Performance Evaluation:
[0138] In order to verify the effectiveness of the proposed method, the results of detection at different distances, under different lighting conditions, and under occlusion conditions are analyzed. Prepare 20 non-training apple images at close distance, medium distance, and long distance, 20 normal light, backlight, and dark light images, and select images with fruit occlusion, branch occlusion, and leaf occlusion. Prepare 20 fruit occlusion, branch occlusion, and leaf occlusion images, see Fig.10 , the improved YOLOv5 can recognize objects that the original YOLOv5 could not recognize under different lighting conditions. Fig.11 , the improved YOLOv5 can recognize objects that the original YOLOv5 cannot recognize under occlusion. The experimental results show that the improved method effectively improves the detection effect compared with the original YOLOv5.
[0139] The embodiment of the present application mainly studies a single-stage method for detecting apples before thinning. In view of the scale changes, complex backgrounds, occlusions, etc. of apples, an attention guidance strategy combining "channel feature extraction and position feature extraction" is used to perform multi-scale detection of apples before thinning. Experimental results show that the method provided in the embodiment of the present application can effectively detect apples in complex environments.
[0140] The embodiment of the present application proposes an attention guidance strategy that combines "channel feature extraction and position feature extraction". In order to accurately realize apple detection before thinning, a deep feature extraction module is integrated. The module is divided into channel feature extraction and position feature information extraction, which can effectively realize feature extraction, obtain contextual information of a series of scales, and enhance the detection ability of the network in complex environments.
[0141] The embodiment of the present application proposes a multi-scale feature detection network in a complex environment. In the original YOLOv5 network, there are three target detection layers in the head end. After feature extraction, the final scales of the feature maps are 80×80, 40×40, and 20×20. In order to better detect apples of different scales, the feature map is upsampled after the feature extraction network of the original network. This can further expand the feature map, which is helpful to extract the feature information of apples of different scales. At the same time, the feature map of size 160×160 extracted by the network is concat fused with the second layer feature map in the Backbone network. After improvement, the model will eventually generate four feature maps of scales of 8×8, 16×16, 32×32, and 64×64, which are sent to different detection layers for multi-scale detection, which helps to improve the detection effect of apples of different sizes.
[0142] Compared with the methods of related technologies, this application has the following beneficial effects:
[0143] (1) Due to factors such as the shape, position and light of the apple, obstruction may occur. The changing external environment will also weaken the features of the apple's surface, causing some features to be inconspicuous. The method provided in the embodiment of the present application improves the problem of decreased detection accuracy or missed detection in complex environments.
[0144] (2) The method provided in the embodiment of the present application not only realizes accurate detection of apples before thinning in various complex environments, but also satisfies the requirement of lightweight detection model and can be moved to terminal applications.
[0145] (3) The method provided in the embodiment of the present application can well balance the detection speed while meeting the detection accuracy.
[0146] In addition, the method provided in the embodiment of the present application is mainly aimed at apples in the early stage of fruit thinning. At this stage, deep learning technology can be applied to many scenarios of apple detection and has achieved good detection results. However, most studies are focused on the detection of apples in the mature stage, and there are few studies on the detection of apples before fruit thinning. Therefore, the design of the network model for apple detection before fruit thinning is implemented in Python 3.7 using the Pytorch framework.
[0147] Smart agriculture is developing rapidly, and investment in smart agriculture is increasing. Through policy support, industrial agglomeration and other measures, the development of smart agriculture has been promoted. The development of smart agriculture is of great significance to the improvement and modernization of agricultural production. It can not only improve agricultural production efficiency and agricultural product quality, but also enhance the sustainability of agricultural production, promote rural economic development, and help farmers increase their income and become rich.
[0148] Apple belongs to the Rosaceae family and has a strong adaptability to climate. It grows between 35-50 degrees north and south latitude and is a major economic tree species. In the process of apple planting and management, the most time-consuming and labor-intensive tasks are apple thinning, spraying, topdressing, and picking. Due to the increase in labor costs and the decrease in the number of skilled workers, the cost of cultivating apples is getting higher and higher. In 2021, the average total cost of apple production in the country was 5,167.78 yuan / mu, up 2.77% from 2020. Among them, the total production cost in the advantaged areas was 5,983.33 yuan / mu, up 2.28% from 2020; the total production cost in other production areas was 4,158.33 yuan / mu, up 4.18% from 2020. In 2021, the average material cost of apple production in China was 2,245.63 yuan per mu, up 4.14% from 2020; the average labor cost was 2,373.44 yuan per mu, up 2.45% from 2020; the average production management and other cost was 554.26 yuan per mu, up 3.45% from 2020. In order to solve the problem of high labor costs in planting, more and more intelligent robots are being developed for apple planting. Intelligent picking robots can automatically identify and pick, which can replace traditional manual picking methods and improve picking efficiency and quality. In addition, robots can also monitor the growth and pest and disease conditions of apples in real time through sensors and image recognition technology, helping fruit farmers to detect problems and take measures in time.
[0149] In order to implement the artificial intelligence-based image processing method of the embodiment of the present application, the embodiment of the present application also provides an artificial intelligence-based image processing device, Fig.12 This is a schematic diagram of the composition structure of an artificial intelligence-based image processing device according to an embodiment of the present application. The artificial intelligence-based image processing device includes: a data acquisition module 501, used to acquire a first image; a first image feature acquisition module 502, used to perform feature extraction processing on the first image at N scales based on a channel dimension and a position dimension, to obtain first image features corresponding to the N scales one by one, where N is a positive integer; a second image feature acquisition module 503, used to perform scale adjustment processing on M first image features, to obtain M second image features, where M is a positive integer less than or equal to N; and a target detection module 504, used to obtain a target detection result of the first image based on the N first image features and the M second image features.
[0150] In some embodiments, the first image feature acquisition module in the artificial intelligence-based image processing device is further configured to perform feature extraction processing on the input of the n-th feature extraction network through the n-th feature extraction network in N cascaded feature extraction networks, and transmit the n-th feature extraction result output by the n-th feature extraction network to the (n + 1)-th feature extraction network to continue feature extraction processing, so as to obtain the (n + 1)-th feature extraction result corresponding to the (n + 1)-th feature extraction network; where n is an integer variable starting from 1 and increasing, and the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th feature extraction network is the first image. When n takes the value of 2 ≤ n < N, the input of the n-th feature extraction network is the (n - 1)-th feature extraction result output by the (n - 1)-th feature extraction network. When n meets the configuration conditions, the feature extraction processing is feature extraction processing based on the channel dimension and the position dimension.
[0151] In some embodiments, the first image feature acquisition module in the artificial intelligence-based image processing device is further configured to perform feature extraction processing based on the channel dimension on the input of the n-th feature extraction network to obtain the channel dimension feature corresponding to the n-th feature extraction network; perform feature extraction processing based on the position dimension on the input of the n-th feature extraction network to obtain the position dimension feature corresponding to the n-th feature extraction network; and combine the n-th channel dimension feature and the n-th position dimension feature to form the n-th feature extraction result.
[0152] In some embodiments, the first image feature acquisition module in the artificial intelligence-based image processing device is further configured to perform max pooling processing and average pooling processing on the input of the n-th feature extraction network respectively to obtain a max pooling result and an average pooling result; perform multi-layer perception processing on the max pooling result and the average pooling result respectively to obtain a first multi-layer perception result corresponding to the max pooling result and a second multi-layer perception result corresponding to the average pooling result; perform fusion processing on the first multi-layer perception result and the second multi-layer perception result to obtain a first fusion result, and determine the channel dimension feature corresponding to the n-th feature extraction network based on the first fusion result and the input of the n-th feature extraction network.
[0153] In some embodiments, the first image feature acquisition module in the artificial intelligence-based image processing device is also used to perform average pooling processing on the first image features in the horizontal direction to obtain horizontal feature information; perform average pooling processing on the first image features in the vertical direction to obtain vertical feature information; fuse the horizontal feature information and the vertical feature information to obtain a second fusion result; expand the second fusion result along the spatial dimension to obtain a horizontal weight in the horizontal direction and a vertical weight in the vertical direction; determine the position dimension feature corresponding to the nth feature extraction network based on the horizontal weight, the vertical weight and the input of the nth feature extraction network.
[0154] In some embodiments, the second image feature acquisition module in the artificial intelligence-based image processing device is also used to perform feature map expansion processing on the M first image features of the scale to obtain M third image features; and fuse the M third image features with the M first image features to obtain the M second image features, wherein the image features involved in the fusion operation have the same initial scale or have different initial scales.
[0155] In some embodiments, the image processing device based on artificial intelligence also includes a training module for obtaining a first image sample and a real detection label of the first image sample; performing the following processing through an initialized image processing model: based on the channel dimension and the position dimension, performing feature extraction processing on the first image sample at N scales to obtain the first image sample features of the first image sample at N scales, where N is a positive integer; performing scale adjustment processing on the first image sample features at M scales to obtain M second image sample features; obtaining the target detection result of the first image sample based on the N first image sample features and the M second image sample features; determining a loss function based on the target detection result of the first image sample and the real detection label of the first image sample, and updating the initialized image processing model based on the loss function to obtain an image processing model for implementing the image processing method based on artificial intelligence. Obtain a second image sample; perform data enhancement processing based on the second image sample to obtain an enhanced image sample set; and obtain the first image sample from the enhanced image sample set.
[0156] In actual application, the first image feature acquisition module 502, the second image feature acquisition module 503 and the target detection module 504 can be implemented by a processor in an image processing device based on artificial intelligence; the data acquisition module 501 can be implemented by a communication interface in an image processing device based on artificial intelligence.
[0157] It should be noted that: the artificial intelligence-based image processing device provided in the above embodiment only uses the division of the above-mentioned program modules as an example when performing image processing. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above.
[0158] Based on the hardware implementation of the above program modules, and in order to implement the image processing method based on artificial intelligence in the embodiment of the present application, the embodiment of the present application also provides an electronic device, Fig.13 The hardware structure diagram of the electronic device of the embodiment of the present application is as follows. The electronic device 600 includes:
[0159] Communication interface 601, capable of exchanging information with other electronic devices;
[0160] The processor 602 is connected to the communication interface 601 to realize information interaction with other electronic devices, and is used to execute the above-mentioned artificial intelligence-based image processing method when running a computer program, and the computer program is stored in the memory 603.
[0161] Specifically, the processor 602 is used to perform feature extraction processing on the first image at N scales based on the channel dimension and the position dimension to obtain first image features corresponding to the N scales one by one, where N is a positive integer; perform scale adjustment processing on M first image features to obtain M second image features, where M is a positive integer less than or equal to N; and obtain a target detection result of the first image based on the N first image features and the M second image features.
[0162] The communication interface 601 is used to acquire a first image.
[0163] It should be noted that the specific processing process of the communication interface 601 and the processor 602 can be understood by referring to the above-mentioned image processing method based on artificial intelligence.
[0164] Of course, in actual application, the various components in the electronic device 600 are coupled together through the bus system 604. It is understandable that the bus system 604 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.13 Various buses are labeled as bus system 604 .
[0165] The memory 603 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device 600. Examples of such data include: any computer program used to operate on the electronic device 600.
[0166] The image processing method based on artificial intelligence provided in the above embodiment of the present application can be applied to the processor 602, or implemented by the processor 602. The processor 602 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above-mentioned image processing method based on artificial intelligence can be completed by the hardware integrated logic circuit or software instructions in the processor 602. The above-mentioned processor 602 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 602 can implement or execute the image processing method, steps and logic block diagram based on artificial intelligence disclosed in the embodiment of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the image processing method based on artificial intelligence disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the memory 603, and the processor 602 reads the information in the memory 603 and completes the steps of the above-mentioned image processing method based on artificial intelligence in combination with its hardware.
[0167] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned artificial intelligence-based image processing method.
[0168] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, a memory 603 storing a computer program, and the computer program can be executed by a processor 602 in an electronic device 600 to complete the steps of the image processing method based on artificial intelligence described in the embodiment of the present application. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, FlashMemory, magnetic surface storage, optical disk, or CD-ROM.
[0169] It should be noted that: "first", "second", "third", etc. are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0170] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0171] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An image processing method based on artificial intelligence, characterized in that: The method includes: Obtain a first image; Based on the channel dimension and the position dimension, perform feature extraction processing on the first image at N scales to obtain first image features corresponding to the N scales, where N is a positive integer; Perform scale adjustment processing on M of the first image features to obtain M second image features, where M is a positive integer less than or equal to N; Based on the N first image features and the M second image features, obtain the object detection result of the first image.
2. The method according to claim 1, characterized in that The performing, based on the channel dimension and the position dimension, feature extraction processing on the first image at N scales to obtain first image features corresponding to the N scales includes: Through the n-th feature extraction network in N cascaded feature extraction networks, perform feature extraction processing on the input of the n-th feature extraction network, and transmit the n-th feature extraction result output by the n-th feature extraction network to the (n + 1)-th feature extraction network to continue feature extraction processing to obtain the (n + 1)-th feature extraction result corresponding to the (n + 1)-th feature extraction network; Wherein, n is an integer variable starting from 1 and increasing, the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the n-th feature extraction network is the first image. When n takes the value of 2 ≤ n < N, the input of the n-th feature extraction network is the (n - 1)-th feature extraction result output by the (n - 1)-th feature extraction network. When n meets the configuration condition, the feature extraction processing is feature extraction processing based on the channel dimension and the position dimension.
3. The method according to claim 2, characterized in that When n meets the configuration condition, the performing, through the n-th feature extraction network in N cascaded feature extraction networks, feature extraction processing on the input of the n-th feature extraction network includes: Perform feature extraction processing based on the channel dimension on the input of the n-th feature extraction network to obtain the channel dimension feature corresponding to the n-th feature extraction network; Perform feature extraction processing based on the position dimension on the input of the n-th feature extraction network to obtain the position dimension feature corresponding to the n-th feature extraction network; Combine the n-th channel dimension feature and the n-th position dimension feature to form the n-th feature extraction result.
4. The method according to claim 3, characterized in that The performing feature extraction processing based on the channel dimension on the input of the n-th feature extraction network to obtain the channel dimension feature corresponding to the n-th feature extraction network includes: Perform max pooling processing and average pooling processing on the input of the n-th feature extraction network respectively to obtain a max pooling result and an average pooling result; Perform multi-layer perceptron processing on the max pooling result and the average pooling result respectively to obtain a first multi-layer perceptron result corresponding to the max pooling result and a second multi-layer perceptron result corresponding to the average pooling result; Perform fusion processing on the first multi-layer perceptron result and the second multi-layer perceptron result to obtain a first fusion result, and based on the first fusion result and the input of the n-th feature extraction network, determine the channel dimension feature corresponding to the n-th feature extraction network.
5. The method according to claim 3, characterized in that: The step of performing feature extraction processing based on the position dimension on the input of the nth feature extraction network to obtain the position dimension feature corresponding to the nth feature extraction network includes: Performing average pooling processing on the first image features in the horizontal direction to obtain horizontal feature information; Performing average pooling processing on the first image features in a vertical direction to obtain vertical feature information; Fusing the horizontal feature information and the vertical feature information to obtain a second fusion result; Expanding the second fusion result along the spatial dimension to obtain a horizontal weight in the horizontal direction and a vertical weight in the vertical direction; Based on the horizontal weight, the vertical weight and the input of the nth feature extraction network, a position dimension feature corresponding to the nth feature extraction network is determined.
6. The method according to claim 1, characterized in that The step of performing a scale adjustment process on the M first image features to obtain M second image features includes: Performing feature map expansion processing on the M first image features of the scale to obtain M third image features; The M third image features are fused with the M first image features to obtain the M second image features.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining a first image sample and a true detection label of the first image sample; The following processing is performed by the initialized image processing model: Based on the channel dimension and the position dimension, performing feature extraction processing on the first image sample at N scales to obtain first image sample features of the first image sample at the N scales; Performing a scale adjustment process on the M first image sample features of the scale to obtain M second image sample features; Based on the N first image sample features and the M second image sample features, obtaining an object detection result of the first image sample; A loss function is determined based on the target detection result of the first image sample and the true detection label of the first image sample, and the initialized image processing model is updated based on the loss function to obtain an image processing model for executing the method described in any one of claims 1 to 6.
8. The method according to claim 7, characterized in that The obtaining of the first image sample comprises: acquiring a second image sample; Performing data enhancement processing based on the second image samples to obtain an enhanced image sample set; The first image sample is obtained from the enhanced image sample set.
9. An image processing device based on artificial intelligence, characterized in that: The device comprises: A data acquisition module, used for acquiring a first image; A first image feature acquisition module, configured to perform feature extraction processing of N scales on the first image based on a channel dimension and a position dimension, to obtain first image features corresponding to the N scales one by one, where N is a positive integer; A second image feature acquisition module, configured to perform scale adjustment processing on the M first image features to obtain M second image features, where M is a positive integer less than or equal to N; The target detection module is used to obtain the target detection result of the first image based on the N first image features and the M second image features.
10. An electronic device, characterized in that: include: A memory for storing computer executable instructions; A processor, configured to implement the method according to any one of claims 1 to 8 when executing computer executable instructions stored in the memory.
11. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the computer executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising computer executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.