Recognition and classification method for distorted rice image

The robust features are extracted through the twin capsule neural network and combined with the VB-ORB algorithm for feature matching, which solves the problems of insufficient recognition capabilities and noise interference in distorted rice image recognition, achieving higher recognition accuracy and robustness.

CN120107804AInactive Publication Date: 2025-06-06ZHEJIANG UNIV CITY COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584801.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient recognition ability, overfitting and noise interference in distorted rice image recognition classification, especially when facing unknown transformations and noises under natural environments, the accuracy and robustness are insufficient.

Method used

The twin capsule neural network was used to extract the robust rice graph features, and the VB-ORB algorithm was used to achieve feature matching on the enhanced feature vector distribution map, overcoming the interference of visual distortion, and capturing the spatial hierarchy relationship and pose phenotype of rice plants.

Benefits of technology

It improves the recognition accuracy and robustness of twisted rice images, reduces the risk of overfitting, and can show higher accuracy and robustness in the face of unknown transformations and noise, achieving better classification and recognition effects of twisted rice images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107804A_ABST
    Figure CN120107804A_ABST
Patent Text Reader

Abstract

The invention relates to an identification and classification method for a distorted rice image. The method comprises the following steps: determining an image region containing rice; extracting image features by using a twin capsule neural network, and obtaining a capsule feature vector distribution diagram; performing contrast enhancement on the capsule feature vector distribution diagram through a single-scale homomorphic filtering algorithm; and through a VB-ORB algorithm based on feature vector calculation, feature matching between the original standard image and the distortion test image is realized on the enhanced feature vector distribution map. The method has the beneficial effects that visual information such as spatial hierarchical relationship and attitude phenotype among different rice plants can be captured, so that the method has higher robustness on optical distortion disturbance from the outside; therefore, the defect that the existing neural network cannot accurately classify and identify the rice image with the visual distortion phenomenon is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image recognition, and more specifically, to a recognition and classification method for distorted rice images. Background Art

[0002] Currently, there are two mainstream methods for achieving recognition and classification of distorted targets: 1) Data augmentation: Data augmentation is a well-known strategy that directly adds explicit transformations to the existing original data domain to generate new samples. The neural network model trained on these samples will learn the features of the images after the transformation, and thus gain robustness to these explicit transformations. Data augmentation methods are easier to understand, easier to implement when the data set to be processed is small, and the effect of increasing the number and diversity of samples is also more intuitive, so they have a very wide range of applications. However, the fidelity of the generated data is limited. The samples generated by some data augmentation methods may be different from the real data, resulting in insufficient recognition of the real data by the model in practical applications. If the distribution of the enhanced data is inconsistent with the original data distribution, the model will overfit the enhanced data and fail to generalize well to the data in the real scene. 2) Embedded Invariant Representation Learning (EIRL): EIRL directly adds constraint rules to the internal loss function of the neural network model, or adjusts the optimization weights to learn the graphic display transformation description, thereby forcing the model to produce invariance to the geometric distortion of the image during the learning optimization process. Compared with the data augmentation method, EIRL has better scalability and practical application effects. However, EIRL usually assumes that the data can be divided into multiple "domains" or "environments", and the model needs to learn invariance between these domains. However, in practical applications, the division of these domains may be difficult to obtain, or the data itself does not fully conform to this division assumption, which affects the generalization ability of the model. Moreover, a key challenge of EIRL is to distinguish causal variables from non-causal variables in the data. In some cases, the data may contain a lot of non-causal information, which may interfere with the model learning the true invariant features. Summary of the invention

[0003] The purpose of the present invention is to address the deficiencies of the prior art and to propose a method for identifying and classifying distorted rice images.

[0004] In a first aspect, a method for identifying and classifying distorted rice images is provided, comprising:

[0005] S1, using the selective search algorithm to preliminarily generate a rectangular calibration frame to determine the image area containing rice;

[0006] S2. Extract image features using a twin capsule neural network to obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of the original standard image and the distorted image at different resolutions through multiple local receptive fields;

[0007] S3, using a single-scale homomorphic filtering algorithm to perform contrast enhancement on the capsule feature vector distribution map and fade irrelevant background to obtain an enhanced feature vector distribution map;

[0008] S4. Through the VB-ORB algorithm based on feature vector calculation, feature matching between the original standard image and the distorted test image is achieved on the enhanced feature vector distribution map, and the distorted rice image is classified and identified according to the feature matching results.

[0009] Preferably, S1 comprises:

[0010] S101, over-segmenting the input image to generate an initial region;

[0011] S102, extracting relevant information of each region, the relevant information including: region size, color histogram and texture feature histogram;

[0012] S103, iteratively merging adjacent regions based on feature similarity to generate candidate regions;

[0013] S104: Perform redundancy screening on the candidate regions to remove duplicate or irregular regions.

[0014] Preferably, in S2, the twin capsule neural network includes a multi-scale convolutional layer, a capsule encoding layer, a link random cutting layer and a decoder.

[0015] Preferably, S2 comprises:

[0016] S201, in the multi-scale convolutional layer, extracting image features through local receptive fields of different sizes;

[0017] S202, in the capsule coding layer, abstracting the image features into feature vectors;

[0018] S203, in the link random cutting layer, randomly shielding part of the feature vectors with Bernoulli distribution to reduce the probability of overfitting;

[0019] S204: Perform image reconstruction in the decoder.

[0020] Preferably, the capsule encoding layer includes an initial capsule layer and an instance capsule layer; S202 includes:

[0021] S2021. In the initial capsule layer, abstract the image features into a first capsule feature vector representing a low-level feature instance;

[0022] S2022: In the instance capsule layer, abstract the first capsule feature vector into a second capsule feature vector representing a high-level feature instance.

[0023] Preferably, S4 includes:

[0024] S401, constructing a scale space pyramid and detecting candidate key dimensions;

[0025] S402, filtering the oscillation response dimension by the principal curvature ratio;

[0026] S403, assigning a binary descriptor to the verified dimension;

[0027] S404, performing feature matching based on the Hamming distance, establishing a correlation between the original image and the distorted image, and determining a classification result of the distorted rice image.

[0028] In a second aspect, a system for identifying and classifying distorted rice images is provided, which is used to execute any of the methods described in the first aspect, including:

[0029] A generation module, used to preliminarily generate a rectangular calibration frame using a selective search algorithm to determine an image area containing rice;

[0030] An extraction module is used to extract image features using a twin capsule neural network and obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of original standard images and distorted images at different resolutions through multiple local receptive fields;

[0031] An enhancement module is used to enhance the contrast of the capsule feature vector distribution map through a single-scale homomorphic filtering algorithm and to fade irrelevant background to obtain an enhanced feature vector distribution map;

[0032] The matching module is used to achieve feature matching between the original standard image and the distorted test image on the enhanced feature vector distribution map through the VB-ORB algorithm based on feature vector calculation, and classify and identify the distorted rice image according to the feature matching result.

[0033] According to a third aspect, a computer storage medium is provided, wherein a computer program is stored in the computer storage medium; when the computer program is executed on a computer, the computer executes any method described in the first aspect.

[0034] In a fourth aspect, an electronic device is provided, including:

[0035] Memory, used to store computer programs;

[0036] A processor is used to execute the computer program to implement any method as described in the first aspect.

[0037] The beneficial effect of the present invention is as follows: the present invention realizes the extraction of robust rice graphic features and the construction of vector distribution through the twin capsule neural network, and then realizes robust feature mapping in the above-mentioned vector distribution through the new algorithm VB-ORB proposed by the present invention, and finally can capture the spatial hierarchical relationship and posture phenotype and other visual information (such as position, angle, scale, etc.) between different rice plants, so as to have stronger robustness to external optical distortion disturbances (such as Gaussian noise, Poisson noise, radial blur), thereby overcoming the problem that the existing neural network cannot accurately classify and identify rice images with visual distortion phenomena, so that when facing unknown transformations and noise, the present invention shows higher accuracy and robustness, reduces the risk of overfitting, and achieves better classification and recognition effect of distorted rice images. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The overall structure diagram of the CapsNetORB model provided for this application;

[0039] Figure 2 A schematic diagram of the overall structure of the twin capsule neural network provided in this application;

[0040] Figure 3 A schematic diagram of the structure of a closed-loop recurrent autoencoder provided in this application;

[0041] Figure 4 Schematic diagram of key point detection of the VB-ORB algorithm provided in this application;

[0042] Figure 5 The original image and the distorted image collected for the verification experiment of this application;

[0043] Figure 6 This is a graph showing the classification accuracy changes of 38 groups of rice images using the CapsNetORB model proposed in the verification experiment of this application;

[0044] Figure 7 Schematic diagram of the recognition rate of 38 groups of plasma rice distorted images by the CapsNetORB model proposed in this application;

[0045] Figure 8 This is a visualization diagram of the feature extraction of rice plants by the twin capsule neural network proposed in the present invention in the verification experiment of this application;

[0046] Fig. 9This is a diagram showing the effect of key point detection and feature association established in the verification experiment of the VB-ORB algorithm proposed in this application. DETAILED DESCRIPTION

[0047] The present invention is further described below in conjunction with embodiments. The description of the following embodiments is only used to help understand the present invention. It should be noted that for ordinary persons in the art, without departing from the principle of the present invention, the present invention can also be modified in some ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

[0048] Embodiment 1:

[0049] In order to overcome the interference of various optical factors in the natural environment, such as various graphic distortion phenomena (for example, compression, fisheye transformation, affine) caused by visual disturbances in the image (for example, scaling, sharpening, overexposure, jitter, relative displacement, etc.), Example 1 of the present application provides a method for accurately identifying and classifying such distorted rice images using a capsule neural network (CapsNet).

[0050] Specifically, the present application provides a CapsNetORB model that is robust to image distortion and a method for identifying and classifying distorted rice images using the model. The model has obvious feature invariance to visual transformations in the image. Figure 1 As shown, the model has two key parts, one is the twin capsule neural network S-CapsNet invented by the present application, and the other is the VB-ORB algorithm based on feature vector calculation invented by the present application. Both algorithms are robust to spatial scale distortion. The former can extract a robust image feature distribution with invariant properties to graphic distortion, and the latter implements a stable mapping from the source domain space to the distortion space on the robust feature distribution. The combination of the two can bring distortion tolerance to the CapsNetORB model, thereby ignoring the interference caused by external visual distortion when identifying and classifying rice images, and improving the recognition accuracy of rice by neural networks in natural environments. In addition, the learning process of CapsNetORB also includes two stages: region of interest positioning and foreground enhancement, which are respectively implemented by the selective search algorithm SSA and the single-scale homomorphic filtering algorithm SSR.

[0051] In addition, the present application provides a method for identifying and classifying distorted rice images, such as Figure 1 As shown, including:

[0052] S1. Use the selective search algorithm to preliminarily generate a rectangular calibration frame to determine the image area containing rice.

[0053] For example, the present application adopts the SSA algorithm as the selective search algorithm, and S1 includes:

[0054] S101, over-segment the input image to generate an initial region.

[0055] Specifically, the image is over-segmented using the graph-based Huttenlocher algorithm, which divides the image into many small regions to obtain initial regions, which serve as the basis for subsequent merging.

[0056] S102, extracting relevant information of each segmented region, wherein the relevant information includes: region size, color histogram, and texture feature histogram.

[0057] S103: Iteratively merge adjacent regions based on feature similarity to generate candidate regions.

[0058] Specifically, all adjacent region pairs (i.e., overlapping or intersecting regions) are found, and then the feature similarity between these regions is calculated based on features such as color, texture, size, and shape, and similar regions are gradually merged. After multiple iterations of merging, the algorithm eventually generates a series of candidate regions and uses these regions as input for target detection. These candidate regions can be used for subsequent target detection or image classification and recognition tasks.

[0059] S104. In order to reduce redundancy and improve efficiency, the algorithm screens the generated candidate regions, such as removing duplicate candidate regions, or removing regions that are too small or irregular in shape.

[0060] S2. Use the twin capsule neural network to extract image features and obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of the original standard image and the distorted image at different resolutions through multiple local receptive fields.

[0061] In S2, the twin capsule neural network includes a multi-scale convolutional layer, a capsule encoding layer, a link random cutting layer and a decoder.

[0062] S2 includes:

[0063] S201. In the multi-scale convolutional layer, extract image features through local receptive fields of different sizes.

[0064] like Figure 2As shown, the present application first adds an additional neural network perception channel based on the original capsule neural network. In each channel, a multi-scale convolution layer is used to replace the conventional convolution layer in the original capsule neural network, and two local receptive fields of different sizes are set to perceive the spatial structure and semantic description of image features, thereby converting the size of the pixel value into the local activity of the descriptor. In CONV-1 and CONV-2, the present application deletes the pooling layer of the original capsule neural network. In order to reduce the feature dimension, the present application replaces the small-sized convolution kernel in the original capsule neural network with a convolution kernel with a larger step size (the step size is greater than or equal to 1) (if the step size is 2, the feature dimension is reduced by 2, and so on). In the first iteration of model training, the output of the multi-scale convolution layer is sent to each capsule unit in the next initial capsule layer with the same probability.

[0065] S202: In the capsule coding layer, abstract the image features into feature vectors.

[0066] The capsule encoding layer includes an initial capsule layer and an instance capsule layer; S202 includes:

[0067] S2021. In the initial capsule layer, the image features are abstracted into a first capsule feature vector representing a low-level feature instance. For example, with respect to a human face, organs such as nose, eyes, and mouth are its low-level feature instances, while the human face is a high-level feature instance.

[0068] S2022: In the instance capsule layer, abstract the first capsule feature vector into a second capsule feature vector representing a high-level feature instance. In the instance capsule layer, the routing protocol algorithm performs parameter update according to the feature vector output by the initial capsule layer.

[0069] S203: In the link random cutting layer, a portion of the feature vectors are randomly shielded using Bernoulli distribution to reduce the probability of overfitting.

[0070] Specifically, in order to eliminate the mutual dependence between capsule units, the present invention adds a link random cut-off layer as a regularization means after the instance capsule layer in the original capsule neural network to make the units independent of each other. However, unlike the convolutional neural network, the output of the entity capsule is a feature vector, rather than a feature map as a scalar. Therefore, in each iteration, the filtering object of the link random cut-off layer should be the entire feature vector, rather than certain dimensions on the vector. In this regard, the present invention has made changes to the calculation rules of the link random cut-off layer, that is, the link random cut-off layer is allowed to treat each feature vector as a whole, thereby ensuring that the direction of the feature vector will not change due to the filtering out of some dimensions by the link random cut-off layer. In this way, in each iterative training, the link random cut-off layer can randomly block a certain proportion of feature vectors based on the Bernoulli distribution, reducing the probability of overfitting.

[0071] S204: Perform image reconstruction in the decoder.

[0072] Specifically, in order to force the entity capsules in the instance capsule layer to extract the instantiation parameters of the image, the present invention adds a decoder after the instance capsule layer of the original capsule neural network to implement the image reconstruction stage. It consists of 3 fully connected layers (see Figure 3 ), which is used to correct the optimization process of the twin capsule neural network by constructing an image reconstruction loss function. This is because if the previous hidden layer has accurately fitted the feature distribution of the image, the image generated by the decoder has a high feature similarity with the original training image. At this time, the feedback loss function value is small. If the previous hidden layer does not mine and extract the image features well, the image generated by the decoder will be very different from the original image. At this time, the returned image reconstruction loss function value is also large, which will encourage the entire model to find a more accurate convergence trend during the training process.

[0073] By adding a decoder, the entire twin capsule neural network model breaks the single serial mode of the original capsule neural network and forms a closed-loop recurrent autoencoder, such as Figure 3 As shown in the figure, all the hidden layers in front of the decoder responsible for image feature extraction (multi-scale convolutional layer, initial capsule layer, instance capsule layer and link random cutting layer) constitute the encoder part of the recurrent autoencoder. The image generation operation of the decoder is equivalent to the inverse operation of the encoder's image feature extraction. The two complete model convergence and image feature extraction in this closed-loop optimization cycle of the recurrent autoencoder.

[0074] like Figure 3 As shown, this application also proposes an optimized loss function for the decoder of the twin capsule neural network. When the feature space of the original training image is given And the capsule feature space extracted by the encoder , the recurrent self-decoder of the twin capsule neural network can solve the mutual mapping between the encoder and the decoder and , Therefore, the image reconstruction loss function shaped by the decoder is to minimize the difference between the generated image and the original image, as shown in the following formula.

[0075]

[0076] Although the added encoder can stimulate the convergence process of the twin capsule neural network, this effect is only auxiliary. In order to prevent the above image reconstruction loss function from affecting the global loss function, the present invention multiplies the above formula by a coefficient of 0.5 to constrain the effect of the reconstruction loss function.

[0077] S3. The capsule feature vector distribution map is contrast enhanced by a single-scale homomorphic filtering algorithm, and irrelevant background is weakened to obtain an enhanced feature vector distribution map.

[0078] S4. Through the VB-ORB algorithm based on feature vector calculation, feature matching between the original standard image and the distorted test image is achieved on the enhanced feature vector distribution map, and the distorted rice image is classified and identified according to the feature matching results.

[0079] Embodiment 2:

[0080] Based on Example 1, Example 2 of the present application provides a more specific method for identifying and classifying distorted rice images, including:

[0081] S1. Use the selective search algorithm to preliminarily generate a rectangular calibration frame to determine the image area containing rice.

[0082] S2. Use the twin capsule neural network to extract image features and obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of the original standard image and the distorted image at different resolutions through multiple local receptive fields.

[0083] S3. The capsule feature vector distribution map is contrast enhanced by a single-scale homomorphic filtering algorithm, and irrelevant background is weakened to obtain an enhanced feature vector distribution map.

[0084] S3 includes:

[0085] S301, a given input image is given by the following formula To process:

[0086]

[0087] In the formula, and They represent the ambient illumination and target reflection components respectively, while the latter is the image enhanced by the single-scale homomorphic filtering algorithm.

[0088] S302, a Gaussian kernel ( ) and the input image Perform convolution operation and get is an approximate value of , therefore, the formula in S301 is rewritten as:

[0089]

[0090] Then, the calculated Each value of is linearly quantized.

[0091]

[0092] After the enhancement In the figure, the grayscale value of the faded background area is between 0 and 127, while the grayscale value of the target area is between 128 and 255, which makes the target area more obvious than the irrelevant background.

[0093] S4. Through the VB-ORB algorithm based on feature vector calculation, feature matching between the original standard image and the distorted test image is achieved on the enhanced feature vector distribution map, and the distorted rice image is classified and identified according to the feature matching results.

[0094] Preferably, S4 includes:

[0095] S401, construct a scale space pyramid and detect candidate key dimensions.

[0096] Specifically, the present application first uses the Difference of Gaussian (DoG) to project a scale space pyramid. Then, with a dimension in the capsule vector as the center, a hypersphere is constructed in the pyramid as the sampling space. The central dimension is then subtracted from the six dimensions located above, below, left, right, in front, and behind the hypersphere to obtain the difference in eigenvalues. If at least four of the absolute values ​​of the difference are greater than a predetermined threshold, the central dimension is temporarily stored as a candidate key point. If multiple candidate key points share the same hypersphere, these key points will be subjected to non-maximum suppression.

[0097] S402. Filter the oscillation response dimension by the principal curvature ratio.

[0098] Specifically, the present invention also proposes a "quality check" stage for candidate dimensions of interest to achieve the fitting of candidate dimensions to nearby dimensions in terms of position, scale, and principal curvature ratio. The present invention divides the candidate dimension by the surrounding dimensions to obtain a ratio, and candidate dimensions with relatively low ratios (and therefore sensitive to visual deformation) are discarded. In the first step, since the Gaussian difference function has a strong edge response at the boundary, the present invention introduces the Hessian matrix to calculate the principal curvature, and discards candidate dimensions with principal curvature ratios below a predetermined threshold.

[0099] S403: Allocate a binary descriptor to the verified dimension.

[0100] Specifically, the present invention uses the BRIEF algorithm to generate a descriptor for each verified dimension of interest, which is a set of bit string descriptors. One of the advantages of BRIEF is that each bit feature has a large variance and the average value is close to 0.5, and the larger variance makes the feature more discriminative. Another advantage is that iterative operations have no dependence on each other. In addition, the generated binary descriptors are also conducive to the implementation of the next feature matching stage.

[0101] S404, performing feature matching based on the Hamming distance, establishing a correlation between the original image and the distorted image, and determining a classification result of the distorted rice image.

[0102] Specifically, the present invention establishes a correspondence based on the above descriptors according to feature similarity. Since BRIEF is a binary model, in the double-cross verification of key dimensions, the present invention selects Hamming distance instead of Euclidean distance to quantify feature similarity. This stage will continue until all dimensions of interest in the original standard image are traversed. Finally, the original image category with the most feature associations is the predicted category of the test image.

[0103] Based on the above ideas, this application gives the necessary steps of the VB-ORB algorithm in detail. For ease of understanding, the following symbols are used to represent the functions and variables in the algorithm: Represents the xth feature dimension in the yth capsule feature vector, using represents the distribution of capsule feature vectors, with k representing the constant multiplication coefficient, represents the scale factor, and Represents a Gaussian function with variable scale, using Represents the scale space region, and The group representing DOG (Octave) is represented by represents the radius of the hypersphere, and and Represent the eigenvalue and its threshold respectively, and use , and Represent the dimensions in the capsule encoding feature vector, the candidate dimensions of interest, and the dimensions of interest after the validity is verified, respectively. represents the absolute value of the eigenvalue difference, size represents the size of the Gaussian kernel, and Represents the Hessian matrix, using and The maximum and minimum eigenvalues ​​of the Hessian matrix are expressed as and Represent the trace and determinant of the Hessian matrix, and use , and Represent the center of gravity, center of mass and sub-sphere of the hypersphere respectively, and use Represents the binary code assignment function, using Represents the descriptor assignment function, using Represents the region of candidate matching dimensions, using represents the Hamming distance.

[0104] The VB-ORB algorithm is as follows:

[0105] enter:

[0106] Output: ; Establish feature associations between the original high-quality image and the distorted image.

[0107] 1. Build a scale space pyramid:

[0108] 1-1. and ( ) Convolution to get (refer to Figure 4 ).

[0109]

[0110] 1-2. Adjacent Subtract (refer to Figure 4 ).

[0111]

[0112] Since the present invention is in two adjacent Definition As the multiplication coefficient, we know The scale factor is , but The scale factor is , similarly, The proportionality factor is .

[0113] 2. Detection of feature dimensions of interest:

[0114] 2-1. is the center of the sphere, and the radius For 4, build a hypersphere from the pyramid (ref. Figure 4 ).

[0115] 2-2. Given , if and only if there are at least 4 When the following formula is satisfied, Considered as (refer to Figure 4 ).

[0116]

[0117] 2-3. If there are multiple Sharing the same hypersphere, all Perform non-maximum suppression and only keep the of .

[0118] 3. Filter oscillation edge response dimension:

[0119] 3-1. Using in a hypersphere ( 1.2, size 9×9) deconvolution , the center of the sphere is .

[0120]

[0121] So as to obtain all of .

[0122]

[0123] 3-2. Use and express of and :

[0124]

[0125] By definition The following ratios can be obtained:

[0126]

[0127] The value is 16 if and only if hour, Accepted as .

[0128] 4. Reshape the sampling coordinate system:

[0129] 4-1. Given all dimensions in a hypersphere , Obtained by the following formula.

[0130]

[0131] Will and Connect and make as the new sampling coordinate system. The direction is .

[0132] 5. Assign descriptors to the verified feature dimensions:

[0133] 5-1. In the rotated coordinate system, a pair of feature dimensions and In the hypersphere (radius , the center of the sphere is ) are randomly sampled. In the two sub-spheres, of and Compared with each other, one The corresponding binary code is obtained by the following formula.

[0134]

[0135]

[0136] 5-2. Continue to randomly select N-1 (N in the present invention is 512) sample dimension groups (for example, A 2 and B 2 , A 3 and B 3 ,… A N and B N ), and repeat step 5-1 for them. Finally, The descriptor can be obtained by the following formula.

[0137]

[0138] 6. Matching and association of feature dimensions of interest:

[0139] 6-1. On an original standard image One As the center, on the distorted image, of Can be generated by KNN.

[0140] 6-2. According to ,exist Determine the distance The two most recent descriptors and .if , then Temporarily save as a candidate matching dimension .

[0141] 6-3.Yes Repeat step 6-2. If Also satisfied , then in and Establish a connection between them.

[0142] 6-4. On this original image, Other Repeat steps 6-1 to 6-3.

[0143] In addition, in order to verify the performance of the CapsNetORB model proposed in the present invention, 12,502 plasma rice growth images in the tillering period were taken at a rice planting base as experimental data, covering 38 different plasma treatment schemes, so the growth conditions of these 38 groups of rice are different. All images are taken horizontally, the image format is JPEG, and each image is a 24-bit color bitmap. In order to simulate several common graphic transformations that images taken under natural conditions are prone to, the collected images are processed with a variety of distortion algorithms, such as motion fuzzy, wide-angle transformation, Gaussian noise, sharpening, etc. Figure 5 As shown. The present invention hopes that the CapsNetORB model can overcome various visual disturbances in the image, correctly extract and distinguish the morphological characteristics of rice plants in different experimental groups, and correctly distinguish the categories to which rice belongs based on the differences in rice phenotypes between the experimental groups. Next, each experimental group is named according to the parameters involved in the plasma treatment of rice, and Table 1 summarizes these parameters according to the treatment order.

[0144] Table 1 Parameters involved in plasma treatment of rice

[0145]

[0146] According to the parameter values ​​and treatment sequence in Table 1, the present invention uses a series of abbreviations to name each experimental group. For example, the abbreviation "3-1-BW-PW-N" means that the plasma treatment object of this group is rice sprouts, the plasma generation gear is 3, the treatment time is 1 hour, the state of planting after treatment is moist, and plasma activated water is given for irrigation later, and no fertilizer is applied.

[0147] For the evaluation of visually distorted rice image classification and recognition, it is incomplete to use only classification accuracy. Therefore, this experiment also uses precision, recall and F1 score as the evaluation criteria for model image classification. The above four indicators are defined as shown in the following four formulas:

[0148]

[0149]

[0150]

[0151]

[0152] Where N represents the total number of test samples. , , and They represent the number of images in the true positive sample group (True Positive, TP), false positive sample group (False Positive, FP), false negative sample group (False Negative, FN) and true negative sample group (True Negative, TN) respectively.

[0153] The experiment first compared the CapsNetORB model with eight other neural network models that also have the ability to resist image distortion, and observed their ability to classify and recognize distorted rice images. They are CapsuleGAN, MS-CapsNet, DR-GAN, SPDA-CNN, Aff-CapsNets, GeoNetM, BCN, and CPL. 70% of the images in the constructed dataset were randomly selected as training samples, and the remaining 30% were used as test samples. Figure 6(a)-(h) show the changing trends of the training accuracy and validation accuracy of the nine models as the training cycle increases. Overall, the CapsNetORB model achieves better classification and recognition results of distorted rice images than other existing models during the training phase.

[0154] Table 2 quantifies the performance indicators of the algorithms in the above training process. The results demonstrate the advantages of CapsNetORB in rice distorted image recognition. Although the training accuracy of CapsNetORB is suppressed by GeoNetM, BCN and CPL, CapsNetORB surpasses all other models in verification accuracy. Specifically, after the model converges, CapsNetORB has the highest verification accuracy of 93.68%, followed by CPL with 93.46%, BCN with 91.28% and GeoNetM with 91.07%. The verification accuracies of DR-GAN and Aff-CapsNet are 90.86% and 90.23% respectively, while the verification accuracies of MS-CapsNet, SPDA-CNN and CapsuleGAN are 88.5%, 85.17% and 84.68% respectively. Among them, DR-GAN and SPDA-CNN both require bounding box and part annotations to assist model training.

[0155] Table 2 Performance comparison of CapsNetORB model and other algorithms in the model training phase

[0156]

[0157] After the model training phase is completed, the CapsNetORB model will be tested comparatively from three aspects: distorted image recognition, feature extraction, and feature matching.

[0158] The first thing that needs to be evaluated is the recognition ability of the CapsNetORB model on distorted rice images. Figure 7The confusion matrix in the figure shows the recognition rate of CapsNetORB for 38 groups of rice images. The ordinate is the true category of the image, and the abscissa is the actual predicted category of the model. The overall distribution of colors (red and orange-red squares) proves that CapsNetORB has a high recognition rate for images in all categories. Among them, the experimental groups with a recognition rate higher than 90% are marked in red fonts, and the experimental groups with an accuracy rate lower than 80% are marked in blue fonts. For the red experimental groups with relatively high recognition rates, their rice has a relatively striking phenotypic appearance (such as the vigorous stems and leaves of 3.5-4-CF-SD-TW and the short plants of NNNSD-TW), making them easier to identify. On the contrary, since the rice growth of the blue experimental groups (such as 3-3-SD-TW-CF and 3.5-3-SD-TW-N) is basically at an average level, there are no significant phenotypic characteristics to attract the attention of the model, so their recognition rates are relatively low. It is also worth noting that not all rice in the red group grew well, because in addition to the healthy plants, CapsNetORB also easily noticed other malnourished rice. Considering the recognition rate and actual rice growth morphology of each experimental group, it is not difficult to find that the plasma treatment schemes of 3-4-SD-TW-CF and 3.5-4-CDS-TW have the best effect on promoting rice growth.

[0159] Furthermore, the embodiment of the present application also conducts an ablation test on the distorted image feature extraction performance of the CapsNetORB model. Since accurate positioning of the region of interest is achieved based on the extracted image features, a robust feature extraction algorithm is essential. In order to test the ability of the twin capsule neural network proposed in the present invention to extract features of distorted images, an ablation experiment was conducted with it and four other CapsNetORB frameworks with different feature extraction algorithms. Specifically, Caps-TripleGAN, MS-CapsNet, CapsuleGAN and standard CapsNet were used to replace the twin capsule neural network in CapsNetORB respectively. The capsule feature vectors learned by the above feature extractor will be sent to the VB-ORB algorithm for feature matching and image category prediction. Therefore, the performance indicators in Table 3 can reflect the characterization capabilities of each feature extractor in disguise.

[0160] Table 3 Comparison of the performance of twin capsule neural network and other algorithms in feature extraction

[0161]

[0162] First, in terms of accuracy, CapsNetORB (that is, the full CapsNetORB) with the twin capsule neural network as the feature extractor achieved the highest 90.18%, followed by Caps-TripleGAN+ VB-ORB with an accuracy of 86.07% and CapsuleGAN+ VB-ORB with an accuracy of 85.17%. On the contrary, the models using MS-CapsNet and standard CapsNet as feature extractors have the lowest accuracy, at 79.87% and 74.18%, respectively. The accuracy test results show that the positive samples correctly identified by the twin capsule neural network + VB-ORB account for the majority of the total number of true positive samples and false positive samples.

[0163] In terms of sensitivity, the combination of Twin Capsule Neural Network + VB-ORB achieved the highest 89.58%, followed by Caps-TripleGAN + VB-ORB and VB-ORB + CapsuleGAN, with sensitivities of 85.81% and 83.82% respectively. The models with MS-CapsNet and standard CapsNet as feature extractors achieved sensitivities of 80.16% and 76.29% respectively. The sensitivity test results show that the combination with Twin Capsule Neural Network as feature extractor has the majority of correctly classified positive samples in all positive images.

[0164] Then the F1 score of the model is tested. As shown in Table 4, the twin capsule neural network + VB-ORB has the highest value of 89.87%, while the MS-CapsNet + VB-ORB combination and the standard CapsNet + VB-ORB combination have the lowest values ​​of 80.01% and 75.22%, respectively. This result verifies that the combination of the twin capsule neural network + VB-ORB can take into account both precision and recall.

[0165] In terms of specificity, the two highest values ​​were obtained by the combination of Twin Capsule Neural Network + VB-ORB (91.43%) and Caps-TripleGAN + VB-ORB (86.11%), followed by CapsuleGAN + VB-ORB with 84.12% and MS-CapsNet + VB-ORB with 83.03%. The standard CapsNet + VB-ORB combination had the lowest specificity of 75.8%.

[0166] Finally, the accuracy results show that compared with the combination of Caps-TripleGAN, MS-CapsNet, CapsuleGAN and standard CapsNet as feature extractors, the complete CapsNetORB achieved 2.66%, 8.37%, 5.85% and 11.4% improvements respectively. The ablation experiment results prove that the twin capsule neural network has excellent image feature learning ability on distorted rice images. This advantage can also be shown by the distribution map of the feature vectors it generates (see Figure 8 ), we can see that the length of the capsule feature vector indicates the possibility of the feature entity existing at that location, and each specific dimension in the vector carries the visual features of that location. Overall, the feature vector length of the rice in the center area is much longer than that of the irrelevant background, which shows that the twin capsule neural network has excellent recognition ability for rice.

[0167] In addition, since the final decision result of CapsNetORB is closely related to the feature matching effect, the feature matching performance of VB-ORB should also be verified. The present invention conducts ablation experiments with classic feature matching algorithms (such as ORB, SIFT, SURF) and recently proposed good feature matching algorithms. Based on the single variable principle, the present invention uses the standard CapsNet to replace the improved twin capsule neural network to provide VB-ORB with a common capsule feature vector, while other algorithms directly perform feature matching on the input test image. The experimental results are shown in Table 4.

[0168] Table 4 Comparison of the performance of VB-ORB and other algorithms in feature matching

[0169]

[0170] In general, the standard CapsNet +VB-ORB combination has the highest score in terms of accuracy, 90.06%, followed by Y.Li with 88.53%. On the contrary, SURF, ORB, and X. Pan rank the last three with accuracies of 73.18%, 77.12%, and 77.67%, respectively.

[0171] In terms of sensitivity, the CapsNet + VB-ORB combination scores higher than the second place Y. Li 1.37% are 89.29% and 87.92% respectively, followed by SuperGlue with 86.6% and A. Baumberg with 83.41%. SIFT and SURF have the lowest scores, 70.64% and 71.64% respectively.

[0172] Next is the F1 score, the top two values ​​(89.67% and 88.22%) are achieved by the standard CapsNet + VB-ORB combination and Y. Li respectively. On the contrary, SIFT and SURF still get the lowest two scores, 70.25% and 72.4% respectively.

[0173] In terms of specificity, the standard CapsNet + VB-ORB scored the highest, followed by Y. Li with 88.73%, A. Baumberg with 87.22%, and SuperGlue with 87.02%. The lowest three specificity scores were 73.37%, 75.38%, and 79.26%, achieved by SIFT, SURF, and ORB, respectively.

[0174] Finally, in terms of accuracy, it can be observed that compared to the second-ranked Y. Li’s 89.18%, the standard CapsNet + VB-ORB achieved an improvement of 2.29%. Fig. 9 The figure shows the key point detection and feature matching effect of VB-ORB algorithm between the original standard image and the distorted test image.

[0175] In summary, during the model training phase, CapsNetORB achieved the highest verification accuracy of 93.68% for the classification and recognition of distorted rice images among all models. During the testing phase, the twin capsule neural network achieved the highest accuracy, sensitivity, F1 score, specificity, and accuracy, which were 90.18%, 89.58%, 89.87%, 91.43%, and 90.81%, respectively. In terms of feature matching, VB-ORB also achieved the highest accuracy, sensitivity, F1 score, specificity, and accuracy, which were 90.06%, 89.29%, 89.67%, 91.02%, and 91.47%, respectively. A series of experimental results prove that CapsNetORB can achieve accurate classification of distorted rice images, and its performance is better than other existing models and algorithms.

[0176] It should be noted that the parts in this embodiment that are the same or similar to those in Embodiment 1 can be referenced to each other and will not be described in detail in this application.

[0177] Embodiment 3:

[0178] Based on Examples 1 and 2, Example 3 of the present application provides a system for identifying and classifying distorted rice images, including:

[0179] A generation module, used to preliminarily generate a rectangular calibration frame using a selective search algorithm to determine an image area containing rice;

[0180] An extraction module is used to extract image features using a twin capsule neural network and obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of original standard images and distorted images at different resolutions through multiple local receptive fields;

[0181] An enhancement module is used to enhance the contrast of the capsule feature vector distribution map through a single-scale homomorphic filtering algorithm and to fade irrelevant background to obtain an enhanced feature vector distribution map;

[0182] The matching module is used to achieve feature matching between the original standard image and the distorted test image on the enhanced feature vector distribution map through the VB-ORB algorithm based on feature vector calculation, and classify and identify the distorted rice image according to the feature matching result.

[0183] It should be noted that the system provided in this embodiment is a system corresponding to the method provided in Embodiments 1 and 2. Therefore, the parts in this embodiment that are the same or similar to Embodiments 1 and 2 can be referenced to each other and will not be repeated in this application.

Claims

1. A method for identifying and classifying distorted rice images, characterized in that: include: S1, using the selective search algorithm to preliminarily generate a rectangular calibration frame to determine the image area containing rice; S2. Extract image features using a twin capsule neural network to obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of the original standard image and the distorted image at different resolutions through multiple local receptive fields; S3, using a single-scale homomorphic filtering algorithm to perform contrast enhancement on the capsule feature vector distribution map and fade irrelevant background to obtain an enhanced feature vector distribution map; S4. Through the VB-ORB algorithm based on feature vector calculation, feature matching between the original standard image and the distorted test image is achieved on the enhanced feature vector distribution map, and the distorted rice image is classified and identified according to the feature matching results.

2. The method for identifying and classifying distorted rice images according to claim 1, characterized in that: S1 includes: S101, over-segmenting the input image to generate an initial region; S102, extracting relevant information of each region, the relevant information including: region size, color histogram and texture feature histogram; S103, iteratively merging adjacent regions based on feature similarity to generate candidate regions; S104: Perform redundancy screening on the candidate regions to remove duplicate or irregular regions.

3. The method for identifying and classifying distorted rice images according to claim 2, characterized in that: In S2, the twin capsule neural network includes a multi-scale convolutional layer, a capsule encoding layer, a link random cutting layer and a decoder.

4. The method for identifying and classifying distorted rice images according to claim 3, characterized in that S2 include: S201, in the multi-scale convolutional layer, extracting image features through local receptive fields of different sizes; S202, in the capsule coding layer, abstracting the image features into feature vectors; S203, in the link random cutting layer, randomly shielding part of the feature vectors with Bernoulli distribution to reduce the probability of overfitting; S204: Perform image reconstruction in the decoder.

5. The method for identifying and classifying distorted rice images according to claim 4, characterized in that: The capsule encoding layer includes an initial capsule layer and an instance capsule layer; S202 includes: S2021. In the initial capsule layer, abstract the image features into a first capsule feature vector representing a low-level feature instance; S2022: In the instance capsule layer, abstract the first capsule feature vector into a second capsule feature vector representing a high-level feature instance.

6. The method for identifying and classifying distorted rice images according to claim 5, characterized in that S4 include: S401, constructing a scale space pyramid and detecting candidate key dimensions; S402, filtering the oscillation response dimension by the principal curvature ratio; S403, assigning a binary descriptor to the verified dimension; S404, performing feature matching based on the Hamming distance, establishing a correlation between the original image and the distorted image, and determining a classification result of the distorted rice image.

7. A recognition and classification system for distorted rice images, characterized in that: Used to perform the method according to any one of claims 1 to 6, comprising: A generation module, used to preliminarily generate a rectangular calibration frame using a selective search algorithm to determine an image area containing rice; An extraction module is used to extract image features using a twin capsule neural network and obtain a capsule feature vector distribution map; the twin capsule neural network learns the feature distribution of original standard images and distorted images at different resolutions through multiple local receptive fields; An enhancement module is used to enhance the contrast of the capsule feature vector distribution map through a single-scale homomorphic filtering algorithm and to fade irrelevant background to obtain an enhanced feature vector distribution map; The matching module is used to achieve feature matching between the original standard image and the distorted test image on the enhanced feature vector distribution map through the VB-ORB algorithm based on feature vector calculation, and classify and identify the distorted rice image according to the feature matching result.

8. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is executed on a computer, the computer executes any one of the methods described in claims 1 to 6.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 6.