Quality rating and classifying method for rice grains
By performing multi-angle image acquisition and preprocessing of rice grains, combined with convolutional neural network and self-attention mechanism, the characteristics of rice grain quality are extracted and rated, and the problem of relying on manual and single standards in the existing technology is solved, and efficient and accurate quality rating is achieved.
Patent Information
- Application Number
- CN202510058004.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the quality rating of rice grains relies on manual assessment to consume labor and have a single standard, so it is impossible to effectively evaluate various defects that affect quality, resulting in unreasonable pricing.
A rice grain quality rating classification method is proposed, which uses multi-angle images for preprocessing and blocking, uses convolutional neural network and self-attention mechanism to extract features, and combines category and position information for quality rating.
It achieves efficient and accurate rating of rice grain quality, reduces the error and inefficiency of manual ratings, can identify various defects that affect quality, and improves the rationality of ratings.
Smart Images

Figure CN119992180A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing, and in particular to a method for grading and classifying the quality of rice grains. Background Art
[0002] In the prior art, the quality assessment of rice grains usually requires manual assessment, which is labor-intensive and relies heavily on the personal experience of the quality assessor.
[0003] In order to solve this problem, some image processing technologies are also used in the prior art to grade rice grains, but the current evaluation standards are relatively simple. For example, the prior art usually uses the size of rice grains as the only feature to screen the quality of rice grains, and is unable to grade other defects that affect the quality of rice grains, resulting in unreasonable pricing of rice grain quality.
[0004] Therefore, how to provide a reasonable rice quality rating and classification method is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In order to solve the technical problem of unreasonable quality grading and classification of rice grains in the prior art, the present invention proposes a quality grading and classification method for rice grains.
[0006] The method for grading and classifying the quality of rice grains proposed by the present invention comprises:
[0007] Acquire multi-angle images of rice grains and perform preprocessing;
[0008] The preprocessed image is divided into blocks, and the position of each image block is encoded during the block division to generate the position information matrix parameters, and the image blocks are input into the convolutional neural network model to obtain the corresponding feature layer;
[0009] Flatten the feature layer into a linear sequence, add a learnable embedding vector before the sequence, and add category information flags and position information matrix parameters to the sequence;
[0010] Feature extraction is performed on the linear sequence, and the importance of each feature layer is paid attention to through the self-attention mechanism, so as to classify the quality rating of rice grains.
[0011] Further, the pre-processing includes:
[0012] Convert multi-angle graphics into corresponding RGB images;
[0013] Perform Gaussian filtering to denoise the RGB image;
[0014] Scale the denoised image;
[0015] Normalize the scaled images and add the dimensions of the training samples.
[0016] Furthermore, the position information matrix parameters include a row index and a column index of the image block in the image, and a position encoding vector.
[0017] Furthermore, the preprocessed image is 244 pixels*244 pixels*3 channels, the size of the convolution kernel is 16*16, and the feature layer obtained by blocking is 14*14*768.
[0018] Furthermore, the Multi-Head Attention structure of the Vit algorithm is used to extract features from linear sequences.
[0019] Furthermore, before feature extraction, layer normalization is performed to convert the input linear sequence data into data with a mean of 0 and a variance of 1.
[0020] Furthermore, before feature extraction, layer normalization is performed to convert the input linear sequence data into data with a mean of 0 and a variance of 1.
[0021] Further, feature extraction of the linear information sequence includes:
[0022] The input information sequence is divided into N heads. The calculation of each head includes generating the query vector matrix (Q), key vector matrix (K), value vector matrix (V), and the corresponding weight matrix (W), and calculating the attention mechanism to capture the relationship between different features;
[0023] The results of N head calculations are stitched together to obtain the feature output of the image corresponding to the rice grain.
[0024] Furthermore, after the feature extraction is completed, the input layer parameters and the feature extraction layer parameters are added to obtain the data of the first residual network;
[0025] Perform layer normalization on the data of the first residual network, and pass the layer normalized data into the multi-layer perceptron;
[0026] Add the data of the multi-layer perceptron to the data of the first residual network to obtain the data of the second residual network;
[0027] The data of the second residual network and the learnable embedding vector are passed into the multi-layer perceptron to complete the classification of the quality rating of the rice grains.
[0028] The present invention utilizes a convolutional neural network to process rice grain images in blocks, adds relevant category and position information, adopts self-attention and multi-head-attention mechanisms, extracts features of eight rice grain quality defects, including broken rice, yellow rice, chalky rice, diseased spots, mixed rice, embryo-retained rice, skin-retained rice and normal rice, adjusts the corresponding parameters of the loss function, and trains the eight rice grain quality categories, so as to achieve efficient and accurate recognition of the quality of rice grains and reduce the errors and inefficiency caused by manual rating. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention is described in detail below with reference to the embodiments and accompanying drawings, wherein:
[0030] Figure 1 It is the main flow chart of the present invention. DETAILED DESCRIPTION
[0031] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] Thus, a feature indicated in this specification will be used to illustrate one of the features of an embodiment of the present invention, rather than implying that each embodiment of the present invention must have the described feature. In addition, it should be noted that this specification describes many features. Although some features can be combined together to illustrate possible system designs, these features can also be used in other combinations that are not explicitly described. Thus, unless otherwise stated, the described combinations are not intended to be limiting.
[0033] The present invention uses a deep learning classification neural network model to rate the quality of rice grains, so as to evaluate the quality of rice grains from various aspects, thereby obtaining reliable and reasonable rating results.
[0034] like Figure 1 As shown, the rice grain quality grading and classification method of the present invention mainly includes the following steps.
[0035] Step 1, obtaining multi-angle images of rice grains and performing preprocessing;
[0036] A device for photographing rice grains can be set up. The device uses a photoelectric sensor to connect 2D industrial cameras and light source control modules at three or more angles in series, and each rice grain is photographed at 360° multi-angles. The rice grains are placed in the device, and multi-angle imaging is performed on each rice grain. The rice grains pass through the working area of the photoelectric sensor, generating light signals. The photoelectric sensor converts the light signals into electrical signals and transmits them to the light source module and the camera. After the light source controller receives the electrical signals, it lights up the corresponding light source after a preset delay. After each camera receives the electrical signals, it starts to expose the rice grains for optical imaging after a preset delay. After the camera converts the analog signals into digital signals that can be processed by the computer, it triggers the corresponding callback function to generate an image storage queue to collect and organize the images of the rice grains. Since it takes a certain amount of time to process and rate the images when imaging rice grains at multiple angles. The use of image storage queues allows the two processes of image acquisition and image processing to proceed in parallel, that is, while continuing to acquire new rice grain images, the existing images in the queue are preprocessed and graded, achieving efficient management of data flow, allowing image data to be stored in the order in which they are generated, avoiding data loss or confusion, ensuring the integrity and orderliness of rice data, and thus improving overall processing efficiency.
[0037] Step 2, divide the preprocessed image into blocks, encode the position of each image block during the block division, generate position information matrix parameters, input the image block into the convolutional neural network model, and obtain the corresponding feature layer;
[0038] First, the image is divided into multiple rectangular blocks of fixed size 32x32 pixels to ensure that each image block contains a similar number of pixels, which is convenient for subsequent feature extraction and processing. Next, by encoding the position of each image block, the position information matrix parameters are generated, including the coordinates of the upper left corner of the block (x, y), the width and height of the block (w, h), and the relative position index of the block relative to the entire image (i, j).
[0039] Step 3: Flatten the feature layer into a linear information sequence, add a learnable embedding vector before the information sequence, and add category information flags and position information matrix parameters to the information sequence;
[0040] In step 3, we obtain the feature data of each image block and add the category label (Cls_Token) based on it. At this point, the construction of the information sequence begins with the extraction of the feature data of each image block and combines it with its corresponding position information matrix parameters. Specifically, the information sequence includes the following:
[0041] Image block feature data: feature representation of each image block obtained through feature extraction algorithm.
[0042] Category label (Cls_Token): The category information flag assigned to each image block for subsequent classification tasks.
[0043] Position information matrix parameters: including the upper left corner coordinates, width, height and relative position index of the block.
[0044] Therefore, the acquisition of information sequence is carried out in step 3, specifically after feature extraction, combining the category label and position information matrix parameters to form a complete information sequence. This process ensures that the information of each image block can be effectively encoded and utilized, providing the necessary data support for subsequent classification and analysis.
[0045] Step 4: Extract features from the linear information sequence and use the self-attention mechanism to focus on the importance of each feature layer, thereby classifying the quality rating of the rice grains.
[0046] The present invention uses a convolutional neural network to perform block processing on an image of rice grains, adds corresponding category and position information, adopts a corresponding algorithm to extract features of eight rice grain quality defects, namely, broken rice, yellow rice, chalky rice, diseased spots, mixed rice, embryo-retained rice, skin-retained rice and normal rice grains, adjusts corresponding parameters of a loss function, and trains the eight rice grain quality categories, so as to achieve efficient and accurate recognition of the quality of rice grains.
[0047] When the collected rice grain image is preprocessed in the above step 1, the following steps are mainly included.
[0048] Step 1.1, conversion process: convert the multi-angle graphics into corresponding RGB images, and convert the original sensor data captured by the camera into a color space to generate a standard red, green, and blue three-channel image.
[0049] Step 1.2, Gaussian filtering denoising: Gaussian filtering denoising is performed on the RGB image. The Gaussian filter smoothes the image through convolution operation, reduces the impact of noise, and retains the main image features.
[0050] Step 1.3, scaling the denoised image. Image scaling is performed using bilinear interpolation and bicubic interpolation methods. In the field of image processing, it is ensured that the quality of the scaled image retains the original details and features as much as possible and reduces distortion and information loss caused by scaling.
[0051] In step 1.4, we normalize the scaled image to ensure that the pixel values of the image are in the range of 0 to 1. The specific operation includes calculating the minimum and maximum values of the image pixels and performing a linear transformation on each pixel value to obtain the normalized image matrix.
[0052] After normalization, the rows and columns of the resulting image matrix have the following meanings:
[0053] Row: represents the height of the image, that is, the number of vertical pixels in the image. Each row corresponds to a row of pixels in the image.
[0054] Column: represents the width of the image, that is, the number of horizontal pixels of the image. Each column corresponds to a column of pixels in the image.
[0055] Therefore, each element of the image matrix corresponds to the normalized value of a specific pixel in the image. The normalized element values range between 0 and 1, where:
[0056] 0 means that the original value of the pixel is equal to the minimum pixel value of the image.
[0057] 1 means that the original value of the pixel is equal to the maximum pixel value of the image.
[0058] A value between 0 and 1 indicates the relative position of the original value of the pixel between the minimum and maximum values.
[0059] In this way, the normalized image matrix not only retains the spatial structure information of the image, but also standardizes the pixel values to a uniform range, which is convenient for subsequent feature extraction and processing. This standardization process helps improve the training efficiency and accuracy of the model because it eliminates the brightness and contrast differences between different images.
[0060] The resulting data after processing is an image with pixel values within the normalized range, usually represented as a matrix of the same size, where each pixel value is within the specified range. For an 8-bit grayscale image, the original range of pixel values is from 0 to 255. After normalization, the pixel values are linearly transformed to the range of 0 to 1.
[0061] For example, consider an 8x8 image matrix, after calculating the minimum and maximum values, the normalized result may be a matrix with the same size as the original image.
[0062] The matrix here is an example of a local element of an image matrix, specifically, it represents the values of certain pixels in the image after normalization. Each element corresponds to the normalized value of a specific pixel in the image, reflecting the relative brightness of the pixel in the original image. It should be noted that this example matrix does not represent all the pixels of the entire image, but only shows how the results of the normalization process are performed in a local range.
[0063] In this way, the normalized image matrix can effectively retain the spatial structure information of the image and provide standardized input for subsequent feature extraction and processing. In addition, the dimension of the training samples has also been expanded. Specifically, the device photographs rice grains from three different angles: front, side, and top, and each rice grain will produce three 500x500 pixel images. These images will serve as the main features of the training samples.
[0064] The dimensions of training samples include the following aspects:
[0065] Image height and width: The dimensions of each image are 500x500 pixels.
[0066] RGB three channels: Each pixel contains information of three channels: red, green and blue, forming a three-dimensional matrix.
[0067] Category label: used to represent the classification information of the rice grain. For example, if the rice grain is classified as "normal", the category label is a vector of length 8 [0,0,0,0,0,0,0,1], where the last element is 1 to represent the category of the rice grain.
[0068] In this case, each element of the category label represents the presence or absence of a specific category, usually represented by one-hot encoding. Specifically: each position in the vector corresponds to a specific category. A value of 1 indicates the presence of the category, and a value of 0 indicates the absence of the category.
[0069] In this way, the training samples contain not only the spatial and color information of the image, but also the classification information, providing rich input data for subsequent model training and feature extraction. The training samples are used to train the deep learning model convolutional neural network to achieve image classification, object detection, and image segmentation tasks.
[0070] In step 1, the collected image is smoothed by Gaussian filtering to reduce the noise generated by the optical module hardware. Then the collected image is losslessly scaled from the size of 2448, 2048, 3 to the size of 244, 244, 3 to standardize the input size of the subsequent image segmentation step. The image storage queue can effectively manage and organize the rice grain images taken from multiple angles. This method can ensure that the images are stored in the order of shooting, avoid data confusion and loss, and improve the stability and reliability of the system. The image storage queue can process image data in parallel, reduce processing delays, and improve the processing efficiency of the system. Through queue management, it can ensure that the multi-angle images of the same rice grain remain consistent during the processing process to avoid confusion. The image storage queue can store images in a predetermined order, which is convenient for subsequent batch processing, analysis and three-dimensional reconstruction. After storage, the images in the image storage queue will be taken out in turn for preprocessing and three-dimensional reconstruction. Preprocessing includes operations such as denoising, enhancement and standardization, which lays a good foundation for subsequent three-dimensional reconstruction and feature extraction. Through the management of the image storage queue, the order and efficiency of image processing can be ensured. The classification module does not randomly block the images. Image segmentation is usually performed according to fixed size and rules, such as dividing the image into fixed-size rectangular blocks of 32x32 pixels. This regularized segmentation method ensures that each image block contains a similar number of pixels, which is convenient for subsequent feature extraction and processing. The position information matrix parameters are generated by encoding the position of the image block. The specific parameters usually include the position of the image block in the original image and the relative position relative to the entire image. These parameters can be obtained in the following ways: Absolute position encoding: directly use the coordinates of the image block in the image. Relative position encoding: use the offset of the image block relative to the center of the image or other reference points. Specific parameters may include: the coordinates of the upper left corner of the block (x, y), the width and height of the block (w, h), the relative position index of the block (i, j), indicating the position of the block in the entire image grid. Information sequence: The information sequence refers to the collection of image data and its position information after segmentation and encoding. In Transformer Encoder, each image block includes its features and position information and is regarded as a sequence element. These elements are sequentially input into Transformer Encoder for feature extraction. Specifically, the information sequence includes the following: Feature data of the image block: the pixel value of each image block or the feature representation after preprocessing. Category label (Cls_Token): information used to identify the category of the image block. Position information matrix parameters: the position information of the image block in the original image, such as coordinates and relative position.
[0071] In the above step 3, the preprocessed image data is divided into blocks, and category labels and position information matrix parameters are added. In a specific embodiment, the corresponding convolution kernel is used to convolve the preprocessed image during the block division to obtain the corresponding feature layer. For example, the preprocessed image is 244 pixels * 244 pixels * 3 channels, the size of the convolution kernel is 16 * 16, and the feature layer obtained by the block division is 14 * 14 * 768, that is, the feature layer is arranged in an array of 14 rows and 14 columns, and the number of layers is 768. On this basis, the obtained feature layer is tiled in height and width dimensions to obtain a 196,768 feature layer. After the tiling is completed, a learnable embedding vector Class_Token is added to the linear image sequence. The learnable embedding vector Class_Token is equivalent to introducing global context information in the entire Patch, so that the entire model considers the characteristics of the entire image and the local area during the training and reasoning process, thereby improving the model's understanding and classification capabilities of the overall semantics of the image.
[0072] After adding the learnable embedding vector Class_Token to the linear information sequence, a 197,768 feature layer is obtained. After adding the embedding vector Class_Token, the position information is added to all feature layers, that is, a 197,768 parameter matrix is generated. This parameter matrix can be trained together with the neural network, so that the entire neural network can distinguish different regions.
[0073] In the present invention, the steps of dividing the preprocessed image data into blocks and adding category labels and position information matrix parameters include: dividing the image into blocks to obtain the corresponding feature layer. In a specific embodiment, the corresponding convolution kernel is used to convolve the preprocessed image during the block division to obtain the corresponding feature layer. For example, the size of the preprocessed image is 244 pixels × 244 pixels × 3 channels, the size of the convolution kernel is 16 × 16, and the feature layer obtained by the block division is 14 × 14 × 768, that is, the feature layer is arranged in an array of 14 rows and 14 columns, and the number of layers is 768. The feature layer is tiled and the embedded vector Class_Token is added. The obtained feature layer is tiled in the height and width dimensions to obtain a 196 × 768 feature layer. After the tiling is completed, the learnable embedded vector Class_Token is added to the linear image sequence. The learnable embedded vector Class_Token is equivalent to introducing global context information in the entire Patch, so that the entire model considers the features of the entire image and the local area at the same time during the training and reasoning process, thereby improving the model's understanding and classification capabilities of the overall semantics of the image. After adding the learnable embedding vector Class_Token to the linear sequence, a 197×768 feature layer is obtained.
[0074] Add position information to the feature layer. After adding the embedding vector Class_Token, add position information to all feature layers to generate a 197×768 parameter matrix. The position information is obtained by calculating the row and column index of the image block in the original image. Specifically, a unique position information code is assigned to each image block, and sine and cosine functions are usually used to generate continuous position codes so that the model can better capture spatial relationships.
[0075] The position information matrix parameters are obtained by calculating the position of the image block in the original image. For each image block, its row and column index in the original image can be used to generate the position information. The specific method is to assign a unique position information code to each image block, which can be determined by the row and column position of the image block in the original image. Sine and cosine functions are usually used to generate continuous position codes so that the model can better capture spatial relationships. For example, if the image is divided into 14x14 blocks, the position information of each block can be determined by its index in the 14x14 grid.
[0076] The position information matrix parameters usually include two main aspects of data:
[0077] Row and column position information: row and column index of each image block.
[0078] Position encoding vector: For ease of neural network processing, row and column indices are usually converted into position encoding vectors. Commonly used methods are sine and cosine position encodings, which convert the index into a vector of fixed dimension and are periodic and continuous. For an image with a block size of 14x14, each image block has a position encoding, and the dimensions of these encoding vectors are consistent with the dimensions of the feature layer.
[0079] The specific parameters usually include the position of the image block in the original image (such as the starting coordinates and size of the block), and the relative position relative to the entire image. These parameters can be obtained in the following ways.
[0080] Absolute position encoding: directly use the coordinates of the image block in the image.
[0081] Relative position encoding: Uses the offset of the image patch relative to the image center or other reference point.
[0082] Specific parameters may include: the coordinates of the upper left corner of the block (x, y), the width and height of the block (w, h), and the relative position index of the block (i, j), which indicates the position of the block in the entire image grid.
[0083] The information sequence refers to the collection of image data and its position information after being divided and encoded. Each image block (including its features and position information) is regarded as a sequence element. These elements are sequentially input into TRANSFORMER ENCOERDE for feature extraction. Specifically, the information sequence includes the following: feature data of the image block, category label (Cls_Token), and position information matrix parameters. The feature data of the image block is the pixel value of each image block or the feature representation after preprocessing. The category label (Cls_Token) is used to identify the information of the image block category. The position information matrix parameters are the position information of the image block in the original image, such as coordinates and relative position.
[0084] In the above step 4, the steps of extracting features from the feature layer sequence obtained in step 3 using the Multi-Head Attention structure of the Vit algorithm include:
[0085] Step 4.1: Split into N Heads for Self-Attention processing. The input feature layer sequence matrix is split into N heads. The calculation of each head includes generating a query vector matrix (Q), a key vector matrix (K), a value vector matrix (V), and the corresponding weight matrix (W), and calculating the attention mechanism to capture the relationship between different features.
[0086] Step 4.2: Multiple Head splicing and feature extraction, splice the outputs of N heads together and perform linear transformation again to obtain a higher-level feature representation. This process allows the model to take into account the feature information of the entire image and the local area at the same time, thereby improving the model's understanding and classification capabilities of the overall semantics of the image.
[0087] Step 4.3: Fully connected association and feature length adjustment, fully connected association is performed on the feature layer sequence after the multi-head attention mechanism, and the length of the feature vector is adjusted, for example, dk=64 is set so that the final output has a suitable feature dimension and representation capability. The Multi-Head Attention structure of the Vit algorithm is used to extract features from the linear sequence. That is, the information sequence obtained in step 2 is passed into the neural network model Transformer Encoder to extract the rice grain features.
[0088] Specifically, the image sequence is divided into N heads, and the feature layer sequence obtained after the preprocessed image data is divided into blocks and feature extraction. Each feature layer represents the characteristics of a local area of the original image. The N heads are divided into N heads by dividing the feature vector into N parts by row. Specifically, assuming that the dimension of the feature vector is d, the dimension of each head is d / N. The role of the N heads is to learn different attention concentration mechanisms in parallel, thereby enhancing the model's ability to capture different features. Each head does not correspond to a specific part of the original image, but to the division and parallel processing of the feature vector.
[0089] Perform Self-Attention processing; concatenate the extracted features; X: Patch+PositionEmbedding sequence matrix, Q: query vector matrix, K: key vector matrix, V: value vector matrix, W: weight matrix corresponding to Q, K, V. Q: query vector matrix, K: key vector matrix, V: value vector matrix and W: weight matrix: These matrices are trained during the learning process of the model. They are used in the Multi-Head Attention structure to calculate the attention score and generate the output of each head. The Q, K, and V matrices are obtained by multiplying the input vector (feature layer sequence matrix) with the corresponding learned weight matrix. Based on the input vector (197, 64), the input vector refers to the feature layer sequence matrix, which has a shape of (197, 768), where 197 is the sequence length and 768 is the dimension of the feature vector. These feature vectors are the linear information sequences obtained in step 2, including the feature information extracted from each image block. Generate the corresponding Q, K, V vector matrices. The input vector refers to the feature layer sequence matrix, whose shape is (197,768), where 197 is the sequence length and 768 is the dimension of the feature vector. These feature vectors are the linear sequences obtained in step 2, including the feature information extracted from each image block. , and perform full connection association. In the Multi-Head Attention structure, the Q, K, and V vector matrices of each head first calculate the attention score, and then obtain the output of each head by weighted summation. The outputs of multiple heads will be spliced together, and then undergo a linear transformation again to generate the final feature representation, dk = 64, which is the length of the feature;
[0090] The query vector Q obtained above is cross-multiplied by the transposed key vector K, divided by the square of dk, that is, 8, and the result is processed by the Softmax function. Then the result is cross-multiplied by the value vector matrix V, and the result is the current Self-Attention feature extraction result.
[0091] The generated image sequence (197, 768) with the embedded vector Class_Token is divided into 12 heads and passed into the Multi-Head Attention module of the Vit algorithm. The feature extraction (Self-Attention) is performed on each head. After the feature extraction is completed, all heads are spliced to obtain the feature output of the image. The output matrix size after segmentation is (197, 12, 64). 197: represents the length of the sequence, that is, the number of each position point or image block in the feature layer sequence. 12: represents the number of heads divided, and each head processes an independent subset of features. 64: represents the dimension of the feature vector in each head, that is, the feature length of each head. This three-dimensional matrix describes the feature representation after processing by the Multi-Head Attention structure. In the Self-Attention processing of each head, the corresponding attention score is generated by processing the query vector (Q) and the key vector (K), and the weight is obtained using the Softmax function. Finally, the weight is applied to the value vector (V) to form the output feature of each head. These output features are independent within each head, but are ultimately concatenated to form the feature output of the entire image, so the size of each Self-Attention feature is (197, 64).
[0092] In step 4 above, in order to address the problem of differences in the proportion of rice grain quality data, the parameters of the loss function in model training are adjusted to increase or decrease the influence of the corresponding rice grain category.
[0093] In one embodiment, before feature extraction, layer normalization is performed to convert the input linear sequence data into data with a mean of 0 and a variance of 1.
[0094] In a further embodiment, after the feature extraction is completed, the input layer parameters and the feature extraction layer parameters are added. When the first residual network processing is performed, the input layer parameters X and the feature layer parameters F obtained by feature extraction are mainly data in the form of tensors or matrices. The specific addition process is as follows: Dimension matching: ensure that the dimensions of X and F are the same so that element-by-element addition operations can be performed. Addition operation: add the corresponding elements at each position, that is, X+F. The addition here is element-by-element, also known as element-by-element addition. This element-by-element addition operation can effectively combine the input layer parameters X and the feature extraction layer parameters F to form the data of the first residual network. The input layer parameters and the feature extraction layer parameters are added to obtain the data of the first residual network. Then, the data of the first residual network is layer normalized, and the layer normalized data is passed into the multilayer perceptron. The data passed through the multilayer perceptron is added to the data of the first residual network to obtain the data of the second residual network. The data of the second residual network and the learnable embedding vector are passed into the multilayer perceptron to complete the classification of the quality rating of the rice grains.
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for grading and classifying the quality of rice grains, characterized in that: include: Acquire multi-angle images of rice grains and perform preprocessing; The preprocessed image is divided into blocks, and the position of each image block is encoded during the block division to generate the position information matrix parameters, and the image blocks are input into the convolutional neural network model to obtain the corresponding feature layer; Flatten the feature layer into a linear information sequence, add a learnable embedding vector before the information sequence, and add category information flags and position information matrix parameters to the information sequence; Feature extraction is performed on the linear information sequence, and the importance of each feature layer is paid attention to through the self-attention mechanism, so as to classify the quality rating of rice grains.
2. The method for grading and classifying rice grain quality according to claim 1, wherein: Preprocessing includes: Convert each multi-angle graphic into a corresponding RGB image; Perform Gaussian filtering to denoise the RGB image; Scale the denoised image; Normalize the scaled images and add the dimensions of the training samples.
3. The method for grading and classifying rice grain quality according to claim 1, wherein: The position information matrix parameters include the row index and column index of the image block in the image, and the position encoding vector.
4. The method for grading and classifying rice grain quality according to claim 3, wherein: The preprocessed image is 244 pixels*244 pixels*3 channels, the size of the convolution kernel is 16*16, and the feature layer obtained by block division is 14*14*768.
5. The method for grading and classifying rice grain quality according to claim 1, wherein: Multi- The Head Attention structure extracts features from linear information sequences.
6. The method for grading and classifying rice grain quality according to claim 5, wherein: Before feature extraction, layer normalization is performed to convert the input linear information sequence data into data with a mean of 0 and a variance of 1.
7. The method for grading and classifying rice grain quality according to claim 5, wherein: Feature extraction of linear information sequences includes: The input information sequence is divided into N heads. The calculation of each head includes generating the query vector matrix (Q), key vector matrix (K), value vector matrix (V), and the corresponding weight matrix (W), and calculating the attention mechanism to capture the relationship between different features; The results of N head calculations are stitched together to obtain the feature output of the image corresponding to the rice grain.
8. The method for grading and classifying rice grain quality according to claim 6, wherein: After feature extraction is completed, the input layer parameters and feature extraction layer parameters are added to obtain the data of the first residual network; Perform layer normalization on the data of the first residual network, and pass the layer normalized data into the multi-layer perceptron; The data passed through the multi-layer perceptron is added to the data of the first residual network to obtain the data of the second residual network; the data of the second residual network and the learnable embedding vector are passed into the multi-layer perceptron to complete the classification of the quality rating of the rice grains.