Wheat grain protein spatial distribution difference identification method and device and storage medium

By improving the U-Net model and K-means++ algorithm, combined with dynamic stratification of radial corrosion, fully automated identification of the spatial distribution of wheat grain proteins is achieved, solving the problem of time-consuming, labor-intensive and insufficient accuracy of methods in the existing technology, and improving the fineness and applicability of the identification.

CN120107966AActive Publication Date: 2025-06-06SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY +1

Patent Information

Application Number
CN202510579735.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In the prior art, when studying the spatial distribution of wheat grain proteins, the method is time-consuming and labor-intensive, insufficient accuracy, and cannot achieve high-throughput research.

Method used

By improving semantic segmentation of U-Net model, dynamic stratification of radial corrosion and component clustering of K-means++ algorithm, fully automated from image preprocessing to protein quantitative analysis is achieved.

Benefits of technology

It reduces the tedious operation of manual depiction, improves the fineness and accuracy of image processing, can accurately reflect the protein distribution gradient, and adapt to different grain morphology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107966A_ABST
    Figure CN120107966A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat grain protein spatial distribution difference identification method. The method comprises the following steps: S1, constructing a semantic segmentation data set; an improved U-Net network is adopted as a model architecture, and preliminary background removal is completed; s2, binarization processing is carried out, and preliminary segmentation is completed; obtaining an endosperm edge contour by using a Sobel edge detection algorithm; determining the center position of an endosperm area; carrying out layer-by-layer corrosion to obtain a radial layered structure image; s3, acquiring a sampling matrix; a K-means + + algorithm is adopted to initialize the clustering center for iteration; and identifying the protein and calculating the space distribution proportion of the protein. According to the method, semantic segmentation of the U-Net model, dynamic layering of radial corrosion and component clustering of the K-means + + algorithm are improved, full automation from image preprocessing to protein quantitative analysis is achieved, tedious operation of manual description is reduced, image processing is more accurate, layering processing can be conducted on the image, and the protein distribution gradient can be accurately reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protein identification, and in particular to a method, a device and a storage medium for identifying spatial distribution differences of wheat grain protein. Background Art

[0002] Achieving high-yield and high-quality wheat production goals simultaneously is an important foundation for meeting the dual needs of food security and improving the quality of life. Different quality special wheats have specific requirements for protein content. For example, strong gluten wheat for bread making requires high protein content, while weak gluten wheat for biscuits requires low protein content.

[0003] However, the quality compliance rate of high-quality wheat in my country is low, which is difficult to meet the demand, and a large amount of it needs to be imported from abroad. In the process of wheat flour making, about 30% of the outer layer of the grain is removed during the grinding process, and the protein content in different parts of the grain varies significantly, and responds differently to cultivation measures such as nitrogen fertilization.

[0004] Therefore, a new idea is provided for the production of special wheat flour. For medium-strong gluten wheat, the selection of varieties whose outer layer protein is not sensitive to nitrogen application can increase the flour protein content and nitrogen fertilizer utilization efficiency; for weak gluten wheat, the selection of varieties whose inner layer protein is not sensitive to nitrogen application can stabilize or reduce the inner endosperm protein content while ensuring yield, thus achieving high yield and quality improvement. Studying the spatial distribution of wheat grain protein is of great significance for improving the milling process, producing various types of special flour, and meeting the diverse needs of food processing.

[0005] However, the current mainstream method for studying the spatial distribution of grain protein is to combine layered grinding with the Kjeldahl method. The layered grinding process is time-consuming and labor-intensive, and it is difficult to grasp the grinding time of each layer. In addition, the Kjeldahl method has many steps, which can easily cause human errors and lead to poor data repeatability. In recent years, sectioning combined with image analysis has become another option for studying the spatial distribution of grain protein, but there is a lack of image analysis software. The protein area is mainly identified by image J, and there are also studies using ARCGIS to identify differences in protein spatial distribution. The common problem of the above image analysis studies is that the outline of the wheat grain endosperm needs to be manually depicted, which is time-consuming and labor-intensive and cannot be applied to high-throughput research. Moreover, when applied to layered research, the number of layers is fixed and cannot be flexible, and the layers are prone to distortion inward. The RGB value is often used to identify the protein staining area, which is not accurate enough.

[0006] In other words, the measures currently taken include at least the following three problems: 1) The mainstream method is to combine layered grinding with Kjeldahl nitrogen determination, which is time-consuming, labor-intensive and inaccurate; 2) When combined with image analysis, the corresponding contours need to be manually drawn, which is time-consuming and labor-intensive and cannot be used for high-throughput research; 3) Lack of flexibility, difficult to process during stratification, and insufficiently accurate recognition results. Summary of the invention

[0007] In response to the above three problems, the purpose of the present invention is to propose a method, device and storage medium for identifying the spatial distribution differences of wheat grain protein. By improving the semantic segmentation of the U-Net model, the dynamic stratification of radial erosion and the component clustering of the K-means++ algorithm, full automation from image preprocessing to protein quantitative analysis is achieved, the tedious operations of manual drawing are reduced, the image processing is more refined and accurate, and the image can be layered to accurately reflect the protein distribution gradient.

[0008] This is achieved through the following technical solutions: Firstly, a method for identifying the spatial distribution difference of wheat grain protein is proposed, which includes the following steps: S1. Prepare multiple cross-sections of mature wheat grains and stain them with resin slices. Use image annotation tools to construct a semantic segmentation dataset with three types of annotations, including endosperm, background, and cavity. After data enhancement, divide them into training set and validation set at a fixed ratio. Use the improved U-Net network as the architecture of the model, embed the channel attention mechanism in the encoder and decoder modules, design a hybrid loss function including cross entropy loss and Dice loss, and then use the AdamW optimizer combined with the cosine annealing learning rate decay strategy for training. After the training is completed, reset the image size of each slice and input it into the trained model after standardization to complete the preliminary background removal and obtain the final foreground mask image. S2, based on the grayscale histogram feature of the blue channel, select a feature threshold for the final foreground mask image in step S1 to perform image binarization processing, first use the im2bw function to assign 0 to the area with a pixel value less than the feature threshold in the binary image, and assign 1 to the area with a pixel value greater than or equal to the feature threshold, then perform denoising through morphological opening operation, complete the preliminary segmentation of the background and grain area, and obtain a preliminary segmented image; continue to use the Sobel edge detection algorithm to detect the preliminary segmented image, and obtain the continuous closed endosperm edge contour of the preliminary segmented image; then use the bwboundaries function to extract the contour coordinates, and determine the center position of the endosperm area based on the contour coordinates; then calculate the Euclidean distance from all edge points on the endosperm edge contour to the corresponding center position, take the maximum value Dmax as the reference radius of radial stratification, divide the radius range into S layers, and erode the preliminary segmented image layer by layer in turn by cyclically calling the imerode function to obtain a radial layered structure image; S3. Use the imread function to load the final foreground mask image in step S1, use the reshape function to flatten and downsample the morphological opening process in step S2, and obtain the sampling matrix Xsample for calculating the initial cluster center; then set the number of clusters K, use the K-means++ algorithm to initialize the cluster center, and then iterate; after the iteration is completed, identify the protein based on the staining spectral characteristics and calculate the protein spatial distribution ratio.

[0009] Preferably, in step S1, there are at least 300 resin slices of cross-sections of wheat grains at maturity, the ratio of training set to validation set is 8:2, and the weight α of the mixed loss function is 0.7; before the encoder and decoder modules are embedded in the channel attention mechanism, the ResNet34 residual block is first introduced in the encoder for multi-scale feature fusion. Multiple slice data can ensure that the morphological diversity of wheat grains is covered as much as possible, and the multi-scale feature fusion facilitates the subsequent channel attention mechanism to improve the recognition accuracy.

[0010] Preferably, before generating the final foreground mask image in step S1, the prediction mask output by the model needs to be subjected to morphological closing operations to connect broken areas, remove isolated noise areas with an area smaller than a fixed pixel, and perform binarization processing in sequence. Morphological closing operations can fill in the broken edges of the endosperm, such as holes caused by uneven staining of slices, and then perform denoising and binarization processing, effectively improving the continuity of the boundary and the accuracy of subsequent recognition.

[0011] Preferably, in step S2, S=9, and the corrosion radius of each layer is r w =S w Dmax,w=1, 2, 3...S, S w is the scale factor of the wth layer; when calling the imerode function to erode each layer layer by layer, the wth layer erosion operation uses a radius of r w The circular structural element is used to generate the annular region mask of the wth layer, and after completing the layer-by-layer erosion, a radial hierarchical structure image from the edge contour of the endosperm to the center position is formed. S=9, divided into 9 annular regions, can adapt to different grain morphologies and facilitate the viewing of protein distribution in different layers.

[0012] Preferably, when the Sobel edge detection algorithm is used for detection in step S2, each preliminary segmented image is first grayed, and the corresponding gradient amplitudes are calculated by Sobel operators in the horizontal and vertical directions, and then the continuous closed endosperm edge contours of each preliminary segmented image are obtained after non-maximum suppression and hysteresis threshold processing. The Sobel operator can highlight the gradient difference between the endosperm area and the background, thereby obtaining a clear endosperm edge contour.

[0013] Preferably, the morphological opening operation in step S2 uses the imopen function of a 3×3 rectangular structure element.

[0014] Preferably, when flattening and downsampling are performed in step S3, the 3×3 rectangular structure is first converted into a MN×3 two-dimensional data matrix through the reshape function, and each row in the two-dimensional data matrix corresponds to the RGB value of a pixel; then the final mask foreground image is downsampled using an alternate row and alternate column sampling method to generate a sampling matrix Xsample. Flattening and downsampling can improve computing efficiency and achieve rapid processing of large-scale images.

[0015] Preferably, after obtaining the sampling matrix Xsample in step S3, the number of clusters K=5 is set, where 0 corresponds to background, 1 corresponds to cavity, 2 corresponds to protein, 3 corresponds to starch, and 4 corresponds to other components; when the K-means++ algorithm is used to initialize the cluster center, the Start parameter in the kmeans function is set to 'plusplus', and the first center is randomly selected from the sampling matrix Xsample, and each subsequent center is selected according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected, and then the similarity between the pixel and the cluster center is measured based on the Euclidean distance, and the allocation stage is first iteratively executed, including calculating the Euclidean distance between each pixel and the 5 centers , x j and a cluster center u K The Euclidean distance between them, c = 1, 2, 3, j is a positive integer, is assigned to the category with the smallest distance, and then the iterative update phase is performed, including recalculating the RGB mean of each type of pixel as the new center, until the maximum number of iterations is 100 or the center change is less than 10 -5 K=5 means that the components are classified into 5 categories, and the K-means++ algorithm is used to disperse the initial centers, thereby avoiding the local optimum of random initialization while taking into account the classification accuracy, so as to ensure the accuracy of overall recognition.

[0016] Secondly, a device is also proposed, including a processor and a memory, the memory storing a computer program, and the processor is configured to execute a method for identifying spatial distribution differences of wheat grain protein as described above when running the computer program.

[0017] In addition, a storage medium is also proposed, in which a computer program is stored. When the computer program runs in a processor, the processor is used to execute the above-mentioned method for identifying spatial distribution differences of wheat grain protein.

[0018] Compared with the prior art, the present invention has the following beneficial effects: The technical solution of the present invention realizes full automation from image preprocessing to protein quantitative analysis by improving the semantic segmentation of the U-Net model, the dynamic stratification of radial erosion and the component clustering of the K-means++ algorithm, reduces the tedious operation of manual drawing, and processes the image more finely and accurately; it also utilizes multi-scale feature fusion and embedded channel attention mechanism to greatly improve the endosperm segmentation accuracy and reduce the boundary segmentation error; and it also performs dynamic radial stratification to solve the traditional stratification problem, thereby adapting to different grain morphologies, and thus having a wide range of applicability; in addition, the image can also be processed in layers to accurately reflect the protein distribution gradient. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of a method for identifying differences in spatial distribution of wheat grain protein; Figure 2 This is a comparison chart of the effects of images in each link after an experiment using a method for identifying differences in spatial distribution of wheat grain protein; Figure 3 This is a schematic diagram of the spatial distribution difference stratification and clustering effect of wheat grain protein. DETAILED DESCRIPTION

[0020] The following will be combined with the attached embodiment of the present invention Figures 1 to 3 , the technical solutions in the embodiments of the present invention are described in detail.

[0021] like Figure 1 The figure shows a flowchart of a method for identifying differences in spatial distribution of wheat grain protein. By removing background interference, dynamic spatial stratification of radial erosion, and component clustering of the K-means++ algorithm, full automation from image preprocessing to protein quantitative analysis is achieved, reducing the tedious operation of manual drawing, greatly improving the accuracy of identification, and solving the stratification problem in identification. The specific steps are as follows: S1. Prepare multiple cross-section resin slices of mature wheat grains and stain them. The number of slices should be at least 300 to ensure that the morphological diversity of wheat grains is covered. Use image annotation tools, such as the Labelme tool, to construct a semantic segmentation dataset with three types of annotations, including endosperm, background, and cavity. After data enhancement, it is divided into training set and validation set at a fixed ratio of 8:2. The 8:2 ratio distribution only uses a small proportion of the dataset for testing and verification, which ensures the accuracy of the training. The data enhancement here uses conventional enhancement methods, such as geometric transformations such as rotation, flipping, scaling, and cropping, and color transformations such as adjusting brightness and hue, in order to expand the database.

[0022] Then, the improved U-Net network is used as the architecture of the model, that is, the improved U-Net model. The ResNet34 residual block is introduced in the encoder corresponding to the model for multi-scale fusion, so as to extract the features of the image at different scales under different convolution kernel sizes and different pooling levels, so as to capture more image detail features and improve clarity. The channel attention mechanism is embedded in the decoder module corresponding to the encoder and the model, and a mixed loss function including cross entropy loss and Dice loss is designed, and the weight of the two is α=0.7, that is, the mixed damage function = 0.7×cross entropy loss + (1-0.7)×Dice loss, to improve the category distinction and boundary segmentation accuracy. In addition, the weight α defaults to 0.7, but it can also be adjusted dynamically to adapt to various different situations of actual wheat grains. Then, the AdamW optimizer is used in combination with the cosine annealing learning rate decay strategy for training, and the average intersection-over-union ratio and other indicators are used for evaluation, which can trigger the early stopping mechanism to avoid overfitting.

[0023] It should be noted that the Labelme tool is an image annotation tool suitable for target detection and semantic segmentation tasks. The improved U-Net model is a convolutional neural network with an encoder-decoder structure, which is suitable for image segmentation. The ResNet34 residual block contains 34 layers, and the model's ability to extract features of different sizes is improved through jump connections. The channel attention mechanism can dynamically allocate the weight α of the channel to meet the requirements of different categories of distinction and boundary segmentation accuracy. The AdamW optimizer is an improved version of the Adam optimizer, which can enhance generalization ability and prevent overfitting. The cosine annealing learning rate decay strategy, in which the learning rate decays from the initial value to the minimum value according to the cosine function curve, and restarts periodically to jump out of the local optimum. The average intersection-over-union ratio is an important indicator for evaluating the segmentation of the model, and can be used to feedback the accuracy of boundary segmentation.

[0024] After the training is completed, the image size of each slice is reset, the image is adjusted to 512×512 pixels and input into the trained model after normalization. Then, the prediction mask output by the model after training is subjected to morphological closing operations to connect the broken areas, remove the isolated noise areas with an area smaller than the fixed pixels, and perform binarization in sequence to complete the preliminary background removal and obtain the preprocessed image, which is also the final foreground mask image.

[0025] In this embodiment, the morphological closing operation can fill the endosperm edge breaks, such as the holes caused by uneven section staining, so that the broken position becomes continuous. For example, a disk with a radius of 3 pixels can be used for closing operations, and the broken areas can be connected through expansion operations. Isolated noise areas with an area smaller than a fixed pixel can be removed. For example, isolated noise areas with an area smaller than 100 pixels can be removed, thereby reducing certain interference and improving the background removal rate. The binarization process here is used to identify the outer contour of the endosperm edge, thereby removing background interference.

[0026] like Figure 2 The figure shows a comparison of the effects of images in each link after the experiment using the wheat grain protein spatial distribution difference recognition method, which is used to verify the experimental effect. Among them, Figure a is a mosaic image of the endosperm stained with protein, which is for the image of the data set constructed by the image annotation tool; Figure b is an image of the endosperm outer contour and cavity identified and marked, which is also an image of the data set constructed by the image annotation tool; Figure c is an image after the background interference is accurately removed, that is, the final foreground mask image; Figure d is an edge detection and centroid positioning image, that is, the image after the subsequent completion of Sobel edge detection and determination of the center position of the endosperm area; Figure e is an image divided into any layers, that is, the subsequent radial hierarchical structure image; Figure f is an image of protein recognition by K-means clustering, that is, the final image after initialization and clustering by the K-means++ algorithm.

[0027] As shown in Figure c, the final foreground mask image in step S1 has clear boundary segmentation, a background removal rate of up to 98.7%, an endosperm boundary positioning error of less than 3 pixels, and a single image processing time of 120ms, providing a high-quality data foundation for the subsequent calculation of protein spatial distribution characteristics.

[0028] S2, based on the blue channel, that is, the B channel of the blue component in the RGB image, based on the grayscale histogram characteristics of the channel, the feature threshold is selected for the final foreground mask image in step S1 to perform image binarization again to obtain a binary image. The image binarization processing in step S2 is the same as the binarization processing method in step S1. Both use the im2bw function to assign 0 to the area in the binary image where the pixel value is less than the feature threshold, corresponding to the background area with a lower blue component in RGB, 0 corresponds to black, and then assign 1 to the area where the pixel value is greater than or equal to the feature threshold, corresponding to the main part of the grain, 1 corresponds to white; then denoise is performed through morphological opening operation, and the preliminary segmentation of the background and grain area is completed to obtain a preliminary segmented image. It should be emphasized that the first binarization processing in step S1 is to remove background interference after identifying the outer contour of the endosperm, targeting the endosperm and the background; while the image binarization processing in step S2, that is, the second binarization processing, is to identify the outer contour of the endosperm for subsequent stratification, and at the same time, targeting the grain and the background.

[0029] In this embodiment, the morphological opening operation may use the imopen function of the 3×3 rectangular structure element. The imopen function is a function used to implement the morphological opening operation in image processing, and is suitable for removing tiny objects in an image and performing noise reduction.

[0030] Then, the Sobel edge detection algorithm is used to detect the preliminary segmented image to obtain the continuous closed endosperm edge contour of the preliminary segmented image. When using the Sobel edge detection algorithm for detection, each preliminary segmented image is first grayed, and the corresponding gradient amplitude is calculated by the Sobel operator in the horizontal direction [-101; -202; -101] and the vertical direction [-1-2-1; 000; 121], and then the continuous closed endosperm edge contour of each preliminary segmented image is obtained after non-maximum suppression and hysteresis threshold processing. The Sobel operator can highlight the gradient difference between the endosperm area and the background, thereby obtaining a clear endosperm edge contour.

[0031] Among them, non-maximum suppression is to check the horizontal and vertical gradient directions for each pixel point, and only retain the pixels with the largest amplitude in the corresponding direction, and suppress the rest to 0. Hysteresis threshold processing is a dynamic threshold method, which is used to distinguish strong edges, weak edges and noise by setting two high and low thresholds, such as a high and low threshold ratio of 2:1; strong edges correspond to clear real edges and are retained; noise is directly discarded; weak edges correspond to between noise and strong edges, and need to be rechecked to see if they are connected to strong edges. If they are connected, they are retained, and if not, they are discarded.

[0032] Furthermore, the bwboundaries function is used to extract the contour coordinates, and the center position of the endosperm area is determined based on the contour coordinates, also called the centroid or geometric center. The calculation formula is: Cx = N ∑ xi , Cy = N ∑ yi ,in N Next, calculate the Euclidean distance from all edge points on the endosperm edge contour to the corresponding center position (Cx, Cy), take the maximum value Dmax as the reference radius of radial stratification, divide the radius range into S layers, corresponding to different radial regions of protein distribution, and erode the preliminary segmented image layer by layer by cyclically calling the imerode function to obtain the overall radial stratified structure image. Figure 2 Figure e in .

[0033] In this embodiment, S can be taken as 9, and the corrosion radius of each layer is r w =S w Dmax,w=1, 2, 3...S, S w is the proportional coefficient of the wth layer, that is, the proportion of the wth layer in the radial direction to the entire radial direction; when calling the imerode function to erode each layer layer by layer, the wth layer erosion operation uses a radius of r w The circular structural element is used to generate the annular region mask of the wth layer, which is the current layer mask minus the previous layer mask. Then, after completing the layer-by-layer erosion, a radial hierarchical structure image from the edge contour of the endosperm to the center position is formed. S=9, divided into 9 annular regions, can adapt to different grain morphologies and facilitate the viewing of protein distribution in different layers.

[0034] It should be noted that the Sobel edge detection algorithm is an edge detection algorithm based on gradient calculation. It calculates the gradient amplitude of the image in the horizontal and vertical directions through the convolution kernel to highlight the edge information. The bwboundaries function can extract the contour coordinates of the target object from the binary image and supports multiple connected area detection. The imerode function is a function of morphological corrosion operation. It shrinks the foreground area of ​​the image through structural elements such as circles and rectangles for segmentation, denoising or boundary extraction.

[0035] S3, performing K-means clustering to identify proteins, including: using the imread function to load the final foreground mask image in step S1, and using the reshape function to flatten and downsample the morphological opening operation process in step S2, for example, first converting the 3×3 rectangular structure into a MN×3 two-dimensional data matrix through the reshape function, each row in the two-dimensional data matrix corresponds to the RGB value of a pixel, and the RGB value range is 0-255; then downsampling the final mask foreground image, for example, a large data volume image of about 262144 pixels, using the alternate row and alternate column sampling method, downsampling the 512×512 pixel image to a 256×256 pixel image, generating a sampling matrix Xsample for subsequent initial cluster center calculation. Flattening and downsampling can improve computing efficiency and achieve rapid processing of large-scale images.

[0036] Then, the number of clusters K is set, and the K-means++ algorithm is used to initialize the cluster centers, and then iterates. After the iteration is completed, proteins are identified based on the staining spectral characteristics and the spatial distribution ratio of proteins is calculated.

[0037] In this embodiment, after obtaining the sampling matrix Xsample, the K-means++ algorithm is used to initialize the cluster center, and then iterate. For example, the number of clusters K=5 is set, where 0 corresponds to the background, 1 corresponds to the cavity, 2 corresponds to protein, 3 corresponds to starch, and 4 corresponds to other components; when the K-means++ algorithm is used to initialize the cluster center, the Start parameter in the kmeans function is set to 'plusplus', and the first center is randomly selected from the sampling matrix Xsample. Each subsequent center is selected according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected, and then the similarity between the pixel and the cluster center is measured based on the Euclidean distance. The allocation stage is first iteratively executed, including calculating the Euclidean distance between each pixel and the 5 centers. , which represents a pixel x j and a cluster center u K The Euclidean distance between them, c = 1, 2, 3, j is a positive integer, is assigned to the category with the smallest distance, and then the iterative update phase is performed, including recalculating the RGB mean of each type of pixel as the new center, until the maximum number of iterations is 100 or the center change is less than 10 -5 K=5 means that the components are classified into 5 categories, and the K-means++ algorithm is used to disperse the initial centers, thereby avoiding the local optimum of random initialization while taking into account the classification accuracy, so as to ensure the accuracy of overall recognition.

[0038] like Figure 3The figure shows a schematic diagram of the stratification and clustering effect of the spatial distribution difference of wheat grain protein. After the clustering is completed, the label mapping is optimized in combination with the characteristics of the stained image: the category with RGB value close to pure black, that is, R / G / B<50, is judged as the background, that is, category 0; the category with high gray value and prominent blue channel, that is, B>200, is judged as cavity, that is, category 1; based on the spectral characteristics of protein dyes, for example, Coomassie Brilliant Blue is used to stain RGB values ​​of R=150-200, G=100-150, B=180-220 to mark proteins, that is, category 2; the remaining two categories are divided into starch and other components according to the grayscale difference, that is, category 3 and category 4.

[0039] Finally, the predicted mask of the endosperm during the previous segmentation of the image is used to filter the background of category 0 and the cavity area of ​​category 1, and the number of pixels of protein, starch and other components in the effective area of ​​the endosperm is counted. N protein , N starch , N other , through the formula P protein= N protein / ( N starch+ N protein+ N other)×100% to calculate the relative protein content.

[0040] Secondly, a device is also proposed, including a processor and a memory, the memory storing a computer program, and the processor is configured to execute a method for identifying spatial distribution differences of wheat grain protein as described above when running the computer program.

[0041] In addition, a storage medium is also proposed, in which a computer program is stored. When the computer program runs in a processor, the processor is used to execute the above-mentioned method for identifying spatial distribution differences of wheat grain protein.

[0042] In summary, the present invention realizes full automation from image preprocessing to protein quantitative analysis by improving the semantic segmentation of the U-Net model, the dynamic stratification of radial erosion and the component clustering of the K-means++ algorithm, reduces the tedious operation of manual drawing, and processes the image more finely and accurately; it also utilizes multi-scale feature fusion and embedded channel attention mechanism to greatly improve the endosperm segmentation accuracy and reduce the boundary segmentation error; and it also performs dynamic radial stratification to solve the traditional stratification problem, thereby adapting to different grain morphologies, and thus having a wide range of applicability; in addition, the image can also be processed in layers to accurately reflect the protein distribution gradient, which is significantly progressive.

[0043] The above embodiments are only for illustrating the technical idea of ​​the present invention, and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for identifying differences in spatial distribution of wheat grain protein, characterized in that: The steps include: S1. Prepare multiple cross-sections of mature wheat grains and stain them with resin slices. Use image annotation tools to construct a semantic segmentation dataset including three types of annotations: endosperm, background, and cavity. After data enhancement, divide them into training set and validation set according to a fixed ratio. The improved U-Net network is used as the model architecture, the channel attention mechanism is embedded in the encoder and decoder modules, a hybrid loss function including cross entropy loss and Dice loss is designed, and then the AdamW optimizer is used in combination with the cosine annealing learning rate decay strategy for training; After the training is completed, the image size of each slice is reset and input into the trained model after normalization to complete the preliminary background removal and obtain the final foreground mask image; S2, based on the grayscale histogram feature of the blue channel, select a feature threshold for the final foreground mask image in step S1 to perform image binarization processing, first use the im2bw function to assign 0 to the area with a pixel value less than the feature threshold in the binary image, and assign 1 to the area with a pixel value greater than or equal to the feature threshold, then perform denoising through morphological opening operation, complete the preliminary segmentation of the background and grain area, and obtain a preliminary segmented image; The Sobel edge detection algorithm is continued to be used to detect the preliminary segmented image to obtain the continuous closed endosperm edge contour of the preliminary segmented image; the bwboundaries function is then used to extract the contour coordinates, and the center position of the endosperm area is determined based on the contour coordinates; Then calculate the Euclidean distance from all edge points on the edge contour of the endosperm to the corresponding center position, take the maximum value Dmax as the reference radius of radial stratification, divide the radius range into S layers, and erode the preliminary segmented image layer by layer by cyclically calling the imerode function to obtain the radial stratified structure image; S3. Use the imread function to load the final foreground mask image in step S1, use the reshape function to flatten and downsample the morphological opening process in step S2, and obtain the sampling matrix Xsample for calculating the initial cluster center; then set the number of clusters K, use the K-means++ algorithm to initialize the cluster center, and then iterate; after the iteration is completed, identify the protein based on the staining spectral characteristics and calculate the protein spatial distribution ratio.

2. A method for identifying differences in spatial distribution of wheat grain protein according to claim 1, characterized in that: In step S1, there are at least 300 resin slices of cross-sections of mature wheat grains, the ratio of training set to validation set is 8:2, and the weight α of the mixed loss function is 0.7; before the encoder and decoder modules are embedded in the channel attention mechanism, the ResNet34 residual block is first introduced in the encoder for multi-scale feature fusion.

3. A method for identifying spatial distribution differences of wheat grain protein according to claim 1, characterized in that: Before generating the final foreground mask image in step S1, the prediction mask output by the model needs to be subjected to morphological closing operations in sequence to connect broken areas, remove isolated noise areas whose areas are smaller than fixed pixels, and perform binarization processing.

4. A method for identifying differences in spatial distribution of wheat grain protein according to claim 1, characterized in that: In step S2, S=9, and the erosion radius of each layer is r w =S w Dmax,w=1, 2, 3...S, S w is the scale factor of the wth layer; when calling the imerode function to erode each layer layer by layer, the wth layer erosion operation uses a radius of r w The circular structure element is used to generate the annular area mask of the wth layer, and after completing the layer-by-layer erosion, a radial hierarchical structure image is formed from the edge contour of the endosperm to the center position.

5. The method for identifying spatial distribution differences of wheat grain protein according to claim 1, characterized in that: When the Sobel edge detection algorithm is used for detection in step S2, each preliminary segmented image is first grayed out, and the corresponding gradient amplitudes are calculated by the Sobel operators in the horizontal and vertical directions respectively. After non-maximum suppression and hysteresis threshold processing, the continuous closed endosperm edge contour of each preliminary segmented image is obtained.

6. A method for identifying differences in spatial distribution of wheat grain protein according to claim 1, characterized in that: The morphological opening operation in step S2 uses the imopen function of a 3×3 rectangular structure element.

7. A method for identifying differences in spatial distribution of wheat grain protein according to claim 6, characterized in that: When flattening and downsampling are performed in step S3, the 3×3 rectangular structure is first converted into a MN×3 two-dimensional data matrix through the reshape function, and each row in the two-dimensional data matrix corresponds to the RGB value of a pixel; then the final mask foreground image is downsampled using the alternate row and alternate column sampling method to generate a sampling matrix Xsample.

8. A method for identifying differences in spatial distribution of wheat grain protein according to claim 7, characterized in that: After obtaining the sampling matrix Xsample in step S3, set the number of clusters K=5, where 0 corresponds to background, 1 corresponds to cavity, 2 corresponds to protein, 3 corresponds to starch, and 4 corresponds to other components; when using the K-means++ algorithm to initialize the cluster center, the Start parameter in the kmeans function is set to 'plusplus', and the first center is randomly selected from the sampling matrix Xsample. Each subsequent center is selected according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected. Then, the similarity between the pixel and the cluster center is measured based on the Euclidean distance. The allocation stage is first iteratively executed, including calculating the Euclidean distance between each pixel and the 5 centers. , x j and a cluster center u K The Euclidean distance between them, c = 1, 2, 3, j is a positive integer, is assigned to the category with the smallest distance, and then the iterative update phase is performed, including recalculating the RGB mean of each type of pixel as the new center, until the maximum number of iterations is 100 or the center change is less than 10 -5 .

9. A device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute a method for identifying spatial distribution differences of wheat grain protein as described in any one of claims 1 to 8 when running the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program. When the computer program runs in the processor, the processor is used to execute the method for identifying spatial distribution differences of wheat grain protein as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Spl16 compositions and methods to increase agronomic performance of plants

    CN103906839A

  • Egg shape index calculation and variety identification method and system

    CN118230025A

  • Image segmentation method combining super-pixels and multi-scale hierarchical feature recognition

    WO2024021413A1

Cited By

  • Corn germplasm resource viability cross-germplasm discrimination method, system, equipment and medium

    CN120651770A

  • Method for monitoring purification process of myrobalan acid extracting solution based on image processing

    CN122176304A

  • A method for monitoring the purification process of chebulagic acid extract based on image processing

    CN122176304B