Method, device and storage medium for identifying differences in the spatial distribution of wheat grain proteins
By improving the combination of U-Net model and K-means++ algorithm, fully automated identification of the spatial distribution of wheat grain proteins is achieved, which solves the problem of time-consuming, labor-intensive and insufficient accuracy in the existing technology, and improves the recognition accuracy and applicability.
Patent Information
- Application Number
- CN202510579735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-07
AI Technical Summary
In the prior art, when identifying the spatial distribution of wheat grain proteins, the method is time-consuming and labor-intensive, insufficient accuracy, poor flexibility and accuracy, and cannot adapt to high-throughput research.
The improved U-Net model is used for semantic segmentation, combined with radial corrosion and K-means++ algorithm for component clustering, and fully automated image preprocessing and protein quantitative analysis are realized, reducing manual delineation operations, and improving recognition accuracy and applicability.
Fully automated identification of the spatial distribution of wheat grain proteins is achieved, the precision and accuracy of image processing is improved, and different grain morphology is adapted to different grain morphology and accurately reflect the protein distribution gradient.
Smart Images

Figure CN120107966B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of protein recognition, and particularly to a method, device, and storage medium for identifying the spatial distribution differences of wheat grain proteins. Background Art
[0002] Simultaneously achieving the goals of high yield and high quality in wheat production is an important basis for meeting the dual needs of food security and improving the quality of life. Different special-purpose wheat varieties have specific requirements for protein content. For example, strong gluten wheat used for bread making requires a high protein content, while weak gluten wheat used for biscuit making requires a low protein content.
[0003] Therefore, it provides a new idea for the production of special-purpose wheat flour. For medium-strong gluten wheat, selecting varieties with insensitive response of outer-layer proteins in grains to nitrogen application can increase the protein content of flour and the nitrogen use efficiency; for weak gluten wheat, selecting varieties with insensitive response of inner-layer proteins in grains to nitrogen application can stabilize or reduce the protein content in the inner endosperm while ensuring the yield, achieving high yield and quality improvement. Studying the spatial distribution of wheat grain proteins is helpful for improving the flour milling process, producing various special-purpose flours, and is of great significance for meeting diverse food processing needs.
[0004] However, the current mainstream method for studying the spatial distribution of grain proteins is through layered flour milling combined with the Kjeldahl method. However, the layered flour milling process is time-consuming and laborious, and it is difficult to grasp the milling time for each layer. Moreover, the Kjeldahl method has many steps and is prone to human error, resulting in poor data repeatability. In recent years, sectioning combined with image analysis has become another option for studying the spatial distribution of grain proteins. However, there is a lack of current image analysis software. Mainly, the protein area is identified through Image J, and there are also studies using ARCGIS to identify the spatial distribution differences of proteins. The common problem of the above image analysis studies is that it is necessary to manually depict the outline of the wheat grain endosperm, which is time-consuming and laborious and cannot be applied to high-throughput research. And when applied to layered research, the number of layers is fixed, and it cannot be flexibly changed. Distortion is likely to occur inside the layers. Identifying the protein staining area mainly distinguishes through RGB values, which is not accurate enough.
[0005] That is to say, at least the following three problems are included in the current measures taken:
[0006] 1) The mainstream method is through layered flour milling combined with the Kjeldahl method, which is time-consuming, laborious, and lacks accuracy;
[0007] 2) When combined with image analysis, it is necessary to manually depict the corresponding outline, which is also time-consuming and laborious and cannot be used for high-throughput research;
[0008] 3) There is insufficient flexibility, it is difficult to process during layering, and the recognition effect is not accurate enough. Summary of the Invention
[0009] In view of the above three problems, the object of the present invention is to propose a method, device and storage medium for identifying the spatial distribution difference of wheat grain proteins. By improving the semantic segmentation of the U-Net model, the dynamic stratification of radial erosion, and the component clustering of the K-means++ algorithm, the full automation from image preprocessing to protein quantitative analysis is realized, the cumbersome operation of manual drawing is reduced, the processing of images is more refined and accurate, and the images can be stratified to accurately reflect the protein distribution gradient.
[0010] It is achieved through the following technical solutions:
[0011] First, a method for identifying the spatial distribution difference of wheat grain proteins is proposed, including the following steps:
[0012] S1. Prepare multiple resin sections of wheat grain cross-sections at the mature stage and stain them. Use an image annotation tool to construct a semantic segmentation dataset including three types of annotations: endosperm, background, and cavity. After data augmentation, divide it into a training set and a validation set according to a fixed ratio; use an improved U-Net network as the model architecture, embed a channel attention mechanism in the encoder and decoder modules, design a hybrid loss function including cross-entropy loss and Dice loss, and then use the AdamW optimizer combined with the cosine annealing learning rate decay strategy for training; after training is completed, reset the image size of each section, and after standardization processing, input it into the trained model to complete the preliminary background removal and obtain the final foreground mask image;
[0013] S2. Based on the gray histogram features of the blue channel, select a feature threshold for image binaryzation processing of the final foreground mask image in step S1. First, use the im2bw function to assign a value of 0 to the area where the pixel value in the binary image is less than the feature threshold, and assign a value of 1 to the area where the pixel value is greater than or equal to the feature threshold. Then, perform denoising through morphological opening operation to complete the preliminary segmentation of the background and grain regions and obtain a preliminary segmentation image; continue to use the Sobel edge detection algorithm to detect the preliminary segmentation image to obtain the continuous and closed endosperm edge contour of the preliminary segmentation image; then use the bwboundaries function to extract the contour coordinates, and determine the center position of the endosperm region based on the contour coordinates; then calculate the Euclidean distance from all edge points on the endosperm edge contour to the corresponding center position, and take the maximum value Dmax as the reference radius for radial stratification. Divide the radius range into S layers evenly, and call the imerode function in a loop to erode the preliminary segmentation image layer by layer to obtain a radial stratification structure image;
[0014] S3. Use the imread function to load the final foreground mask image in step S1, and flatten and downsample the morphological opening operation process in step S2 through the reshape function to obtain the sampling matrix Xsample for calculating the initial cluster centers. Then set the number of clusters K, initialize the cluster centers using the K-means++ algorithm, and perform iterations. After the iterations are completed, identify proteins based on the staining spectral characteristics and calculate the proportion of the protein spatial distribution.
[0015] Preferably, in step S1, there are at least 300 resin sections of the cross-section of mature wheat grains. The ratio of the training set to the validation set is 8:2, and the weight α of the mixed loss function is 0.7. Before embedding the channel attention mechanism in the encoder and decoder modules, first introduce the ResNet34 residual block in the encoder for multi-scale feature fusion. Multiple slice data can ensure covering the morphological diversity of wheat grains as much as possible, and multi-scale feature fusion facilitates the subsequent operation of the channel attention mechanism and improves the recognition accuracy.
[0016] Preferably, before generating the final foreground mask image in step S1, it is also necessary to perform morphological closing operations on the predicted mask output by the model to connect the broken regions, remove isolated noise regions with an area smaller than a fixed number of pixels, and binaryzation processing. The morphological closing operation can fill the breaks at the endosperm edge, such as holes caused by uneven slice staining, etc., and then perform denoising and binaryzation processing, effectively improving the continuity of the boundary and the accuracy of subsequent recognition.
[0017] Preferably, in step S2, take S = 9, and the erosion radius of each layer is r w = S w ·Dmax, w = 1, 2, 3... S, S w is the proportionality coefficient of the w-th layer; when calling the imerode function to perform layer-by-layer erosion on each layer, the erosion operation of the w-th layer uses a circular structuring element with a radius of r w to generate the annular region mask of the w-th layer, and form a radial layered structure image from the endosperm edge contour to the center position after completing the layer-by-layer erosion. S = 9, divided into 9 annular regions, which can adapt to different grain morphologies and are also convenient for viewing the protein distribution in different layers.
[0018] Preferably, when using the Sobel edge detection algorithm for detection in step S2, first grayscale each preliminary segmentation image, calculate the corresponding gradient magnitudes respectively through the Sobel operators in the horizontal and vertical directions, and then obtain the continuous and closed endosperm edge contours of each preliminary segmentation image after non-maximum suppression and hysteresis threshold processing. The Sobel operator can highlight the gradient difference between the endosperm region and the background, thus obtaining a clear endosperm edge contour.
[0019] Preferably, the morphological opening operation in step S2 uses the imopen function with a 3×3 rectangular structuring element.
[0020] Preferably, when performing flattening and downsampling in step S3, first convert the 3×3 rectangular structure into a two-dimensional data matrix of MN×3 through the reshape function, where each row in the two-dimensional data matrix corresponds to the RGB values of a pixel point; then perform downsampling on the final masked foreground image using the interlaced sampling method to generate the sampling matrix Xsample. Flattening and downsampling can improve the calculation efficiency and achieve fast processing of large-scale images.
[0021] Preferably, after obtaining the sampling matrix Xsample in step S3, set the number of clusters K = 5, where 0 corresponds to the background, 1 corresponds to the cavity, 2 corresponds to the protein, 3 corresponds to the starch, and 4 corresponds to other components; when initializing the cluster centers using the K-means++ algorithm, set the Start parameter in the kmeans function to 'plusplus', randomly select the first center from the sampling matrix Xsample, and then select each subsequent center according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected. Then, based on the Euclidean distance to measure the similarity between the pixel and the cluster center, first iteratively execute the assignment stage, including calculating the Euclidean distance , x j and a cluster center u K between them, c = 1, 2, 3, j is a positive integer, and assign it to the category with the smallest distance. Then execute the iterative update stage, including recalculating the RGB mean of each class of pixels as the new center until the maximum number of iterations is 100 times or the center change amount is less than 10 −5 . K = 5 means that the component classification is 5 categories. Combining with the K-means++ algorithm to disperse the initial centers can avoid the local optimum of random initialization while taking into account the classification accuracy, so as to ensure the overall recognition accuracy.
[0022] Secondly, a device is also proposed, including a processor and a memory. The memory stores a computer program, and the processor is configured to execute a method for identifying the spatial distribution difference of wheat grain proteins as described above when running the computer program.
[0023] In addition, a storage medium is also proposed. A computer program is stored in the storage medium, and when the computer program runs in the processor, the processor is used to execute a method for identifying the spatial distribution difference of wheat grain proteins as described above.
[0024] The beneficial effects of the present invention compared with the prior art are:
[0025] The technical solution of the present invention realizes full automation from image preprocessing to protein quantitative analysis by improving the semantic segmentation of the U-Net model, the dynamic layering of radial erosion, and the component clustering of the K-means++ algorithm, reducing the cumbersome operations of manual drawing, and making the image processing more refined and accurate; it also uses multi-scale feature fusion and embedded channel attention mechanism to greatly improve the endosperm segmentation accuracy and reduce the boundary segmentation error; moreover, dynamic radial layering is carried out to solve the traditional layering problem, so as to adapt to different grain morphologies, thus having wide applicability; in addition, the image can be layered to accurately reflect the protein distribution gradient. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flowchart of a method for identifying the spatial distribution difference of wheat grain proteins;
[0027] Figure 2 It is a comparison chart of the effects of images in each link after an experiment using the method for identifying the spatial distribution difference of wheat grain proteins;
[0028] Figure 3 It is a schematic diagram of the layering and clustering effects of the spatial distribution difference of wheat grain proteins. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be described in detail in conjunction with the attached Figures 1 to 3 drawings in the embodiments of the present invention.
[0030] As Figure 1 shown, it is a flowchart of a method for identifying the spatial distribution difference of wheat grain proteins. Through removing background interference, dynamic spatial layering of radial erosion, and component clustering of the K-means++ algorithm, full-automatic operation from image preprocessing to protein quantitative analysis is realized, reducing the cumbersome operations of manual drawing, and also greatly improving the recognition accuracy and solving the layering problem in recognition. The following specific steps are adopted:
[0031] S1. Prepare multiple resin sections of wheat grain cross-sections at the mature stage and stain them. The number of sections is at least 300 to ensure covering the morphological diversity of wheat grains. Use an image annotation tool, such as the Labelme tool, to construct a semantic segmentation data set including three types of annotations: endosperm, background, and cavity. After data augmentation, it is divided into a training set and a validation set according to a fixed ratio of 8:2. The 8:2 ratio allocation uses only a small proportion of the data set for test verification to ensure the accuracy of training. The data augmentation here adopts conventional augmentation methods, such as geometric transformations like rotation, flipping, scaling, and cropping, and color transformations like adjusting brightness and hue, aiming to expand the database.
[0032] Then, an improved U-Net network is adopted as the architecture of the model, that is, the U-Net model is improved. ResNet34 residual blocks are introduced into the encoder corresponding to this model for multi-scale fusion, so as to extract features of the image at different scales under different convolutional kernel sizes, different pooling levels, etc., in order to capture more detailed features of the image and improve clarity. A channel attention mechanism is embedded in the encoder and the decoder module corresponding to the model. A hybrid loss function including cross-entropy loss and Dice loss is designed and the weight of both is α = 0.7, that is, the hybrid loss function = 0.7×cross-entropy loss+(1 - 0.7)×Dice loss, to improve the class discrimination and boundary segmentation accuracy. In addition, the weight α is defaulted to 0.7, but it can also be dynamically adjusted to adapt to various different situations of actual wheat grains. Then, the AdamW optimizer is combined with the cosine annealing learning rate decay strategy for training, and evaluated by indicators such as the mean intersection over union, which can trigger the early stopping mechanism to avoid overfitting.
[0033] It should be noted that the Labelme tool is an image annotation tool applicable to object detection and semantic segmentation tasks. The improved U-Net model is a convolutional neural network with an encoder-decoder structure and is applicable to image segmentation. The ResNet34 residual block contains 34 layers and improves the model's ability to extract features of different sizes through skip connections. The channel attention mechanism can dynamically allocate the weight α of the channels, so as to meet the requirements of different class discrimination and boundary segmentation accuracy. The AdamW optimizer is an improved version of the Adam optimizer, which can enhance the generalization ability and prevent overfitting. The cosine annealing learning rate decay strategy, whose learning rate decays from the initial value to the minimum value according to the cosine function curve and restarts periodically to jump out of the local optimum. The mean intersection over union is an important indicator for evaluating the segmentation of the model and can be used to feedback the accuracy of boundary segmentation.
[0034] After training is completed, the image size of each slice is reset, the image is adjusted to 512×512 pixels and after being normalized, it is input into the trained model. Then, for the predicted mask output after the model is trained, morphological closing operations are performed in sequence to connect broken regions, remove isolated noise regions with an area smaller than a fixed number of pixels, and binaryzation processing are carried out to complete the preliminary background removal and obtain the preprocessed image, which is also the final foreground mask image.
[0035] In this embodiment, morphological closing operation can fill the breaks at the edge of the endosperm, such as the holes caused by uneven slice staining, etc., so that the broken positions become continuous. For example, a disk with a radius of 3 pixels can be used for the closing operation, and through the dilation operation, the broken areas can be connected. Isolated noise regions with an area smaller than a fixed number of pixels are removed. For example, isolated noise regions smaller than 100 pixels can be removed, thereby reducing certain interference and improving the background removal rate. The binarization process here is used to identify the outer contour of the endosperm edge, so as to remove background interference.
[0036] As Figure 2 shown, it is a comparison diagram of the images of each link after an experiment using the wheat grain protein spatial distribution difference recognition method to verify the experimental effect. Among them, Figure a is the spliced image of the protein-stained endosperm read, which is for the image of the dataset constructed by the image annotation tool; Figure b is the image of identifying the outer contour and cavity of the endosperm and marking them, which is also for the image of the dataset constructed by the image annotation tool; Figure c is the image after accurately removing background interference, that is, the final foreground mask image; Figure d is the image of edge detection and centroid positioning, that is, the image after subsequent Sobel edge detection and determining the center position of the endosperm region; Figure e is the image divided into any level, that is, the subsequent radial hierarchical structure image; Figure f is the image of K-means clustering to identify proteins, that is, the final image after initialization and clustering by the K-means++ algorithm.
[0037] Combined with Figure c, the final foreground mask image in step S1 has clear boundary segmentation, a background removal rate as high as 98.7%, an endosperm boundary positioning error of less than 3 pixels, and the processing time of a single image can also be controlled within 120 ms, providing a high-quality data basis for the subsequent calculation of protein spatial distribution characteristics.
[0038] S2. Based on the blue channel, that is, the B channel of the blue component in the RGB image, and based on the grayscale histogram features of this channel, the feature threshold is selected for the final foreground mask image in step S1 to perform image binarization again, obtaining a binary image. The image binarization process in step S2 and the binarization process in step S1 are the same. First, the im2bw function is used to assign the area where the pixel value in the binary image is less than the feature threshold to 0, corresponding to the background area with a lower blue component in RGB, 0 corresponding to black. Then, the area where the pixel value is greater than or equal to the feature threshold is assigned to 1, corresponding to the main part of the grain, 1 corresponding to white. Then, denoising is performed through morphological opening operation to complete the preliminary segmentation of the background and grain areas, obtaining a preliminary segmentation image. It should be emphasized that the first binarization process in step S1 is to remove background interference after identifying the outer contour of the endosperm, targeting the endosperm and the background; while the image binarization process in step S2, that is, the second binarization process, is to identify the outer contour of the endosperm for subsequent layering, and at the same time, it targets the grain and the background.
[0039] In this embodiment, the morphological opening operation can use the imopen function with a 3×3 rectangular structuring element. The imopen function is a function used to implement morphological opening operation in image processing, which is suitable for removing small objects in the image and performing denoising.
[0040] Then, continue to use the Sobel edge detection algorithm to detect the preliminary segmentation image to obtain the continuous closed endosperm edge contour of the preliminary segmentation image. When using the Sobel edge detection algorithm for detection, first, each preliminary segmentation image is grayscale processed. The gradient magnitudes are calculated respectively through the Sobel operators in the horizontal direction [-1 0 1; -2 -0 -2; -1 0 1] and the vertical direction [-1 -2 -1; 0 0 0; 1 2 1]. Then, after non-maximum suppression and hysteresis threshold processing, the continuous closed endosperm edge contour of each preliminary segmentation image is obtained. The Sobel operator can highlight the gradient difference between the endosperm area and the background, thus obtaining a clear endosperm edge contour.
[0041] Among them, non-maximum suppression is to check the horizontal and vertical gradient directions for each pixel point, and only retain the pixel with the largest magnitude in the corresponding direction, and the rest are suppressed to 0. The hysteresis threshold processing is a dynamic threshold method. By setting two thresholds, high and low, such as the high-low threshold ratio is 2:1, it is used to distinguish strong edges, weak edges, and noise; strong edges correspond to clear real edges and are retained; noise is directly discarded; weak edges are between noise and strong edges and need to be rechecked whether they are connected to strong edges. If connected, they are retained; if not connected, they are discarded.
[0042] Further, continue to use the bwboundaries function to extract the contour coordinates, and determine the central position of the endosperm region based on the contour coordinates, also known as the centroid or geometric center. The calculation formula is Cx = N ∑ xi , Cy = N ∑ yi ,where N is the number of contour pixels. Then, calculate the Euclidean distance from all edge points on the endosperm edge contour to the corresponding central position (Cx, Cy), and take the maximum value Dmax as the reference radius for radial stratification. Divide the radius range into S layers, corresponding to different radial regions of protein distribution. By repeatedly calling the imerode function, the preliminary segmentation image is eroded layer by layer to obtain the overall radial stratification structure image. Refer to Figure 2 Figure e in
[0043] In this embodiment, S = 9 can be taken, and the erosion radius of each layer is r w =S w ·Dmax, w = 1, 2, 3... S, S w is the proportionality coefficient of the w-th layer, that is, the proportion of the w-th layer in the radial direction to the overall radial direction; when calling the imerode function to perform layer-by-layer erosion on each layer, the erosion operation of the w-th layer uses a circular structuring element with a radius of r w to generate the annular region mask of the w-th layer, that is, the current layer mask minus the previous layer mask. Then, after completing the layer-by-layer erosion, a radial stratification structure image from the endosperm edge contour to the central position is formed. S = 9, divided into 9 annular regions, which can adapt to different grain morphologies and is also convenient for viewing the protein distribution in different layers.
[0044] It should be noted that the Sobel edge detection algorithm is an edge detection algorithm based on gradient calculation. It calculates the gradient amplitude of the image in the horizontal and vertical directions through a convolution kernel to highlight the edge information. The bwboundaries function can extract the contour coordinates of the target object from a binary image and supports the detection of multiple connected regions. The imerode function is a function for morphological erosion operation. It shrinks the foreground region of the image through structuring elements such as circles and rectangles and is used for segmentation, denoising, or boundary extraction.
[0045] S3. Perform K-means clustering to identify proteins, including: loading the final foreground mask image in step S1 using the imread function, flattening and downsampling the process of morphological opening operation in step S2 through the reshape function. For example, first convert the 3×3 rectangular structure into a two-dimensional data matrix of MN×3 through the reshape function. Each row in the two-dimensional data matrix corresponds to the RGB value of a pixel point, and the RGB value range is 0-255. Then, for the final mask foreground image, such as a large data volume image of about 262,144 pixels, use the interlaced and inter-column sampling method for downsampling, downsample the 512×512 pixel image to a 256×256 pixel image, and generate the sampling matrix Xsample for subsequent calculation of the initial cluster centers. Flattening and downsampling can improve the calculation efficiency and achieve fast processing of large-scale images.
[0046] Then, set the number of clusters K, initialize the cluster centers using the K-means++ algorithm, and then perform iteration. After the iteration is completed, identify proteins based on the staining spectral characteristics and calculate the proportion of the protein spatial distribution.
[0047] In this embodiment, after obtaining the sampling matrix Xsample, initialize the cluster centers using the K-means++ algorithm and then perform iteration. For example, set the number of clusters K = 5, where 0 corresponds to the background, 1 corresponds to the cavity, 2 corresponds to the protein, 3 corresponds to the starch, and 4 corresponds to other components. When initializing the cluster centers using the K-means++ algorithm, set the Start parameter in the kmeans function to 'plusplus', randomly select the first center from the sampling matrix Xsample, and then select each subsequent center according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected. Then, based on the Euclidean distance to measure the similarity between the pixel and the cluster center, first iteratively execute the assignment stage, including calculating the Euclidean distance between each pixel point and the 5 centers , this formula represents a pixel point x j and a cluster center u K The Euclidean distance between them, c = 1, 2, 3, j is a positive integer, and it is assigned to the category with the smallest distance. Then execute the iterative update stage, including recalculating the RGB mean of each class of pixels as the new center until the maximum number of iterations is 100 times or the center change amount is less than 10 −5 . K = 5 means that the component classification is 5 categories. Combining with the K-means++ algorithm to disperse the initial centers can avoid the local optimum of random initialization while taking into account the classification accuracy, so as to ensure the overall recognition accuracy.
[0048] Such as Figure 3The figure shows a schematic diagram of the stratification and clustering effect of the spatial distribution difference of wheat grain protein. After the clustering is completed, the label mapping is optimized in combination with the characteristics of the stained image: the category with RGB value close to pure black, that is, R / G / B<50, is judged as the background, that is, category 0; the category with high gray value and prominent blue channel, that is, B>200, is judged as cavity, that is, category 1; based on the spectral characteristics of protein dyes, for example, Coomassie Brilliant Blue is used to stain RGB values of R=150-200, G=100-150, B=180-220 to mark proteins, that is, category 2; the remaining two categories are divided into starch and other components according to the grayscale difference, that is, category 3 and category 4.
[0049] Finally, the predicted mask of the endosperm during the previous segmentation of the image is used to filter the background of category 0 and the cavity area of category 1, and the number of pixels of protein, starch and other components in the effective area of the endosperm is counted. N protein , N starch , N other , through the formula P protein= N protein / ( N starch+ N protein+ N other)×100% to calculate the relative protein content.
[0050] Secondly, a device is also proposed, including a processor and a memory, the memory storing a computer program, and the processor is configured to execute a method for identifying spatial distribution differences of wheat grain protein as described above when running the computer program.
[0051] In addition, a storage medium is also proposed, in which a computer program is stored. When the computer program runs in a processor, the processor is used to execute the above-mentioned method for identifying spatial distribution differences of wheat grain protein.
[0052] In summary, the present invention realizes full automation from image preprocessing to protein quantitative analysis by improving the semantic segmentation of the U-Net model, the dynamic stratification of radial erosion and the component clustering of the K-means++ algorithm, reduces the tedious operation of manual depiction, and processes the image more finely and accurately; it also utilizes multi-scale feature fusion and embedded channel attention mechanism to greatly improve the endosperm segmentation accuracy and reduce the boundary segmentation error; and it also performs dynamic radial stratification to solve the traditional stratification problem, thereby adapting to different grain morphologies, and thus having a wide range of applicability; in addition, the image can also be processed in layers to accurately reflect the protein distribution gradient, which is significantly progressive.
[0053] The above embodiments are only for illustrating the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the present invention.
Claims
1. A method for identifying the difference in the spatial distribution of wheat grain proteins, characterized in that, It includes the following steps: S1. Prepare multiple resin sections of the cross-section of mature wheat grains and stain them. Use an image annotation tool to construct a semantic segmentation dataset including three types of annotations: endosperm, background, and cavity. After data augmentation, divide it into a training set and a validation set according to a fixed ratio; Adopt an improved U-Net network as the architecture of the model. Embed a channel attention mechanism in the encoder and decoder modules. Design a hybrid loss function including cross-entropy loss and Dice loss, and then use the AdamW optimizer combined with the cosine annealing learning rate decay strategy for training; After training is completed, reset the image size of each section, and after normalization processing, input it into the trained model to complete the preliminary background removal and obtain the final foreground mask image; S2. Based on the gray histogram features of the blue channel, select a feature threshold for image binarization of the final foreground mask image in step S1. First, use the im2bw function to assign a value of 0 to the area where the pixel value in the binarized image is less than the feature threshold, and assign a value of 1 to the area where the pixel value is greater than or equal to the feature threshold. Then, perform denoising through morphological opening operation to complete the preliminary segmentation of the background and grain regions and obtain a preliminary segmentation image; Continue to use the Sobel edge detection algorithm to detect the preliminary segmentation image to obtain the continuous and closed endosperm edge contour of the preliminary segmentation image; then use the bwboundaries function to extract the contour coordinates, and determine the central position of the endosperm region based on the contour coordinates; Then calculate the Euclidean distance from all edge points on the endosperm edge contour to the corresponding central position, and take the maximum value Dmax as the reference radius for radial stratification. Divide the radius range into S layers evenly, and sequentially perform layer-by-layer erosion on the preliminary segmentation image by circularly calling the imerode function to obtain a radially stratified structure image; S3. Use the imread function to load the final foreground mask image in step S1, and flatten and downsample the process of the morphological opening operation in step S2 through the reshape function to obtain a sampling matrix Xsample for calculating the initial cluster center; then set the number of clusters K, use the K-means++ algorithm to initialize the cluster center, and then perform iteration; after the iteration is completed, identify proteins based on the staining spectral characteristics and calculate the proportion of the protein spatial distribution.
2. The method for identifying the spatial distribution difference of wheat grain proteins according to claim 1, characterized in that In step S1, there are at least 300 resin sections of the cross-section of mature wheat grains, the ratio of the training set to the validation set is 8:2, and the weight α of the hybrid loss function is 0.7; before embedding the channel attention mechanism in the encoder and decoder modules, first introduce a ResNet34 residual block in the encoder for multi-scale feature fusion.
3. The method for identifying the spatial distribution difference of wheat grain proteins according to claim 1, wherein Before generating the final foreground mask image in step S1, it is also necessary to sequentially perform morphological closing operation to connect the broken regions, remove the isolated noise regions with an area less than a fixed number of pixels, and binarization processing on the predicted mask output by the model.
4. A method for identifying the spatial distribution difference of wheat grain proteins according to claim 1, characterized in that In step S2, take S = 9, and the etching radius of each layer is r w = S w ·Dmax, w = 1, 2, 3... S, S w is the proportionality coefficient of the w-th layer; when calling the imerode function to perform layer-by-layer etching on each layer, the etching operation of the w-th layer uses a circular structuring element with a radius of r w to generate the annular region mask of the w-th layer, and after completing the layer-by-layer etching, a radially stratified structure image from the endosperm edge contour to the center position is formed.
5. A method for identifying the spatial distribution difference of wheat grain proteins according to claim 1, characterized in that When performing detection using the Sobel edge detection algorithm in step S2, first, grayscale processing is performed on each preliminary segmented image. The corresponding gradient magnitudes are calculated separately through Sobel operators in the horizontal and vertical directions. After non-maximum suppression and hysteresis threshold processing, continuous closed endosperm edge contours of each preliminary segmented image are obtained.
6. The method for identifying the spatial distribution difference of wheat grain proteins according to claim 1, wherein, The morphological opening operation in step S2 uses the imopen function with a 3×3 rectangular structuring element.
7. A method for identifying the spatial distribution difference of wheat grain proteins according to claim 6, characterized in that, When performing flattening and downsampling in step S3, first, the 3×3 rectangular structure is converted into a two-dimensional data matrix of MN×3 through the reshape function. Each row in the two-dimensional data matrix corresponds to the RGB values of a pixel point. Then, the final masked foreground image is downsampled using an interlaced sampling method to generate the sampling matrix Xsample.
8. The method for identifying the difference in the spatial distribution of wheat grain proteins according to claim 7, wherein, After obtaining the sampling matrix Xsample in step S3, set the number of clusters K = 5, where 0 corresponds to the background, 1 corresponds to the cavity, 2 corresponds to the protein, 3 corresponds to the starch, and 4 corresponds to other components; when initializing the cluster centers using the K-means++ algorithm, set the Start parameter in the kmeans function to 'plusplus', randomly select the first center from the sampling matrix Xsample, and then select each subsequent center according to the proportional probability of the square of the distance from the data point to the nearest selected center until 5 initial centers are selected. Then, based on the Euclidean distance to measure the similarity between the pixel and the cluster center, first iteratively execute the assignment phase, including calculating the Euclidean distance between each pixel point and the 5 centers , x j and a cluster center u K The Euclidean distance between them, c = 1, 2, 3, j is a positive integer, and it is assigned to the category with the smallest distance. Then, execute the iterative update phase, including recalculating the RGB mean of each class of pixels as the new center until the maximum number of iterations is 100 times or the center change amount is less than 10 -5 .
9. A device, characterized in that, It includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute a method for identifying differences in the spatial distribution of wheat kernel proteins as described in any one of claims 1 to 8 when running the computer program.
10. A storage medium, characterized in that, A computer program is stored in a storage medium. When the computer program runs in a processor, the processor is used to execute a method for identifying differences in the spatial distribution of wheat kernel proteins as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Spl16 compositions and methods to increase agronomic performance of plants
CN103906839A
Egg shape index calculation and variety identification method and system
CN118230025A