Cotton root semantic segmentation method based on deep learning
By constructing a root semantic segmentation model based on the attention mechanism, the problem of insufficient generalization ability in cotton root image segmentation is solved, high-precision and fine-grained cotton root segmentation is achieved, and stable segmentation effects and reliable root analysis data are provided.
Patent Information
- Application Number
- CN202510695215.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have limited generalization capabilities in cotton root image segmentation and unsatisfactory processing results. Traditional methods damage plants, and high-tech equipment is expensive and complex. Commonly used models lack segmentation accuracy under complex backgrounds.
A root semantic segmentation model based on the attention mechanism is constructed. The transformer network and CNN network are used to extract multi-scale features. The FPN structure and the AGL mechanism are combined to achieve fine-grained segmentation through cross-grafting of multi-scale features. The attention module is introduced to improve the model's sensitivity to fine root features.
It achieves accurate segmentation of cotton roots in complex backgrounds, improves the generalization and robustness of the model, reduces missed and mis-segmentation phenomena, and provides high-precision root parameter results.
Smart Images

Figure CN120635443A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crop root morphology detection, and in particular to a cotton root semantic segmentation method based on deep learning. Background Art
[0002] Cotton is a deep-rooted crop, and its roots are the primary organ for water and nutrient absorption and transport. Scientific monitoring and management of root growth directly impacts cotton plant growth and pest resistance, boll yield and quality, and reproductive breeding. During its peak root growth phase, the taproot grows up to 2.5 cm daily, reaching depths exceeding 1 meter, and lateral roots can extend up to 50 cm laterally. Deep tillage during this phase effectively controls nutrient growth, promotes root development, and prolongs the root's functional life. During its peak root absorption phase, the root network is essentially established, absorption capacity reaches peak levels, but rooting capacity declines. Strengthening water and fertilizer management during this phase can maintain root vitality and delay root aging. Therefore, dynamic monitoring and analysis of cotton plant root growth holds important research and application implications for cotton growth management and breeding.
[0003] Currently, research techniques and methods for studying plant root systems remain at the forefront and a hot topic in related fields both domestically and internationally. However, traditional methods, such as excavation, drilling, and root cutting, can cause varying degrees of damage to plants. Aeroponics, hydroponics, and paper-based culture methods are suitable for laboratory environments and are difficult to apply to cotton fields. Meanwhile, advanced technologies are being applied to non-destructive root studies. For example, Keyes et al. used synchrotron X-ray tomography (SRXCT) in 2017 to achieve three-dimensional root images accurate to 1 μm, and Pflugfelder et al. used magnetic resonance imaging (MRI) to accurately model soybean roots. However, these methods share common challenges: expensive and complex testing equipment, and high soil moisture content can lead to significant errors and even hinder imaging. Micro-root canal analysis, combining a small color camera with a 360° multi-layer rotating scanner, can capture dynamic changes in root morphology. However, this technique is limited by equipment installation and sampling time, which can affect root growth paths and image clarity.
[0004] Furthermore, the background in cotton root images varies during growth. Given the slenderness, small size, and high background noise of root objects, commonly used image semantic segmentation models such as Mask R-CNN and DeepLab struggle to achieve the required segmentation accuracy. To address this, Dr. He Kaiming integrated the ResNet residual architecture with an FPN feature detection network, constructing a Mask R-CNN model using a Faster RCNN backbone and a deconvolutional mask branch. This model achieves pixel-level semantic segmentation of image objects, achieving an average precision (AP) of approximately 38%. With the emergence and continuous improvements of the one-stage detection model YOLO, the YOLO v8 model currently achieves an average precision (AP) exceeding 50% for semantic segmentation of image objects. However, the datasets used for these detection models are mostly large objects in everyday scenes, where the proportion of interior pixels is higher than that of outline pixels. However, the characteristics of crop root images are the opposite: outline pixels are higher and interior pixels are relatively lower. Therefore, simply applying these classic models to transfer learning of cotton root objects is difficult. Xie Chenxi's team at Beihang University and Pengcheng Laboratory proposed a one-stage Pyramid Grafting Network (PGNet) for salient object detection (SOD). This model extracts and grafts features from images of different resolutions. An attention-guided loss (AGL) supervises the generated attention matrix, helping the network better interact with attention from different models. The model achieved promising results in detecting the contours of steel structures and feathers. Chen Guowei et al. proposed a high-precision natural image matting network, PPMatting, which consists of a semantic context branch (SCB) and a high-resolution detail branch (HRDB). The interaction between the two branches improves the network's ability to perceive detail, outperforming other models in extracting the root details of a person's hair. Su Jiayi et al. proposed a semantic segmentation network for grape roots based on a modified U-Net architecture. Yan Jingkun et al. added a global attention mechanism module to the OCRNet network to strengthen the focus on root targets, achieving good results in the semantic segmentation of micro-root canal images.
[0005] While the Transformer-based Segmenter model, the PGNet (Pyramid Grafting Network) model, and the HRDB (High Resolution Detail Branch)-based PP-Matting model have all achieved good evaluation scores in the field of high-precision segmentation of detailed images, they all share a common problem: they rely on experimental datasets and have limited generalization capabilities. This results in poor performance for image objects such as cotton roots, which are primarily composed of edges and branch outlines. Summary of the Invention
[0006] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a cotton root semantic segmentation method based on deep learning. By constructing a higher-precision and finer-grained semantic segmentation model suitable for cotton roots, it can achieve accurate segmentation of cotton roots in complex backgrounds, effectively overcoming the defects of the existing technology such as limited generalization ability and unsatisfactory processing effect.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A cotton root semantic segmentation method based on deep learning includes the following steps: Step 1: Acquire a cotton root system image using a root system image acquisition device; Step 2: Preprocess the obtained cotton root images, annotate the image objects, and construct training and test sets; Step 3: Establish a root semantic segmentation model based on the attention mechanism; The root semantic segmentation model includes a first feature extraction module, a second feature extraction module, an attention module, a feature detection module, and a feature fusion module, wherein the first feature extraction module is used to extract multi-scale regional features of the main root from the cotton root system image; the second feature extraction module is used to extract multi-scale regional features of lateral roots and root hairs from the cotton root system image; the attention module is used to use the attention mechanism to process the bottom-level feature map in the multi-scale feature map output by the first feature extraction module and the second feature extraction module; the feature detection module is used to perform feature detection on the regional features of different scales output by the first feature extraction module, the second feature extraction module, and the attention module; the feature fusion module is used to splice and fuse the feature detection module and the output regional features by cross-grafting of multi-scale features to obtain a semantic segmentation result; Step 4: Use the training set to train the root semantic segmentation model; Step 5: Use the test set to test the trained root semantic segmentation model to obtain a trained root semantic segmentation model; Step 6: Use the trained root semantic segmentation model to process the root image to be segmented to obtain the cotton root semantic segmentation result.
[0008] Furthermore, the root image acquisition device includes an import-type inclined tube, a control module, an image collector, a cleaning module and an image processing module based on edge computing; the control module is used to generate control instructions; the import-type inclined tube is used to collect cotton roots; the image collector is used to obtain cotton root images according to the control instructions, and the cleaning module is used to clean the import-type inclined tube according to the control instructions; the image processing module is used to process and upload the cotton root images using edge computing according to the control instructions.
[0009] Furthermore, a claw-type positioning structure is provided on the inlet-type inclined tube.
[0010] Furthermore, in step 2, the cotton root system image is preprocessed including an edge cropping process, a median filtering process, a data cleaning process, and a size normalization process; in step 2, the image target is labeled using the Lableme labeling tool, and after labeling, the labeling parameters are stored in a json file.
[0011] Furthermore, the first feature extraction module adopts a transformer network, the second feature extraction module adopts a CNN network, and the feature detection module adopts an FPN network.
[0012] Furthermore, the attention module includes a first feature fusion unit, a first activation function unit, a first convolution unit, a second activation function unit, a first multiplication processing unit, a second multiplication processing unit, a second convolution unit, a second feature fusion unit, a third activation function unit and a third feature fusion unit; The first feature fusion unit adds the bottom-level feature map output by the first feature extraction module and the bottom-level feature map output by the second feature extraction module, and then inputs the sum into the first activation function unit. The first activation function unit strengthens the root features in the feature map and then inputs the sum into the first convolution unit. The first convolution unit performs convolution processing on the feature map and then inputs the sum into the second activation function unit. The second activation function unit suppresses the background response in the feature map and then inputs the sum into the first multiplication processing unit. The first multiplication processing unit multiplies the sum with the bottom-level feature map output by the first feature extraction module and then outputs the sum. The second multiplication processing unit multiplies the lowest-level feature map output by the first feature extraction module and the lowest-level feature map output by the second feature extraction module, and then inputs the result into the second convolution unit. The second convolution unit convolves the feature map and then inputs the result into the second feature fusion unit. The second feature fusion unit adds the result and the lowest-level feature map output by the second feature extraction module, and then inputs the result into the third activation function unit. The third activation function unit outputs the result after strengthening the root features in the feature map. The third feature fusion unit performs addition processing on the feature map output by the first multiplication processing unit and the feature map output by the third activation function unit.
[0013] Furthermore, the first activation function unit and the third activation function unit both use ReLU activation function, and the second activation function unit uses Sigmoid activation function.
[0014] Furthermore, the loss function during the training of the root semantic segmentation model is: L fl =-α(1- y '* y )*γ*log( y '), Among them, α and γ represent weight coefficients, y' is the model output result, and y is the model input result.
[0015] Furthermore, in step 5, when the trained root semantic segmentation model is tested using the test set, the test indicators include average pixel accuracy, mean absolute error, and F-score index. If the test result of the trained root semantic segmentation model does not meet the standard, return to step 4 for retraining.
[0016] Furthermore, the average pixel accuracy of the trained root semantic segmentation model is not less than 85%, the average absolute error is not higher than 15%, and the F-score index is not lower than 90%.
[0017] The remarkable effects of the present invention are: 1. Strong generalization ability: The present invention establishes a root semantic segmentation model based on the attention mechanism. Not only does it not require artificially designed features for guidance, the model itself can learn the required features from the training data and make reasonable use of them. Therefore, it has better generalization ability and can perform stably even in complex scenes. Moreover, by introducing the attention mechanism, the model's sensitivity to small root features is improved, and accurate segmentation of cotton roots under complex backgrounds is achieved. By adding an attention module, the model focuses on more important features and suppresses unnecessary features. The model not only performs better, but also has stronger robustness to noise input. Experiments have shown that the visual self-attention mechanism can not only optimize the segmentation boundaries, but also reduce the phenomena of missed segmentation and mis-segmentation.
[0018] (2) The present invention improves the network's feature discrimination and selection capabilities through the attention module, thereby alleviating the problem of reduced accuracy caused by interference features. In addition, the adopted attention structure can focus on more important channel features, reducing network complexity.
[0019] (3) The present invention constructs an end-to-end root image segmentation framework that can perform accurate segmentation without losing too much detail information, significantly improving the accuracy of the final root parameter results and providing reliable reference information for subsequent root calculation and analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of the method of the present invention; Figure 2 is a structural diagram of the root system image acquisition device; Figure 3 It is a structural diagram of the introduction type inclined tube; Figure 4 This is the left view of the imported inclined tube; Figure 5 This is a schematic diagram of the working state of the imported inclined tube; Figure 6 This is a schematic diagram of cotton root system image data processing; Figure 7 is a network structure diagram of the root semantic segmentation model; Figure 8 is a block diagram of the attention module; Figure 9 It is a cotton root system image obtained by semantic segmentation. DETAILED DESCRIPTION
[0021] The specific implementation manner and working principle of the present invention will be further described in detail below with reference to the accompanying drawings.
[0022] like Figure 1As shown in the figure, a cotton root semantic segmentation method based on deep learning is shown in the figure. The specific steps are as follows: Step 1: Acquire a cotton root system image using a root system image acquisition device; This embodiment studies the distribution of cotton roots in the soil during cotton growth in Xinjiang, and combines the research results and professional experience of a team of agricultural experts from Shihezi University and a team of experts in field robot design and development from the University of Lincoln in the United Kingdom to design an image acquisition device that can be used for long-term, all-round measurement of cotton roots. The specific structure is as follows: Figure 2 As shown, the root image acquisition device includes an import-type inclined tube, a control module, an image collector, a cleaning module and an image processing module based on edge computing; the control module is used to generate control instructions; the import-type inclined tube is used to collect cotton root systems; the image collector uses a high-resolution satellite camera to obtain cotton root system images according to control instructions. Since the cotton root system can reach 1m in length and 50cm in side diameter during the vigorous period, the depth and angle of the image acquisition point of the image collector are controlled by the control unit; the cleaning module is used to clean the import-type inclined tube according to the control instructions to ensure the long-term use of the equipment and reduce the image noise introduced by dust particles attached to the tube wall; in order to reduce the load on the system central server, the image processing module is used to process and upload the cotton root system image according to the control instructions using edge computing.
[0023] This embodiment is aimed at the sandy soil environment and irrigation operation of Xinjiang cotton fields. In order to reduce the analysis error caused by misaligned shooting, a claw-type positioning structure is provided on the inlet inclined tube, such as Figure 3-Figure 5 As shown, the claw-type positioning structure includes an oblique tube body 1, a mounting ring 2, a mounting groove 3, a fixing claw 5, a connecting rod 7 and a fixing block 8. The mounting ring 2 slides on the end of the oblique tube body 1 close to the ground surface, and a plurality of mounting grooves 3 are fixed on the outer wall of the mounting ring 2. The fixing claws 5 and the mounting grooves 3 are arranged in a one-to-one correspondence. One end of each fixing claw 5 is arranged in the mounting groove 3 and is rotatably connected to the mounting groove 3 through a rotating shaft. The middle part of the fixing claw 5 and one end of the connecting rod 7 are rotatably connected through a first pin shaft, and the other end of the connecting rod 7 and the fixing block 8 are rotatably connected through a second pin shaft. The fixing block 8 is fixed to the outer wall of the oblique tube body 1.
[0024] Based on the above structure, it can be known that in the process of moving the mounting ring 2 along the inclined tube body 1 to move it toward the fixed block 8, the fixing claw 5 will rotate around the rotating shaft under the action of the connecting rod 7, and at the same time, the connecting rod 7 will also rotate around the second pin shaft driven by the fixing claw 5, so that the multiple fixing claws 5 are opened away from the end of the mounting groove 3 and away from the outer wall of the inclined tube body 1, so that when the imported inclined tube is used to collect root system images, the inclined tube body 1 can be fixed in the soil, so that its position remains relatively fixed, thereby reducing the analysis error caused by misplaced shooting.
[0025] Furthermore, the plurality of fixing grooves 3 are evenly distributed circumferentially on the outer wall of the mounting ring 2 , so that the inclined tube body 1 can be fixed in position in all directions by the fixing claws 5 , thereby improving convenience in use.
[0026] from Figure 3 and Figure 5 It can also be seen that a pin 4 is independently installed on the mounting ring 2, and a plurality of insertion holes 6 are opened on the outer wall of the inclined tube body 1. Through the cooperation of the pin 4 and the insertion holes 6, the position of the moving mounting ring 2 can be fixed, thereby improving the stability of the position and posture of the fixing claw 5 during root system image acquisition, further improving the position stability of the imported inclined tube, and further reducing the analysis error caused by misplaced shooting.
[0027] Step 2: Preprocess the obtained cotton root images, annotate the image objects, and construct training and test sets; In the embodiment of the present invention, about 50,000 cotton root system images were collected as a training data set for the model. The cotton root system images were subjected to edge cropping, median filtering, data cleaning and size normalization. The size of the source image was determined by the setting of the built-in lens of the micro-root tube, such as Figure 6 The original image shown in (a) has a resolution of 2550*2273. Considering the computational complexity and training speed of the system, we set the resolution of the standard processed image to 640*640 and divide the original image into 4*4 sub-images, as shown in the following example: Figure 6 (b) (edges are padded with 0 pixels). Apply the Labelme annotation tool to annotate the image objects, as shown in Figure 6 (c). Save the annotation parameters into the json file, such as Figure 6 (d).
[0028] In this example, the steps to construct the training set and test set are as follows: When the root system patches in the root system image are sparsely distributed, the steps are as follows: Step A1: Obtain the boundaries of each image patch, and expand a buffer zone of random size around the boundaries. Step A2: rasterizing the image patch range after the buffer zone is expanded to obtain a mask of the root system within the image patch range after the buffer zone is expanded; Step A3: Slide a fixed-size window on the mask and calculate the ratio between the foreground and background. Step A4: when the ratio is greater than a set threshold, the image and the data of the area on the mask are cropped as the image and label; Step A5: Integrate the cropped data and labels to obtain a training sample set; Step A6: randomly generate a training set and a test set based on the training sample set; The above process only rasterizes the area near the spots. When the spots are sparsely distributed, it can avoid generating a mask that takes up a large amount of space for the entire image, reduce the number of sliding windows, and improve the sample extraction speed.
[0029] When the root system spots in the root system image are densely distributed, the steps are: Step B1: rasterizing the root system image vector data to obtain a mask covering the entire image; Step B2: Use a fixed-size sliding window to slide simultaneously on the root image and mask to extract data as images and labels; Step B3: Integrate the cropped data and labels to obtain a training sample set; Step B4: Randomly generate a training set and a test set based on the training sample set.
[0030] Through the above-mentioned sample set preparation method, the sample set can be purified, and erroneous samples with inaccurate labels can be automatically screened out, thereby improving the purity of correct samples in the sample set and significantly reducing the cost of preparing the sample set.
[0031] Step 3: Establish a root semantic segmentation model based on the attention mechanism; The cotton root target images are different from the daily target detection images. They are collected by the high-precision lens built into the positioning oblique root canal buried in the cotton field soil. The original collected image data is as follows Figure 6 shown.
[0032] from Figure 6Cotton root images exhibit complex and noisy backgrounds, slender and widely extended objects, variable and irregular root shapes, multiple contour features, varying sizes of main roots, lateral roots, and root fibers, and a small number of pixels occupied along the diameter. Common semantic segmentation detection models such as Mask R-CNN and DeepLab are unable to meet the required detection accuracy for root systems. The Transformer-based Segmenter model, the PGNet (Pyramid Grafting Network) model, and the HRDB (High-Resolution Detail Branch)-based PP-Matting model have all achieved impressive evaluation scores in the field of high-precision segmentation of detail images. PGNet-UH (trained on the UH dataset) achieved a MAE of 0.020 and an F-score of 0.945. PP-Matting achieved an MSE of 0.005 on the high-precision Composition-1k images and 0.009 on the low-precision Distinctions-646 training set. However, a common problem is that they rely on experimental datasets and have limited model generalization capabilities. For image targets such as cotton roots, which are basically composed of edges and branch outlines, the processing effect is not ideal.
[0033] In response to the above characteristics, this example establishes a high-precision small target segmentation model, PCSegRoot, suitable for high-resolution cotton root images. This model uses a transformer network and a CNN network to extract regional features of the main root, lateral roots, and root hairs, respectively. It employs an FPN structure to detect features of regional targets at different scales, and uses multi-scale feature cross-grafting to combine and fuse branch features. The AGL mechanism improves the model's sensitivity to small features, resulting in a pyramidal, cross-model root semantic segmentation model based on an attention mechanism. Specifically: like Figure 7As shown, the root semantic segmentation model in this example includes a first feature extraction module, a second feature extraction module, an attention module, a feature detection module, and a feature fusion module, wherein the first feature extraction module is used to extract multi-scale regional features of the main root from the cotton root system image; the second feature extraction module is used to extract multi-scale regional features of lateral roots and root hairs from the cotton root system image; the attention module is used to use the attention mechanism to process the bottom-level feature map in the multi-scale feature map output by the first feature extraction module and the second feature extraction module; the feature detection module is used to perform feature detection on the regional features of different scales output by the first feature extraction module, the second feature extraction module, and the attention module; the feature fusion module is used to splice and fuse the feature detection module and the output regional features by cross-grafting of multi-scale features to obtain a semantic segmentation result; The root semantic segmentation model based on the attention mechanism established in the present invention not only does not require artificially designed features for guidance, but the model itself can learn the required features from the training data and make reasonable use of them. Therefore, it has better generalization ability and can perform stably even when facing complex scenes. Moreover, by introducing the attention mechanism, the model's sensitivity to fine root features is improved, and accurate segmentation of cotton roots under complex backgrounds is achieved.
[0034] In this embodiment of the present invention, the structure of the attention module is as follows: Figure 8 As shown, it includes a first feature fusion unit, a first activation function unit, a first convolution unit, a second activation function unit, a first multiplication processing unit, a second multiplication processing unit, a second convolution unit, a second feature fusion unit, a third activation function unit and a third feature fusion unit; The first feature fusion unit adds the bottom-level feature map T1 output by the first feature extraction module and the bottom-level feature map T2 output by the second feature extraction module, and then inputs the sum into the first activation function unit. The first activation function unit strengthens the root features in the feature map and then inputs the feature map into the first convolution unit. The first convolution unit performs convolution processing on the feature map and then inputs the sum into the second activation function unit. The second activation function unit suppresses the background response in the feature map and then inputs the sum into the first multiplication processing unit. The first multiplication processing unit multiplies the sum with the bottom-level feature map T1 output by the first feature extraction module and then outputs the sum. The second multiplication processing unit multiplies the bottom feature map T1 output by the first feature extraction module and the bottom feature map T2 output by the second feature extraction module, and then inputs the result into the second convolution unit. The second convolution unit performs convolution processing on the feature map and then inputs the result into the second feature fusion unit. The second feature fusion unit adds the result and the bottom feature map T2 output by the second feature extraction module, and then inputs the result into the third activation function unit. The third activation function unit outputs the result after strengthening the root features in the feature map. The third feature fusion unit adds the feature map output by the first multiplication processing unit and the feature map output by the third activation function unit and outputs a feature map T.
[0035] In this example, the first activation function unit and the third activation function unit both use the ReLU activation function, the second activation function unit uses the Sigmoid activation function, and the first convolution unit and the second convolution unit are both 1×1×1 convolutions.
[0036] The root semantic segmentation model described in this embodiment improves the sensitivity of the model to small root features by introducing the attention mechanism, and realizes the accurate segmentation of cotton roots under complex backgrounds. By adding the attention module, the model focuses on more important features and suppresses unnecessary features. The model not only performs better, but also has stronger robustness to noise input. Experiments have shown that the visual self-attention mechanism can not only optimize the boundary of segmentation, but also reduce the phenomenon of missed segmentation and mis-segmentation. At the same time, the attention module improves the feature discrimination and selection ability of the network, thereby alleviating the problem of reduced accuracy caused by interference features, and the adopted attention structure can focus on more important channel features, reducing the complexity of the network.
[0037] Step 4: Use the training set to train the root semantic segmentation model; In this example, the loss function for training the root semantic segmentation model is: L fl =-α(1- y '* y )*γ*log( y '), Among them, α and γ represent weight coefficients, y' is the model output result, and y is the model input result.
[0038] Step 5: Use the test set to test the trained root semantic segmentation model to obtain a trained root semantic segmentation model. If the test result of the trained root semantic segmentation model does not meet the standard, return to step 4 for retraining. The trained root semantic segmentation model should achieve the following segmentation results on high-resolution images (2550*2273, slightly different depending on the acquisition device and parameter settings): mPA (Mean Pixel Accuracy) no less than 85%, MAE (Mean Absolute Error) no more than 15%, and F-score index no less than 90%.
[0039] Step 6: Use the trained root semantic segmentation model to process the root image to be segmented and obtain Figure 9 The cotton root semantic segmentation results are shown.
[0040] In summary, the present invention establishes a root semantic segmentation model with higher precision and finer granularity based on the attention mechanism. It not only does not require artificially designed features for guidance, but the model itself can learn the required features from the training data and make reasonable use of them. Therefore, it has better generalization ability and can perform stably even in complex scenes. It improves the model's sensitivity to fine root features and realizes accurate segmentation of cotton roots under complex backgrounds, thereby providing data support for root detection and analysis.
[0041] The technical solution provided by the present invention is introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A cotton root semantic segmentation method based on deep learning, characterized in that: The steps include: Step 1: Acquire a cotton root system image using a root system image acquisition device; Step 2: Preprocess the obtained cotton root images, annotate the image objects, and construct training and test sets; Step 3: Establish a root semantic segmentation model based on the attention mechanism; The root semantic segmentation model includes a first feature extraction module, a second feature extraction module, an attention module, a feature detection module, and a feature fusion module, wherein the first feature extraction module is used to extract multi-scale regional features of the main root from the cotton root system image; the second feature extraction module is used to extract multi-scale regional features of lateral roots and root hairs from the cotton root system image; the attention module is used to use the attention mechanism to process the bottom-level feature map in the multi-scale feature map output by the first feature extraction module and the second feature extraction module; the feature detection module is used to perform feature detection on the regional features of different scales output by the first feature extraction module, the second feature extraction module, and the attention module; the feature fusion module is used to splice and fuse the feature detection module and the output regional features by cross-grafting of multi-scale features to obtain a semantic segmentation result; Step 4: Use the training set to train the root semantic segmentation model; Step 5: Use the test set to test the trained root semantic segmentation model to obtain a trained root semantic segmentation model; Step 6: Use the trained root semantic segmentation model to process the root image to be segmented to obtain the cotton root semantic segmentation result.
2. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: The root system image acquisition device includes an import-type inclined tube, a control module, an image collector, a cleaning module and an image processing module based on edge computing; the control module is used to generate control instructions; the import-type inclined tube is used to collect cotton root systems; the image collector is used to obtain cotton root system images according to the control instructions, and the cleaning module is used to clean the import-type inclined tube according to the control instructions; the image processing module is used to process and upload cotton root system images using edge computing according to the control instructions.
3. The cotton root semantic segmentation method based on deep learning according to claim 2, characterized in that: The introduction-type inclined tube is provided with a claw-type positioning structure.
4. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: In step 2, the cotton root system image is preprocessed, including an edge clipping process, a median filtering process, a data cleaning process, and a size normalization process; in step 2, the image target is labeled using the Lableme annotation tool, and the annotation parameters are stored in a json file after annotation.
5. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: The first feature extraction module adopts the transformer network, the second feature extraction module adopts the CNN network, and the feature detection module adopts the FPN network.
6. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: The attention module includes a first feature fusion unit, a first activation function unit, a first convolution unit, a second activation function unit, a first multiplication processing unit, a second multiplication processing unit, a second convolution unit, a second feature fusion unit, a third activation function unit and a third feature fusion unit; The first feature fusion unit adds the bottom-level feature map output by the first feature extraction module and the bottom-level feature map output by the second feature extraction module, and then inputs the sum into the first activation function unit. The first activation function unit strengthens the root features in the feature map and then inputs the sum into the first convolution unit. The first convolution unit performs convolution processing on the feature map and then inputs the sum into the second activation function unit. The second activation function unit suppresses the background response in the feature map and then inputs the sum into the first multiplication processing unit. The first multiplication processing unit multiplies the sum with the bottom-level feature map output by the first feature extraction module and then outputs the sum. The second multiplication processing unit multiplies the lowest-level feature map output by the first feature extraction module and the lowest-level feature map output by the second feature extraction module, and then inputs the result into the second convolution unit. The second convolution unit convolves the feature map and then inputs the result into the second feature fusion unit. The second feature fusion unit adds the result and the lowest-level feature map output by the second feature extraction module, and then inputs the result into the third activation function unit. The third activation function unit outputs the result after strengthening the root features in the feature map. The third feature fusion unit performs addition processing on the feature map output by the first multiplication processing unit and the feature map output by the third activation function unit.
7. The cotton root semantic segmentation method based on deep learning according to claim 6, characterized in that: The first activation function unit and the third activation function unit both use the ReLU activation function, and the second activation function unit uses the Sigmoid activation function.
8. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: The loss function during the training of the root semantic segmentation model is: L fl =-α(1- y '* y )*γ*log( y '), Among them, α and γ represent weight coefficients, y' is the model output result, and y is the model input result.
9. The cotton root semantic segmentation method based on deep learning according to claim 1, characterized in that: In step 5, when the trained root semantic segmentation model is tested using the test set, the test indicators include average pixel accuracy, mean absolute error, and F-score index. If the test results of the trained root semantic segmentation model do not meet the standards, return to step 4 for retraining.
10. The cotton root semantic segmentation method based on deep learning according to claim 9, characterized in that: The trained root semantic segmentation model has an average pixel accuracy of no less than 85%, an average absolute error of no more than 15%, and an F-score index of no less than 90%.
Citation Information
Cited By
Cotton radicle image segmentation method, and root length and root coarse measurement method and system
CN121639728A