Cotton data processing method and device based on unmanned aerial vehicle and electronic equipment
By collecting cotton field data using drones equipped with RGB and lidar sensors, and combining various machine learning algorithms and gene editing technologies, the problem of low processing efficiency of cotton plant height gene data was solved, enabling rapid and accurate localization of cotton plant height-related genes and simplifying the breeding process.
Patent Information
- Application Number
- CN202411817469.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing technologies have low data processing efficiency for cotton plant height genes, making it impossible to quickly and accurately improve the target material.
By using drones equipped with RGB and LiDAR sensors to collect cotton field data, and combining various machine learning algorithms and gene editing technologies, cotton plant height data is automatically processed to achieve precise canopy boundary calibration and gene localization.
It improves the accuracy and efficiency of cotton plant height gene data processing, enabling rapid and accurate acquisition of plant height-related genes and simplifying the breeding process.
Smart Images

Figure CN119785241B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biotechnology, and in particular to a cotton data processing method and device based on a UAV and an electronic device. BACKGROUND
[0002] Cotton is one of the most important economic crops in the world and is the main producer of natural fibers for the global textile industry. Cotton plant type, as a comprehensive agronomic trait, refers to the three-dimensional structure of the aboveground part of cotton. Cotton plant type is a key factor affecting cotton yield and mechanical harvesting efficiency. Improving cotton plant type can accelerate the breeding process, effectively improve cotton yield, and improve fiber quality, thereby improving the economic benefits of cotton planting. The main cultivation characteristics of cotton at present are dwarf, dense planting and early maturity, among which dwarf refers to plant height.
[0003] Plant height is one of the most important factors in cotton plant type, and optimizing plant height can increase planting density, thereby improving the efficiency of mechanical harvesting of cotton and yield. By introducing dwarf genes in crops, plant height can be changed to significantly improve yield. However, through traditional breeding methods such as hybridization, it is necessary to spend a lot of time in the field for the selection and cultivation of hybrid combinations, which cannot quickly and accurately improve the target material. By manually measuring cotton plant height data in the field and correlating with genotype data, cotton plant height genes are identified and cotton materials are improved, which can shorten the improvement process of the target material. However, the existing technology has low data processing efficiency for cotton plant height genes.
[0004] Embodiments of the application
[0005] The purpose of the embodiments of the present application is to provide a cotton data processing method and device based on a UAV and an electronic device to solve the technical problem of low data processing efficiency for cotton plant height genes.
[0006] In a first aspect, the present application provides a cotton data processing method based on a UAV, the UAV flies above a cotton field, and the UAV is loaded with an RGB sensor and a laser radar sensor; the method comprises:
[0007] acquiring RGB image data of the cotton field through the RGB sensor and acquiring laser point cloud data of the cotton field through the laser radar sensor;
[0008] The RGB image data is aligned and spliced to obtain a first processing result. The RGB image data is automatically batch-processed for recognition according to the first processing result to obtain a first recognition result. The first soil part and the first background part in the RGB image data are divided according to the first recognition result, and all first pixel points of the green vegetation part in the RGB image data are retained. A first group canopy range boundary is created for each material in the cotton group in the RGB image data according to the first pixel points, the first soil part and the first background part. A digital orthographic image map is constructed and a first digital surface model (DSM) is generated according to the first group canopy range boundary. First plant height data is obtained based on the first digital surface model (DSM).
[0009] The laser point cloud data is classified and denoised to obtain a second processing result. The laser point cloud data is divided into ground point type point cloud data and non-ground point type point cloud data according to the second processing result. The laser point cloud data is automatically batch-processed for recognition according to the second processing result to obtain a second recognition result. The second soil part and the second background part in the laser point cloud data are divided according to the second recognition result, and all second pixel points of the green vegetation part in the laser point cloud data are retained. A second group canopy range boundary is created for each material in the cotton group in the laser point cloud data according to the second pixel points, the second soil part and the second background part. A digital elevation model (DEM) is generated based on the ground point type point cloud data and the second group canopy range boundary by structure from motion (SfM). A second digital surface model (DSM) is generated based on the non-ground point type point cloud data and the second group canopy range boundary by SfM. Second plant height data is obtained based on the difference between the digital elevation model (DEM) and the second digital surface model (DSM).
[0010] The field measurement plant height data of part of the plant individuals in the cotton field is obtained. The field measurement plant height data is obtained through a measurement process from the ground to the top of the stem of the plant individual. The measurement time of the measurement process is less than a preset time period from the acquisition time of the RGB image data and the laser point cloud data.
[0011] A plurality of different machine learning algorithms are used to train and correct the first plant height data, the second plant height data and the field measurement plant height data, respectively, to obtain final first plant height data and final second plant height data.
[0012] Filtering the cotton population genomic variation data SNPs in the cotton field to obtain target cotton population genomic variation data SNPs with actual quality and accuracy greater than that before filtering, calculating the best linear unbiased estimate BLUE based on the final first plant height data and the final second plant height data, and performing GWAS analysis on the best linear unbiased estimate BLUE and the target cotton population genomic variation data SNPs to obtain an analysis result, and obtaining plant height gene data based on the analysis result;
[0013] Extracting the coding sequence of the plant height related gene located based on the plant height gene data, amplifying the plant height gene using specific primers according to the coding sequence to obtain an amplified fragment, inserting the amplified fragment into a designated target vector to obtain a target gene overexpression vector, and generating gene overexpression new material data based on the target gene overexpression vector.
[0014] Determining the sgRNA sequence of the plant height related gene based on the plant height gene data, inserting the sgRNA sequence into a designated gene knockout vector to obtain a target gene knockout vector, and generating gene editing new material data based on the target gene knockout vector.
[0015] In one possible implementation, the training and correction of the first plant height data, the second plant height data and the field measured plant height data using a plurality of different machine learning algorithms to obtain the final first plant height data and the final second plant height data include:
[0016] Detecting the extraction accuracy of the plant height data using a plurality of different machine learning algorithms through at least one of the determination coefficient R 2 , the mean absolute error MAE and the root mean square error RMSE to obtain a comprehensive detection result.
[0017] Selecting the best machine learning model result according to the comprehensive detection result, and determining the first plant height data corresponding to the best machine learning model result as the final first plant height data and the second plant height data corresponding to the best machine learning model result as the final second plant height data; wherein the final first plant height data and the final second plant height data are used in the process of the GWAS analysis.
[0018] In one possible implementation, the determination coefficient R 2 is greater and the mean absolute error MAE and the root mean square error RMSE are smaller, indicating that the model accuracy of data training through the machine learning algorithm is higher.
[0019] In one possible implementation, the transformation of the plant height related genes is performed by using a genotype-free high-efficiency genetic transformation system SAMT for cotton, which uses cotton seed apical meristem stem cells as explants, combines Agrobacterium and ultrasonic treatment, integrates foreign vectors into stem cells to induce stem cells to produce adventitious buds, and uses zeocin to inhibit the growth of main buds to induce the generation of axillary buds.
[0020] In one possible implementation, the script is used to divide the soil and cotton population canopy by color segmentation through different color spectrums of the soil and the cotton population canopy, and the boundaries of the cotton population canopy are automatically identified and demarcated according to the color spectrum.
[0021] In one possible implementation, the first plant height data is obtained based on the first digital surface model (DSM).
[0022] The difference between the upper percentile and the lower percentile based on the first digital surface model (DSM) is determined as the plant height value extracted from the RGB image data, and the first plant height data is obtained.
[0023] In one possible implementation, the plurality of different machine learning algorithms include any of the following:
[0024] Random forest regression algorithm, K-nearest neighbor algorithm, Huber regression algorithm, least angle regression algorithm, light gradient boosting machine algorithm, decision tree regression algorithm, extra tree regression algorithm, orthogonal matching pursuit algorithm, Bayesian ridge regression algorithm, AdaBoost regression algorithm, CatBoost regression algorithm, and ridge regression algorithm.
[0025] In a second aspect, the present application provides a cotton data processing device based on a UAV, wherein the UAV flies above a cotton field, and the UAV is loaded with an RGB sensor and a laser radar sensor; the device comprises:
[0026] The acquisition module is configured to acquire RGB image data of the cotton field through the RGB sensor and laser point cloud data of the cotton field through the laser radar sensor.
[0027] The first processing module is configured to perform alignment and splicing processing on the RGB image data to obtain a first processing result, automatically perform batch identification on the RGB image data according to the first processing result by running a script to obtain a first identification result, divide a first soil part and a first background part in the RGB image data according to the first identification result, retain all first pixel points of a green vegetation part in the RGB image data, create a first group canopy range boundary for each material in a cotton group in the RGB image data according to the first pixel points, the first soil part and the first background part, construct a digital orthographic image map according to the first group canopy range boundary, generate a first digital surface model (DSM), and obtain first plant height data based on the first digital surface model (DSM).
[0028] The second processing module is configured to perform classification and denoising processing on the laser point cloud data to obtain a second processing result, divide ground point type point cloud data and non-ground point type point cloud data from the laser point cloud data according to the second processing result, automatically perform batch identification on the laser point cloud data according to the second processing result by running a script to obtain a second identification result, divide a second soil part and a second background part in the laser point cloud data according to the second identification result, retain all second pixel points of a green vegetation part in the laser point cloud data, create a second group canopy range boundary for each material in a cotton group in the laser point cloud data according to the second pixel points, the second soil part and the second background part, generate a digital elevation model (DEM) based on the ground point type point cloud data and the second group canopy range boundary by structure from motion (SfM), generate a second digital surface model (DSM) based on the non-ground point type point cloud data and the second group canopy range boundary by SfM, and obtain second plant height data based on the digital elevation model (DEM) and the second digital surface model (DSM).
[0029] The acquisition module is configured to acquire in-situ measured plant height data of part of plant individuals in the cotton field, wherein the in-situ measured plant height data is obtained through a measurement process from the ground to the top of the stem of the plant individual, and the measurement time of the measurement process and the acquisition time of the RGB image data and the laser point cloud data differ by a time period less than a preset time period.
[0030] The training module is configured to train and verify the first plant height data, the second plant height data and the in-situ measured plant height data by using a plurality of different machine learning algorithms to obtain final first plant height data and final second plant height data.
[0031] The analysis module is configured to filter the cotton population genomic variation data SNPs in the cotton field to obtain target cotton population genomic variation data SNPs with actual quality and accuracy greater than that before filtering, calculate a best linear unbiased estimate (BLUE) based on the final first plant height data and the final second plant height data, perform GWAS analysis on the BLUE and the target cotton population genomic variation data SNPs to obtain an analysis result, and obtain plant height gene data based on the analysis result.
[0032] The first construction module is configured to extract a coding sequence of a located plant height related gene based on the plant height gene data, amplify the plant height gene based on the coding sequence to obtain an amplified fragment, insert the amplified fragment into a specified target vector to obtain a target gene overexpression vector, and generate gene overexpression new material data based on the target gene overexpression vector.
[0033] The second construction module is configured to determine an sgRNA sequence of a plant height related gene based on the plant height gene data, insert the sgRNA sequence into a specified gene knockout vector to obtain a target gene knockout vector, and generate gene editing new material data based on the target gene knockout vector.
[0034] In a third aspect, the present application further provides an electronic device, including a memory and a processor, the memory stores a computer program which can be run on the processor, and the processor implements the method of the first aspect when executing the computer program.
[0035] In a fourth aspect, the present application further provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions make the processor execute the method of the first aspect when being called and run by the processor.
[0036] The present application brings the following beneficial effects:
[0037] The application provides a cotton data processing method and device based on a UAV and an electronic device. The RGB image data of a cotton field is collected by an RGB sensor carried on the UAV, the laser point cloud data of the cotton field is collected by a laser radar sensor carried on the UAV, the RGB image data is aligned and spliced to obtain a first processing result, the RGB image data is automatically identified in batches according to the first processing result by running a script to obtain a first identification result, the first soil part and the first background part in the RGB image data are divided according to the first identification result, and all first pixel points of the green vegetation part in the RGB image data are retained, the first group canopy range boundary is created for each material in the cotton group in the RGB image data according to the first pixel points, the first soil part and the first background part, the digital orthographic image is constructed according to the first group canopy range boundary, the first digital surface model (DSM) is generated, and the first plant height data is obtained based on the first digital surface model (DSM), the laser point cloud data is classified and denoised to obtain a second processing result, the laser point cloud data is divided into ground point type point cloud data and non-ground point type point cloud data according to the second processing result, the laser point cloud data is automatically identified in batches according to the second processing result by running a script to obtain a second identification result, the second soil part and the second background part in the laser point cloud data are divided according to the second identification result, and all second pixel points of the green vegetation part in the laser point cloud data are retained, the second group canopy range boundary is created for each material in the cotton group in the laser point cloud data according to the second pixel points, the second soil part and the second background part, the digital elevation model (DEM) is generated based on the ground point type point cloud data by structure from motion (SfM) using the second group canopy range boundary, the second digital surface model (DSM) is generated based on the non-ground point type point cloud data by SfM using the second group canopy range boundary, and the second plant height data is obtained based on the difference between the digital elevation model (DEM) and the second digital surface model (DSM), the in-situ measured plant height data of part of the plant individuals in the cotton field is obtained, wherein the in-situ measured plant height data is obtained through a measurement process from the ground to the top of the stem of the plant individual, the measurement time of the measurement process and the collection time of the RGB image data and the laser point cloud data differ by a time period less than a preset time period, the first plant height data, the second plant height data and the in-situ measured plant height data are trained and corrected using a plurality of different machine learning algorithms to obtain final first plant height data and final second plant height data, the target cotton group genomic variation data SNPs are filtered to obtain data with actual quality and accuracy greater than that before filtering, the best linear unbiased estimate (BLUE) is calculated based on the final first plant height data and the final second plant height data, and the best linear unbiased estimate (BLUE) and the target cotton group genomic variation data SNPs are subjected to GWAS analysis to obtain an analysis result, and plant height gene data is obtained based on the analysis result.Based on the plant height gene data, the coding sequence of the plant height related gene positioned is extracted, the plant height gene is amplified by using specific primers according to the coding sequence, the amplification fragment is obtained, the amplification fragment is inserted into a specified target vector, a target gene overexpression vector is obtained, and gene overexpression new material data is generated based on the target gene overexpression vector, the sgRNA sequence of the plant height related gene is determined based on the plant height gene data, the sgRNA sequence is inserted into a specified gene knockout vector, a target gene knockout vector is obtained, and gene editing new material data is generated based on the target gene knockout vector, in the scheme, RGB and laser radar sensors are comprehensively used to obtain cotton multi-growth period images and point cloud data, cotton population canopy boundaries are batch recognized and automatically calibrated through scripts, a plurality of machine learning models are screened to obtain the most suitable model, plant height data is associated with population genotype data, and gene overexpression and gene editing materials are verified, so that the canopy boundary calibration can be automatically completed in batches, the combination of the two sensors and the multi-model verification data is more accurate, the plant height related genes can be accurately identified by combining the multi-period phenotype data, the accuracy and efficiency are high, and the operation is simple and feasible, and the data processing efficiency of the cotton plant height gene is improved.
[0038] In order to make the above objectives, characteristics and advantages of the present application more apparent and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings without creative labor on the basis of these drawings.
[0040] Figure 1 A flowchart of the cotton data processing method based on the unmanned aerial vehicle provided by the embodiments of the present application is shown in the figure.
[0041] Figure 2 Another flowchart of the cotton data processing method based on the unmanned aerial vehicle provided by the embodiments of the present application is shown in the figure.
[0042] Figure 3 In the cotton data processing method based on the unmanned aerial vehicle provided by the embodiments of the present application, an example of material planting and unmanned aerial vehicle platform selection is shown in the figure.
[0043] Figure 4 In the cotton data processing method based on the unmanned aerial vehicle provided by the embodiments of the present application, an example of the obtained cotton canopy automatic calibration result is shown in the figure.
[0044] Figure 5An example of the result obtained by using machine learning on the RGB and lidar data in the UAV-based cotton data processing method provided in the embodiments of the present application;
[0045] Figure 6 An example of the result obtained by performing GWAS on the RGB and lidar data and the genotype data in the UAV-based cotton data processing method provided in the embodiments of the present application;
[0046] Figure 7 An example of the result obtained by performing overexpression and gene editing on the GhPH_UAV1 gene in the UAV-based cotton data processing method provided in the embodiments of the present application;
[0047] Figure 8 A structural schematic diagram of a UAV-based cotton data processing device provided in the embodiments of the present application is shown.
[0048] Figure 9 A structural schematic diagram of an electronic device provided in the embodiments of the present application is shown. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in conjunction with the accompanying drawings, obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0050] The terms "comprise" and "have" and any variations thereof mentioned in the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units is not limited to the listed steps or units, but optionally further comprises other steps or units not listed, or optionally further comprises other steps or units inherent to the process, method, product, or device.
[0051] Currently, the rapid development of genomics has improved the efficiency of crop breeding. By performing whole genome association analysis (GWAS) on multiple periods of plant height data and genotypic data, genes controlling plant height are identified and new materials with improved plant height can be quickly and accurately obtained through gene editing, etc., thereby shortening the breeding process. Because the plant height of cotton changes rapidly during growth, it is crucial to quickly obtain plant height data of a large population for GWAS analysis. Traditional field plant height data collection relies on manual measurement, which has significant drawbacks such as low throughput, high cost, and strong subjectivity. Obtaining phenotypic data through indoor phenotyping platforms improves throughput to some extent, but the high cost of these facilities and their inability to simulate field growth environments greatly limit their application. High-throughput plant height phenotypic data acquisition and automated batch analysis in the field remain a technical bottleneck.
[0052] The rapid development of unmanned aerial vehicle technology and imaging analysis technology has opened up new ways for high-throughput measurement of crop field phenotypes, providing accurate phenotypic data for GWAS analysis. Compared with traditional phenotypic analysis methods, high-throughput phenotypic analysis can dynamically observe the traits of large sample populations at different growth stages and provide more objective trait characterization based on image or point cloud data. However, there are still the following bottlenecks in using unmanned aerial vehicle platforms to acquire and analyze cotton plant height data in the field: (1) The canopy of the cotton population changes in images at different periods, and manual calibration is extremely inefficient, making it impossible to achieve automated and precise boundary calibration; (2) Plant height data obtained by a single sensor is one-sided and cannot form mutual verification of different data sets; (3) When training the plant height obtained by the unmanned aerial vehicle and the plant height measured by the human, a comprehensive machine learning model should be selected for comparison to obtain the most suitable model data; (4) The genes located by GWAS using the plant height data obtained by the unmanned aerial vehicle are not subjected to overexpression and gene editing, and cotton plant height gene overexpression and gene editing materials are not successfully obtained. Therefore, the existing technology has low efficiency in processing cotton plant height gene data, and it is currently impossible to quickly and accurately improve the target material.
[0053] Based on this, the embodiments of the present application provide a cotton data processing method and device based on an unmanned aerial vehicle and an electronic device, which can solve the technical problem of low efficiency in processing cotton plant height gene data.
[0054] The embodiments of the present application will be further described below with reference to the accompanying drawings.
[0055] Figure 1 A flowchart of a cotton data processing method based on an unmanned aerial vehicle is provided. The unmanned aerial vehicle flies above the cotton field, and the RGB sensor and the laser radar sensor are carried on the unmanned aerial vehicle. As shown in Figure 1 The method comprises:
[0056] Step S110, collecting RGB image data of the cotton field through the RGB sensor and collecting laser point cloud data of the cotton field through the laser radar sensor.
[0057] As an optional implementation, as shown in Figure 2 , a UAV can be used to carry an RGB sensor and a laser radar sensor respectively, and fly over the cotton field at a specific growth period of cotton to collect high-definition image data
[0058] For example, 419 materials were planted in the experimental field (cotton field), and the distribution of the 419 materials is shown in Figure 3 A, the UAV platform for image acquisition is Mavic 3E carrying an RGB sensor and Matrice 350RTK carrying a laser radar sensor. Both UAVs are equipped with advanced integrated real-time kinematics (RTK) GNSS modules, which can provide positioning data with a positioning accuracy of 1 cm + 1 ppm (horizontal) and 1.5 cm + 1 ppm (vertical) Figure 3 B). The UAVs flew on June 11, June 18, June 26, July 3 and July 8, 2023, during which the cotton plants were in the rapid growth stage. On the day of flight, some plant individuals were selected, and the artificial measurement of plant height data was obtained by measuring from the ground to the top of the stem using a ruler.
[0059] Step S120, aligning and stitching the RGB image data to obtain a first processing result, automatically recognizing the RGB image data in batches according to the first processing result by running a script to obtain a first recognition result, dividing a first soil part and a first background part in the RGB image data according to the first recognition result, and retaining all first pixel points of the green vegetation part in the RGB image data, creating a first group canopy range boundary for each material in the cotton group in the RGB image data according to the first pixel points, the first soil part and the first background part, constructing a digital orthographic image map according to the first group canopy range boundary and generating a first digital surface model (DSM), and obtaining first plant height data based on the first digital surface model (DSM).
[0060] For example, the 419 cotton materials formed 419 different canopy structures at each growth period. The conventional method of manual calibration is not only inefficient, but also cannot achieve accurate calibration of the canopy edge. The code can automatically recognize and batch process the image, automatically recognize and segment according to the difference in color and space between the soil and the canopy, so as to create an accurate boundary for each material in the cotton group, minimize the impact of the soil, and maximize the retention of group canopy data. The effect is shown in Figure 4 .
[0061] As shown in Figure 2The obtained RGB photos of the five growth stages are aligned and spliced. Further, software DJI Terra V4.2.2 can be used to build a digital orthographic image map (DOM) and generate a digital surface model (DSM). The difference between the first percentile and the 95th percentile is selected as the plant height value extracted from the unmanned aerial vehicle RGB image, and finally the plant height data of 419 materials at five growth stages is obtained.
[0062] As an optional implementation, a script can be used to divide the soil and cotton population canopy by color segmentation according to the different color spectra of the soil and the cotton population canopy, and the boundary of the cotton population canopy is automatically identified and calibrated according to the color spectrum. The script code uses the different color spectra of the soil and the cotton canopy to separate the soil and the canopy by color segmentation, so that the boundary of the cotton population canopy can be automatically and accurately identified and calibrated according to the color. Through this step, the plant height of each material can be calculated by subsequent RGB and laser radar data, so as to obtain accurate cotton plant height data.
[0063] The script "Boundary demarcation.py" is run through PYTHON V3.8 to automatically and batch-identify the image, complete the conversion from a color image to a gray-scale image, separate the soil and the background, retain all pixel points of the green vegetation, and thus realize the creation of accurate boundaries for each material in the cotton population, minimize the influence of the soil, and maximize the retention of the population canopy range data.
[0064] Further, the script "Boundary demarcation.py" in the embodiments of the present application is specifically as follows:
[0065] import cv2
[0066] import numpy as np
[0067] Boundary1=cv2.imread('Cotton.tiff')
[0068] Boundary2=np.array(Boundary1,dtype=np.float32) / 255.0
[0069] (b,g,r)=cv2.split(Boundary2)
[0070] Gray=2.4*g-b-r
[0071] (minVal, maxVal, minLoc, maxLoc) = cv2.minMaxLoc(Gray)
[0072] Gray_u8 = np.array((gray - minVal) / (maxVal - minVal) * 255, dtype=np.uint8)
[0073] (thresh, th) = cv2.threshold(Gray_u8, -1.0, 255, CV2.THRESH OTSU)
[0074] (b8, g8, r8) = cv2.split(Boundary1)
[0075] Color_pictures = cv2.merge([b8 & th, g8 & th, r8 & th])
[0076] cv2.imshow("Boundary1", Color_pictures)
[0077] It should be noted that the above code Boundary demarcation.py creates accurate boundaries for each material in the cotton population is programmed based on the following methods: 1, import the required library: cv2 (OpenCV library) and numpy; 2, read the image file 'RGB.tiff' taken by visible light using the `imread` function of the OpenCV library; 3, convert the read image data to a NumPy array and normalize it to 0-1 for further processing; 4, use the `split` function of the OpenCV library to split the image into B (blue), G (green), and R (red) color channels; 5, calculate the gray image. Here, a specific formula is used: `gray=2.4*g-b-r`, which is to multiply the value of the green channel by 2.4, then subtract the values of the blue and red channels to get the gray image; 6, use the `minMaxLoc` function of the OpenCV library to find the minimum and maximum values in the gray image, as well as their positions; 7, normalize the gray image to the range of 0-255 for threshold processing. Here, a specific formula is used: `(gray-minVal) / (maxVal-minVal)*255`, which is to subtract the minimum value (minVal) from the gray (gray) image, then divide by the difference between the maximum value (maxVal) and the minimum value (minVal), and finally multiply by 255 to get the gray image in the range of 0-255; 8, use the `threshold` function of the OpenCV library to perform threshold processing on the gray image. Here, the Otsu algorithm is used to automatically calculate the threshold value; 9, process the B, G, and R channels of the color image with the threshold-processed gray image to get a new color image; 10, use the `merge` function of the OpenCV library to merge the new color image into one image; 12, use the `imshow` function of the OpenCV library to display the processed image.
[0078] In one possible implementation, software DJI Terra V4.2.2 can be used to align and stitch the large number of RGB photos obtained, construct a digital orthophoto map, and generate a digital surface model (DSM).
[0079] The above process of obtaining first plant height data based on the first digital surface model DSM can specifically include the following steps: determining the difference between the upper and lower percentiles based on the first digital surface model DSM as the plant height value extracted from the RGB image data, to obtain the first plant height data.
[0080] Considering that the soil condition changes with human activities and rainfall, the difference between the upper and lower percentiles is selected as the plant height value extracted from the UAV RGB image. Generally, the minimum first and second percentiles and the maximum 95-98 percentiles are selected as the ground main stem bottom and top end height respectively, and in the embodiment of the application, the difference between the second percentile and the 98th percentile is selected as the plant height value extracted from the UAV RGB image.
[0081] In step S130, the laser point cloud data is classified and denoised to obtain a second processing result. According to the second processing result, the laser point cloud data is divided into ground point type point cloud data and non-ground point type point cloud data. According to the second processing result, the laser point cloud data is automatically batch-processed by running a script to obtain a second identification result. According to the second identification result, a second soil part and a second background part are divided in the laser point cloud data, and all second pixel points of the green vegetation part in the laser point cloud data are retained. According to the second pixel points, the second soil part and the second background part, a second population canopy range boundary is created for each material in the cotton population in the laser point cloud data. Based on the ground point type point cloud data, a digital elevation model (DEM) is generated by structure from motion (SfM) using the second population canopy range boundary. Based on the non-ground point type point cloud data, a second digital surface model (DSM) is generated by SfM using the second population canopy range boundary. Based on the difference between the digital elevation model (DEM) and the second digital surface model (DSM), second plant height data is obtained.
[0082] For example, the software DJI Terra V4.2.2 is used to coarsely classify the obtained laser point cloud data into ground points and non-ground points, as shown in FIG. 2. Figure 2 As shown in FIG. 2, the point cloud data is denoised to reduce noise in the point cloud. The ground point point cloud data is generated into a digital elevation model (DEM) by structure from motion (SfM) technology, and the non-ground point point cloud data is generated into a digital surface model (DSM). For example, the plant height data of 419 materials in 5 periods is calculated by the difference between the DEM and the DSM.
[0083] As a possible implementation, the cotton plant height obtained by the lidar is calculated by the difference between the DSM and the DEM.
[0084] In step S140, the field measurement plant height data of part of the plant individuals in the cotton field is obtained.
[0085] The in-situ measured plant height data is obtained by a measuring process from the ground to the top of the stem of the individual plant, and the measuring time of the measuring process is less than a preset time period from the acquisition time of the RGB image data and the laser point cloud data. For example, as shown in FIG. 8, on the day of the flight of the UAV, some individual plants are selected, and the artificial measured plant height data is obtained by manually measuring from the ground to the top of the stem using a ruler. For another example, the above-mentioned measuring process of the plant height is performed immediately after the acquisition of the UAV, so that the plant height does not change in a short time, thereby ensuring the consistency of the data. Figure 2
[0086] In step S150, a plurality of different machine learning algorithms are used to train and correct the first plant height data, the second plant height data and the in-situ measured plant height data, respectively, to obtain final first plant height data and final second plant height data.
[0087] The plurality of different machine learning algorithms can include any one or more of the following: random forest regression algorithm, K-nearest neighbor algorithm, Huber regression algorithm, least angle regression algorithm, light gradient boosting machine algorithm, decision tree regression algorithm, extra tree regression algorithm, orthogonal matching pursuit algorithm, Bayesian ridge regression algorithm, AdaBoost regression algorithm, CatBoost regression algorithm, and ridge regression algorithm.
[0088] In practical applications, different machine learning methods have different effects. For example, the random forest regression algorithm has a high accuracy, but the training time is long; the K-nearest neighbor algorithm has a high accuracy, but the training time is long; the Huber regression algorithm has a high accuracy, but the training time is long; the least angle regression algorithm has a high accuracy, but the training time is long; the light gradient boosting machine algorithm has a high accuracy, but the training time is long; the decision tree regression algorithm has a high accuracy, but the training time is long; the extra tree regression algorithm has a high accuracy, but the training time is long; the orthogonal matching pursuit algorithm has a high accuracy, but the training time is long; the Bayesian ridge regression algorithm has a high accuracy, but the training time is long; the AdaBoost regression algorithm has a high accuracy, but the training time is long; the CatBoost regression algorithm has a high accuracy, but the training time is long; and the ridge regression algorithm has a high accuracy, but the training time is long. Figure 2 As shown, the following 12 machine learning algorithms are used: Random Forest Regression (RFR), K-Nearest Neighbors (KNN), Huber Regression, Least Angle Regression, Light Gradient Boosting Machine (LightGBM), Decision Tree Regressor, Extra Trees Regressor, Orthogonal Matching Pursuit, Bayesian Ridge Regressor, AdaBoost Regressor, CatBoost Regressor, and Ridge Regressor, to train and validate the plant height values extracted from RGB and lidar with the manually measured plant height values. The extraction accuracy of plant height is evaluated using three indicators: coefficient of determination (R 2 ), mean absolute error (MAE), and root mean square error (RMSE). The best model result is selected as the final plant height data according to the comprehensive results for subsequent GWAS analysis.
[0089] As an optional implementation, the extraction accuracy of plant height data is detected using at least one of the coefficient of determination R 2 , mean absolute error MAE, and root mean square error RMSE using a plurality of different machine learning algorithms, and a comprehensive detection result is obtained. The best machine learning model result is selected according to the comprehensive detection result, and the first plant height data corresponding to the best machine learning model result is determined as the final first plant height data, and the second plant height data corresponding to the best machine learning model result is determined as the final second plant height data. The final first plant height data and the final second plant height data are used in the process of GWAS analysis.
[0090] The coefficient of determination R 2 is greater and the mean absolute error MAE and the root mean square error RMSE are smaller, indicating that the model accuracy of data training by the machine learning algorithm is higher. The 12 machine learning algorithms are used to train and correct the plant height data obtained by RGB and lidar with the manually measured data, to obtain the accuracy evaluation of plant height prediction, and finally generate a precision test result set. The coefficient of determination (R 2The greater the R2, the greater the average absolute error (MAE) and the root mean square error (RMSE) are, representing the better the model accuracy.
[0091] For example, 14 machine learning algorithms are used to train and validate 2095 plant height values extracted from RGB and lidar of 5 periods and manually measured plant height values. In the model construction of RGB data, it is found that the model constructed by the least angle regression algorithm has the highest accuracy, the largest R2(0.914), the smallest MAE(0.042) and RMSE(0.055) Figure 5 A). In the model construction of lidar data, it is found that the model constructed by the Huber regression algorithm has the highest accuracy, the largest R2(0.934), the smallest MAE(0.040) and RMSE(0.050) Figure 5 B). The final first plant height data and the final second plant height data obtained by the two best models will be used for subsequent GWAS analysis and verification.
[0092] In step S160, the cotton population genomic variation data SNPs in the cotton field are filtered to obtain target cotton population genomic variation data SNPs with actual quality and accuracy greater than before filtering, the best linear unbiased estimate BLUE is calculated based on the final first plant height data and the final second plant height data, and the GWAS analysis is performed on the best linear unbiased estimate BLUE and the target cotton population genomic variation data SNPs to obtain an analysis result, and the plant height gene data is obtained based on the analysis result.
[0093] It should be noted that the process of calculating the best linear unbiased estimate BLUE is a common step in GWAS analysis, that is, the phenotype data is first converted into this form to be analyzed together with the genotype.
[0094] As an optional implementation, the SNP screening condition is set to MAF>0.05 and the missing rate is <10%. The significance threshold of GWAS is set to -log10(p)=5. In this step, the cotton population genomic variation data (SNPs) are filtered to obtain high-quality SNPs (MAF>0.05, missing rate<10%). The best linear unbiased estimate (BLUE) of the cotton population plant height data obtained by the unmanned aerial vehicle platform is calculated, and the GWAS analysis is performed on the SNPs, and the significance threshold of the correlation is set to -log10(p)=5, and the plant height related genes are obtained in the significant interval.
[0095] For example, as Figure 2As shown, GWAS analysis was performed on all five time points of 419 cotton materials using laser radar-based plant height data, and five peaks on chromosomes A01 (~ 3.75-3.88 Mb), A02 (~ 43.8-44.5 Mb), A07 (~ 27.8-28.2 Mb), A10 (~ 110.7-111.1 Mb) and D11 (~ 10.8-11.0 Mb) were found to be significantly related to plant height Figure 6 A) The same results were obtained based on RGB analysis Figure 6 B). The two genes located on chromosome A01 were consistent with the cotton plant height genes GhUBP15 (Ye et al., 2023) and GhPH1 (Tian et al., 2024) identified using UAV RGB phenotype data and artificial measurement data, respectively, indicating that the results obtained by the method of the present application are comprehensive and reliable. A new gene identified from chromosome D11 may be related to cotton plant height, named GhPH_UAV1.
[0096] Step S170, based on the plant height gene data, the coding sequence of the located plant height related gene is extracted, and the plant height gene is amplified using specific primers according to the coding sequence, to obtain an amplified fragment, and the amplified fragment is inserted into a specified target vector to obtain a target gene overexpression vector, and gene overexpression new material data is generated based on the target gene overexpression vector.
[0097] For example, the coding sequence of the located plant height related gene is extracted, and specific primers are designed to amplify the gene. The amplified fragment is inserted into a 35S promoter-driven pCAMBIA-2300 vector to construct an overexpression vector, as shown in Figure 2 As shown, gene overexpression new materials are created. For example, the coding sequence of the GhPH_UAV1 gene is extracted, and specific primers are designed to amplify the gene. The primer information is shown in Table 1 below, i.e. GhPH_UAV1 gene primer information:
[0098]
[0099] Further, the transformation of the plant height related genes all uses the cotton high-efficiency genetic transformation system SAMT without genotype restriction. The cotton high-efficiency genetic transformation system SAMT is used to integrate the exogenous vector into the stem cells to induce the stem cells to produce adventitious buds, and the spectinomycin is used to inhibit the growth of the main buds to induce the generation of axillary buds. Specifically, the transformation of the plant height genes all uses the cotton high-efficiency genetic transformation system (SAMT) without genotype restriction, the stem cells of the cotton seed top meristem are used as explants, the exogenous vector is integrated into the stem cells by combining the Agrobacterium and ultrasonic treatment, and then the stem cells are induced to produce adventitious buds, and the spectinomycin is used to inhibit the growth of the main buds to induce the generation of axillary buds, thereby effectively reducing the probability of chimeras. The obtained overexpression transgenic materials and gene editing mutant materials can be stably inherited, and there is no phenomenon of sterile transgenic regenerants and no abnormal seedlings.
[0100] In step S180, the sgRNA sequence of the plant height related gene is determined based on the plant height gene data, the sgRNA sequence is inserted into a designated gene knockout vector, a target gene knockout vector is obtained, and gene editing new material data is generated based on the target gene knockout vector.
[0101] As a possible implementation, the sgRNA of the plant height related gene can be generated, and the sgRNA sequence is inserted into the CRISPR / Cas9 gene knockout vector specially developed for Gossypium hirsutum to construct a gene knockout vector, as shown in Figure 2 , to create gene editing new materials.
[0102] In order to further illustrate the role of GhPH_UAV1 in the PH of cotton, the application embodiment uses CRISPR / Cas9 technology to create a cotton material with GhPH_UAV1 gene knockout (GhPH_UAV1-Cas9) and a cotton material with GhPH_UAV1 gene overexpression (GhPH_UAV1-OE). qRT-PCR analysis shows that the expression amount of the GhPH_UAV1 gene in the GhPH_UAV1-OE material is significantly higher than that in the normal plant height cotton material ZM24, proving that the GhPH_UAV1 gene overexpression experiment is successfully implemented Figure 7 A). Further, compared with ZM24, the plant height of the GhPH_UAV1-OE material is increased by 16.48%, and the plant height of the GhPH_UAV1-Cas9 material is reduced by 21.44% Figure 7 B-C). Further, the application embodiment observes the cell development of the two materials, and compared with ZM24, the cell length of the GhPH_UAV1-OE material is significantly increased, and the cell length of the GhPH_UAV1-Cas9 material is significantly reduced Figure 7D-E). The above results directly prove that the GhPH_UAV1 gene identified by the method of the embodiment of the application has a significant role in positively regulating cotton plant height.
[0103] In the embodiment of the application, RGB and laser radar sensors are comprehensively used to obtain cotton multi-growth-period images and point cloud data, cotton population canopy boundaries are batch-recognized and automatically calibrated through scripts, the most suitable model is screened through multiple machine learning, plant height data is associated with population genotype data, and gene overexpression and gene editing materials are verified to realize that canopy boundary calibration can be automatically and batch-completed, the data are more accurate through the combination of the two sensors and multiple model verification data, the plant height related genes can be accurately identified through the combination of multi-period phenotype data, the accuracy and efficiency are high, and the operation is simple and feasible, and the data processing efficiency of cotton plant height genes is improved.
[0104] Through the use of two sensors to obtain field plant height, comparison and verification for subsequent GWAS analysis can be performed, and the phenotype data are more reliable. Moreover, canopy boundary calibration can be automatically and batch-completed through code, and the efficiency is high. Furthermore, 14 kinds of machine learning methods are used for modeling, the most suitable model is selected, and the phenotype data are more accurate. In addition, the embodiment of the application can perfectly and smoothly combine the obtaining and analysis of plant height data of a large population in the field by a UAV at multiple growth periods with the identification of genes and verification of gene functions through GWAS, accurately connect phenotype and genotype data, and to a great extent, make up for the deficiencies of time-consuming and labor-intensive manual investigation of phenotype traits in the field, subjective errors, lack of process batch analysis, and the like. The embodiment of the application can play an important role in the process of obtaining and analyzing multi-growth-period phenotype traits of a large population in the field.
[0105] The method provided by the embodiment of the application can be used as a method for revealing the genetic basis of cotton plant height through multi-sensor measurement of a UAV. The method helps researchers to no longer manually measure the plant height of a large population in the field, but only needs to operate the UAV to collect images at a specific period, and then batch-processes all population samples through the constructed plant height model to accurately obtain plant height data, which is directly used for association analysis with genotype data, to obtain plant height regulating genes, and to create new plant height directional improvement materials through gene editing, thereby providing accurate and valuable gene and material resources for modern molecular breeding.
[0106] It should be noted that, unless otherwise specified, the experimental methods used in the embodiments of the application are conventional methods. Unless otherwise specified, the materials, reagents, and the like used in the embodiments of the application can be obtained from commercial channels.
[0107] Figure 8A structural diagram of a cotton data processing device based on a UAV is provided. The UAV flies above a cotton field, and an RGB sensor and a laser radar sensor are carried on the UAV. As shown in Figure 8 The cotton data processing device based on the UAV 800 includes:
[0108] The acquisition module 801 is configured to acquire RGB image data of the cotton field through the RGB sensor and laser point cloud data of the cotton field through the laser radar sensor.
[0109] The first processing module 802 is configured to perform alignment and splicing processing on the RGB image data to obtain a first processing result, automatically batch-recognize the RGB image data according to the first processing result through a running script to obtain a first recognition result, divide a first soil part and a first background part in the RGB image data according to the first recognition result, and retain all first pixel points of a green vegetation part in the RGB image data, create a first group canopy range boundary for each material in a cotton group in the RGB image data according to the first pixel points, the first soil part and the first background part, construct a digital orthographic image map and generate a first digital surface model (DSM) according to the first group canopy range boundary, and obtain first plant height data based on the first digital surface model (DSM).
[0110] The second processing module 803 is configured to perform classification and denoising processing on the laser point cloud data to obtain a second processing result, divide the laser point cloud data into ground point type point cloud data and non-ground point type point cloud data according to the second processing result, automatically batch-recognize the laser point cloud data according to the second processing result through a running script to obtain a second recognition result, divide a second soil part and a second background part in the laser point cloud data according to the second recognition result, and retain all second pixel points of a green vegetation part in the laser point cloud data, create a second group canopy range boundary for each material in a cotton group in the laser point cloud data according to the second pixel points, the second soil part and the second background part, generate a digital elevation model (DEM) through structure from motion (SfM) based on the ground point type point cloud data and the second group canopy range boundary, generate a second digital surface model (DSM) through SfM based on the non-ground point type point cloud data and the second group canopy range boundary, and obtain second plant height data based on the digital elevation model (DEM) and the second digital surface model (DSM).
[0111] The acquisition module 804 is configured to acquire field measurement plant height data of part of plant individuals in the cotton field; wherein the field measurement plant height data is obtained through a measurement process from the ground to the top of the stem of the plant individual, and a measurement time of the measurement process is less than a preset time period from a collection time of the RGB image data and the laser point cloud data.
[0112] The training module 805 is configured to train and verify the first plant height data, the second plant height data and the field measurement plant height data respectively by using a plurality of different machine learning algorithms, to obtain final first plant height data and final second plant height data.
[0113] The analysis module 806 is configured to filter cotton population genomic variation data SNPs in the cotton field to obtain target cotton population genomic variation data SNPs with actual quality and accuracy greater than those before filtering, calculate a best linear unbiased estimate BLUE based on the final first plant height data and the final second plant height data, and perform GWAS analysis on the best linear unbiased estimate BLUE and the target cotton population genomic variation data SNPs to obtain an analysis result, and obtain plant height gene data based on the analysis result.
[0114] The first construction module 807 is configured to extract a coding sequence of a plant height related gene located based on the plant height gene data, amplify the plant height gene by using specific primers according to the coding sequence to obtain an amplified fragment, and insert the amplified fragment into a specified target vector to obtain a target gene overexpression vector, and generate gene overexpression new material data based on the target gene overexpression vector.
[0115] The second construction module 808 is configured to determine an sgRNA sequence of a plant height related gene based on the plant height gene data, insert the sgRNA sequence into a specified gene knockout vector to obtain a target gene knockout vector, and generate gene editing new material data based on the target gene knockout vector.
[0116] The cotton data processing device based on the unmanned aerial vehicle provided in the embodiments of the present application has the same technical features as the cotton data processing method based on the unmanned aerial vehicle provided in the above embodiments, and can solve the same technical problems and achieve the same technical effects.
[0117] The electronic device provided in the embodiments of the present application, as shown in Figure 9 The electronic device 900 includes a processor 902 and a memory 901, the memory stores a computer program executable on the processor, and the processor executes the computer program to implement the steps of the method provided in the above embodiments.
[0118] Referring toFigure 9 The electronic device further includes a bus 903 and a communication interface 904, the processor 902, the communication interface 904 and the memory 901 are connected through the bus 903; the processor 902 is used to execute the executable modules stored in the memory 901, such as computer programs.
[0119] The memory 901 can contain a high-speed random access memory (RAM) and can also include a non-volatile memory such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 904 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0120] The bus 903 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 Only one bidirectional arrow is used in the above description, but it does not mean that there is only one bus or only one type of bus.
[0121] The memory 901 is used to store programs, and the processor 902 executes the programs after receiving execution instructions. The method executed by the device defined by the process disclosed in any embodiment of the present application can be applied to the processor 902 or realized by the processor 902.
[0122] The processor 902 can be an integrated circuit chip having a processing capability for signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 902 or the instruction in the form of software. The processor 902 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 901, and the processor 902 reads the information in the memory 901, and combines the hardware to complete the steps of the above method.
[0123] Corresponding to the above unmanned aerial vehicle based cotton data processing method, the embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores computer executable instructions, when the processor calls and runs the computer executable instructions, the computer executable instructions make the processor run the steps of the above unmanned aerial vehicle based cotton data processing method.
[0124] The unmanned aerial vehicle based cotton data processing device provided by the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided by the embodiment of the present application has the same implementation principle and technical effect as the foregoing method embodiments. For the sake of brevity, the part of the device embodiment not mentioned in the foregoing method embodiments can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can be referred to the corresponding process in the foregoing method embodiments, which will not be described here.
[0125] In the embodiments of the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. The embodiments described above are merely specific implementation manners of the present application, and for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, communication interfaces, or a combination of other forms, which can be electric, mechanical, or in other forms.
[0126] For another example, the flowcharts and block diagrams in the drawings show the possible implementation architectures, functions and operations of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders from that shown in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0127] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0128] In addition, each functional unit in the embodiments of the present application can be integrated into one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated into one unit.
[0129] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the unmanned aerial vehicle-based cotton data processing method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0130] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings, in addition, the terms "first", "second", "third" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0131] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit them, the protection scope of the present application is not limited thereto, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features thereof; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A cotton data processing method based on unmanned aerial vehicles (UAVs), characterized in that, The drone flies over the cotton field and is equipped with an RGB sensor and a lidar sensor; the method includes: The RGB image data of the cotton field is acquired by the RGB sensor, and the laser point cloud data of the cotton field is acquired by the lidar sensor. The RGB image data is aligned and stitched to obtain a first processing result. Based on the first processing result, the RGB image data is automatically batch-recognized by running a script to obtain a first recognition result. Based on the first recognition result, the first soil part and the first background part in the RGB image data are divided, and all first pixels of the green vegetation part in the RGB image data are retained. Based on the first pixels, the first soil part and the first background part, a first population canopy range boundary is created for each material in the cotton population in the RGB image data. A digital orthophoto map is constructed based on the first population canopy range boundary and a first digital surface model (DSM) is generated. Based on the first digital surface model (DSM), the first plant height data is obtained. The laser point cloud data is classified and denoised to obtain a second processing result. Based on the second processing result, the laser point cloud data is divided into ground point type point cloud data and non-ground point type point cloud data. Based on the second processing result, the laser point cloud data is automatically batch-identified by running a script to obtain a second identification result. Based on the second identification result, the laser point cloud data is divided into a second soil part and a second background part, and all second pixels of the green vegetation part in the laser point cloud data are retained. Based on the second pixels, the second soil part, and the second background part, a second population canopy range boundary is created for each material in the cotton population in the laser point cloud data. Based on the ground point type point cloud data, a digital elevation model (DEM) is generated using the second population canopy range boundary through structural motion (SfM). Based on the non-ground point type point cloud data, a second digital surface model (DSM) is generated using the second population canopy range boundary through SfM. Based on the difference between the digital elevation model (DEM) and the second digital surface model (DSM), the second plant height data is obtained. The field measurement data of plant height of some individual plants in the cotton field is obtained; wherein the field measurement data of plant height is obtained by measuring from the ground to the top of the stem of the individual plant, and the time difference between the measurement time of the measurement process and the acquisition time of the RGB image data and the laser point cloud data is less than a preset time period. The first plant height data, the second plant height data, and the field-measured plant height data are trained and corrected using various different machine learning algorithms to obtain the final first plant height data and the final second plant height data. The SNPs of genomic variation data of cotton population in the cotton field are filtered to obtain SNPs of genomic variation data of target cotton population with actual quality and accuracy greater than those before filtering. The best linear unbiased estimator BLUE is calculated based on the final first plant height data and the final second plant height data. GWAS analysis is performed on the best linear unbiased estimator BLUE and the SNPs of genomic variation data of target cotton population to obtain the analysis results. Based on the analysis results, plant height gene data is obtained. Based on the plant height gene data, the coding sequence of the plant height-related gene is extracted and located. The plant height gene is amplified using specific primers according to the coding sequence to obtain an amplified fragment. The amplified fragment is inserted into a specified target vector to obtain a target gene overexpression vector. New gene overexpression material data is generated based on the target gene overexpression vector. Based on the plant height gene data, the sgRNA sequence of the plant height-related gene is determined, the sgRNA sequence is inserted into the specified gene knockout vector to obtain the target gene knockout vector, and new gene editing material data is generated based on the target gene knockout vector.
2. The method according to claim 1, characterized in that, The process of using multiple different machine learning algorithms to train and correct the first plant height data, the second plant height data, and the field-measured plant height data to obtain the final first plant height data and the final second plant height data includes: Using various machine learning algorithms to obtain the coefficient of determination R 2 The extraction accuracy of the plant height data is tested by at least several indicators, including mean absolute error (MAE) and root mean square error (RMSE), to obtain a comprehensive test result. The optimal machine learning model result is selected based on the comprehensive detection results, and the first plant height data corresponding to the optimal machine learning model result is determined as the final first plant height data, and the second plant height data corresponding to the optimal machine learning model result is determined as the final second plant height data; wherein, the final first plant height data and the final second plant height data are used in the GWAS analysis process.
3. The method according to claim 2, characterized in that, The determination coefficient R 2 The larger the value and the smaller the mean absolute error (MAE) and root mean square error (RMSE), the higher the accuracy of the model trained on the data using the machine learning algorithm.
4. The method according to claim 1, characterized in that, The transformation of the plant height-related genes all utilized the genotype-free, high-efficiency cotton genetic transformation system SAMT. The high-efficiency cotton genetic transformation system SAMT uses cotton seed apical meristem stem cells as explants, combined with Agrobacterium and ultrasonic treatment, to integrate exogenous vectors into stem cells to induce the stem cells to produce adventitious buds, and uses zizomycin to inhibit the growth of the main bud to induce the production of axillary buds.
5. The method according to claim 1, characterized in that, The script is used to divide the soil and cotton canopy by color segmentation using the different color spectra of the soil and cotton canopy. The boundaries of the cotton canopy are automatically identified and marked based on the color spectra.
6. The method according to claim 1, characterized in that, The process of obtaining the first plant height data based on the first digital surface model (DSM) includes: Based on the first digital land surface model (DSM), the difference between the upper percentile and the lower percentile is determined as the plant height value extracted from the RGB image data, thus obtaining the first plant height data.
7. The method according to claim 1, characterized in that, The various machine learning algorithms include any of the following: Random Forest Regression Algorithm, K-Nearest Neighbors Algorithm, Huber Regression Algorithm, Least Angle Regression Algorithm, Lightweight Gradient Boosting Machine Algorithm, Decision Tree Regression Algorithm, Extra Tree Regression Algorithm, Orthogonal Matching Pursuit Algorithm, Bayesian Ridge Regression Algorithm, AdaBoost Regression Algorithm, CatBoost Regression Algorithm, and Ridge Regression Algorithm.
8. A cotton data processing device based on unmanned aerial vehicles (UAVs), characterized in that, The drone flies over the cotton field and is equipped with an RGB sensor and a lidar sensor; the device includes: The acquisition module is used to acquire RGB image data of the cotton field through the RGB sensor and acquire laser point cloud data of the cotton field through the lidar sensor. The first processing module is used to align and stitch the RGB image data to obtain a first processing result. Based on the first processing result, the RGB image data is automatically batch-recognized by running a script to obtain a first recognition result. Based on the first recognition result, the first soil part and the first background part in the RGB image data are divided, and all first pixels of the green vegetation part in the RGB image data are retained. Based on the first pixels, the first soil part and the first background part, a first population canopy range boundary is created for each material in the cotton population in the RGB image data. Based on the first population canopy range boundary, a digital orthophoto map is constructed and a first digital surface model (DSM) is generated. Based on the first digital surface model (DSM), the first plant height data is obtained. The second processing module is used to classify and denoise the laser point cloud data to obtain a second processing result. Based on the second processing result, the laser point cloud data is divided into ground point type point cloud data and non-ground point type point cloud data. Based on the second processing result, the laser point cloud data is automatically batch-identified by running a script to obtain a second identification result. Based on the second identification result, the laser point cloud data is divided into a second soil part and a second background part, and all second pixels of the green vegetation part in the laser point cloud data are retained. Based on the second pixels, the second soil part, and the second background part, a second population canopy range boundary is created for each material in the cotton population in the laser point cloud data. Based on the ground point type point cloud data, a digital elevation model (DEM) is generated using the second population canopy range boundary through structural motion (SfM). Based on the non-ground point type point cloud data, a second digital surface model (DSM) is generated using the second population canopy range boundary through SfM. Based on the digital elevation model (DEM) and the second digital surface model (DSM), a second plant height data is obtained. The acquisition module is used to acquire the actual measured plant height data of some individual plants in the cotton field; wherein, the actual measured plant height data is obtained by measuring from the ground to the top of the stem of the individual plant, and the time difference between the measurement time of the measurement process and the acquisition time of the RGB image data and the laser point cloud data is less than a preset time period. The training module is used to train and validate the first plant height data, the second plant height data, and the field-measured plant height data using a variety of different machine learning algorithms, so as to obtain the final first plant height data and the final second plant height data. The analysis module is used to filter the SNPs of the cotton population genome variation data in the cotton field to obtain the target cotton population genome variation data SNPs whose actual quality and accuracy are greater than those before filtering. Based on the final first plant height data and the final second plant height data, the optimal linear unbiased estimator BLUE is calculated, and GWAS analysis is performed on the optimal linear unbiased estimator BLUE and the target cotton population genome variation data SNPs to obtain the analysis results. Based on the analysis results, the plant height gene data is obtained. The first construction module is used to extract the coding sequence of the plant height-related gene located based on the plant height gene data, amplify the plant height gene using specific primers according to the coding sequence to obtain an amplified fragment, insert the amplified fragment into a specified target vector to obtain a target gene overexpression vector, and generate gene overexpression new material data based on the target gene overexpression vector. The second construction module is used to determine the sgRNA sequence of the plant height-related gene based on the plant height gene data, insert the sgRNA sequence into the specified gene knockout vector to obtain the target gene knockout vector, and generate gene editing new material data based on the target gene knockout vector.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Wide-ridge crop plant height and row spacing automatic extraction method based on laser radar
CN113487636A
Flue-cured tobacco plant height measuring method and device based on unmanned aerial vehicle cooperation and electronic equipment
CN116563362A