A multi-source remote sensing data fusion method based on linear regression feature augmentation

By combining multispectral, RGB, thermal infrared, and 3D structure features through a linear regression-based feature amplification method, new training features are generated, which solves the problem of feature imbalance in drone multi-sensor data fusion and achieves more accurate plant phenotypic evaluation and more comprehensive data analysis.

CN116740517BActive Publication Date: 2025-10-24CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310698170.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-10-24
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

Existing UAV multi-sensor data fusion methods are single and ignore the correlation and differences between different sensors, resulting in feature imbalance or redundancy, failing to fully utilize the advantages of sensors and achieving comprehensive and accurate plant phenotypic evaluation.

Method used

A feature amplification method based on linear regression is adopted to generate new training features through multiple linear regression and cross-validation, enhance data diversity and model performance, and combine the features of multispectral, RGB, thermal infrared and 3D structure for fusion.

Benefits of technology

It improves the accuracy of plant phenotypic assessment and the robustness of the model, enables more comprehensive data fusion and analysis, and enhances the computational efficiency and noise resistance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740517B_ABST
    Figure CN116740517B_ABST
Patent Text Reader

Abstract

The application discloses a multi-source remote sensing data fusion method based on linear regression feature expansion, and belongs to the field of remote sensing data fusion.The application randomly fuses different types of features from multispectral, RGB, thermal infrared and 3D structure by a linear regression and cross-validation method and generates new features, obtains a large number of feature sets combining the information of various sensors after multiple cycles, and effectively improves the performance and robustness of a training model. Experiments show that the method can obtain higher plant phenotype evaluation accuracy than traditional data fusion methods. The method can improve the expression capacity of data and the performance of a model, can realize more comprehensive remote sensing data fusion and analysis, has the characteristics of small calculation amount, high model precision and strong noise resistance, and can provide more high-quality data support and decision analysis for quantitative remote sensing application fields.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of remote sensing data fusion, and particularly relates to a multi-source remote sensing data fusion method based on linear regression feature expansion. BACKGROUND

[0002] Due to the continuous increase in population, changes in human dietary structure, and the extensive use of biofuels in various industries, the global demand for food production has increased significantly. In order to meet the needs of human consumption on the current resource basis, it is necessary to cultivate high-yield crop varieties. With the rapid development of genomics and high-throughput sequencing, traditional phenotyping methods have been unable to meet the needs of today's breeding, and accelerating high-throughput phenomics research is particularly important. By using remote sensing, computer technology can quickly and accurately analyze plant phenotypes, which can help researchers better understand complex genetic traits and accelerate the progress of breeding and precision agriculture, and has become an important research field in agronomy.

[0003] In recent years, unmanned aerial vehicle remote sensing technology has been increasingly applied in plant phenotype evaluation. Compared with satellite remote sensing, unmanned aerial vehicles can carry high-resolution imaging sensors and obtain more detailed image data than satellite remote sensing, enabling more accurate feature extraction and classification. In addition, unmanned aerial vehicles can flexibly adjust flight routes, altitudes, and angles according to needs to obtain more abundant data information. Unmanned aerial vehicles can quickly take off for remote sensing data collection when needed, avoiding satellite data delays and the influence of weather and other factors, enabling immediate data collection and processing. In plant phenotype evaluation, high-resolution vegetation image data obtained by remote sensing sensors carried by unmanned aerial vehicles can be used to monitor and analyze vegetation coverage, chlorophyll content, and growth status, thereby enabling rapid, efficient, and accurate extraction and analysis of plant phenotype characteristics.

[0004] The commonly used sensors in the study of plant phenotyping using unmanned aerial vehicle multi-sensor data fusion include hyperspectral imaging sensors, multispectral imaging sensors, thermal infrared sensors, RGB sensors, and three-dimensional laser radars, etc. These sensors can obtain different types of data, such as vegetation index, plant height, canopy structure, temperature, and terrain information. By fusing these data, a comprehensive analysis of the growth status of crops can be achieved, thereby more accurately estimating the yield of crops. Through unmanned aerial vehicle multi-sensor data fusion, plant phenotyping information can be more comprehensively and quickly obtained, and support can be provided for agricultural production, ecological research, etc. However, the current fusion method of different sensor features is single, and the features extracted by different sensors are simply stacked. Although this is a simple method, it ignores the correlation and difference between different sensors. The features of different sensors may have different scales, dynamic ranges, and statistical distributions, and simple stacking may lead to unbalanced or redundant features. In addition, simple stacking does not consider the complementarity and weight distribution between different sensors, and may not fully utilize the advantages of different sensors. In order to overcome these limitations, more advanced multi-source remote sensing data fusion methods need to be explored. SUMMARY

[0005] The present application is aimed at the problem of single data fusion method, and provides a multi-source remote sensing data fusion method based on linear regression feature expansion for crop phenotyping evaluation. This method can improve the expression ability of data and the performance of the model, and at the same time can realize more comprehensive remote sensing data fusion and analysis, with the characteristics of small amount of calculation, high model precision, strong anti-noise ability, etc., which can provide better data support and decision analysis for quantitative remote sensing application field.

[0006] To achieve the above purpose, the technical solutions adopted by the present application are as follows:

[0007] A multi-source remote sensing data fusion method based on linear regression feature expansion, the method comprises the following steps:

[0008] Step 1: Data acquisition and processing: using unmanned aerial vehicle to obtain different source orthophoto data or laser radar data, in addition, using unmanned aerial vehicle to carry camera to obtain plant images at different angles on crop canopy, and reconstructing 3D structure of corn canopy; in this way, multi-source and multi-angle data can be obtained, and the robustness of subsequent model can be enhanced.

[0009] Step 2: Feature extraction: according to the characteristics of different sensors, spectral features, color features, structural features and texture features are extracted from various images; by obtaining different source remote sensing data and plant image data, various information sources can be comprehensively utilized, the richness and diversity of data can be improved, and the growth status of plants can be more comprehensively described.

[0010] Step 3: Feature expansion

[0011] Feature augmentation refers to a method of generating new training features to augment the training dataset by a series of transformations on the original data. Feature augmentation can increase the diversity of data without increasing the acquisition of new data, thus improving the performance of machine learning models. The feature augmentation method based on linear regression and the validation process proposed in this study are as follows:

[0012] (1) Divide the dataset into training and test sets using a certain ratio;

[0013] (2) On the training set, randomly select i features from m different sources of feature sets, a total of mi features, and select the same features on the test set;

[0014] (3) Combine the mi features on the training set with the plant trait value, denoted as D; then use multiple linear regression (MLR) to perform 10-fold cross-validation on D; the out-of-sample prediction results generated during cross-validation are used as new training features; train the MLR model using all the data of D, and test it on the mi features of the test set; the output prediction results are used as new test features;

[0015] (4) Repeat step (3) n times to generate n new features in the training and test datasets; since the mi features selected in each iteration are different, the features generated between different iterations are independent of each other; combine these n new features with all the original features in the training and test datasets to form a new training set and a new test set. This method combines the original features with the plant trait value through linear regression and cross-validation to generate new training and test features, increasing the diversity and quantity of data.

[0016] Further, the step one is specifically: first, use a UAV to obtain multispectral, RGB, and thermal infrared orthophotos; in addition, use a UAV equipped with an RGB camera and cross-over circular oblique photography technology to obtain plant images at different angles 4 meters above the crop canopy; use the images obtained by cross-over circular oblique photography to reconstruct the 3D structure of the corn canopy in industrial software. This step is to obtain multiple types of data according to the growth characteristics of plants to enhance the robustness of the subsequent model construction.

[0017] Further, in step one, during the orthophoto acquisition process, the flight height of each sensor is set to 20-50 meters, the heading and lateral overlap is 70%-85%, and 5-30 ground control points are evenly distributed in the entire data acquisition area, and a differential global navigation satellite system is used for measurement to further improve the accuracy of image processing.

[0018] Further, in step one, in the cross-circumferential oblique photography, 90% in-circle and 60% out-of-circle overlap are set. Setting higher in-circle and out-of-circle overlap can ensure more overlapping areas between adjacent images, thereby improving the coverage of data acquisition. This is very important for obtaining comprehensive plant canopy information, especially in the process of three-dimensional reconstruction and feature extraction.

[0019] Further, in step one, the processing of orthophoto is carried out in industrial software, including image geolocation, import of ground control point coordinates, image alignment, establishment of dense point cloud, establishment of digital surface model and orthophoto, and radiation correction.

[0020] Further, in step two, according to the characteristics of each sensor, the single-band pixel brightness value, coverage, vegetation index, plant height and texture feature are extracted from the RGB image; the single-band reflectivity, vegetation index and texture feature are extracted from the multi-spectral image; the normalized canopy temperature and texture feature are extracted from the thermal infrared image; and the plant height, coverage and plant area index are extracted from the 3D structure.

[0021] Further, in step three, considering the number of characteristics of each sensor, i is set to 5 and n is set to 200 to verify the data fusion effect of the feature amplification method. The number of each type of feature can be adjusted.

[0022] Further, the method further comprises step four: model training and verification: the amplified training features and the original feature set are respectively input into the machine learning regression algorithm, the accuracy is tested by 5-fold cross-validation, and the plant trait evaluation accuracy of the amplified feature set and the original feature set is compared.

[0023] The beneficial effects of the present application relative to the prior art are: the present application randomly fuses different types of features from multi-spectral, RGB, thermal infrared and 3D structure by linear regression and cross-validation method and generates new features, and a large number of feature sets combining information of multiple sensors are obtained after multiple cycles, which effectively improves the performance and robustness of the training model. Experiments show that this method can obtain higher plant phenotype evaluation accuracy than traditional data fusion methods. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 is the overall flowchart of the present application;

[0025] Figure 2 is a 10-fold cross-validation diagram in the feature amplification process in the present application;

[0026] Figure 3 is a precision comparison diagram of the traditional data fusion method and the feature amplification method when evaluating the corn leaf area index in the large bell period.

[0027] Figure 4 Figure for accuracy comparison between traditional data fusion method and feature augmentation method when evaluating corn leaf area index at milk stage;

[0028] Figure 5 Figure for accuracy comparison between traditional data fusion method and feature augmentation method when evaluating corn biomass at tasseling stage;

[0029] Figure 6 Figure for accuracy comparison between traditional data fusion method and feature augmentation method when evaluating corn biomass at milk stage. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the protection scope of the present application.

[0031] Example 1:

[0032] Step 1, use unmanned aerial vehicle vertical photography to obtain multispectral, RGB and thermal infrared 30-meter images; in addition, use unmanned aerial vehicle equipped with RGB camera and cross-ring tilt photography technology to obtain plant images at different angles above the crop canopy at 4 meters and generate canopy point cloud. On the day of unmanned aerial vehicle flight task execution, ground leaf area index and biomass data are measured.

[0033] Step 2, use industrial software to process images of each sensor to obtain orthographic images of multispectral, RGB and thermal infrared sensors, and obtain the canopy 3D structure reconstructed from cross-ring tilt photography measurement data.

[0034] Step 3, according to the characteristics of each sensor, single-band pixel brightness value, coverage, vegetation index, plant height, texture features are extracted from RGB images; single-band reflectivity, vegetation index, texture features are extracted from multispectral images; normalized canopy temperature, texture features are extracted from thermal infrared images; plant height, coverage, plant area index, canopy volume are extracted from 3D structure.

[0035] (1) Single-band reflectivity refers to the ability of an object's surface to reflect incident light within a specific wavelength range. It is a unitless ratio that represents the proportional relationship between the light intensity reflected by the object's surface and the incident light intensity.

[0036] (2) Pixel brightness value represents the brightness intensity of each pixel in the image, which can be used to describe the overall brightness distribution of the image.

[0037] (3) Coverage refers to the proportion of pixels occupied by a specific object or feature in an image. It can be used to quantify the distribution range and density of a certain object or feature in the image.

[0038] (4) Vegetation index is an index used to evaluate the degree of vegetation growth by comparing the reflectivity of light at different wavelengths in vegetation and non-vegetation areas. Commonly used vegetation indices include the normalized difference vegetation index and the green vegetation index.

[0039] (5) Plant height refers to the height of a plant, which is obtained through image analysis or other measurement methods. In remote sensing images, specific algorithms and techniques can be used to estimate plant height.

[0040] (6) Texture features describe the texture or surface features of different regions in an image. They are calculated based on the degree of change between the gray values or color distributions of pixels. Texture features can be used for image classification, object recognition, and image segmentation tasks to provide detailed information about different regions in the image. Commonly used texture features include the gray level co-occurrence matrix and the local binary pattern.

[0041] (7) Normalized difference vegetation index (NDVI) is an index used to evaluate the degree of vegetation growth by comparing the reflectivity of light at different wavelengths in vegetation and non-vegetation areas. Commonly used vegetation indices include the normalized difference vegetation index and the green vegetation index.

[0042] (8) Plant area index (PAI) refers to the total leaf area of a plant that intersects with a unit of horizontal surface area in the vertical direction. It is a parameter that describes the plant community and can be used to evaluate the density and growth status of vegetation.

[0043] (9) Crown volume refers to the three-dimensional space volume occupied by the plant canopy. It represents the distribution range and volume size of the plant community in the vertical and horizontal directions.

[0044] Step 4,

[0045] (1) Divide the data into training set and test set in the ratio of 4:1;

[0046] (2) In the training set, randomly select 5 features from the multispectral, RGB, thermal infrared, and 3D feature sets, respectively, for a total of 20 features, and select the same features in the test set;

[0047] (3) Combine the 20 features in the training set with the plant trait values column by column, denoted as D; then use multivariate linear regression to perform 10-fold cross-validation on D (see Figure 1 for 10-fold cross-validation schematic); the out-of-sample prediction results generated during the cross-validation process are used as new training features; use all the data of D to train the MLR model, and test it on the 20 features of the test set; the output prediction results are used as new test features; Figure 1 ​

[0048] (4) Step (3) is repeated 200 times, and 200 new features are generated in the training and test data sets; since the 20 features selected in each iteration are different, the features generated between different iterations are independent of each other; these 400 new features are combined with all the original features in the training and test data sets by column to form a new training set and a new test set. For specific procedures, refer to Figure 2 .

[0049] Step 5, input the new training set generated in step 4 into random forest, lasso regression and K nearest neighbor regression to construct models, and verify the model performance on the new test set. 5 independent verifications are used to test the feature enhancement effect.

[0050] When evaluating the leaf area index of corn at the large horn stage using multi-source data fusion, the average R 2 of each algorithm of the conventional data fusion method is 0.70, and the average R 2 of each algorithm of the feature augmentation method is 0.76. Figure 3 )

[0051] When evaluating the leaf area index of corn at the milk stage using multi-source data fusion, the average R 2 of each algorithm of the conventional data fusion method is 0.33, and the average R 2 of each algorithm of the feature augmentation method is 0.42. Figure 4 )

[0052] When evaluating the biomass of corn at the large horn stage using multi-source data fusion, the average R 2 of each algorithm of the conventional data fusion method is 0.60, and the average R 2 of each algorithm of the feature augmentation method is 0.65. Figure 5 )

[0053] When evaluating the biomass of corn at the milk stage using multi-source data fusion, the average R 2 of each algorithm of the conventional data fusion method is 0.51, and the average R 2 of each algorithm of the feature augmentation method is 0.58. Figure 6 )

Claims

1. A multi-source remote sensing data fusion method based on linear regression-based feature augmentation, characterized in that: The method is: Step one: data acquisition and processing: using unmanned aerial vehicles to obtain different sources of orthographic image data or laser radar data, in addition, using unmanned aerial vehicles to carry cameras to obtain different angle plant images on the crop canopy, and reconstructing the 3D structure of the corn canopy; Step two: feature extraction: according to the characteristics of different sensors, the spectral features, color features, structural features and texture features are extracted from various images; According to the characteristics of each sensor, the single band pixel brightness value, coverage, vegetation index, plant height and texture feature are extracted from the RGB image; The single band reflectivity, vegetation index and texture feature are extracted from the multispectral image; The normalized canopy temperature and texture feature are extracted from the thermal infrared image; The plant height, coverage, plant area index and canopy volume are extracted from the 3D structure; Step three: feature amplification (1) a certain proportion is used to divide the data set into training set and test set; (2) On the training set, randomly select m features from each of i different source feature sets, for a total of mi features, and select the same features on the test set. ( 3) The training set mi The characteristics are combined with the plant trait value and expressed as D ; Then use multiple linear regression to perform 10-fold cross validation on D; The out-of-sample prediction results generated in the cross-validation process are used as new training features; Using D All data of the training set is used to train the MLR model and tested on the test set of mi The output prediction is used as a new test feature; (4) Step (3) is repeated n Next, new features are generated in the training and test data sets n one new feature; Since the features selected in each iteration mi are different, the resulting features from each iteration are independent of one another; this n new feature is combined with all the original features in the training and test datasets to form a new training set and a new test set.

2. The multi-source remote sensing data fusion method based on linear regression feature augmentation of claim 1, characterized in that: The step one is specifically: first, using unmanned aerial vehicles to obtain multispectral, RGB and thermal infrared orthographic images; In addition, using unmanned aerial vehicles carrying RGB cameras and cross-ring tilt photography technology to obtain plant images at different angles above the crop canopy; The 3D structure of the corn canopy is reconstructed in the industrial software by using the images obtained by cross-ring tilt photography.

3. The multi-source remote sensing data fusion method based on linear regression feature augmentation of claim 2, characterized in that: In step one, during the orthographic image acquisition process, the flight height of each sensor is set to 20-50 meters, the heading and lateral overlap is 70%-85%, and 5-30 ground control points are evenly distributed in the entire data acquisition area, and the differential global navigation satellite system is used for measurement.

4. The multi-source remote sensing data fusion method based on linear regression feature augmentation according to claim 2 or 3, characterized in that: In step one, in the cross-ring tilt photography, the in-circle overlap is set to 90% and the out-of-circle overlap is set to 60%.

5. The multi-source remote sensing data fusion method based on linear regression feature augmentation of claim 2, wherein: In step one, the processing of orthographic images is carried out in industrial software, including image geolocation, importing ground control point coordinates, aligning images, establishing dense point cloud, establishing digital surface model and establishing orthographic image, and radiation correction.

6. The multi-source remote sensing data fusion method based on linear regression feature augmentation of claim 1, wherein: In step three, considering the number of features of each sensor, set i 5, n 200 to verify the data fusion effect of the feature expansion method.

7. The multi-source remote sensing data fusion method based on linear regression feature augmentation of claim 1, 2, 3, 5 or 6, characterized in that: The method further comprises step four: model training and verification: input the amplified training features and the original feature set into the machine learning regression algorithm respectively, test its accuracy by 5-fold cross-validation, and compare the plant trait evaluation accuracy of the amplified feature set and the original feature set.

Citation Information

Patent Citations

  • Video spatio-temporal feature optimization method and system based on quality attention mechanism, electronic equipment and storage medium

    CN115243031A

  • Corn aboveground biomass estimation method and system

    CN116187478A