Sugarcane growth stage monitoring method and system
Through the multimodal fusion monitoring model combined with sugarcane images, point clouds and environmental data, the problem of low accuracy of sugarcane growth stage monitoring is solved, and efficient and accurate monitoring of sugarcane growth stage is achieved.
Patent Information
- Application Number
- CN202510381823.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The accuracy of monitoring of sugarcane growth stages is low, resulting in waste of sugarcane resources and missing the best agricultural operation window.
A multimodal fusion monitoring model is adopted, and the image data, point cloud data and environmental data of sugarcane are collected, and the multimodal fusion monitoring model is input after preprocessing. The appearance characteristics, morphological characteristics, physiological characteristics and environmental characteristics of sugarcane are comprehensively analyzed to determine the growth stage of sugarcane.
It significantly improves the efficiency and accuracy of sugarcane growth stage monitoring, reduces the problem of poor subjectivity and timeliness of manual monitoring, and provides more accurate growth stage information.
Smart Images

Figure CN120279323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural planting, and particularly relates to a method and system for monitoring the growth stages of sugarcane. Background Art
[0002] Sugarcane is an important cash crop with extremely wide uses. Sugarcane can be used as a renewable energy alternative to produce bioethanol, reducing the consumption of traditional fossil fuels. The cellulose extracted from sugarcane can also be used in industries such as papermaking and textile. Moreover, the stalks of sugarcane can be used as the core raw material for the sugar industry.
[0003] With the increasing demand for agricultural intelligence, the refined management of sugarcane planting has become a key direction for industrial upgrading. Among them, the accurate monitoring of the growth stages of sugarcane has significant economic value for agricultural production. Specifically, sugarcane has significant differences in water and fertilizer requirements, environmental adaptability, and pest and disease risks at different growth stages. For example: during the tillering stage, the supply of nitrogen fertilizer needs to be strengthened, while during the mature stage, water control is required to increase sugar accumulation. Therefore, in order to ensure the yield of sugarcane, it is necessary to monitor the growth stages of sugarcane in real time to obtain the growth information of sugarcane at each growth stage.
[0004] Traditional methods for monitoring the growth stages of sugarcane mainly rely on manual monitoring, that is, by manually observing and collecting the growth data of sugarcane, and identifying the growth stages of sugarcane based on professional knowledge and empirical knowledge. However, these knowledges require a long learning time and practice. Moreover, the methods relying on manual monitoring have problems of strong subjectivity and poor timeliness, which are likely to lead to waste of sugarcane resources or missing the best window for agricultural operations. In this way, it results in slow efficiency in monitoring the growth stages of sugarcane, being time-consuming and laborious, and there will also be large errors in the monitoring results, affecting the accuracy of monitoring the growth stages of sugarcane. Summary of the Invention
[0005] The technical problem to be solved by the present invention is the low accuracy of monitoring the growth stages of sugarcane.
[0006] To solve the above technical problem, the present invention provides a method and system for monitoring the growth stages of sugarcane. The specific technical solutions are as follows:
[0007] In a first aspect, the present invention provides a method for monitoring the growth stage of sugarcane, including: First, collect first environmental data and first hyperspectral image data of a target area, as well as first image data and first three-dimensional point cloud data of the sugarcane planted in the target area. Then, perform a first preprocessing on the first image data to obtain second image data, perform a second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, perform a third preprocessing on the first environmental data to obtain second environmental data, and perform a fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data. Finally, input the second image data, second three-dimensional point cloud data, second environmental data, and second hyperspectral image data into a multi-modal fusion monitoring model, and output the monitoring result of the growth stage of the sugarcane and the corresponding location information. Among them, the multi-modal fusion monitoring model is used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of the sugarcane according to the second image data, second three-dimensional point cloud data, second environmental data, and second hyperspectral image data, and perform a fusion analysis on the shape feature data, morphological feature data, physiological feature data, and environmental feature data to determine the monitoring result of the growth stage of the sugarcane and the corresponding location information; the monitoring result of the growth stage is one of the following: germination stage, seedling stage, tillering stage, elongation stage, and maturity stage.
[0008] In this method, first, collect first environmental data and first hyperspectral image data of a target area, as well as first image data and first three-dimensional point cloud data of the sugarcane. Then, perform preprocessing on the collected data respectively. Finally, input the second image data, second three-dimensional point cloud data, second environmental data, and second hyperspectral image data after preprocessing into a multi-modal fusion monitoring model, and output the monitoring result of the growth stage of the sugarcane and the corresponding location information. In this way, through the multi-modal fusion monitoring model, the growth stage of the sugarcane can be comprehensively analyzed and determined from four dimensions of the shape feature, morphological feature, physiological feature, and environmental feature of the sugarcane, so as to significantly improve the efficiency and accuracy of monitoring the growth stage of the sugarcane.
[0009] In combination with the first aspect, in an alternative implementation, the above multi-modal fusion monitoring model includes: an image feature extraction network, a point cloud feature extraction network, an environmental feature extraction network, a spectral feature extraction network, a first fusion layer, a second fusion layer, a third fusion layer, and an output layer. Among them, the image feature extraction network and the point cloud feature extraction network are respectively connected to the first fusion layer, the environmental feature extraction network and the spectral feature extraction network are respectively connected to the second fusion layer, the first fusion layer and the second fusion layer are respectively connected to the third fusion layer, and the third fusion layer is connected to the output layer. Specifically, the image feature extraction network is used to extract features from the second image data and output shape feature data; the point cloud feature extraction network is used to extract features from the second three-dimensional point cloud data and output morphological feature data; the environmental feature extraction network is used to extract features from the second environmental data and output environmental feature data; the spectral feature extraction network is used to extract features from the second hyperspectral image data and output physiological feature data; the first fusion layer is used to perform feature fusion on the shape feature data and the morphological feature data and output first feature fusion data; the second fusion layer is used to perform feature fusion on the environmental feature data and the physiological feature data and output second feature fusion data; the third fusion layer is used to perform feature fusion on the first fusion data and the second fusion data and output target feature fusion data; the output layer is used to classify according to the target feature fusion data and output the growth stage monitoring result of the sugarcane and the corresponding location information.
[0010] In this implementation, the multi-modal fusion monitoring model adopts a classification feature extraction and hierarchical hybrid fusion mechanism. In this way, the local feature interaction ability is retained, and the global efficiency is optimized, thereby effectively improving the efficiency and accuracy of the multi-modal fusion monitoring model to determine the growth stage monitoring result of the sugarcane.
[0011] In combination with the first aspect, in an alternative implementation, the above first fusion layer includes: a first cross-modal attention fusion module. The first cross-modal attention fusion module is used to perform feature interaction fusion on the shape feature data and the morphological feature data to determine the first feature fusion data.
[0012] In this implementation, the first fusion layer can calculate the feature interaction weight through the multi-head attention mechanism of the first cross-modal attention fusion module to avoid feature dilution caused by simple splicing and achieve cross-modal fine-grained strongly correlated feature fusion. In this way, the fusion effect of the first fusion layer on the shape feature data and the morphological feature data can be improved, so that the output layer can accurately output the growth stage monitoring result.
[0013] In combination with the first aspect, in an alternative implementation, the above-mentioned second fusion layer includes: an adaptive weight generator module. The adaptive weight generator module is used to adaptively and dynamically allocate the weights of the environmental feature data and the physiological feature data during the process of feature fusion of the environmental feature data and the physiological feature data.
[0014] In this implementation, the second fusion layer can dynamically allocate the weights of the environmental feature data and the physiological feature data through the adaptive weight generator module, and dynamically adjust the global contributions of the environmental feature data and the physiological feature data to improve the adaptability to complex sugarcane growth environments. Moreover, the computational overhead of the adaptive weight generator module is relatively low, so that the processing efficiency and accuracy of the multi-modal fusion monitoring model can be improved.
[0015] In combination with the first aspect, in an alternative implementation, the above-mentioned third fusion layer includes: a second cross-modal attention fusion module. The second cross-modal attention fusion module is used to perform feature cross-fusion on the first fusion data and the second fusion data to determine the target feature fusion data.
[0016] In this implementation, the third fusion layer can perform re-feature interaction fusion on the first fusion data and the second fusion data through the second cross-modal attention fusion module, that is, integrate the features of each modal data, fuse the global multi-modal feature data, and generate a more discriminative comprehensive representation feature number. In this way, the collaborative expression of the multi-modal feature data can be strengthened, and the classification robustness of the multi-modal fusion monitoring model in complex environments can be further improved.
[0017] In combination with the first aspect, in an alternative implementation, the image feature extraction network is a Residual ResNet network. The point cloud feature extraction network is: a PointNet++ network, or a RandLA-Net network. The environmental feature extraction network includes: a Long Short-Term Memory LSTM network and a multi-layer fully connected layer. The spectral feature extraction network includes: a three-dimensional convolutional neural network 3DCNN and a Spectral Transformer network.
[0018] In this implementation, the above specific feature extraction network can effectively extract the shape feature data, morphological feature data, environmental feature data and physiological feature data of sugarcane, and improve the accuracy of feature extraction.
[0019] In combination with the first aspect, in an alternative implementation, the above-mentioned output layer includes: a fusion feature input layer, a fully connected layer, a regularization and overfitting prevention module, an attention enhancement module, and a monitoring result determination layer. The loss function of the output layer is: a weighted cross-entropy loss function. Among them, the monitoring result determination layer includes a Softmax function.
[0020] In this implementation manner, the output layer can effectively suppress overfitting, accelerate convergence, and improve the generalization performance of the model by integrating the feature input layer, the fully connected layer, the regularization and overfitting prevention module, the attention enhancement module, and the monitoring result determination layer. Moreover, it can dynamically adjust the class weights to prevent the multi-modal fusion monitoring model from being biased towards the majority class, thereby improving the accuracy of the output layer in determining the growth stage monitoring result.
[0021] Combined with the first aspect, in an alternative implementation manner, the above first preprocessing method includes one or more of the following: image denoising, region of interest (ROI) extraction, color and contrast adjustment, image enhancement, cropping and alignment, geometric correction, and pixel value normalization processing. The second preprocessing method includes one or more of the following: point cloud registration, point cloud filtering, point cloud scale and position normalization processing, and point cloud segmentation. The third preprocessing method includes one or more of the following: data denoising, filling missing values, removing outliers, time feature extraction, and environmental data normalization processing. The fourth preprocessing method includes one or more of the following: cropping processing, smoothing processing, continuum removal processing, and standardization processing.
[0022] In this implementation manner, the above specific preprocessing method can improve the accuracy of the multi-modal fusion monitoring model in extracting the corresponding data of the shape features, morphological features, physiological features, and environmental features of sugarcane, thereby improving the accuracy of sugarcane growth stage monitoring.
[0023] Combined with the first aspect, in an alternative implementation manner, the above shape feature data is used to characterize one or more of the following features of sugarcane: sugarcane height, tiller number, germination number, leaf number, leaf size, leaf texture, and leaf color distribution. The morphological feature data is used to characterize one or more of the following features of sugarcane: sugarcane height, stem diameter, tiller number, germination number, leaf number, leaf size, and leaf texture. The environmental feature data is used to characterize one or more of the following features of the target area: environmental temperature, light intensity, air humidity, and soil humidity. The physiological feature data is used to characterize one or more of the following features of sugarcane: chlorophyll content and Brix value.
[0024] In this implementation manner, based on the above feature data, the multi-modal fusion monitoring model can accurately determine the growth stage of sugarcane.
[0025] Second aspect, the present invention provides a sugarcane growth stage monitoring system, which includes: a collection module, a data preprocessing module, and a monitoring module. Among them, the collection module is used to collect the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane planted in the target area. The data preprocessing module is used to perform a first preprocessing on the first image data to obtain second image data, perform a second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, perform a third preprocessing on the first environmental data to obtain second environmental data, and perform a fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data. The monitoring module is used to input the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data into the multi-modal fusion monitoring model, and output the growth stage monitoring result of the sugarcane and the corresponding position information; wherein, the multi-modal fusion monitoring model is used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of the sugarcane according to the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data, and perform a fusion analysis on the shape feature data, morphological feature data, physiological feature data, and environmental feature data to determine the growth stage monitoring result of the sugarcane and the corresponding position information; the growth stage monitoring result is one of the following: germination stage, seedling stage, tillering stage, elongation stage, maturity stage.
[0026] Third aspect, the present invention provides an electronic device, which includes: a memory, one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is enabled to execute the method according to the first aspect and any one of its alternative methods as described above.
[0027] Fourth aspect, the present invention provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are run on an electronic device, the electronic device is enabled to execute the method according to the first aspect and any one of its alternative methods as described above.
[0028] It can be understood that the beneficial effects that can be achieved by the sugarcane growth stage monitoring system in the second aspect, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect can refer to the beneficial effects in the first aspect and any one of its possible design manners, which will not be elaborated here. Description of the Drawings
[0029] Figure 1 It is a schematic flowchart of the sugarcane growth stage monitoring method provided by the embodiment of the present application;
[0030] Figure 2 It is a schematic structural diagram of the multi-modal fusion monitoring model provided by the embodiment of the present application;
[0031] Figure 3 This is a schematic structural diagram of the sugarcane growth stage monitoring system provided by the embodiments of the present application. Specific embodiments
[0032] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0033] Sugarcane is an important cash crop with extremely wide uses. Sugarcane can be used as a renewable energy alternative to produce bioethanol, reducing the consumption of traditional fossil fuels. The cellulose extracted from sugarcane can also be used in industries such as papermaking and textiles. Moreover, the stalks of sugarcane can also be used as the core raw material in the sugar industry.
[0034] With the increasing demand for agricultural intelligence, the refined management of sugarcane planting has become a key direction for industrial upgrading. Among them, the accurate monitoring of the sugarcane growth stage has significant economic value for agricultural production. Specifically, sugarcane has significant differences in water and fertilizer requirements, environmental adaptability, and pest and disease risks at different growth stages. For example: during the tillering stage, nitrogen fertilizer supply needs to be strengthened, while during the mature stage, water control is required to increase sugar accumulation. Therefore, in order to ensure the yield of sugarcane, it is necessary to monitor the growth stage of sugarcane in real time and obtain the growth information of sugarcane at each growth stage.
[0035] Traditional methods for monitoring the sugarcane growth stage mainly rely on manual monitoring, that is, by manually observing and collecting the growth data of sugarcane, and identifying the growth stage of sugarcane based on professional knowledge and empirical knowledge. However, these knowledges require a long learning time and practice. Moreover, the method relying on manual monitoring has problems of strong subjectivity and poor timeliness, which are likely to lead to waste of sugarcane resources or missed best agricultural operation windows. In this way, the efficiency of monitoring the sugarcane growth stage is slow, time-consuming and laborious, and there will also be large errors in the monitoring results, affecting the accuracy of monitoring the sugarcane growth stage.
[0036] To solve the above problems, an embodiment of the present invention provides a method and system for monitoring the growth stage of sugarcane. The method first collects environmental data and hyperspectral image data of the target area, as well as image data and three-dimensional point cloud data of the sugarcane planted in the target area. After preprocessing, it is input into a multi-modal fusion monitoring model, and the monitoring result of the growth stage of the sugarcane is output. In this way, through the multi-modal fusion monitoring model, the growth stage of the sugarcane can be comprehensively analyzed and determined from four dimensions: the appearance characteristics, morphological characteristics, physiological characteristics, and environmental characteristics of the sugarcane, avoiding the limitations of a single data source, and the method of machine learning using the multi-modal fusion monitoring model can effectively improve the efficiency and accuracy of monitoring the growth stage of the sugarcane.
[0037] Next, in conjunction with the accompanying drawings, the solution provided by the embodiment of the present invention will be introduced.
[0038] Specifically, referring to Figure 1 , which is a schematic flowchart of the method for monitoring the growth stage of sugarcane provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps S101 - S103:
[0039] S101. Collect the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane planted in the target area.
[0040] In an embodiment of the present application, in order to improve the accuracy of monitoring the growth stage of sugarcane, features that are highly relevant to determining the growth stage of sugarcane are selected, namely, the appearance characteristics, morphological characteristics, physiological characteristics, and environmental characteristics of the sugarcane, and the growth stage of the sugarcane is comprehensively determined from these four feature dimensions, so as to ensure that the accuracy of monitoring the growth stage of sugarcane can be effectively improved.
[0041] To extract the appearance characteristics, morphological characteristics, physiological characteristics, and environmental characteristics of the sugarcane, it is first necessary to collect the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane planted in the target area. Among them, one or more sugarcanes can be planted in the target area.
[0042] Specifically, the first environmental data can be used to characterize the temperature, soil humidity, etc. of the environment in the target area. The first hyperspectral image data can be used to characterize the spectral characteristics and spatial information of the sugarcane in the target area. The first image data can be used to characterize the appearance characteristics such as the color and morphology of the sugarcane. The first three-dimensional point cloud data can be used to characterize the morphological characteristics such as the shape and stem thickness of the sugarcane.
[0043] Exemplarily, the first image data can be a high-resolution (e.g., resolution 1024×1024) RGB-format image. The first three-dimensional point cloud data can be 100,000 points / frame point cloud data, specifically including XYZ coordinates and reflection intensity. The first environmental data can be sensor time-series data such as temperature and soil humidity. The first hyperspectral image data can be hyperspectral image data including 256 bands and a spatial resolution of 0.1 m.
[0044] In some embodiments, the first environmental data can be collected in real time by environmental sensors (e.g., light sensors, temperature and humidity sensors, soil humidity sensors, etc.) arranged in the target area.
[0045] The first hyperspectral image data can be obtained by an unmanned aerial vehicle (UAV) equipped with a hyperspectral camera for aerial photography of the target area. Exemplarily, the wavelength range of the hyperspectral camera can be 400 - 2500 nm. Based on the reflection characteristics of sugarcane leaves in different bands (such as the red edge and near infrared), physiological characteristics such as the chlorophyll content and Brix value of sugarcane can be inverted according to the first hyperspectral image data.
[0046] The first image data of sugarcane can be collected by a high-resolution visible light camera, and the data format of the first image data can be an RGB image. Specifically, the acquisition range of the first image data can cover the top, side, and root areas of sugarcane.
[0047] The first three-dimensional point cloud data of sugarcane can be collected by a lidar (LiDAR) or a three-dimensional laser scanner.
[0048] S102. Perform a first preprocessing on the first image data to obtain second image data, perform a second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, perform a third preprocessing on the first environmental data to obtain second environmental data, and perform a fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data.
[0049] Next, perform data preprocessing on the first environmental data, first hyperspectral image data, first image data, and first three-dimensional point cloud data collected in S101 respectively, remove the noise in the collected data, perform data processing according to the application requirements corresponding to each collected data, and finally convert it into the input format of the multi-modal fusion monitoring model. In this way, the accuracy of the multi-modal fusion monitoring model in extracting data corresponding to the external shape features, morphological features, physiological features, and environmental features of sugarcane can be improved, and thus the accuracy of sugarcane growth stage monitoring can be improved.
[0050] In some embodiments, the above-mentioned first preprocessing method may include one or more of the following: image denoising, region of interest (ROI) extraction, color and contrast adjustment, image enhancement, cropping and alignment, geometric correction, and pixel value normalization processing.
[0051] Specifically, image denoising can remove the noise in the first image data to eliminate sensor noise or environmental interference. Exemplarily, the following methods can be used for image denoising: Gaussian filtering method, median filtering method, wavelet transform method. ROI extraction can focus on the region in the first image data that represents sugarcane. Exemplarily, the following methods can be used for ROI extraction: threshold segmentation method, edge detection method, etc. Color and contrast adjustment can make the first image data meet the requirements for standardized input to the multi-modal fusion monitoring model. Image enhancement can highlight the characteristic details of sugarcane in the first image data and weaken the interference. Cropping and alignment can perform size normalization processing on the first image data to unify the size or perspective of the images. Geometric correction can correct the lens distortion of the first image data. Pixel value normalization processing can normalize the pixel values in the first image data to accelerate the convergence of the multi-modal fusion monitoring model and improve the generalization ability.
[0052] The above-mentioned second preprocessing method includes one or more of the following: point cloud registration, point cloud filtering, point cloud scale and position normalization processing, and point cloud segmentation.
[0053] Specifically, point cloud registration can fuse the point cloud data from different perspectives in the first three-dimensional point cloud data and unify them into the same coordinate system to construct complete three-dimensional point cloud data. Point cloud filtering can perform denoising processing on the first three-dimensional point cloud data to remove outliers, or downsample to reduce the data volume and improve the processing efficiency. Point cloud scale and position normalization processing can eliminate the differences in scale and position in the first three-dimensional point cloud data to achieve standardized input of the three-dimensional point cloud data and adapt to the multi-modal fusion monitoring model. Point cloud segmentation can separate the point cloud data used to represent sugarcane in the first three-dimensional point cloud data and specifically process this part of the point cloud data representing sugarcane to further improve the processing efficiency.
[0054] The above-mentioned third preprocessing method includes one or more of the following: data denoising, filling missing values, deleting outliers, time feature extraction, and environmental data normalization processing.
[0055] Specifically, data denoising can perform denoising processing on the first environmental data to eliminate the noise and interference of the acquisition sensor and improve the smoothness of the time-series data. Filling missing values can supplement the missing data in the first environmental data to ensure the continuity of the environmental data. Exemplarily, the methods for filling missing values can include: linear interpolation method, mean filling method, and KNN imputation method. In one implementation, the average value can be calculated based on the environmental data within a preset neighborhood range centered on the position of the missing value, and this average value can be used as the filling value for the missing value. Removing outliers can remove the abnormal data in the first environmental data to prevent the multi-modal fusion monitoring model from being misled by the outliers and affecting the accuracy of the model output results. Time feature extraction can extract the periodic trend information of the data in the first environmental data to enhance the time-series semantic expression of the environmental data. Normalization processing of the environmental data can eliminate the dimensional difference of the first environmental data to improve the stability of the convergence of the multi-modal fusion monitoring model.
[0056] The above fourth preprocessing method includes one or more of the following: cropping processing, smoothing processing, continuum removal processing, and standardization processing.
[0057] Specifically, cropping processing can intercept the sugarcane area in the first hyperspectral image data, reduce redundant data, and thus can reduce the computational complexity of feature extraction by the multi-modal fusion monitoring model. Smoothing processing can suppress the spectral and spatial noise in the first hyperspectral image data, improve the signal-to-noise ratio of the first hyperspectral image data, and retain details. Continuum removal processing can eliminate the background or substrate signal in the first hyperspectral image data to further highlight the spectral characteristics of sugarcane. Standardization processing can unify the data distribution of different bands in the first hyperspectral image data to enhance the robustness of the multi-modal fusion monitoring model.
[0058] S103. Input the second image data, the second 3D point cloud data, the second environmental data, and the second hyperspectral image data into the multi-modal fusion monitoring model, and output the monitoring results of the growth stage of sugarcane and the corresponding position information.
[0059] Finally, the second image data, the second 3D point cloud data, the second environmental data, and the second hyperspectral image data can be input into the multi-modal fusion monitoring model. Through the multi-modal fusion monitoring model, multi-modal data feature extraction and fusion analysis are performed to determine the monitoring results of the growth stage of sugarcane and the corresponding position information. Among them, the monitoring results of the growth stage are one of the following: germination stage, seedling stage, tillering stage, elongation stage, and maturity stage. The position information is used to represent the position of the sugarcane in the target area. In this way, it is convenient for users to quickly locate the target sugarcane and its growth stage, and further facilitate the agricultural management of sugarcane by users (such as zoned fertilization, irrigation, etc.).
[0060] In the embodiments of the present application, the multimodal fusion monitoring model can be used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of sugarcane based on the second image data, second 3D point cloud data, second environmental data, and second hyperspectral image data, and perform fusion analysis on the shape feature data, morphological feature data, physiological feature data, and environmental feature data to determine the growth stage monitoring result of sugarcane and the corresponding location information.
[0061] Among them, the multimodal fusion monitoring model is a model trained according to sample data. Specifically, the sample data includes: image sample data, 3D point cloud sample data, environmental sample data, and hyperspectral image sample data corresponding to five types of growth stage labels (i.e., the growth stage labels are: germination stage, seedling stage, tillering stage, elongation stage, and maturity stage).
[0062] Using the sugarcane growth stage monitoring method provided in the above embodiments of the present application, first, collect the first environmental data and first hyperspectral image data of the target area, as well as the first image data and first 3D point cloud data of the sugarcane. Then, preprocess the collected data respectively. Finally, input the preprocessed second image data, second 3D point cloud data, second environmental data, and second hyperspectral image data into the multimodal fusion monitoring model to output the growth stage monitoring result of sugarcane and the corresponding location information. In this way, through the multimodal fusion monitoring model, the growth stage of sugarcane can be comprehensively analyzed and determined from four dimensions of the shape features, morphological features, physiological features, and environmental features of sugarcane, thereby significantly improving the efficiency and accuracy of sugarcane growth stage monitoring.
[0063] In some embodiments, Figure 2 is a schematic structural diagram of the multimodal fusion monitoring model provided in the embodiments of the present application, as Figure 2 shown, the above multimodal fusion monitoring model 200 includes: an image feature extraction network 201, a point cloud feature extraction network 202, an environmental feature extraction network 203, a spectral feature extraction network 204, a first fusion layer 205, a second fusion layer 206, a third fusion layer 207, and an output layer 208.
[0064] Among them, the image feature extraction network 201 and the point cloud feature extraction network 202 are respectively connected to the first fusion layer 205, the environmental feature extraction network 203 and the spectral feature extraction network 204 are respectively connected to the second fusion layer 206, the first fusion layer 201 and the second fusion layer 206 are respectively connected to the third fusion layer 207, and the third fusion layer 207 is connected to the output layer 208.
[0065] Specifically, the image feature extraction network can be used to extract features from the second image data and output shape feature data. The point cloud feature extraction network can be used to extract features from the second 3D point cloud data and output morphological feature data. The environmental feature extraction network can be used to extract features from the second environmental data and output environmental feature data. The spectral feature extraction network can be used to extract features from the second hyperspectral image data and output physiological feature data.
[0066] The first fusion layer can be used to fuse the shape feature data and the morphological feature data and output the first feature fusion data. The second fusion layer can be used to fuse the environmental feature data and the physiological feature data and output the second feature fusion data. The third fusion layer can be used to fuse the first fusion data and the second fusion data and output the target feature fusion data. The output layer can be used to classify according to the target feature fusion data and output the monitoring result of the growth stage of sugarcane and the corresponding position information.
[0067] In one implementation, the first fusion layer fuses the shape feature data and the morphological feature data to obtain the first feature fusion data, which specifically may include: feature alignment and standardization, feature projection, attention weight assignment, and residual connection fusion. Among them, feature alignment includes: spatial alignment and temporal alignment.
[0068] Specifically, the expression of the first feature fusion data can be:
[0069] F1 = LayerNorm(F img + Attention(Q, K, V));
[0070]
[0071] Q = W q ·F img ;
[0072] K = W k ·F pc ;
[0073] V = W v ·F pc ;
[0074] Among them, F1 represents the first feature fusion data, LayerNorm() represents the normalization process, F img represents the shape feature data, F pcrepresents morphological feature data, Attention(Q, K, V) represents attention weights, softmax() represents the normalization exponential function, Q represents the query of the self-attention mechanism in the first fusion layer, K represents the key of the self-attention mechanism in the first fusion layer, V represents the value of the self-attention mechanism in the first fusion layer, and W q represents the projection matrix corresponding to Q, and W k represents the projection matrix corresponding to K, and W v represents the projection matrix corresponding to V.
[0075] The second fusion layer performs feature fusion on environmental feature data and physiological feature data to obtain the second feature fusion data, which specifically includes: dynamic weight calculation and weighted fusion.
[0076] Specifically, the expression of the second feature fusion data can be:
[0077] F2 = α · F env + (1 - α) · F spec ;
[0078] α = σ(W α · [F env ; F spec + b α ), α ∈ [0, 1];
[0079] where F2 represents the second feature fusion data, σ represents the Sigmoid function, F env represents environmental feature data, F spec represents physiological feature data, W α represents the first learning parameter matrix, b α represents the second learning parameter matrix, and α represents the weight coefficient corresponding to F env .
[0080] The third fusion layer performs feature fusion on the first fusion data and the second fusion data to obtain the target feature fusion data, which specifically includes: cross-attention calculation and fusion output.
[0081] Specifically, the expression of the target feature fusion data can be:
[0082] F final = LayerNorm(F1 + Attention(Q ′ , K ′ , V ′ ))
[0083] Q ′ = W q ′ · F1;
[0084] K ′ = W k ′ ·F2;
[0085] V ′ = W v ′ ·F2;
[0086] Wherein, F final represents the target feature fusion data, and Q ′ represents the query of the self-attention mechanism in the third fusion layer, K ′ represents the key of the self-attention mechanism in the third fusion layer, and V ′ represents the value of the self-attention mechanism in the third fusion layer, and W q ′ represents the projection matrix corresponding to Q ′ ; W k ′ represents the projection matrix corresponding to K ′ ; W v ′ represents the projection matrix corresponding to V ′ ;
[0087] In some embodiments, the shape feature data can be used to characterize one or more of the following features of sugarcane: sugarcane height, tiller number, germination number, leaf number, leaf size, leaf texture, leaf color distribution. The morphological feature data can be used to characterize one or more of the following features of sugarcane: sugarcane height, stem diameter, tiller number, germination number, leaf number, leaf size, leaf texture. The environmental feature data can be used to characterize one or more of the following features of the target area: environmental temperature, light intensity, air humidity, soil humidity. The physiological feature data can be used to characterize one or more of the following features of sugarcane: chlorophyll content, Brix value.
[0088] Wherein, the sugarcane features characterized by the shape feature data and the morphological feature data may have overlapping features. For example, tiller number, germination number, leaf number, leaf size, leaf texture, etc. In this way, after the shape feature data and the morphological feature data are subjected to feature fusion, the fused feature data (i.e., the first fusion data) can more accurately characterize the features of sugarcane, and further improve the accuracy of sugarcane growth stage monitoring.
[0089] In the embodiments of the present application, as Figure 2The multi-modal fusion monitoring model shown adopts a classification feature extraction and hierarchical hybrid fusion mechanism. Specifically, first, the shape feature data, morphological feature data, environmental feature data, and physiological feature data of sugarcane are extracted through four extraction networks respectively. Then, the shape feature data and morphological feature data that are highly correlated with the growth stage of sugarcane, have obvious features, and have similar or overlapping characterization features are subjected to feature fusion to achieve fine-grained strong correlation fusion, obtaining the first fusion data. Next, the environmental feature data and physiological feature data are subjected to feature fusion to achieve coarse-grained weak correlation fusion, obtaining the second fusion data. Further, the first fusion data and the second fusion data are subjected to feature fusion to integrate the features of each modal data, obtaining the target feature fusion data. Finally, the output layer performs classification processing based on the target feature fusion data and outputs the monitoring results of the growth stage of sugarcane and the corresponding position information.
[0090] In this way, the classification feature extraction and hierarchical hybrid fusion mechanism of the multi-modal fusion monitoring model not only retains the local feature interaction ability but also realizes the global efficiency optimization, thus effectively improving the efficiency and accuracy of the multi-modal fusion monitoring model in determining the monitoring results of the growth stage of sugarcane.
[0091] In some embodiments, the first fusion layer may include: a first cross-modal attention fusion module. The first cross-modal attention fusion module can be used to perform feature interaction fusion on the shape feature data and the morphological feature data to determine the first feature fusion data.
[0092] Specifically, the first cross-modal attention fusion module is a cross-modal attention fusion module based on Transformer. The first fusion layer can calculate the feature interaction weights through the multi-head attention mechanism of the first cross-modal attention fusion module, that is, dynamically calculate the interaction weights of the shape feature data and the morphological feature data (such as the correlation between leaf texture and the three-dimensional shape of the stem), and can capture the local detailed correlation between the shape features and morphological features of sugarcane to avoid feature dilution caused by simple splicing, realizing cross-modal fine-grained strong correlation feature fusion. At the same time, it can also solve the problem of insufficient single-view feature information of sugarcane. In this way, it can improve the fusion effect of the first fusion layer on the shape feature data and the morphological feature data, facilitating the output layer to accurately output the monitoring results of the growth stage.
[0093] In some embodiments, the second fusion layer includes: an adaptive weight generator module. The adaptive weight generator module can be used to adaptively and dynamically allocate the weights of the environmental feature data and the physiological feature data during the process of feature fusion of the environmental feature data and the physiological feature data.
[0094] Specifically, the second fusion layer can dynamically allocate the weights of the environmental feature data and the physiological feature data through the adaptive weight generator module, and dynamically adjust the global contributions of the environmental feature data and the physiological feature data to improve the adaptability to the complex sugarcane growth environment. In this way, the problem that the traditional fixed-weight fusion is insensitive to scene changes can be solved to ensure that the growth stage monitoring results are more in line with the actual growth characteristics of sugarcane and improve the accuracy of the growth stage monitoring results. Moreover, the calculation cost of the adaptive weight generator module is relatively low, which can improve the processing efficiency of the multi-modal fusion monitoring model.
[0095] In some embodiments, the third fusion layer includes: a second cross-modal attention fusion module. The second cross-modal attention fusion module can be used to perform feature cross-fusion on the first fusion data and the second fusion data to determine the target feature fusion data.
[0096] Specifically, the second cross-modal attention fusion module is a cross-modal attention fusion module based on Transformer. The third fusion layer can perform another feature cross-fusion on the first fusion data and the second fusion data through the second cross-modal attention fusion module, that is, synthesize the features of each modal data, fuse the global multi-modal feature data, and generate a more discriminative comprehensive representation feature number, that is, the target feature fusion data. In this way, the co-expression of multi-modal feature data can be strengthened, and the classification robustness of the multi-modal fusion monitoring model in a complex environment can be further improved.
[0097] In the embodiments of the present application, the multi-modal fusion monitoring model can achieve hierarchical hybrid fusion of the appearance feature data, the morphological feature data, the environmental feature data, and the physiological feature data through the first fusion layer, the second fusion layer, and the third fusion layer. Combining the characteristics of the feature data to be fused and the application requirements, the first fusion layer, the second fusion layer, and the third fusion layer respectively adopt different fusion modules (i.e., the corresponding first cross-modal attention fusion module, adaptive weight generator module, second cross-modal attention fusion module) to implement different fusion methods. In this way, the features of each modal data can be effectively synthesized, the accuracy of the fused target feature fusion data in expressing the sugarcane features can be improved, and further the accuracy of the multi-modal fusion monitoring model in determining the growth stage monitoring results can be improved.
[0098] In some embodiments, the above image feature extraction network is a Residual ResNet network. The point cloud feature extraction network is: PointNet++ network, or, RandLA-Net network. The environmental feature extraction network includes: a long short-term memory LSTM network and a multi-layer fully connected layer. The spectral feature extraction network includes: a three-dimensional convolutional neural network 3DCNN and a SpectralTransformer network.
[0099] Specifically, the ResNet network can achieve hierarchical feature extraction of images, that is, extract features such as edges and textures in the shallow layer and capture semantic features in the deep layer. In this way, the accuracy of image feature data extraction can be significantly improved.
[0100] The PointNet++ network is a hierarchical feature learning architecture that can extract global or local features through symmetric functions (such as max pooling). Specifically, it can finely process the second three-dimensional point cloud data through local geometric modeling and permutation invariance to achieve feature extraction and determine the morphological feature data. The RandLA-Net network is based on random sampling and efficient local feature aggregation, which can process large-scale point cloud data and aggregate neighborhood features through an attention mechanism to achieve neighborhood point cloud data perception.
[0101] The long short-term memory (LSTM) network and the multi-layer fully connected layer can, based on temporal modeling and feature space mapping, capture environmental features that depend on time in the environment based on the second environmental data, and can project high-dimensional features into a low-dimensional decision space, thereby improving the accuracy of extracting environmental feature data.
[0102] The 3D convolutional neural network (3DCNN) and the Spectral Transformer network can achieve functions such as joint feature extraction, multi-band modeling, and dynamic weight allocation, that is, they can simultaneously capture the spatial neighborhood and spectral correlation in the second hyperspectral image data, process the continuous band features of the second hyperspectral image data, and adaptively focus on key spectral intervals, thereby effectively performing spectral analysis and inversion on the second hyperspectral image data, and estimating and extracting physiological feature data of sugarcane such as chlorophyll content and sugar content (Brix value).
[0103] In some embodiments, the output layer includes: a fused feature input layer, a fully connected layer, a regularization and overfitting prevention module, an attention enhancement module, and a monitoring result determination layer. The loss function of the output layer is: a weighted cross-entropy loss function. Among them, the monitoring result determination layer includes a Softmax function.
[0104] Specifically, the fused feature input layer is used to receive the target feature fusion data output by the third fusion layer. The fully connected layer can be used to map the target feature fusion data to the category space. The regularization and overfitting prevention (Dropout) module can effectively suppress overfitting, accelerate convergence, and stabilize the training process; moreover, it can ensure that the multi-modal fusion monitoring model still maintains high generalization performance in sparse data scenarios (such as newly planted areas). The attention enhancement module can perform secondary weighting on the target feature fusion data before classification, so that the multi-modal fusion monitoring model focuses on the key information in the target feature fusion data. The monitoring result determination layer is used to classify the output data of the attention enhancement module to determine and output the growth stage monitoring result. Specifically, the monitoring result determination layer can first perform a classification head mapping. Then, the monitoring result determination layer can determine the probability distribution of the growth stage through the Softmax function, and use the growth stage corresponding to the highest probability as the growth stage monitoring result.
[0105] Among them, a weighted cross-entropy loss function is introduced in the output layer, which can dynamically adjust the class weights for the problem of unbalanced samples in the sugarcane growth stage (such as the amount of data in the mature stage is much larger than that in the germination stage), avoid the multi-modal fusion monitoring model from biasing towards the majority class, and thus improve the accuracy of the multi-modal fusion monitoring model in determining the growth stage monitoring result.
[0106] In some embodiments, the sugarcane growth stage monitoring method provided by this application further includes: establishing a three-dimensional model of the sugarcane based on the second three-dimensional point cloud data. After the multi-modal fusion monitoring model outputs the growth stage monitoring result of the sugarcane and the corresponding position information, not only can the growth stage monitoring result and the corresponding position information of the sugarcane be displayed, but also the three-dimensional model corresponding to the sugarcane can be presented. In this way, the visualization information of the sugarcane growth stage can be further enriched, facilitating users to intuitively understand the growth stage of the sugarcane through the three-dimensional model of the sugarcane.
[0107] In some embodiments, the embodiments of this application further provide a sugarcane growth stage monitoring system. Figure 3 As shown in the structural schematic diagram of the sugarcane growth stage monitoring system provided by the embodiments of this application, Figure 3 as shown, the sugarcane growth stage monitoring system 300 includes: a collection module 301, a data preprocessing module 302, and a monitoring module 303.
[0108] Among them, the collection module 301 can be used to collect the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane planted in the target area.
[0109] In some embodiments, the acquisition module 301 may include: an environmental sensor, a drone device equipped with a hyperspectral camera, a high-resolution visible light camera, a lidar, a three-dimensional laser scanner, etc. Among them, the environmental sensor can be used to collect first environmental data, the drone device equipped with a hyperspectral camera can be used to collect first hyperspectral image data, the high-resolution visible light camera can be used to collect first image data, and the lidar and the three-dimensional laser scanner can be used to collect first three-dimensional point cloud data.
[0110] The data preprocessing module 302 can be used to perform a first preprocessing on the first image data to obtain second image data, a second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, a third preprocessing on the first environmental data to obtain second environmental data, and a fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data.
[0111] The monitoring module 303 can be used to input the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data into a multi-modal fusion monitoring model, and output the growth stage monitoring result of the sugarcane and the corresponding location information; among them, the multi-modal fusion monitoring model is used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of the sugarcane according to the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data, and perform a fusion analysis on the shape feature data, the morphological feature data, the physiological feature data, and the environmental feature data to determine the growth stage monitoring result of the sugarcane and the corresponding location information. Among them, the growth stage monitoring result is one of the following: germination stage, seedling stage, tillering stage, elongation stage, maturity stage.
[0112] By using the sugarcane growth stage monitoring system provided in the above embodiments of the present application, the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane, can be collected through the acquisition module. The collected data can be respectively preprocessed through the data preprocessing module. Finally, through the monitoring module, the preprocessed second image data, second three-dimensional point cloud data, second environmental data, and second hyperspectral image data can be input into the multi-modal fusion monitoring model, and the growth stage monitoring result of the sugarcane and the corresponding location information can be output. In this way, through the multi-modal fusion monitoring model, the growth stage of the sugarcane can be comprehensively analyzed and determined from the four dimensions of the shape feature, morphological feature, physiological feature, and environmental feature of the sugarcane, so as to significantly improve the efficiency and accuracy of the sugarcane growth stage monitoring.
[0113] In some embodiments, the embodiments of the present invention further provide an electronic device, including: a memory and one or more processors; the memory is coupled to the processors; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processors, the electronic device is caused to execute the relevant steps of the sugarcane growth stage monitoring method in the above method embodiments.
[0114] In some embodiments, the embodiments of the present invention further provide a computer-readable storage medium, including computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the relevant steps of the sugarcane growth stage monitoring method in the above method embodiments.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0116] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0117] In the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", "connection", "fixation" and other terms should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0118] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0119] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.
Claims
1. A method for monitoring the growth stage of sugarcane, characterized in that, Including: Collecting first environmental data and first hyperspectral image data of a target area, as well as first image data and first three-dimensional point cloud data of sugarcane planted in the target area; Performing first preprocessing on the first image data to obtain second image data, performing second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, performing third preprocessing on the first environmental data to obtain second environmental data, and performing fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data; Inputting the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data into a multi-modal fusion monitoring model, and outputting the growth stage monitoring result of the sugarcane and the corresponding position information; Wherein, the multi-modal fusion monitoring model is used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of the sugarcane according to the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data, and perform fusion analysis on the shape feature data, the morphological feature data, the physiological feature data, and the environmental feature data to determine the growth stage monitoring result of the sugarcane and the corresponding position information; the growth stage monitoring result is one of the following: germination stage, seedling stage, tillering stage, elongation stage, maturity stage.
2. The method according to claim 1, wherein The multi-modal fusion monitoring model includes: an image feature extraction network, a point cloud feature extraction network, an environmental feature extraction network, a spectral feature extraction network, a first fusion layer, a second fusion layer, a third fusion layer, and an output layer; Wherein, the image feature extraction network and the point cloud feature extraction network are respectively connected to the first fusion layer, the environmental feature extraction network and the spectral feature extraction network are respectively connected to the second fusion layer, the first fusion layer and the second fusion layer are respectively connected to the third fusion layer, and the third fusion layer is connected to the output layer; The image feature extraction network is used to extract features from the second image data and output the shape feature data; the point cloud feature extraction network is used to extract features from the second three-dimensional point cloud data and output the morphological feature data; the environmental feature extraction network is used to extract features from the second environmental data and output the environmental feature data; the spectral feature extraction network is used to extract features from the second hyperspectral image data and output the physiological feature data; the first fusion layer is used to perform feature fusion on the shape feature data and the morphological feature data and output first feature fusion data; the second fusion layer is used to perform feature fusion on the environmental feature data and the physiological feature data and output second feature fusion data; the third fusion layer is used to perform feature fusion on the first fusion data and the second fusion data and output target feature fusion data; the output layer is used to classify according to the target feature fusion data and output the growth stage monitoring result of the sugarcane and the corresponding position information.
3. The method according to claim 2, characterized in that, The first fusion layer includes: a first cross-modal attention fusion module; The first cross-modal attention fusion module is used to perform feature interaction and fusion on the shape feature data and the morphological feature data to determine the first feature fusion data.
4. The method according to claim 2 or 3, characterized in that, The second fusion layer includes: an adaptive weight generator module; The adaptive weight generator module is used to adaptively and dynamically allocate the weights of the environmental feature data and the physiological feature data in the process of feature fusion of the environmental feature data and the physiological feature data.
5. The method according to claim 4, wherein The third fusion layer includes: a second cross-modal attention fusion module; The second cross-modal attention fusion module is used to perform feature cross-fusion on the first fusion data and the second fusion data to determine the target feature fusion data.
6. The method according to claim 2, wherein The image feature extraction network is: a residual ResNet network; The point cloud feature extraction network is: a PointNet++ network, or a RandLA-Net network; The environmental feature extraction network includes: a long short-term memory LSTM network and a multi-layer fully connected layer; The spectral feature extraction network includes: a three-dimensional convolutional neural network 3DCNN and a Spectral Transformer network.
7. The method according to claim 2, characterized in that The output layer includes: a fused feature input layer, a fully connected layer, a regularization and overfitting prevention module, an attention enhancement module, and a monitoring result determination layer; the loss function of the output layer is: a weighted cross-entropy loss function; Wherein, the monitoring result determination layer includes a Softmax function.
8. The method according to claim 1, wherein The method of the first preprocessing includes one or more of the following: image denoising, region of interest ROI extraction, color and contrast adjustment, image enhancement, cropping and alignment, geometric correction, pixel value normalization processing; The method of the second preprocessing includes one or more of the following: point cloud registration, point cloud filtering, point cloud scale and position normalization processing, point cloud segmentation; The method of the third preprocessing includes one or more of the following: data denoising, filling missing values, deleting outliers, time feature extraction, environmental data normalization processing; The method of the fourth preprocessing includes one or more of the following: cropping processing, smoothing processing, continuum removal processing, and standardization processing.
9. The method according to claim 1, wherein The shape feature data is used to characterize one or more of the following features of the sugarcane: sugarcane height, tiller number, germination number, leaf number, leaf size, leaf texture, leaf color distribution; The morphological feature data is used to characterize one or more of the following features of the sugarcane: sugarcane height, stem thickness, tiller number, germination number, leaf number, leaf size, leaf texture; The environmental feature data is used to characterize one or more of the following features of the target area: environmental temperature, light intensity, air humidity, soil humidity; The physiological feature data is used to characterize one or more of the following features of the sugarcane: chlorophyll content, Brix value.
10. A sugarcane growth stage monitoring system, characterized in that, Includes: An acquisition module, a data preprocessing module, and a monitoring module; The acquisition module is used to acquire the first environmental data and the first hyperspectral image data of the target area, as well as the first image data and the first three-dimensional point cloud data of the sugarcane planted in the target area; The data preprocessing module is used to perform a first preprocessing on the first image data to obtain second image data, perform a second preprocessing on the three-dimensional point cloud data to obtain second three-dimensional point cloud data, perform a third preprocessing on the first environmental data to obtain second environmental data, and perform a fourth preprocessing on the first hyperspectral image data to obtain second hyperspectral image data; The monitoring module is used to input the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data into a multi-modal fusion monitoring model, and output the growth stage monitoring result of the sugarcane and the corresponding position information; wherein, the multi-modal fusion monitoring model is used to extract the shape feature data, morphological feature data, physiological feature data, and environmental feature data of the sugarcane according to the second image data, the second three-dimensional point cloud data, the second environmental data, and the second hyperspectral image data, and perform a fusion analysis on the shape feature data, the morphological feature data, the physiological feature data, and the environmental feature data to determine the growth stage monitoring result of the sugarcane and the corresponding position information; the growth stage monitoring result is one of the following: germination stage, seedling stage, tillering stage, elongation stage, and maturity stage.