Crop Full Growth Cycle Recognition Method and System Based on Deep Learning
Through multispectral image sequence processing and multi-task deep learning model, the limitations of crop growth cycle identification in the existing technology are solved, a comprehensive assessment of growth stage and health status is achieved, and the automation level and decision-making reliability of agricultural management are improved.
Patent Information
- Application Number
- CN202510608239.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The prior art relies on a single visible light band image in crop growth cycle identification, making it difficult to capture physiological state changes under multi-spectral channels, resulting in limited identification of growth stages, and lack of quantitative assessment of crop health status, making it difficult to guide precise agricultural operations, and resource waste or management lag.
Multi-spectral image sequences are obtained through image acquisition equipment, standardized lighting adjustment processing is performed, crop area segmentation is performed, morphology, texture and spectral response features are extracted using multi-task deep learning model, cross-stage correlation analysis is performed, and growth stage category labels, health status scores and environmental adaptability index are output.
Multi-dimensional analysis of crop growth cycles is achieved, timing consistency and accuracy of growth stage judgments are improved, quantitative assessment of growth health status is provided, and precise agricultural management decisions are supported.
Smart Images

Figure CN120125918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular, to a method and system for identifying the entire growth cycle of crops based on deep learning. Background Art
[0002] In modern agriculture, the intelligent identification of the crop growth cycle is the core link of precision agricultural management, and is widely used in crop health monitoring, growth stage determination, and environmental regulation decision-making. Existing technologies usually obtain visible light images of crops by deploying image acquisition devices, extract image features based on convolutional neural networks, and combine classification models to output crop growth stage labels. For example, a support vector machine model is trained using the color and morphological features in RGB images to classify crops into stages such as germination, growth, and maturity; or a pre-trained deep learning model is used to perform regression analysis on phenotypic parameters such as leaf area and stem height to indirectly infer the crop growth state.
[0003] However, the above methods have significant limitations in practical applications: First, traditional methods rely on image data in a single visible light band and are difficult to capture the physiological state changes of crops in multi-spectral channels (such as chlorophyll content, water stress response), resulting in the discrimination of growth stages by the model being limited to apparent morphological features and unable to correlate with deep indicators such as environmental adaptability. Second, the lighting conditions in the crop growth environment are complex and variable. Existing technologies usually use global histogram equalization or fixed parameter correction to process images, unable to distinguish overexposed and underexposed areas, resulting in the loss of local details and an increase in the misjudgment rate of the model under strong light or shadow interference. In addition, traditional models only output the classification results of growth stages and lack a quantitative assessment of the crop health status (such as nutrient absorption efficiency, pest and disease resistance), making it difficult to guide precise farming operations such as irrigation and fertilization, resulting in resource waste or management lag. In addition, the processing of various modal features (such as morphology, texture, spectrum) in existing technologies is isolated from each other, and no cross-modal temporal correlation mechanism is established. For example, when judging the growth stage only by morphological features, it may be confused with the morphological changes in the normal growth stage due to environmental stress; while relying solely on spectral features to evaluate the health status is easily interfered by the reflectance deviation caused by light fluctuations. Such feature fragmentation problems further limit the model's ability to analyze the complex phenotypic evolution of the entire crop growth cycle. Therefore, there is an urgent need for a method to improve the recognition accuracy and agricultural management efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for identifying the entire growth cycle of crops based on deep learning. The embodiments of the present invention are implemented as follows:
[0005] In a first aspect, an embodiment of the present invention provides a method for identifying the entire growth cycle of crops based on deep learning. The method includes: obtaining a multi-spectral image sequence of a target crop within a continuous time period through an image acquisition device, and performing standardized illumination adjustment processing on each image frame in the multi-spectral image sequence to obtain a set of standard illumination images corresponding to the multi-spectral image sequence; performing crop region segmentation processing on each image in the set of standard illumination images, extracting local feature regions related to the target crop, and generating a set of standardized crop images containing the local feature regions; inputting each image in the set of standardized crop images into a pre-trained multi-task deep learning model, and respectively extracting a morphological feature vector, a texture distribution vector, and a spectral response vector in the standardized crop image through parallel convolutional branches in the multi-task deep learning model; performing cross-stage correlation analysis in the fully connected layer of the multi-task deep learning model according to the combined features of the morphological feature vector, the texture distribution vector, and the spectral response vector, and outputting a multi-task classification result corresponding to the growth cycle of the target crop; the multi-task classification result includes a growth stage category label, a growth health status score, and an environmental adaptability index; generating a stage identification report corresponding to the entire growth cycle of the target crop according to the growth stage category label in the multi-task classification result, and associatively storing the growth health status score and the environmental adaptability index in a crop management database.
[0006] In a second aspect, an embodiment of the present invention provides a computer system, including: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and are configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method described above is implemented.
[0007] The beneficial effects included in the present invention at least are:
[0008] The crop full growth cycle recognition method based on deep learning provided by the present invention obtains a multi-spectral image sequence of a target crop within a continuous time period, and performs standardized illumination adjustment processing on the multi-spectral image sequence to generate a set of standard illumination images; performs crop area segmentation processing on the images in the set of standard illumination images, extracts local feature regions of the target crop, and generates a set of standardized crop images; inputs the standardized crop images into a multi-task deep learning model, and extracts a morphological feature vector, a texture distribution vector, and a spectral response vector through parallel convolutional branches; outputs a multi-task classification result including a growth stage category label, a health status score, and an environmental adaptability index based on cross-stage correlation analysis; finally generates a stage recognition report and associates and stores it in a crop management database. In this way, the multi-spectral image sequence can comprehensively capture the phenotypic characteristics of the crop under different spectral channels, and the standardized illumination adjustment processing effectively eliminates the interference of environmental illumination differences on the image quality, ensuring the robustness of subsequent feature extraction; the morphological feature vector captures the structural evolution law of the crop stem and leaves through multi-scale convolution, the texture distribution vector quantifies the degradation or health recovery mode of the leaf surface texture with the growth stage, and the spectral response vector fuses the physiological state sensitive information of multiple spectral channels. The three cooperate to enhance the multi-dimensional analysis ability of the model for the crop growth state. At the same time, cross-stage correlation analysis combines time series modeling to associate the growth trend changes within a continuous time period with the real-time classification results, improving the temporal consistency of growth stage discrimination. The introduction of the growth health status score and the environmental adaptability index not only provides a quantitative evaluation index for the crop physiological state, but also directly maps to the dynamic regulation strategies of agricultural operations such as irrigation and fertilization through comparative analysis with environmental parameters and standard growth ranges, forming a closed-loop management link from data collection, intelligent recognition to precise control. Finally, the associated storage of the stage recognition report and the crop management database provides a data basis for historical data backtracking, growth trend prediction, and strategy effect verification, significantly improving the automation level and decision-making reliability of crop full life cycle management. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a flowchart of a crop full growth cycle recognition method based on deep learning provided by an embodiment of the present invention.
[0010] Figure 2 is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] In an embodiment of the present invention, the execution subject of the crop full growth cycle recognition method based on deep learning is a computer system, including but not limited to servers, personal computers, laptops, tablets, smart phones, etc., and is not specifically limited. As Figure 1 shown, the crop full growth cycle recognition method based on deep learning includes the following steps:
[0012] Step S100: Obtain a multi-spectral image sequence of the target crop within a continuous time period through an image acquisition device, and perform standardized illumination adjustment processing on each image frame in the multi-spectral image sequence to obtain a set of standard illumination images corresponding to the multi-spectral image sequence.
[0013] In the embodiments of the present invention, the image acquisition device may refer to an optical imaging device equipped with a multi-spectral sensor, which has the ability to synchronously collect multi-spectral channels such as visible light, near-infrared, and short-wave infrared, and can perform periodic image acquisition on the target crop at fixed time intervals within a continuous time period to generate a multi-spectral image sequence composed of images at multiple time points. Each image frame in the multi-spectral image sequence contains the reflectance information of the target crop in different spectral bands at the same geographical location. Among them, the visible light band (400 - 700 nm) is used to characterize the leaf color and morphological characteristics, the near-infrared band (700 - 1300 nm) is used to detect the chlorophyll content and water distribution, and the short-wave infrared band (1300 - 2500 nm) reflects the cell structure changes and disease characteristics. The standardized illumination adjustment processing refers to addressing the problem of abnormal brightness distribution caused by environmental illumination fluctuations in the multi-spectral image sequence. By dynamically adjusting the pixel brightness values of each spectral channel, the influence of overexposed and underexposed areas is eliminated, and the illumination conditions of all image frames are unified to a preset standard range. Specifically, during implementation, first, extract the brightness distribution histogram of each image frame in different spectral channels. By setting the standard illumination range threshold (for example, the brightness value of the visible light channel is limited to [20, 235]), identify the overexposed areas (brightness values higher than the upper limit) and underexposed areas (brightness values lower than the lower limit) that exceed the threshold range. Subsequently, respectively use the dynamic contrast compression algorithm (for overexposed areas) and the adaptive gamma correction algorithm (for underexposed areas) for local brightness correction to generate intermediate adjustment images. Further, perform spatial illumination gradient analysis on the intermediate adjustment images through an illumination equalization model constructed by a convolutional neural network to generate an illumination compensation coefficient matrix matching the image size, and perform per-pixel brightness compensation on each pixel point, and finally output standardized image frames that meet the standard illumination range threshold. The set of standard illumination images is composed of all standardized image frames arranged in the order of acquisition time, and the illumination distribution consistency of each of its image frames is significantly improved, providing a stable data basis for subsequent crop feature analysis.
[0014] Step S200: Perform crop region segmentation processing on each image in the set of standard illumination images, extract the local feature regions related to the target crop, and generate a set of standardized crop images containing the local feature regions.
[0015] In the embodiments of the present invention, the crop area segmentation process is, for example, a process of accurately separating the target crop stem and leaf area from the background area (such as soil, weeds or facility structures) in a standard illumination image. Specifically, during implementation, anisotropic diffusion filtering processing can be first performed on each standardized image frame. This algorithm calculates the image gradient direction and amplitude, and performs smoothing filtering along the edge tangent direction, while suppressing high-frequency noise (such as speckle noise caused by uneven illumination) and retaining the edge details of the crop stems and leaves, generating a preprocessed image after noise reduction. On this basis, a multi-scale edge detection algorithm (such as multi-resolution Canny edge detection based on a Gaussian pyramid) is used to analyze the preprocessed image, and edge gradient amplitude and direction information are extracted at different spatial scales, generating a multi-scale edge feature map. The multi-scale edge feature map is pixel-level fused with the channel components (including the luminance channel L, the green-red opponent channel a, and the blue-yellow opponent channel b) of the preprocessed image in the Lab color space to enhance the color contrast difference between the crop green area and the background, generating a fused color-edge feature map. Further, based on a superpixel segmentation algorithm (such as the SLIC algorithm), the fused color-edge feature map is divided into multiple homogeneous regions with uniform color and texture distributions, and the color consistency (pixel color variance within the region), texture continuity (gray-level co-occurrence matrix energy value), and edge closure (coincidence rate of the region boundary and the multi-scale edge feature map) of each region are calculated, generating a regional similarity metric value. According to the regional similarity metric value, an iterative merging operation (such as hierarchical clustering based on a greedy strategy) is performed on adjacent regions until a preset merging termination condition is met (such as the number of regions is reduced to the actual tiller number of the target crop), generating an initial segmentation mask containing candidate crop regions. Subsequently, a morphological closing operation (such as dilation-erosion operation using a circular structuring element) is performed on the initial segmentation mask to fill the region holes caused by leaf occlusion or light reflection, and an edge-sensitive region growing algorithm (such as seed point expansion based on gradient amplitude constraint) is used to connect the broken crop contour boundaries, generating an optimized accurate segmentation mask. Finally, according to the accurate segmentation mask, a local feature region containing only the target crop stem and leaf structure is cropped from the preprocessed image after noise reduction, and it is scaled to a unified size (such as 256×256 pixels), generating a standardized crop image set. The standardized crop image set eliminates background interference and focuses on the morphological and physiological characteristics of the target crop, providing a standardized input for subsequent multi-task feature extraction.
[0016] Step S300: Input each image in the standardized crop image set into a pre-trained multi-task deep learning model, and respectively extract the morphological feature vector, texture distribution vector, and spectral response vector in the standardized crop image through the parallel convolutional branches in the multi-task deep learning model.
[0017] In the embodiments of the present invention, the multi-task deep learning model can be a neural network architecture integrating multi-branch feature extraction and joint training, which synchronously extracts different feature modalities of crop phenotypes through parallel convolutional branches. The morphological feature vector is used to characterize the stem height, leaf expansion degree and overall plant type structure of the target crop, and its extraction process is realized through the first convolutional branch. This convolutional branch includes multiple dilated convolutional layers (for example, a three-level stack with dilation rates of 2, 4, and 8), and captures morphological features at different scales (such as the ratio of local stem diameter to global plant height) by expanding the receptive field of the convolutional kernel. The texture distribution vector is used to describe the surface texture of the leaf, stomatal density and vein distribution pattern, and its extraction process is realized through the second convolutional branch. This convolutional branch can be built with a Histogram of Oriented Gradients (HOG) filtering layer, and generates a feature code related to the leaf texture directionality by calculating the statistical histogram of the gradient direction in the local area of the image. The spectral response vector is used to quantify the reflection characteristics of the crop under different spectral channels, and its extraction process is realized through the third convolutional branch. This convolutional branch includes a spectral attention mechanism layer, which performs weighted fusion on the multi-spectral channel feature maps through learnable weight coefficients (for example, assigning higher weights to the near-infrared channel to enhance chlorophyll sensitivity), and generates a spectral response vector related to the physiological state of the crop. In specific implementation, the standardized crop images are respectively input into three parallel convolutional branches: the first convolutional branch gradually extracts hierarchical features of the stem morphology through multi-level dilated convolutional operations, and finally outputs a morphological feature vector with a dimension of 512; the second convolutional branch calculates the directional gradient distribution of the leaf area through the HOG filtering layer, and generates a texture distribution vector with a dimension of 256 after being compressed by the fully connected layer; the third convolutional branch dynamically adjusts the weights of each channel through the spectral attention mechanism, and generates a spectral response vector with a dimension of 128 after global average pooling. The three feature vectors comprehensively characterize the crop growth state from the morphological, texture and spectral dimensions, providing multi-modal data support for subsequent cross-stage correlation analysis.
[0018] Step S400: According to the combined features of the morphological feature vector, the texture distribution vector and the spectral response vector, perform cross-stage correlation analysis in the fully connected layer of the multi-task deep learning model, and output a multi-task classification result corresponding to the growth cycle of the target crop.
[0019] In the embodiments of the present invention, cross-stage correlation analysis can be used to model the temporal series dependence of multi-modal features, so as to reveal the dynamic association between the evolution of crop growth stages and the changes in phenotypic characteristics. In specific implementation, first, the morphological feature vector, texture distribution vector, and spectral response vector are concatenated to generate a combined feature vector with a dimension of 896, and it is input into the time series analysis module in the fully connected layer. This module is composed of a bidirectional long short-term memory network (Bi-LSTM), which synchronously processes the combined feature vectors in consecutive time periods through forward and backward LSTM units to capture the feature change trends at historical and future time points (such as the phased acceleration of leaf unfolding rate or the linear growth of stem height). Based on the growth trend change features output by the Bi-LSTM, the fully connected layer calculates three task outputs respectively: the first classification probability distribution of the growth stage category label (such as germination stage, tillering stage, heading stage, maturity stage), the second regression value of the growth health status score (a continuous value within the range of 0-100), and the third regression value of the environmental adaptability index (a normalized value within the range of 0-1). For the first classification probability distribution, it is normalized through the Softmax function, and the category with the highest probability is selected as the final growth stage category label; for the second regression value and the third regression value, the dynamic range scaling technology (such as linear interpolation or Sigmoid function) is used to map them to a preset numerical interval to ensure the comparability and interpretability of the growth health status score and the environmental adaptability index. In the multi-task classification result, the growth stage category label is used to identify the current physiological development stage of the crop, the growth health status score comprehensively reflects the leaf chlorophyll content, stem strength, and disease infection degree, and the environmental adaptability index quantifies the response efficiency of the crop to temperature, humidity, and light conditions. Through multi-task collaborative training and feature sharing, the model can optimize the accuracy of classification and regression tasks simultaneously, and significantly improve the comprehensiveness and reliability of growth cycle recognition.
[0020] Step S500: Generate a stage recognition report corresponding to the full growth cycle of the target crop according to the growth stage category label in the multi-task classification result, and store the growth health status score and the environmental adaptability index in association with the crop management database.
[0021] In the embodiments of the present invention, the stage identification report is, for example, a comprehensive document integrating growth stage category labels, historical stage evolution paths, and recommended agronomic measures. Specifically, in implementation, the standard growth parameter ranges corresponding to the current growth stage category labels can be retrieved from a preset crop growth stage knowledge base (for example, the leaf expansion threshold during the tillering stage is 30-50 cm, the stem height range is 40-60 cm, and the chlorophyll content reference value is 35-50 SPAD), and the growth health status score and environmental adaptability index in the multi-task classification results are compared and analyzed with the standard ranges. If the growth health status score is lower than the leaf expansion threshold of the current stage (for example, a score of 60 corresponds to an expansion of 28 cm, which is lower than the lower threshold of 30 cm), it is determined that there is an abnormal nutrient absorption in the target crop (such as nitrogen deficiency or poor root development), triggering the generation of fertilization plan adjustment parameters; if the environmental adaptability index deviates from the fitness threshold corresponding to the stem height range (for example, an index of 0.3 is lower than the threshold of 0.5), it is determined that there is an abnormal environmental response (such as insufficient irrigation or temperature stress), triggering the formulation of an irrigation frequency correction coefficient and a shading period optimization strategy. The stage identification report includes the above diagnostic results and the corresponding set of remedial measures (for example, applying 15 kg / acre of urea, shortening the irrigation interval from 7 days to 5 days, and adding a sunshade net to cover for 4 hours per day), and converts the fertilization plan adjustment parameters and the irrigation frequency correction coefficient into a device-executable control instruction set (such as the nitrogen injection amount setting value of an intelligent water and fertilizer integrated machine, the valve opening duration of a drip irrigation system) through an IoT gateway to real-time regulate the growth environment of the target crop area. At the same time, the growth health status score and the environmental adaptability index are stored in specific fields of the crop management database according to the timestamp (such as the health_score and adaptation_index columns in a MySQL table), supporting historical data backtracking and trend analysis. The crop management database is designed with a relational architecture, and is associated with the stage identification report, environmental sensor data, and device execution logs through foreign keys, providing full-link data support for crop growth optimization.
[0022] As an implementation manner, in step S100, perform a standardized light adjustment process on each image frame in the multi-spectral image sequence to obtain a standard light image set corresponding to the multi-spectral image sequence, which may specifically include:
[0023] Step S110: For each image frame in the multi-spectral image sequence, extract the brightness distribution histogram of the image frame in different spectral channels, and determine the overexposed area and underexposed area in the brightness distribution histogram that exceed the standard light range threshold according to the preset standard light range threshold.
[0024] In an embodiment of the present invention, it is assumed that the crop is Red Fuji. The multi-spectral image sequence is, for example, a set of multi-band optical images of the Red Fuji apple tree canopy during the growth period continuously collected by an image acquisition device equipped with a multi-spectral sensor (such as an airborne hyperspectral camera or a fixed farmland monitoring system). Each image frame contains synchronous imaging data of the visible light channel (400 - 700 nm), the near-infrared channel (700 - 1300 nm), and the short-wave infrared channel (1300 - 2500 nm). The brightness distribution histogram is a statistical distribution graph of the pixel brightness values (ranging from 0 to 255) within each spectral channel, reflecting the proportion of the number of pixels in different brightness intervals in the image frame. Specifically, during implementation, first, the brightness distribution histograms of the red (R), green (G), and blue (B) sub-channels are extracted from the visible light channel image of the Red Fuji apple tree canopy to quantify the leaf color saturation and the fruit coloring state; the single-channel brightness distribution histogram is extracted from the near-infrared channel image to characterize the leaf water content and the chlorophyll absorption characteristics; the single-channel brightness distribution histogram is extracted from the short-wave infrared channel image to detect the cell structure integrity and potential disease areas. The standard illumination range threshold is preset according to the illumination characteristics of the Red Fuji apple planting environment. For example, the thresholds for the R / G / B sub-channels of the visible light channel are set to [25 - 230], [30 - 225], [20 - 235] respectively, the threshold for the near-infrared channel is set to [15 - 240], and the threshold for the short-wave infrared channel is set to [10 - 245]. This threshold range ensures that the reflection characteristics of the leaves and fruits in different spectral channels are within the resolvable range. Based on the brightness distribution histogram, by traversing the pixel brightness values of each spectral channel, the pixel regions with brightness values higher than the upper limit of the corresponding channel threshold are marked as overexposed regions (such as the pixel group with a brightness value ≥ 230 in the visible light R channel), and the pixel regions with brightness values lower than the lower limit of the threshold are marked as underexposed regions (such as the pixel group with a brightness value ≤ 10 in the short-wave infrared channel). Taking the visible light image collected during the fruit ripening period of the Red Fuji apple as an example, the overexposed region may correspond to the fruit surface reflection region caused by strong direct sunlight, and the underexposed region may correspond to the leaf region covered by the tree canopy shadow. Both need to eliminate the interference of abnormal illumination on the extraction of morphological and physiological characteristics through subsequent processing.
[0025] Step S120: Perform a dynamic contrast compression algorithm on the overexposed region and an adaptive gamma correction algorithm on the underexposed region to generate an intermediate adjusted image corresponding to the image frame.
[0026] In the embodiments of the present invention, the dynamic contrast compression algorithm is, for example, for high-brightness pixels in overexposed areas, to reduce the extreme brightness values through local contrast limitation and non-linear mapping to prevent detail loss. Specifically, in implementation, for the overexposed area (such as the fruit surface specular reflection area) in the Red Fuji apple image, calculate the average brightness value of its neighboring pixels, and dynamically adjust the contrast compression intensity based on this mean value: if the brightness difference around the overexposed area is large (such as the junction between the fruit and the leaf), then use a piecewise linear function to compress the original brightness value [230 - 255] to the interval [200 - 230] to retain the highlight details; if the surrounding brightness is relatively uniform (such as the continuous specular reflection area in the center of the fruit), then use a logarithmic function to map the brightness value [230 - 255] to the interval [180 - 210] to suppress the overexposure effect. The adaptive gamma correction algorithm refers to dynamically adjusting the gamma value (γ) according to the local brightness distribution for low-brightness pixels in underexposed areas to improve the visibility of dark details. Taking the shaded area of the Red Fuji apple tree crown as an example, first calculate the average brightness value of the pixels in the underexposed area (for example, the average value is 15). If this average value is lower than the lower limit of the channel threshold (for example, the B-channel threshold is 20), then calculate the adaptive gamma value based on the formula γ = log(mean) / log(target mean) (for example, when the target mean is set to 50, γ ≈ 0.7), and enhance the leaf texture and vein structure by raising the original brightness value [0 - 20] to the interval [20 - 80] through power-law transformation. In the intermediate adjustment image after the above processing, the brightness of the fruit surface specular reflection patches in the overexposed area is reduced to an analyzable range, the details of the shaded leaves in the underexposed area are enhanced, and at the same time, the natural color transition between the Red Fuji apple fruits and leaves is retained, laying a foundation for subsequent global illumination equalization.
[0027] Step S130: Input the intermediate adjustment image into the illumination equalization model, extract the spatial illumination gradient map of the intermediate adjustment image through the convolutional layer in the illumination equalization model, and generate a corresponding illumination compensation coefficient matrix according to the spatial illumination gradient map.
[0028] In the embodiments of the present invention, the light equalization model represents an end-to-end light correction model constructed based on a convolutional neural network, which learns the spatial light distribution pattern of the image through multiple convolutional operations. The spatial light gradient map is a two-dimensional matrix reflecting the brightness change rate and direction of each pixel point in the intermediate adjusted image, and is synthesized by calculating the gradient amplitudes in the horizontal direction (Sobel_x operator) and the vertical direction (Sobel_y operator). Specifically in implementation, the intermediate adjusted image of the Red Fuji apple is input into the first convolutional layer of the light equalization model (such as a 3×3 convolutional kernel, a stride of 1, and a padding of 1) to extract shallow edge features; subsequently, the second convolutional layer (a 5×5 convolutional kernel, a stride of 2, and a padding of 2) is used to capture large-scale light distribution features, and the channel dimension of the output feature map is the same as the number of spectral channels of the input image. Based on the feature map, the light gradient amplitude of each pixel position is calculated to generate a spatial light gradient map, where the area with a high gradient amplitude corresponds to the sudden change area of the canopy light (such as the junction of the fruit and the leaf or the edge of the branch projection), and the area with a low gradient amplitude corresponds to the uniform light area (such as the surface of the mature fruit). Further, the spatial light gradient map is converted into a light compensation coefficient matrix through a fully connected layer, and each element value in this matrix is a floating point number within the range of 0.8 - 1.2, which is used to indicate the brightness compensation intensity of the corresponding pixel point: the compensation coefficient in the area with a high gradient amplitude (such as the fruit edge) is close to 1.0 to maintain details, and the compensation coefficient in the area with a low gradient amplitude (such as the uniformly illuminated leaf) is 1.1 - 1.2 to improve the overall brightness consistency. Taking the canopy image of the Red Fuji apple as an example, the compensation coefficient of the leaf area with sufficient light at the top of the canopy is set to 1.15 due to the low gradient amplitude; the compensation coefficient of the fruit in the side branch shadow area is set to 0.95 due to the high gradient amplitude, so as to achieve global light equalization while suppressing local overcompensation.
[0029] Step S140: Perform per-pixel brightness compensation on each pixel point in the intermediate adjusted image according to the light compensation coefficient matrix to generate a standardized image frame that meets the standard light range threshold.
[0030] In the embodiment of the present invention, the per-pixel brightness compensation means performing a dot product operation on the illumination compensation coefficient matrix and the corresponding pixel brightness value of the intermediate adjustment image to achieve pixel-level brightness adjustment. During specific implementation, each spectral channel of the intermediate adjustment image of Red Fuji apples is processed independently: for the visible light R channel, each element in the compensation coefficient matrix is multiplied by the pixel value of the R channel. For example, a pixel with an original R value of 200 is adjusted to 220 under the action of a compensation coefficient of 1.1; for the near-infrared channel, a pixel with an original value of 180 is adjusted to 176.4 (rounded to 176) under the action of a compensation coefficient of 0.98. This process ensures that the brightness values of each channel strictly fall within the preset standard illumination range threshold after compensation (for example, the value range of the visible light R channel after adjustment is [25 - 230]). Taking the image of Red Fuji apple fruits as an example, after compensation, the R channel value of the reflective area on the fruit surface drops from 228 to 215 (lower than the threshold upper limit of 230), the B channel value of the shaded leaves rises from 18 to 22 (higher than the threshold lower limit of 20), and at the same time, the near-infrared channel value of the leaves in the middle of the tree crown drops from 235 to 229 (within the threshold range). The finally generated standardized image frame eliminates the brightness abnormality caused by illumination fluctuations, and significantly improves the characterization consistency of the fruit color, leaf texture, and branch structure of Red Fuji apples in different spectral channels, providing a standardized input for subsequent crop area segmentation.
[0031] Step S150: Summarize all the standardized image frames to generate a standard illumination image set.
[0032] In the embodiment of the present invention, the standard illumination image set is a set of standardized image frames arranged in chronological order, and the illumination condition, color fidelity, and detail clarity of each image frame meet the preset technical specifications. During specific implementation, the standardized image frames at multiple time points (such as the germination period, flowering period, fruit swelling period, and maturity period) continuously collected during the whole growth cycle of Red Fuji apples are named according to the timestamp and stored in a structured database to form an illumination standardized image sequence that can be traced in the time dimension. For example, after processing the image frame collected on April 10, 2023, during the germination period, the R / G / B values of the tender buds in the visible light channel are stably in the range of [30 - 210] / [40 - 200] / [25 - 195]; in the image frame collected on September 20, 2023, during the maturity period, the R channel values of the fruits are evenly distributed in the range of [180 - 225], accurately reflecting the coloring degree of Red Fuji apples. The standard illumination image set ensures the comparability of the morphological and physiological characteristics at different growth stages by eliminating the chronological illumination differences, provides cross-cycle consistent data support for the multi-task deep learning model, and further improves the accuracy of identifying the growth stages and evaluating the health status of Red Fuji apples.
[0033] As an implementation manner, in step S200, each image in the standard illumination image set is subjected to crop region segmentation processing to extract local feature regions related to the target crop, and a standardized crop image set including the local feature regions is generated. Specifically, it may include:
[0034] Step S210: Perform anisotropic diffusion filtering on each standardized image frame in the standard illumination image set to suppress high-frequency noise in the standardized image frame and retain crop edge details, and generate a preprocessed image after noise reduction.
[0035] In the embodiment of the present invention, the anisotropic diffusion filtering is a non-linear image smoothing technology based on partial differential equations, which protects the target crop edge structure while suppressing noise by adaptively adjusting the diffusion intensity. Specifically, when implementing, for the standardized image frame of the Red Fuji apple crown (such as the RGB image of leaves and fruits in the visible light channel), first calculate the gradient magnitude of each pixel point in the image, and dynamically adjust the diffusion coefficient according to the gradient magnitude: in the flat area (such as the uniformly colored area on the fruit surface), the gradient magnitude is low, and high-intensity diffusion is used to eliminate high-frequency noise (such as sensor noise or light reflection spots); in the edge area (such as the serrated contour of the leaf or the branch boundary), the gradient magnitude is high, and low-intensity diffusion is used to retain details. Taking the Red Fuji apple leaf image as an example, during the filtering process, small noise points (such as dust or water droplet reflections) on the leaf surface are smoothed out, while the edge sharpness of the leaf veins and the fruit stalk depression is maintained. After iterative diffusion calculation (for example, the number of iterations is set to 5 times, and the time step is 0.15), a preprocessed image after noise reduction is generated, whose signal-to-noise ratio is significantly improved and the integrity of the morphological features is guaranteed, providing high-quality input for subsequent edge detection and region segmentation.
[0036] Step S220: Apply a multi-scale edge detection algorithm to the preprocessed image after noise reduction to extract edge gradient magnitude and direction information at different resolutions respectively, and generate a multi-scale edge feature map related to the crop stem and leaf contour.
[0037] In the embodiment of the present invention, the multi-scale edge detection algorithm is a technology that constructs multi-resolution image layers through a Gaussian pyramid and independently performs edge gradient analysis on each layer. Specifically, when implementing, perform Gaussian pyramid downsampling on the preprocessed image of the Red Fuji apple (for example, generate a scale of the original Figure 1Three - layer images of 1 / 2, 1 / 4, and 1 / 8), the Sobel operator is applied at each scale layer to calculate the horizontal and vertical gradients, and the gradient magnitude map and direction map are synthesized. At the highest resolution layer (the original image scale), the gradient magnitude map accurately captures the fine edges of the leaves (such as leaf margin serrations and insect - eaten holes); at the low - resolution layer (1 / 8 scale), the gradient magnitude map highlights the large - scale contours (such as the branching direction of the main trunk or the overall shape of the fruit). Through up - sampling and weighted fusion (such as linear interpolation combined with amplitude - threshold weighting), the multi - scale gradient information is integrated into a multi - scale edge feature map. Taking the image of the fruit enlargement period of Red Fuji apples as an example, this feature map simultaneously retains the fine edges at the fruit stalk connection (from the high - resolution layer) and the overall contour of the tree crown (from the low - resolution layer), avoiding edge breaks or redundant noise caused by single - scale detection, and laying a data foundation for subsequent color - edge feature fusion.
[0038] Step S230: Fuse the multi - scale edge feature map with the channel components of the pre - processed image after noise reduction in the Lab color space to enhance the color contrast difference between the crop area and the background, and generate a fused color - edge feature map.
[0039] In the embodiment of the present invention, the Lab color space is composed of a luminance channel (L), a green - red opponent channel (a), and a blue - yellow opponent channel (b). Its color gamut range is wider than that of the RGB space, and it is more suitable for characterizing the natural color differences of Red Fuji apple leaves and fruits. Specifically, during implementation, the pre - processed image after noise reduction is converted from the RGB space to the Lab space, and the three channel components of L, a, and b are extracted. At the same time, the multi - scale edge feature map is normalized to the [0, 1] interval as the fourth channel and pixel - level superposition fusion is performed with the Lab channels. During the fusion process, the high - response area of the a channel (corresponding to the green area of the Red Fuji apple leaves) and the high - amplitude area of the edge feature map (corresponding to the stem and leaf contours) highly coincide spatially. By channel weighting (such as the weight of the a channel is 0.6 and the weight of the edge feature map is 0.4), the contrast between the crop target and the background (such as brown soil or white reflective film) is enhanced. Taking the image of Red Fuji apples during the flowering period as an example, in the fused feature map, the response of the white petals in the b channel is superimposed on the calyx contour of the edge feature map, forming a high - contrast area; while the response of the green leaves in the a channel is superimposed on the vein edges, further highlighting the morphological boundaries. The fused color - edge feature map integrates color and spatial features, significantly improving the segmentability of crop targets under complex backgrounds.
[0040] Step S240: Based on the pixel clustering distribution in the fused color - edge feature map, use the super - pixel segmentation algorithm to divide the fused color - edge feature map into multiple homogeneous regions, and calculate the regional similarity metric value according to the color consistency, texture continuity, and edge closure degree of each region.
[0041] In the embodiments of the present invention, the superpixel segmentation algorithm divides an image into continuous regions with consistent visual perception, and its boundaries coincide with the natural contours of crop stems and leaves. Specifically, when implemented, the simple linear iterative clustering (SLIC) algorithm is used to process the fused color-edge feature map of Red Fuji apples: the number of clustering centers is initialized to 500 (set based on the image size and target complexity), and the distance between pixels and clustering centers is calculated in a 5-dimensional feature space (L, a, b, edge amplitude, edge direction), and the clustering is iteratively updated until convergence. The generated superpixel regions need to meet the conditions of color consistency (the variance of the Lab channels within the region ≤ 10), texture continuity (the contrast of the gray-level co-occurrence matrix ≤ 5), and edge closure (the coincidence rate of the region boundary and the fused edge feature map ≥ 80%). Taking the image of Red Fuji apples at the mature stage as an example, a single superpixel may cover 1-2 complete leaves or the surface area of a single fruit, while the soil background is divided into large-scale superpixels in a loose manner. The regional similarity metric value is calculated through a weighted formula (for example, the color similarity weight is 0.5, the texture similarity weight is 0.3, and the edge closure weight is 0.2), which is used to quantify the mergibility between adjacent superpixels and provide a basis for subsequent iterative merging.
[0042] Step S250: Perform an iterative merging operation on adjacent regions according to the regional similarity metric value until a preset merging termination condition is met, and generate an initial segmentation mask including candidate crop regions.
[0043] In the embodiments of the present invention, the iterative merging operation is a hierarchical clustering process based on a greedy strategy, which aggregates crop targets by gradually merging regions with the highest similarity. Specifically, when implemented, a region adjacency graph is established and the similarity metric values of all adjacent region pairs are calculated, and region pairs with similarity higher than a threshold (for example, 0.85) are preferentially merged. After each merge, the region attributes (color mean, texture statistics, and boundary coordinates) are updated, and the similarity with adjacent regions is recalculated. The merging termination condition is set to one of the following two: 1) the number of remaining regions drops to a preset target (for example, 50, corresponding to the main structural units of the Red Fuji apple tree crown); 2) the maximum regional similarity metric value is lower than the threshold of 0.7. Taking the image of Red Fuji apples during the fruit thinning period as an example, after the initial superpixels are merged, the scattered young fruit regions are aggregated into coherent fruit clusters, while the mis-segmented branch shadow regions are retained as independent backgrounds due to low similarity. In the generated initial segmentation mask, the candidate crop regions are marked in the form of a binary image, and its boundary initially fits the main structure of the tree crown, but there are still region holes and breaks caused by leaf occlusion or light reflection.
[0044] Step S260: Perform a morphological closing operation on the initial segmentation mask to fill the region holes caused by uneven illumination or leaf occlusion, and use an edge-sensitive region growing algorithm to connect the broken crop contour boundaries to generate an optimized accurate segmentation mask.
[0045] In the embodiments of the present invention, morphological closing is a composite operation of dilation followed by erosion, which is used to fill small-scale holes and smooth irregular boundaries. A circular structuring element (radius 5 pixels) is used to dilate the initial segmentation mask of Red Fuji apples, so that small holes in the fruit area (such as leaf gaps with a diameter ≤ 10 pixels) are filled with adjacent foreground pixels; subsequently, an erosion operation of the same size is performed to restore the original main body contour. On this basis, an edge-sensitive region growing algorithm is applied: using the mask boundary pixels as seed points, region expansion is performed according to the conditions of gradient direction consistency (the difference in gradient directions of adjacent pixels ≤ 15°) and amplitude continuity (the change rate of gradient amplitude ≤ 20%) to connect the broken boundaries caused by missed edge detection. Taking the dense leaf area of Red Fuji apples as an example, after the closing operation fills the gaps between leaves, the region growing algorithm extends along the gradient direction of the veins to connect the fragmented leaf areas into complete semantic units. In the optimized accurate segmentation mask, the boundary of the crop area is highly consistent with the true physiological structure, and the background misdetection rate is less than 3%, meeting the requirements of accurate feature extraction.
[0046] Step S270: Crop a local feature area containing only the stem and leaf structure of the target crop from the preprocessed image after noise reduction according to the optimized accurate segmentation mask, and scale the local feature area to a unified size to generate a standardized crop image set.
[0047] In the embodiments of the present invention, the local feature area is a set of pixels within the minimum bounding rectangle corresponding to the accurate segmentation mask, which completely contains the target structures such as the stems, leaves, flowers, or fruits of the Red Fuji apple crown. Specifically, during implementation, all connected regions in the mask are traversed, and the coordinates of their bounding rectangles are calculated (for example, the upper left corner point (x min , y min ) and the lower right corner point (x max , y max )), and the rectangular area is intercepted from the preprocessed image after noise reduction and the original resolution is retained. To ensure the uniformity of the input data specifications, the intercepted image is scaled to a fixed size (such as 512×512 pixels), and the bicubic interpolation algorithm is used to maintain the clarity of leaf textures and fruit details. Taking images of Red Fuji apples at different growth stages as an example, after the young shoot area in the germination stage is cropped and scaled, the tip shape and scale structure of the buds are completely retained; the fruit area in the mature stage fully presents the coloring area and fruit shape index. The generated standardized crop image set eliminates background interference, focuses on the phenotypic characteristics of the target crop, and provides input data with uniform size, illumination, and spatial distribution for the multi-task deep learning model, significantly improving the generalization ability and classification accuracy of the model.
[0048] As an implementation method, in step S300, the morphological feature vectors, texture distribution vectors, and spectral response vectors in the standardized crop images are respectively extracted through the parallel convolutional branches in the multi-task deep learning model, including:
[0049] Step S310: Input the standardized crop images into the first convolutional branch, the second convolutional branch, and the third convolutional branch of the multi-task deep learning model respectively; wherein, the first convolutional branch includes multiple dilated convolutional layers for capturing morphological features at different scales in the standardized crop images; the second convolutional branch includes a histogram of oriented gradients (HOG) filtering layer for extracting the texture distribution pattern of the leaf edges in the standardized crop images; the third convolutional branch includes a spectral attention mechanism layer for weighted fusion of the response features of the standardized crop images in different spectral channels.
[0050] In the embodiment of the present invention, the parallel convolutional backbone architecture of the multi-task deep learning model realizes multi-dimensional characterization of the phenotypes of Red Fuji apples through a heterogeneous feature extraction module. The first convolutional branch uses dilated convolutional layers (such as a three-level stack with dilation rates of 2, 4, and 8 respectively) to expand the receptive field of the convolutional kernel, so as to capture multi-scale morphological features of the stem thickness, leaf expansion degree, and overall plant type of Red Fuji apples without increasing the number of parameters. For example, the convolutional layer with a dilation rate of 2 focuses on local stem texture details (such as internode length and epidermal wrinkles), while the convolutional layer with a dilation rate of 8 integrates the overall canopy contour information (such as canopy projection area and branching angle). The second convolutional branch incorporates a histogram of oriented gradients (HOG) filtering layer. By dividing the image into cell units (such as 8×8 pixels) and counting the distribution histogram of gradient directions within each unit (such as divided into 9 direction intervals), the texture direction consistency of the leaf edges of Red Fuji apples (such as vein direction and serration arrangement pattern) is quantified. The spectral attention mechanism layer of the third convolutional branch dynamically adjusts the contribution degrees of different spectral channels through a learnable weight matrix (such as a 3×3 convolutional kernel to generate a channel attention map). For example, a higher weight is assigned to the near-infrared channel to enhance the sensitivity to leaf water stress, while suppressing the interference of fruit coloring differences in the visible light channel on the interpretation of physiological states. Taking the image of Red Fuji apples during the flowering period as an example, the standardized crop images are synchronously input into the three branches: the first convolutional branch extracts the flower bud distribution density and inflorescence morphology, the second convolutional branch quantifies the regularity of the petal edge texture, and the third convolutional branch fuses multi-spectral data to evaluate the physiological activity of the flowers, thereby realizing the parallel extraction and complementary enhancement of morphological, texture, and spectral features.
[0051] Step S320: Perform multi-scale feature extraction on the standardized crop images through the dilated convolutional layers in the first convolutional branch to generate a morphological feature vector related to the crop stem structure.
[0052] In the embodiments of the present invention, the morphological feature vector is a quantitative descriptor of the plant type structure of Red Fuji apples abstracted through hierarchical convolution operations. Specifically, during implementation, the standardized crop images are sequentially processed through three levels of dilated convolutional layers in the first convolution branch: the first layer (for example, with a dilation rate of 2) outputs a feature map that focuses on microscopic morphology (such as the diameter at the fruit stalk connection and the angle between the base of the leaf), the second layer (for example, with a dilation rate of 4) extracts mesoscopic features (such as the bifurcation angle between the main branch and the lateral branch), and the third layer (for example, with a dilation rate of 8) captures macroscopic features (such as the ratio of the tree crown height to the crown width). After each level of convolutional layer, batch normalization and the ReLU activation function are connected to suppress internal covariate shift and enhance the non-linear expression ability. Finally, the three-level feature maps are compressed into a 512-dimensional vector through global average pooling, where the high-response dimensions correspond to the key morphological indicators of Red Fuji apples (for example, the 128th dimension represents the straightness of the stem, and the 256th dimension reflects the leaf density). Taking the image of Red Fuji apples during the fruit thinning period as an example, the numerical differences in specific dimensions of the morphological feature vector can distinguish normal fruit clusters (compact spherical distribution) from overgrown fruit clusters (loose and elongated morphology), providing a structural basis for growth stage classification.
[0053] Step S330: Perform edge direction statistical analysis on the standardized crop image through the histogram of oriented gradients filtering layer in the second convolution branch to generate a texture distribution vector related to the surface texture of the crop leaves.
[0054] In the embodiments of the present invention, the texture distribution vector is an encoding of the leaf surface pattern based on gradient direction statistics. The histogram of oriented gradients filtering layer divides the Red Fuji apple leaf image into overlapping block regions (such as 16×16 pixel blocks with a step size of 8 pixels), and each block contains 4 cell units (8×8 pixels). Within each unit, the pixel gradient direction (for example, divided into 9 intervals from 0° to 180°) is calculated and the histogram is statistically analyzed. Through concatenating the unit histograms within the block and L2-Hys normalization (re-normalization after truncating the threshold at 0.2), texture features with rotational invariance are generated. Taking the leaves of Red Fuji apples during the mature period as an example, the gradient direction histogram of healthy leaves shows a bimodal distribution in the 45° and 135° directions (corresponding to parallel leaf veins and perpendicular lateral veins), while the diseased leaves show a multi-modal disordered distribution due to vein breaks. After dimensionality reduction through a fully connected layer, the texture distribution vector is encoded in 256 dimensions to represent the texture complexity of the leaves under different physiological states. For example, high-dimensional values reflect the disordered speckled gradient directions caused by apple leaf rust, and low-dimensional values correspond to the direction consistency of healthy leaves.
[0055] Step S340: Calculate the weight coefficients of different spectral channels through the spectral attention mechanism layer in the third convolution branch, and perform channel weighted fusion on the multi-spectral feature map according to the weight coefficients to generate a spectral response vector related to the physiological state of the crop.
[0056] In the embodiments of the present invention, the spectral response vector is a comprehensive physiological activity index of Red Fuji apples that integrates multi-spectral information. The spectral attention mechanism layer first obtains the global features of each channel through global average pooling (for example, the pooling value of the visible light R channel reflects the fruit coloring degree, and the pooling value of the near-infrared channel is related to the chlorophyll content). Subsequently, two fully connected networks (with the hidden layer dimension being 1 / 4 of the number of channels) are used to generate channel attention weights. The weight coefficients are normalized to the 0-1 interval through the Sigmoid function to guide the weighted fusion of the feature map channels. Taking the image of Red Fuji apples during the swelling period as an example, the weight of the near-infrared channel may be increased to 0.9 (reflecting the activity of water transportation), the weight of the short-wave infrared channel is reduced to 0.3 (because the cell structure changes are not significant during this stage), and the weight of the visible light B channel is set to 0.6 to balance the fruit color and background interference. The weighted fused feature map is compressed into a 128-dimensional spectral response vector through spatial pyramid pooling. The high value of a specific dimension (such as the 64th dimension) corresponds to the peak photosynthesis efficiency, providing a quantitative basis for the health status score.
[0057] As an implementation manner, in step S400, cross-stage correlation analysis is performed in the fully connected layer of the multi-task deep learning model, and multi-task classification results corresponding to the growth cycle of the target crop are output, which may specifically include:
[0058] Step S410: Concatenate the morphological feature vector, the texture distribution vector, and the spectral response vector to generate a combined feature vector, and input the combined feature vector into the time series analysis module in the fully connected layer.
[0059] In the embodiments of the present invention, the combined feature vector can be constructed, for example, by concatenating morphological (such as 512 dimensions), texture (such as 256 dimensions), and spectral (such as 128 dimensions) features to form an 896-dimensional joint representation, integrating multi-source information on the growth state of Red Fuji apples. The time series analysis module receives a sequence of combined feature vectors for a continuous time period (for example, collected daily from the germination period to the maturity period), and forms an input tensor with a time step of T (dimension T×896) after sorting by timestamp. Taking the annual growth cycle of Red Fuji apples as an example, T is set to 365 corresponding to daily data. The module processes local time series data through a sliding window (for example, the window size is 30 days and the step size is 1 day) to capture the feature evolution law (such as the S-shaped curve of the fruit swelling rate changing with the accumulated temperature). The high-dimensional fusion of the combined feature vectors ensures that the model simultaneously perceives the synergistic effects of plant type structure changes (morphological vector), leaf senescence process (texture vector), and photosynthetic efficiency fluctuations (spectral vector), providing a complete data basis for cross-stage correlation analysis.
[0060] Step S420: Perform time-dependence modeling on the combined feature vector through a bidirectional long short-term memory network in the time series analysis module to capture the growth trend change features in the continuous time period of the standardized crop image set.
[0061] In the embodiments of the present invention, the Bidirectional Long Short-Term Memory Network (Bi-LSTM) can learn the context dependencies of historical and future time steps through forward and backward LSTM units respectively. Each LSTM unit includes an input gate, a forget gate, an output gate, and a cell state, dynamically adjusting the information flow to model the temporal dynamics of Red Fuji apple growth. The forward LSTM processes the combined feature vectors sequentially from time step 1 to T, capturing the progressive feature evolution from the budding stage to the maturity stage (such as the cumulative effect from inflorescence differentiation to fruit coloring); the backward LSTM processes in reverse order from time step T to 1, identifying the retrospective associations from the maturity stage to the budding stage (such as the dependence of fruit quality on the accumulated temperature during the flowering period). The two-way outputs are concatenated to form a 1024-dimensional (512 dimensions for each direction) temporal feature vector, where the high-response dimensions may correspond to the turning points of key phenological periods (for example, the sudden increase in the 256th dimension during the flowering period reflects the photoperiod sensitivity). Taking the flower bud differentiation stage of Red Fuji apple as an example, Bi-LSTM captures the delayed effect of low-temperature accumulation on flower bud quality, generating the growth trend change characteristics with time causality.
[0062] Step S430: According to the growth trend change characteristics, calculate the first classification probability distribution corresponding to the growth stage category label, the second regression value corresponding to the growth health status score, and the third regression value corresponding to the environmental adaptability index in the fully connected layer respectively
[0063] In the embodiments of the present invention, the fully connected layer can realize the parallel prediction of growth parameters through a multi-task learning head. For the growth stage classification task (such as budding stage, leaf expansion stage, flowering stage, fruit development stage, maturity stage, dormancy stage), the first classification probability distribution calculates the logical values of each category through the fully connected layer before Softmax (such as input 1024 dimensions, output 6 dimensions); for the health status score regression task, the second regression value is mapped to the 0-100 interval through a linear layer (such as input 1024 dimensions, output 1 dimension) to quantify the overall physiological state of Red Fuji apple (for example, a score of 80 indicates normal chlorophyll content and no visible diseases); the environmental adaptability index regression task generates a 0-1 value through another linear layer (such as input 1024 dimensions, output 1 dimension) to evaluate the tolerance of the plant to temperature and humidity fluctuations (for example, an index of 0.7 indicates maintaining normal metabolism within a 10°C day-night temperature difference). Taking the color change stage of Red Fuji apple as an example, the probability of the "maturity stage" category in the first classification probability distribution reaches 0.92, the second regression value of 85 reflects that the fruit sugar content meets the standard, and the third regression value of 0.65 indicates the need to guard against the decline in adaptability caused by sunburn.
[0064] Step S440: Perform Softmax normalization processing on the first classification probability distribution to determine the growth stage category label; at the same time, perform dynamic range scaling on the second regression value and the third regression value, so that the growth health status score and the environmental adaptability index are mapped to a preset numerical interval.
[0065] In the embodiment of the present invention, the Softmax function converts the first classification logical value into a probability distribution, and selects the category corresponding to the maximum probability as the growth stage category label (for example, if the probability of the germination stage is 0.15, the flowering stage is 0.73, and the maturity stage is 0.12, then it is determined as the flowering stage). The dynamic range scaling adopts a linear transformation formula: health status score = original regression value × 25 + 50, mapping the [-2, 2] interval output by the model to [0, 100]; the environmental adaptability index is compressed to [0, 1] through the Sigmoid function, and the formula is 1 / (1 + exp(-original regression value)). Taking the drought-affected Red Fuji apple as an example, when the original regression value of the model is -0.8, the health score drops to 30 after scaling (preset threshold warning), and the environmental adaptability index is 0.33 (lower than the safety threshold of 0.5), triggering the generation of an irrigation control instruction.
[0066] Step S450: Output the growth stage category label, the growth health status score, and the environmental adaptability index as the multi-task classification result.
[0067] In the embodiment of the present invention, the multi-task classification result can be integrated and output in a structured data format (such as a JSON object), including discrete category labels and continuous regression values. The typical output for the flowering stage of Red Fuji apples is: the growth stage category label "flowering stage" (such as encoded 03), the growth health status score of 78 (such as corresponding to good integrity of flower organs), and the environmental adaptability index of 0.82 (such as reflecting the ideal pollination conditions of suitable temperature and high humidity). This result is synchronously stored in the time series record table of the crop management database and associated with the agronomic operation logs (such as the time of thinning flowers and the amount of fertilizer application) to support the closed-loop verification of the growth optimization strategy. Through the multi-task collaborative optimization mechanism, the model meets the agricultural application-level accuracy requirements in terms of classification accuracy (such as the F1-score for flowering stage recognition is 94%), scoring error (such as the MAE for health status is 2.3), and index correlation (such as the R² for environmental adaptability is 0.89).
[0068] As an implementation method, in the above step S500, a stage recognition report corresponding to the entire growth cycle of the target crop is generated, which may specifically include:
[0069] Step S510: According to the growth stage category label in the multi-task classification result, retrieve the standard growth parameter range corresponding to the current growth stage from the preset crop growth stage knowledge base; the standard growth parameter range includes the leaf expansion threshold, the stem height interval, and the chlorophyll content reference value.
[0070] In the embodiments of the present invention, the crop growth stage knowledge base can be a structured database that stores the standard physiological parameters and agronomic indicators of each growth stage of Red Fuji apples. Based on the growth stage category label in the multi-task classification result (for example, the code "03" corresponds to the flowering period), the system retrieves the standard growth parameter range of this stage from the "Red Fuji apple growth stage parameter table" in the knowledge base through an SQL query statement. The leaf expansion threshold refers to the minimum expansion length that the leaves should reach in the current stage (for example, the leaf expansion threshold in the flowering period is 25 cm), which is used to determine whether the photosynthesis area meets the standard; the stem height range refers to the reasonable height range from the base of the main branch to the apical bud (for example, the stem height range in the flowering period is 40 - 60 cm), which is used to evaluate the rationality of the tree structure; the chlorophyll content reference value is quantified in SPAD units (for example, the reference value in the flowering period is 35 - 50 SPAD), which is used to diagnose the nitrogen nutrient level of the leaves. Taking the germination period of Red Fuji apples as an example, the retrieval results are a leaf expansion threshold of 5 - 10 cm, a stem height range of 20 - 30 cm, and a chlorophyll content reference value of 25 - 40 SPAD. The system compares the above parameters with the real-time monitoring data to generate a basis for analyzing the deviation degree of the growth state.
[0071] Step S520: Compare the growth health status score in the multi-task classification result with the leaf expansion threshold in the standard growth parameter range. If the growth health status score is lower than the leaf expansion threshold, it is determined that the target crop has an abnormal nutrient absorption index.
[0072] In the embodiments of the present invention, the growth health status score can convert the regression value output by the model into an actual physical quantity through linear mapping (for example, a score of 70 corresponds to a leaf expansion of 23 cm, and a score of 90 corresponds to an expansion of 28 cm). The system calculates the relative deviation between the growth health status score mapping value and this threshold according to the leaf expansion threshold corresponding to the current growth stage category label (for example, the threshold in the flowering period is 25 cm): If the mapping value is lower than the lower limit of the threshold (for example, a score of 65 corresponds to a leaf expansion of 22 cm, which is lower than the threshold of 25 cm), the determination of abnormal nutrient absorption is triggered. Specifically, during implementation, the insufficient leaf expansion in the flowering period of Red Fuji apples may be caused by restricted root development or an imbalance in the soil nitrogen-phosphorus ratio. The system determines the specific abnormal type (for example, a score of 62 corresponds to nitrogen-deficient nutrient absorption abnormality) by associating with the abnormal pattern library in the knowledge base (for example, "low leaf expansion + normal chlorophyll" corresponds to water stress, and "low leaf expansion + low chlorophyll" corresponds to nitrogen deficiency symptoms), generates a diagnostic conclusion, and records it in the abnormal index log.
[0073] Step S530: Conduct a correlation analysis between the environmental adaptability index in the multi-task classification result and the stem height range in the standard growth parameter range. If the environmental adaptability index deviates from the fitness threshold corresponding to the stem height range, it is determined that the target crop has an abnormal environmental response index.
[0074] In the embodiments of the present invention, the environmental adaptability index and the fitness threshold of the stem height range are associated through a preset fitting curve. For example, the stem height range during the flowering period of Red Fuji apples is 40 - 60 cm, corresponding to the theoretical threshold of the environmental adaptability index of 0.5 - 0.8 (the index 0.5 represents the minimum adaptation requirement when the stem height is 40 cm, and 0.8 corresponds to the ideal state when the stem height is 60 cm). When the real-time environmental adaptability index is lower than 0.5 (for example, the index is 0.45), it is determined that the stem height has not reached the lower limit (for example, the measured height is 38 cm), and there is an abnormal environmental response; if the index is higher than 0.8 (for example, the index is 0.85), the stem height may exceed the upper limit (the measured height is 65 cm), indicating the risk of excessive vegetative growth. When making an abnormal determination, the system synchronously analyzes the data of environmental sensors (such as soil humidity, daily average temperature), determines the abnormal cause (for example, the index 0.45 may be associated with root hypoxia caused by continuous rainfall), and writes it into the environmental response abnormal index report to provide a basis for matching remedial measures.
[0075] Step S540: Based on the abnormal nutrient absorption index and the abnormal environmental response index, match the corresponding set of remedial measures from the crop growth stage knowledge base; the set of remedial measures includes fertilization plan adjustment parameters, irrigation frequency correction coefficients, and light intensity compensation strategies.
[0076] In the embodiments of the present invention, the set of remedial measures can be retrieved and generated from the "Red Fuji Apple Abnormal Response Strategy Table" in the knowledge base through a decision tree algorithm. The input parameters are the type of abnormal nutrient absorption (such as nitrogen deficiency type) and the type of abnormal environmental response (such as insufficient stem height + excessive humidity). The system traverses the conditional branches of the strategy table to match the remedial entries that meet all abnormal constraints. The fertilization plan adjustment parameters include the additional amount of nitrogen fertilizer (such as 10 kg / acre of urea), the adjustment of the phosphorus-potassium ratio (such as adjusting from 1:0.5:1 to 1:0.8:1); the irrigation frequency correction coefficient is dynamically calculated according to the soil humidity deviation (for example, the current humidity of 85% exceeds the threshold upper limit of 75%, and the coefficient is set to 0.7, extending the original 7-day irrigation cycle to 10 days); the light intensity compensation strategy coordinately adjusts the canopy light environment through the sunshade net coverage rate (such as 30%) and the on-time duration of supplementary lights (such as increasing 2 hours per day). Taking the example of Red Fuji apples encountering low temperature during the flowering period, the set of remedial measures may include spraying antifreeze (concentration 0.5%), delaying the flower thinning operation (postponing for 5 days), and increasing the inter-row film mulching (raising the ground temperature by 2°C) to ensure the normal development of flower organs.
[0077] Step S550: Generate a stage identification report based on the growth stage category label and the set of remedial measures in the multi-task classification result, convert the fertilization plan adjustment parameters and irrigation frequency correction coefficients in the stage identification report into a control instruction set, and transmit it to the intelligent irrigation device and fertilization device deployed in the target crop area through the IoT gateway, triggering the intelligent irrigation device and fertilization device to execute the operations corresponding to the control instruction set.
[0078] In the embodiment of the present invention, the stage identification report can be encapsulated in XML or JSON format, including the growth stage category label (such as "flowering stage"), the list of abnormal indicators (such as "abnormal nitrogen-deficient nutrient absorption", "abnormal environmental response with insufficient stem height"), and the details of remedial measures (such as "top-dress 10 kg / acre of urea", "adjust the irrigation cycle to 10 days"). The control instruction set is encapsulated into an instruction code that can be parsed by the device through the MQTT protocol. For example, the fertilization instruction "FERT_CMD{N:10,P:8,K:10}" means adjusting the application rates of nitrogen, phosphorus, and potassium to 10 kg, 8 kg, and 10 kg / acre; the irrigation instruction "IRRIG_CMD{CYCLE:240,DURATION:30}" means setting the irrigation cycle to 240 hours (10 days) and the single irrigation duration to 30 minutes. The instruction set is sent to the intelligent water and fertilizer integrated machine and drip irrigation controller in the target area through the IoT gateway (such as LoRaWAN gateway or 5G base station), triggering the device to perform quantitative fertilization and precise irrigation.
[0079] As an implementation manner, the method provided by the embodiment of the present invention further includes the process of training the multi-task deep learning model, which may specifically include:
[0080] Step S10: Obtain a historical crop growth cycle data set, which includes multi-spectral images of multiple crop samples at different growth stages, corresponding growth stage annotations, health status scores, and environmental parameter records.
[0081] In the embodiments of the present invention, the historical crop growth cycle dataset can be constructed by long-term monitoring of the Red Fuji apple planting area, covering the phenotypic and environmental data of each phenological stage within the complete growth cycle. The multi-spectral images in the dataset are collected by an airborne hyperspectral camera on a drone at fixed time intervals (e.g., noon every day), covering the visible light (400 - 700 nm), near-infrared (700 - 1300 nm), and short-wave infrared (1300 - 2500 nm) bands. Each image frame is attached with longitude and latitude coordinates and the acquisition timestamp. The growth stage annotation is divided according to the phenological period standards formulated by agricultural experts, including 7 categories: germination stage, leaf expansion stage, flowering stage, fruit swelling stage, color change stage, maturity stage, and dormancy stage. 3 independent annotators perform cross-validation on the same image to ensure annotation consistency. The health status score is comprehensively calculated based on the leaf chlorophyll content (SPAD value), fruit sugar content (Brix value), and the proportion of diseased area. The scoring range is 0 - 100, and it is measured on-site by a handheld detection device and associated with the corresponding image. The environmental parameter records include temperature, humidity, light intensity, rainfall collected by a weather station, and pH value, nitrogen, phosphorus, and potassium content monitored by soil sensors. They are stored at an hourly granularity and aligned with the image timestamp. Taking the color change stage of Red Fuji apples as an example, the dataset contains 5000 multi-spectral images of this stage, labeled as "color change stage", with the health status scores distributed in the range of 75 - 92. The environmental parameter records show that the average daily temperature is 20 - 25 °C and the soil humidity is 60 - 70%, constituting the basic data units for model training.
[0082] Step S20: Perform standardized illumination adjustment processing and crop area segmentation processing on each multi-spectral image in the historical crop growth cycle dataset to generate a training standardized crop image set.
[0083] In the embodiments of the present invention, the standardized illumination adjustment processing follows the aforementioned operation method to correct the illumination anomalies of the Red Fuji apple multi-spectral images. Specifically, when implementing, for each image frame in the historical dataset, the brightness histogram of each spectral channel is extracted. Based on the preset standard illumination range threshold (e.g., the near-infrared channel threshold [15 - 240]), overexposed and underexposed areas are identified, and dynamic contrast compression and adaptive gamma correction are respectively applied. Then, a standardized image is generated through an illumination equalization model. The crop area segmentation processing adopts the process described in steps S110 - S150. High-frequency noise is suppressed through anisotropic diffusion filtering, and multi-scale edge features and Lab color space information are fused. Finally, an accurate mask of the Red Fuji apple tree crown is extracted and cropped to 512 × 512 pixels. For example, after processing the original image of the flowering stage collected in 2019, the brightness of the overexposed petal area drops from 250 to 230, and the brightness of the underexposed leaf area rises from 18 to 25. After segmentation, the image only retains the inflorescence and leaf structures, eliminating the interference of the soil background, and forming training samples with unified size and illumination conditions.
[0084] Step S30: Construct an initial multi-task deep learning model, which includes parallel convolutional branches, a fully connected layer, and a time series analysis module.
[0085] In the embodiment of the present invention, the initial multi-task deep learning model can adopt a heterogeneous multi-branch architecture design. In the parallel convolutional branches, the first convolutional branch is configured with 3 dilated convolutional layers with increasing dilation rates (such as 2, 4, 8) for extracting multi-scale features of the stem morphology of Red Fuji apples; the second convolutional branch integrates a histogram of oriented gradients filtering layer, sets the cell unit size to 8×8 pixels, and has 9 direction intervals, which is specifically used to quantify the direction distribution of leaf textures; the third convolutional branch introduces a spectral attention mechanism layer, generates a channel weight vector through a fully connected network, and dynamically fuses multi-spectral features. The fully connected layer contains 1024 neurons, receives the concatenated multi-modal feature vectors (512-dimensional morphology + 256-dimensional texture + 128-dimensional spectrum), and prevents overfitting through a Dropout layer (ratio 0.5). The time series analysis module is composed of a stacked bidirectional long short-term memory network (Bi-LSTM), with 256 LSTM units in each layer, and the time step is set to 30 to match the monthly cycle growth law. The model output layer is divided into three task heads: a growth stage classification head (Softmax activation), a health status regression head (linear activation), and an environmental adaptability regression head (Sigmoid activation), corresponding to a 7-class probability distribution, a 0-100 score, and a 0-1 index respectively.
[0086] Step S40: Input the standardized crop image set for training into the initial multi-task deep learning model, and optimize the model parameters through a joint loss function; the joint loss function includes the weighted sum of the growth stage classification loss, the health status regression loss, and the environmental adaptability regression loss.
[0087] In the embodiment of the present invention, the joint loss function is defined as L for example total =α·L class +β·L health +γ·L adapt, where α, β, and γ are dynamic weight coefficients. The classification loss for the growth stage uses the cross-entropy loss function to calculate the difference between the predicted probability distribution and the true annotation. For example, the true label of a Red Fuji apple flowering period sample is the one-hot vector [0, 0, 1, 0, 0, 0, 0], and the loss value corresponding to the model prediction probability [0.02, 0.15, 0.68, 0.10, 0.05, 0.00, 0.00] is -log(0.68). The regression loss for the health status uses the smooth L1 loss function to perform a robust calculation on the absolute error between the predicted score and the true score. For example, the loss between the true score of 85 and the predicted value of 82 is |85 - 82| = 3. The regression loss for the environmental adaptability uses the mean square error function. For example, the loss between the true index of 0.7 (the theoretical value calculated based on environmental parameters) and the predicted value of 0.65 is (0.7 - 0.65) 2 = 0.0025. During the training process, the initial weights are set to α = 1.0, β = 0.5, and γ = 0.5, and are dynamically adjusted according to the gradient magnitudes of each loss term after each round of iteration: If the gradient norm of the classification loss is more than twice that of the regression loss, then α is reduced to 0.8 and β and γ are increased to 0.6, and vice versa for the reverse adjustment. The optimizer selects AdamW, with an initial learning rate of 3e-4, a batch size of 32, and gradually converges during 100 rounds of training.
[0088] Step S50: Generate a pre-trained multi-task deep learning model according to the optimized model parameters, and verify the multi-task classification accuracy of the model in the test dataset.
[0089] In the embodiment of the present invention, the test dataset may include 2,000 groups of samples from the independent planting area of Red Fuji apples, without data augmentation processing of the training set data. The evaluation metrics for multi-task classification accuracy include, for example: the classification accuracy of the growth stage (Top-1 Accuracy), the mean absolute error (MAE) of the health status score, and the coefficient of determination (R²) of the environmental adaptability index. For example, the model achieves a classification accuracy of 93.2% (the recognition accuracy of the flowering period is 96.5%) on the test set, the health status MAE is 2.8 (the scoring error is ±2.8 points), and the environmental adaptability R² is 0.86 (the predicted value and the theoretical value are highly linearly correlated). Performance verification analyzes specific category biases through a confusion matrix. For example, the proportion of samples in the color-changing period misjudged as the mature period is less than 3%, indicating that the model has a strong ability to distinguish color gradient features. After the final model parameters are solidified, they are deployed to edge computing devices to support real-time processing of Red Fuji apple monitoring data.
[0090] As an implementation, the method provided in the embodiment of the present invention further includes a process of performing data augmentation processing on the standardized crop image set for training, which may specifically include:
[0091] Step S21: Randomly apply rotation transformation, mirror flipping transformation, and scale scaling transformation to each image in the training standardized crop image set to generate enhanced image samples.
[0092] In the embodiments of the present invention, data augmentation aims to improve the generalization ability of the model to the pose and scale changes of Red Fuji apples. The rotation transformation takes the center of the image as the axis and randomly rotates within the angle range of ±30°, simulating different aerial photography angles of drones; the mirror flipping includes horizontal and vertical directions, with the probability set to 50% each, enhancing the adaptability to the symmetric structure of fruit trees; the scale scaling randomly adjusts the image size to 80 - 120% of the original image, and then crops it back to 512×512 pixels, simulating the observation perspectives at different distances. For example, a close-up image of a fruit collected at noon, after being rotated by 25°, flipped horizontally, and scaled to 90%, the position of the fruit stalk moves from the lower left to the upper right, and the fruit size slightly shrinks, but the texture and spectral features remain intact, effectively expanding the diversity of training samples.
[0093] Step S22: Perform color jitter processing on the local regions in the enhanced image samples to simulate the phenotypic changes of crops under different lighting conditions, and obtain jitter-enhanced image samples.
[0094] In the embodiments of the present invention, color jitter can be, for example, simulated by adjusting the hue (H), saturation (S), and value (V) channels in the HSV color space for light fluctuations. Randomly apply ±15° hue shift (simulating the color temperature change of morning and evening light), ±30% saturation change (such as simulating the fading effect of strong light), and ±20% value change (simulating the brightness fluctuation caused by cloud occlusion) to the Red Fuji apple image. For example, the enhanced leaf image shows a yellow-green hue when the hue is +10°, the vein contrast decreases when the saturation is -20%, and the highlight area is overexposed when the value is +15%, forcing the model to learn color invariance features. The jitter parameters are randomly sampled according to the Gaussian distribution to ensure that the enhancement effect is natural and conforms to agronomic laws.
[0095] Step S23: Combine the jitter-enhanced image samples with the original training standardized crop image set to generate an extended training data set.
[0096] In the embodiments of the present invention, the original training set contains 12,000 images of Red Fuji apples at each growth stage. After rotation transformation, mirror flipping, and scale scaling, 12,000 enhanced samples are generated respectively, and color jitter further generates 12,000 jitter samples. After merging, the total amount of the extended data set reaches 48,000 images. Data expansion significantly alleviates the problem of class imbalance. For example, the number of samples in the dormant period in the original data is only 800, which is increased to 3,200 after enhancement, and the proportion increases from 6.7% to 6.8%, reducing the overfitting risk of the model to the majority class (such as the fruit swelling period). The extended data set is divided into training-validation-test subsets according to 7:2:1 to ensure distribution consistency.
[0097] Step S24: During the training process, according to the quantity distribution of samples at different growth stages in the expanded training dataset, dynamically adjust the batch sampling weights to make the learning balance of the model for each growth stage category.
[0098] In the embodiment of the present invention, the dynamic sampling weight mechanism can be realized in two steps. For example, first calculate the initial weights based on the inverse frequency of the sample quantity to suppress oversampling of the majority class; secondly, dynamically fine-tune the weights according to the accuracy of each category during the training process to strengthen the learning of difficult examples. For example, the number of samples in the germination stage, flowering stage, and maturity stage in the expanded dataset are 4000, 6000, and 5000 respectively. The initial sampling weights are calculated as 1 / 4000, 1 / 6000, and 1 / 5000. After normalization, the weight coefficients are 0.31, 0.21, and 0.25. During the mid-training, it is detected that the classification accuracy of the flowering stage reaches 95% while that of the germination stage is only 82%, then the weight adjustment is triggered: the weight of the germination stage is increased to 0.35, and the weight of the flowering stage is decreased to 0.18, forcing the model to increase the learning frequency of samples in the germination stage. The weight update frequency is set to every 5 epochs, and the adjustment amplitude is smoothed through a sliding window to avoid training oscillations.
[0099] As an implementation manner, in the above step S24, dynamically adjusting the batch sampling weights may specifically include:
[0100] Step S241: Count the number of samples corresponding to each growth stage category in the expanded training dataset, and calculate the proportion of the number of samples in each growth stage category.
[0101] In the embodiment of the present invention, the sample quantity statistics can be realized by traversing the annotation file of the dataset. The number of samples in the 7 growth stages of Red Fuji apples are 4000 in the germination stage, 3500 in the leaf expansion stage, 6000 in the flowering stage, 5500 in the fruit expansion stage, 5000 in the color conversion stage, 5000 in the maturity stage, and 3200 in the dormancy stage, with a total of 32,200 images. The proportion calculation is 12.4% in the germination stage, 10.9% in the leaf expansion stage, 18.6% in the flowering stage, 17.1% in the fruit expansion stage, 15.5% in the color conversion stage, 15.5% in the maturity stage, and 9.9% in the dormancy stage, reflecting the characteristics that the original data collection focuses on the flowering to maturity stages. The statistical results are stored as a hash table for subsequent weight calculation calls.
[0102] Step S242: Determine the initial sampling weights of each growth stage category according to the reciprocal of the proportion of the number of samples, and perform normalization processing on the initial sampling weights.
[0103] In the embodiments of the present invention, the initial sampling weight calculation can be the reciprocal of the proportion of each category. For example, the weight of the budding stage = 1 / 0.124 ≈ 8.06, the weight of the leaf - unfolding stage = 1 / 0.109 ≈ 9.17, and so on. To prevent the interference of extreme values, the original weights are truncated (the upper limit is 10 times the mean value, and the lower limit is 0.1 times the mean value), and then normalized by the Softmax function so that the sum is 1. After processing, the weight of the budding stage drops from 8.06 to 0.146, the weight of the dormancy stage is adjusted from 10.10 to 0.183, and the weight of the flowering stage is adjusted from 5.38 to 0.097, ensuring that the minority class (dormancy stage) has a higher sampling probability and the majority class (flowering stage) has its weight appropriately reduced.
[0104] Step S243: In each round of training iteration, according to the classification accuracy of the model for each growth stage category in the current training stage, the initial sampling weight is dynamically corrected.
[0105] In the embodiments of the present invention, the classification accuracy can be measured by the F1 - score of each category in the validation set. For example, after the 20th round of training, the F1 - score of the budding stage is 82%, the F1 - score of the flowering stage is 96%, and the F1 - score of the dormancy stage is 75%. Then, the relative ratio of the classification accuracy of each category is calculated as 82:96:75. The dynamic correction formula is: new weight = initial weight × (benchmark accuracy / category accuracy), where the benchmark accuracy takes the average value of all categories (for example, 84.3%). The weight of the budding stage is corrected to 0.146×(84.3 / 82)=0.150, the weight of the dormancy stage is corrected to 0.183×(84.3 / 75)=0.207, and the weight of the flowering stage is corrected to 0.097×(84.3 / 96)=0.085, so that the weights of the low - accuracy categories are increased and the weights of the high - accuracy categories are decreased.
[0106] Step S244: If the classification accuracy of any growth stage category is lower than the preset threshold, then increase the sampling weight corresponding to any growth stage category; otherwise, maintain or reduce the sampling weight.
[0107] For example, the preset classification accuracy threshold in the embodiments of the present invention can be set to 80%. When the F1 - score of a certain category is lower than the threshold for 3 consecutive rounds, an emergency increase in weight is triggered. For example, the F1 - scores of the dormancy stage in the 25th - 27th rounds are 78%, 77%, and 76% respectively, then its weight is increased from 0.207 to 0.250, with an increase of 20%. On the contrary, if the accuracy of a certain category continuously exceeds 95% for 5 rounds, the weight is decreased by 10%. This mechanism prevents the model from being under - fitted to the minority class and at the same time avoids over - fitting of the majority class.
[0108] Step S245: Extract batch training samples from the extended training dataset according to the corrected sampling weights.
[0109] In the embodiments of the present invention, batch sampling can adopt a stratified sampling strategy. Each batch contains 32 samples, and the number of samples for each category is allocated according to the corrected weights. For example, if the corrected weights are 0.150 for the germination stage, 0.085 for the flowering stage, and 0.250 for the dormancy stage, then the number of samples in the germination stage in the batch = 32×0.150≈5, the flowering stage≈3, the dormancy stage≈8, and the remaining categories are allocated proportionally. The sampling process is implemented through a weighted random number generator to ensure that the sample distribution of each category in each round of training conforms to the current weight configuration. This dynamic mechanism increases the frequency of the samples in the dormancy stage of Red Fuji apples within an epoch by 2.3 times, significantly improving the feature learning effect of the model for this stage.
[0110] As an implementation manner, in step S40, optimizing the model parameters through a joint loss function may include the following steps:
[0111] Step S41: Measure the difference between the growth stage category label predicted by the model and the true annotation through a cross-entropy loss function as the growth stage classification loss.
[0112] In the embodiments of the present invention, the cross-entropy loss function is used to quantify the prediction error of the multi-task deep learning model for the growth stage classification of Red Fuji apples. Specifically, when implemented, the growth stage classification head of the model outputs a 7-dimensional probability distribution (corresponding to the germination stage, leaf expansion stage, flowering stage, fruit expansion stage, color conversion stage, maturity stage, dormancy stage), and the true annotation is a one-hot encoded category label. Taking the samples in the flowering stage of Red Fuji apples as an example, the true label is [0,0,1,0,0,0,0], and the model prediction probability distribution is [0.02,0.15,0.68,0.10,0.05,0.00,0.00]. The cross-entropy loss is calculated as -log(0.68)=0.385. When this loss value is backpropagated, it drives the update of the model parameters to improve the prediction confidence of the flowering stage category and at the same time suppress the misjudgment probability of other categories. For the class imbalance problem existing in the dataset (such as fewer samples in the dormancy stage), the cross-entropy loss alleviates the overfitting tendency of the model to the majority class through a weighted strategy (the class weight is inversely proportional to the number of samples).
[0113] Step S42: Measure the difference between the growth health status score predicted by the model and the true score through a smooth L1 loss function as the health status regression loss.
[0114] In the embodiments of the present invention, the smooth L1 loss function is used to evaluate the regression accuracy of the model for the health status score of Red Fuji apples, which combines the robustness of the L1 loss and the smoothness of the L2 loss. For example, the formula is defined as: if the absolute error |y pred -y true |<1, the loss is ; otherwise it is |y pred -ytrue |-0.5. For example, when the true health score is 85 (calculated comprehensively based on a chlorophyll content of 45 SPAD, a fruit sugar content of 14 Brix, and a disease area of 2%), and the model predicted value is 82, the absolute error 3 > 1, and the loss value is 3 - 0.5 = 2.5. This loss function has strong fault tolerance for outliers (such as score mutations caused by sensor failures), and at the same time ensures gradient smoothness within a small error range, promoting the stable convergence of the model. In the health score regression task during the color transition period of Red Fuji apples, the smooth L1 loss effectively handles local score deviations caused by light reflection, improving the overall regression accuracy.
[0115] Step S43: Measure the difference between the environmental adaptability index predicted by the model and the theoretical adaptability index calculated based on environmental parameter records through the mean squared error loss function as the environmental adaptability regression loss.
[0116] In the embodiments of the present invention, the mean squared error loss function is used to optimize the regression accuracy of the environmental adaptability index of Red Fuji apples. The theoretical adaptability index is calculated through a multiple linear regression model of environmental parameters (such as daily average temperature, soil humidity, and light duration) and physiological indicators (such as transpiration rate and photosynthetic efficiency), and is normalized to the 0 - 1 interval. For example, the theoretical index 0.75 corresponds to the ideal conditions of a daily average temperature of 22 °C, a humidity of 65%, and 12 hours of light. When the model predicted value is 0.70, the mean squared error loss is (0.75 - 0.70)^2 = 0.0025. This loss function emphasizes sensitivity to subtle environmental fluctuations, forcing the model to learn the adaptation decay pattern under extreme conditions such as sudden drops in temperature or continuous drought. In the training samples during the budding period of Red Fuji apples, the mean squared error loss preferentially optimizes the prediction of the adaptability index under low temperature stress (<10 °C), ensuring that the model captures the non-linear relationship between budding delay and environmental factors.
[0117] Step S44: Dynamically adjust the weight coefficients of each loss term according to the historical gradient magnitudes of the growth stage classification loss, the health status regression loss, and the environmental adaptability regression loss, so that the optimization speeds of each task during the training process are balanced.
[0118] In the embodiments of the present invention, the dynamic weight adjustment mechanism realizes the balanced optimization of multi-task learning by monitoring the gradient norms of each loss term (i.e., the L2 norm of the gradient vector). After each round of training iteration, calculate the gradient norms g cls 、g health 、g adapt of the growth stage classification loss, the health status regression loss, and the environmental adaptability regression loss, and compare them with the historical moving averages μ cls 、μ health 、μ adapt (decay factor 0.9). The adjustment factor α t = g cls / μcls , β t = g health / μ health , γ t = g adapt / μ adapt . If α t > 1.2 (the classification gradient is significantly higher than the historical level), then reduce the classification loss weight α from 1.0 to 0.8 to prevent the classification task from dominating the parameter update; if γ t <0.8 (the environmental adaptation gradient remains low), then increase γ from 0.5 to 0.7 to enhance the learning intensity of this task. This mechanism ensures that during the mid - training stage of the Red Fuji apple multi - task model, the gradient contribution degrees of the classification and regression tasks are maintained in a proportionally balanced range (for example, classification: health: environment ≈ 1:0.5:0.5), avoiding overfitting or underfitting of a single task.
[0119] Step S45: Update the parameters of the initial multi - task deep learning model through the backpropagation algorithm until the joint loss function converges to a preset threshold.
[0120] In the embodiments of the present invention, the backpropagation algorithm can calculate the gradient of the joint loss function with respect to the model parameters based on the chain rule and update the parameters through an optimizer (such as AdamW). Taking the weight matrix W of the fully - connected layer of the Red Fuji apple multi - task model as an example, its gradient is the weighted sum of three parts: . The learning rate adopts, for example, the cosine annealing strategy, with an initial value of 3e - 4, a minimum value of 1e - 5, and a period set to 50 rounds to promote the model to jump out of the local optimum. The training termination condition is, for example, that the decrease amplitude of the joint loss function on the validation set is less than 1e - 4 for 10 consecutive rounds, or the total number of training rounds reaches 200.
[0121] As an implementation method, in step S44, dynamically adjusting the weight coefficients of each loss term may include the following steps:
[0122] Step S441: In each round of training iteration, calculate the gradient norms of the classification loss during the growth stage, the regression loss of the health status, and the regression loss of the environmental adaptation ability with respect to the model parameters respectively.
[0123] In the embodiments of the present invention, the gradient norm calculation can be implemented through an automatic differentiation tool (such as autograd of PyTorch). For the classification loss L cls , calculate its gradient i with respect to each trainable parameter θ , and obtain the L2 norm . Similarly, calculate and . Taking the weight parameters of the third convolutional layer of the Red Fuji apple model as an example, the classification gradient norm in a certain round of iteration , the health regression gradient norm , the environmental adaptation gradient norm . The gradient norm reflects the influence intensity of each loss term on parameter update. The classification task dominates at this layer, and multi-task learning needs to be balanced by weight adjustment.
[0124] Step S442: Determine the current adjustment factor of each loss term according to the ratio of the gradient norm corresponding to each loss term to its historical average value.
[0125] In the embodiment of the present invention, the historical average value is updated by exponential moving average (EMA), and the decay factor is set to 0.9. For example, the historical mean μ of the classification gradient cls is initialized to 1.0, and the current gradient norm g cls = 1.24. After update, μ cls = 0.9×1.0 + 0.1×1.24 = 1.024. The adjustment factor k cls = g cls / μ cls = 1.24 / 1.024 ≈ 1.21. Similarly, calculate k health = 0.87 / 0.95 ≈ 0.92, k adapt = 0.53 / 0.60 ≈ 0.88. An adjustment factor greater than 1 indicates that the current gradient is higher than the historical level, and the weight of this task needs to be reduced; less than 1 means increasing the weight. This mechanism enables the Red Fuji apple model to adaptively balance the learning speed of each task during training and avoid optimization bias caused by task difficulty differences.
[0126] Step S443: If the current adjustment factor of any loss term is greater than the preset upper bound threshold, reduce the weight coefficient of any loss term; if the current adjustment factor is less than the preset lower bound threshold, increase the weight coefficient of any loss term.
[0127] Suppose the preset upper bound threshold in the embodiment of the present invention is 1.2 and the lower bound threshold is 0.8. For example, k cls = 1.21 > 1.2, and the classification loss weight α is reduced from 1.0 to 0.8×1.0 = 0.8; k health = 0.92 ∈ [0.8, 1.2], and the health regression weight β remains 0.5; k adapt = 0.88 < 0.8, and the environmental adaptation weight γ is increased from 0.5 to 0.5×1.2 = 0.6. The weight adjustment amplitude is controlled by a multiplicative factor (such as 0.8 or 1.2) to prevent violent fluctuations from affecting training stability. This strategy ensures that during the training stage dominated by flowering period samples, the Red Fuji apple model appropriately suppresses the weight of the classification task, enhances the attention to the health status and environmental adaptation tasks, and improves the balance of multi-task prediction.
[0128] Step S444: Recalculate the weighted sum of the joint loss function according to the adjusted weight coefficients, and perform the next round of training iteration
[0129] In the embodiment of the present invention, for example, assume that the adjusted weight coefficients are α = 0.8, β = 0.5, and γ = 0.6, then the joint loss function is updated to L t otal = 0.8·L cls + 0.5·L health + 0.6·L adapt . In the next round of training, after the forward propagation of the batch data of Red Fuji apples (such as 32 samples), calculate the losses of the three tasks and sum them up with weights, and then generate new parameter gradients through backpropagation. After multiple rounds of dynamic adjustment, the model gradually converges to a state where the gradient norms of the three tasks are balanced (for example, k cls ≈ 1.0, k health ≈ 1.0, k adapt ≈ 1.0), and finally realize the collaborative optimization of the identification of the growth stage, health assessment and environmental adaptation prediction of Red Fuji apples.
[0130] It can be understood that in the calculation of the fusion of variables with different dimensions, in order to overcome the dimension difference, those skilled in the art can perform reasonable dimension elimination operations by means of conventional means, such as normalization, standardization operations, etc. The embodiments of the present invention do not elaborate on related operations too much.
[0131] In an optional derivative embodiment, after step S500 of generating a stage identification report corresponding to the entire growth cycle of the target crop, the method provided by the embodiment of the present invention may further include:
[0132] Step S600: Extract the historical growth and health status scores and environmental adaptability indexes from the crop management database, arrange the historical growth and health status scores and environmental adaptability indexes in a time series, and generate a crop growth trend data set.
[0133] In the embodiments of the present invention, the crop management database is designed with a relational architecture, for example. Its "Growth Record Table of Red Fuji Apples" includes a timestamp (accurate to the minute), geographical coordinates, a growth health status score (0 - 100), an environmental adaptability index (0 - 1), and associated environmental sensor data fields. Target crop records are filtered by SQL query statements according to a time range (e.g., the last three years) and geographical location (such as the apple planting belt at 35° north latitude), and score and index data of a continuous time series are extracted. For example, the data of the Red Fuji apple orchard numbered PF_Orchard_2023 contains 365 daily-granularity records. The growth health status score gradually increases from 45 at the budding stage to 92 at the maturity stage, and the environmental adaptability index varies in the range of 0.5 - 0.9 with seasonal temperature and humidity fluctuations. After data cleaning, a structured crop growth trend dataset is generated. Each row record of it contains the date, score, index, and standardized time encoding (such as day of year DOY), providing a standardized input for subsequent periodic analysis.
[0134] Step S700: Perform a periodic pattern analysis on the crop growth trend dataset to identify the growth rate fluctuation characteristics and environmental response lag duration of the target crop within a preset observation period.
[0135] In the embodiments of the present invention, the periodic pattern analysis adopts a method combining Fourier transform and autocorrelation function, for example. Perform spectral analysis on the sequence of the growth health status score of Red Fuji apples to extract the dominant periods (such as the annual period of 365 days, the semi-growth period of 180 days) and their harmonic components, and calculate the amplitude and phase corresponding to each period. The growth rate fluctuation characteristics are quantified by the variance of the first-order difference sequence. For example, the difference variance during the color-changing period (DOY210 - 240) is 15.2, which is significantly higher than 8.7 during the swelling period (DOY150 - 180), reflecting the drastic change in physiological activity during this stage. The environmental response lag duration is determined by cross-correlation analysis: perform sliding window cross-correlation calculation on the environmental adaptability index sequence and meteorological data (such as daily average temperature). It is found that the response lag of Red Fuji apples to temperature increase is 3 - 5 days (the offset of the cross-correlation peak), while the response lag to precipitation events is 1 - 2 days. This feature indicates that irrigation strategies need to be adjusted 3 days after continuous high temperature to relieve the heat stress effect.
[0136] Step S800: Based on the growth rate fluctuation characteristics and environmental response lag duration, load the growth kinetic parameters corresponding to the target crop variety into the time series prediction model to predict the expected growth stage evolution path within a future time window.
[0137] In the embodiments of the present invention, the time series prediction model adopts, for example, an autoregressive integrated moving average model with external variables (ARIMAX). Its growth kinetic parameters include the accumulated temperature requirement unique to the Red Fuji apple variety (for example, the effective accumulated temperature during the flowering period ≥ 850℃·d), photoperiod sensitivity (daily light ≥ 12 hours induces flower bud differentiation), and water use efficiency (water consumption per gram of dry matter is 400 - 500 mL). The model inputs are historical growth trend data, real-time environmental parameters, and meteorological forecast data after lag adjustment (such as the temperature and rainfall probability in the next 15 days). Taking the prediction of the budding period in 2024 as an example, after loading the parameters, the model outputs the growth stage evolution path for the next 30 days: DOY80 - 110 is the budding period, and the health score rises from 40 to 65; DOY111 - 140 is the leaf expansion period, and the score reaches 75; at the same time, it is predicted that the color change period is delayed by 5 days (the original DOY210 is extended to 215) due to insufficient accumulated temperature caused by the low temperature in April. The prediction results are presented in the form of a confidence interval (such as 80% confidence level), supporting risk-sensitive decision-making.
[0138] Step S900: Determine whether it is necessary to adjust the crop management strategy according to the deviation between the expected growth stage evolution path and the growth health status score in the current multi-task classification result; if the deviation exceeds the preset tolerance threshold, generate a set of dynamic regulation strategies including irrigation volume correction parameters, fertilization ratio adjustment coefficients, and shading period optimization.
[0139] In the embodiments of the present invention, the deviation is calculated as the absolute difference between the predicted score and the measured score divided by the width of the prediction interval. For example, the current measured score is 70 (DOY100), and the predicted score interval is [65, 75]. The deviation = |70 - 70| / (75 - 65) = 0 < the tolerance threshold of 0.3, and it is determined that no adjustment is required; if the measured score is 62 (deviation 0.8 > 0.3), then the strategy generation is triggered. The set of dynamic regulation strategies matches the rules in the knowledge base through a decision tree: the irrigation volume correction parameter is calculated according to the soil moisture deviation (for example, if the current humidity of 60% is lower than the target of 70%, the irrigation volume needs to be increased by 20%); the fertilization ratio adjustment coefficient is based on the detection result of the leaf nitrogen content (for example, if the N content of 2.5% is lower than the threshold of 3%, the nitrogen fertilizer ratio is increased by 15%); the shading period optimization is based on the light-temperature ratio (for example, when the sunshine hours / average daily temperature > 1.2, the sunshade net is enabled to cover from 10:00 to 14:00 every day). The set of strategies is sorted by priority (such as irrigation > fertilization > shading) to ensure the priority allocation of key resources.
[0140] Step S1000: Mark the set of dynamic regulation strategies in association with the stage identification report and update it to the strategy execution record table in the crop management database for subsequent crop management terminal devices to call according to the priority order.
[0141] In the embodiments of the present invention, the associated tag can be implemented through a unique transaction ID. For example, the policy ID PF_Adj_20240517_001 is linked to the stage identification report IDPF_Report_20240517. The newly added fields in the policy execution record table include policy type (irrigation / fertilization / lighting), parameter value, effective time, and expected impact period. For example, the irrigation volume correction parameter "+20%" is associated with the transaction ID, and it is expected to increase the soil humidity to 70% in the next 7 days. The crop management terminal device (such as an intelligent irrigation controller) polls the policy execution record table through the REST API, parses the instructions in the order of priority, and ensures that the operations are triggered within the specified time window. The historical policy execution data supports retrospective analysis to optimize the subsequent decision-making logic.
[0142] As an implementation manner, after generating the dynamic regulation policy set in step S900, the method provided by the embodiments of the present invention may further include the following steps:
[0143] Step S901: Generate a first control instruction set bound to the geographical coordinates of the target crop area according to the irrigation volume correction parameter in the dynamic regulation policy set, and send the first control instruction set to the intelligent irrigation system deployed in the corresponding area; the intelligent irrigation system adjusts the nozzle flow rate and irrigation time interval according to the first control instruction set.
[0144] In the embodiments of the present invention, the first control instruction set is encapsulated in the JSON-LD format and contains geographical coordinates (such as 34.5°N, 118.2°E), irrigation volume correction parameter (+20%), execution time window (06:00-08:00 every day), and device ID list (such as Zone1_Valve3). After parsing the instructions, the intelligent irrigation system drives the solenoid valve to extend the single irrigation duration from 30 minutes to 36 minutes, and the nozzle flow rate increases from 20 L / min to 24 L / min. The pressure sensor real-time feeds back the pipe network data. If the water pressure is abnormal (such as below 200 kPa), an abnormal handling protocol is triggered (such as switching to a standby pump). The instruction execution log (such as actual irrigation volume, energy consumption) is transmitted back to the crop management database for deviation analysis and device health monitoring.
[0145] Step S902: Calculate the mixing weights of different fertilizer types according to the fertilization ratio adjustment coefficient in the dynamic regulation policy set, generate a second control instruction set, and send the second control instruction set to the intelligent fertilizer mixing device; the intelligent fertilizer mixing device dynamically mixes nitrogen, phosphorus, and potassium fertilizers according to the second control instruction set and transports them to the root area of the target crop.
[0146] In the embodiments of the present invention, the second control instruction set is based on the fertilization ratio adjustment coefficient (such as Adjust the ratio from 1:0.5:1 to 1:0.8:1) to generate the ratio parameters executable by the production equipment. After receiving the instruction, the intelligent fertilization device controls the stepping motor to adjust the opening of the fertilizer storage tank valve: the flow rate of the nitrogen fertilizer tank maintains the reference value, the flow rate of the phosphate fertilizer tank increases by 60% (from 5 L / h to 8 L / h), and the flow rate of the potassium fertilizer tank remains unchanged. After the fertilizer solution is evenly diluted in the mixing chamber, it is transported to the root area of Red Fuji apple trees through the drip irrigation belt. The conductivity sensor monitors the concentration of the fertilizer solution in real time. If the deviation exceeds 5% (for example, the target EC is 2.5 mS / cm and the measured value is 2.7 mS / cm), the dynamic compensation algorithm is triggered to adjust the valve opening to ensure accurate fertilizer supply.
[0147] Step S903: Collect the new multi-spectral images of the crops in real time after executing the first control instruction set and the second control instruction set, and extract the updated growth health status score and environmental adaptability index through the multi-task deep learning model.
[0148] In the embodiment of the present invention, the new multi-spectral images are collected 24 hours after the strategy is executed to ensure that the regulation measures are effective. The airborne hyperspectral camera of the unmanned aerial vehicle flies according to the preset route to obtain the visible light-near infrared images (400-1300 nm) of the Red Fuji apple canopy. After the images are adjusted by standardized illumination and segmented by crop area, they are input into the multi-task deep learning model. For example, the health score after regulation rises from 62 to 68, and the environmental adaptability index rises from 0.55 to 0.63, indicating that the irrigation and fertilization adjustments are initially effective. The data collection frequency can be dynamically set according to the strategy influence period (for example, the shading strategy is collected once every 6 hours) to achieve high-frequency status monitoring.
[0149] Step S904: Compare the updated growth health status score and environmental adaptability index with the expected growth stage evolution path. If the deviation amount drops within the tolerance threshold, mark the current dynamic regulation strategy set as an effective strategy; otherwise, re-trigger the extraction of the crop growth trend data set and the generation process of the dynamic regulation strategy set until the deviation amount meets the preset conditions.
[0150] In the embodiment of the present invention, for example, the deviation amount threshold is set to 0.2 (allowing a 20% deviation from the prediction interval). For example, the deviation amount between the measured score of 68 and the predicted value of 70 after regulation is |68 - 70| / 10 = 0.2, which meets the threshold requirement, and the strategy is marked as effective. If the measured score is still 63 (deviation amount 0.7), re-execute steps S600 - S900: extract the latest trend data, identify the current fluctuation characteristics, and generate a correction strategy (such as additional foliar fertilizer spraying). The iterative optimization is performed at most 3 times. If the standard is still not met, an artificial intervention alarm is triggered.
[0151] Step S905: Integrate the final effective dynamic regulation strategy set, its corresponding multi-spectral images, and the updated multi-task classification results into a closed-loop control log and store it in the strategy verification archive in the crop management database.
[0152] In an embodiment of the present invention, the closed-loop control log includes, for example, a policy version number, an execution timeline, a multispectral image snapshot, and a score change curve. For example, the policy PF_Adj_20240517_001 is associated with the comparison of visible light images before and after regulation (the leaf color changes from yellow-green to dark green), the trend of the health score (62 → 68 → 73), and the change in the environmental index (0.55 → 0.63 → 0.70). After being subjected to the SHA-256 hashing operation, the log is stored in the blockchain evidence storage module to ensure the immutability of the data. The policy verification archive supports multi-dimensional queries (such as filtering by geographical location, crop variety, and policy type), provides a reference for intelligent decision-making in subsequent similar scenarios, and continuously optimizes the precision management efficiency of the entire growth cycle of Red Fuji apples.
[0153] Please refer to Figure 2 , an embodiment of the present invention further provides a computer system. The computer system 100 includes: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation to the embodiments of the present invention.
[0154] An embodiment of the present invention provides a computer system. The computer system in the embodiment of the present invention includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and are configured to be executed by one or more processors. When the one or more programs are executed by the processor, the methods described above in the embodiments of the present invention are implemented.
Claims
1. A method for identifying the entire growth cycle of crops based on deep learning, characterized in that, The method includes: Obtaining a multi - spectral image sequence of a target crop within a continuous time period through an image acquisition device, and performing standardized illumination adjustment processing on each image frame in the multi - spectral image sequence to obtain a set of standard illumination images corresponding to the multi - spectral image sequence; Performing crop region segmentation processing on each image in the set of standard illumination images, extracting local feature regions related to the target crop, and generating a set of standardized crop images containing the local feature regions; Inputting each image in the set of standardized crop images into a pre - trained multi - task deep learning model, and respectively extracting a morphological feature vector, a texture distribution vector, and a spectral response vector in the standardized crop image through parallel convolutional branches in the multi - task deep learning model; Concatenating the morphological feature vector, the texture distribution vector, and the spectral response vector to generate a combined feature vector, and inputting the combined feature vector into a time - series analysis module in a fully - connected layer; performing time - dependence modeling on the combined feature vector through a bidirectional long short - term memory network in the time - series analysis module to capture the growth trend change features within a continuous time period in the set of standardized crop images; according to the growth trend change features, respectively calculating a first classification probability distribution corresponding to a growth stage category label, a second regression value corresponding to a growth health status score, and a third regression value corresponding to an environmental adaptability index in the fully - connected layer; performing Softmax normalization processing on the first classification probability distribution to determine the growth stage category label; simultaneously performing dynamic range scaling on the second regression value and the third regression value so that the growth health status score and the environmental adaptability index are mapped to a preset numerical interval; outputting the growth stage category label, the growth health status score, and the environmental adaptability index as multi - task classification results; Generating a stage recognition report corresponding to the full growth cycle of the target crop according to the growth stage category label in the multi - task classification result, and associatively storing the growth health status score and the environmental adaptability index in a crop management database.
2. The method according to claim 1, wherein The performing standardized illumination adjustment processing on each image frame in the multi - spectral image sequence to obtain a set of standard illumination images corresponding to the multi - spectral image sequence includes: For each image frame in the multi - spectral image sequence, extracting the brightness distribution histogram of the image frame in different spectral channels, and determining the over - exposed region and the under - exposed region in the brightness distribution histogram that exceed the preset standard illumination range threshold; Performing a dynamic contrast compression algorithm on the over - exposed region and simultaneously performing an adaptive gamma correction algorithm on the under - exposed region to generate an intermediate adjustment image corresponding to the image frame; Inputting the intermediate adjustment image into an illumination equalization model, extracting the spatial illumination gradient map of the intermediate adjustment image through the convolutional layer in the illumination equalization model, and generating a corresponding illumination compensation coefficient matrix according to the spatial illumination gradient map; Perform per-pixel brightness compensation on each pixel in the intermediate adjusted image according to the illumination compensation coefficient matrix to generate a standardized image frame that meets the standardized illumination range threshold; Aggregate all the standardized image frames to generate the standardized illumination image set.
3. The method according to claim 2, wherein The extracting the morphological feature vector, texture distribution vector, and spectral response vector from the standardized crop image through the parallel convolutional branches in the multi-task deep learning model respectively includes: Input the standardized crop image into the first convolutional branch, the second convolutional branch, and the third convolutional branch of the multi-task deep learning model respectively; wherein, the first convolutional branch includes multiple dilated convolutional layers for capturing morphological features at different scales in the standardized crop image; the second convolutional branch includes a histogram of oriented gradients filtering layer for extracting the texture distribution pattern of the leaf edges in the standardized crop image; the third convolutional branch includes a spectral attention mechanism layer for weighted fusion of the response features of the standardized crop image in different spectral channels; Perform multi-scale feature extraction on the standardized crop image through the dilated convolutional layers in the first convolutional branch to generate a morphological feature vector related to the crop stem structure; Perform edge direction statistical analysis on the standardized crop image through the histogram of oriented gradients filtering layer in the second convolutional branch to generate a texture distribution vector related to the crop leaf surface texture; Calculate the weight coefficients of different spectral channels through the spectral attention mechanism layer in the third convolutional branch, and perform channel weighted fusion on the multi-spectral feature map according to the weight coefficients to generate a spectral response vector related to the crop physiological state.
4. The method according to claim 1, wherein The method further includes the process of training the multi-task deep learning model, including: Obtain a historical crop growth cycle data set, which includes multi-spectral images of multiple crop samples at different growth stages, corresponding growth stage annotations, health status scores, and environmental parameter records; Perform the standardized illumination adjustment process and crop region segmentation process on each multi-spectral image in the historical crop growth cycle data set to generate a training standardized crop image set; Construct an initial multi-task deep learning model, which includes the parallel convolutional branches, a fully connected layer, and a time series analysis module; Input the training standardized crop image set into the initial multi-task deep learning model, and optimize the model parameters through a joint loss function; the joint loss function includes a weighted sum of a growth stage classification loss, a health status regression loss, and an environmental adaptability regression loss; Generate the pre-trained multi-task deep learning model according to the optimized model parameters, and verify the multi-task classification accuracy of the model in the test data set.
5. The method according to claim 4, characterized in that The optimizing the model parameters through the joint loss function includes: Measure the difference between the growth stage class label predicted by the model and the true annotation through a cross-entropy loss function as the growth stage classification loss; Measure the difference between the growth health status score predicted by the model and the true score through the smooth L1 loss function, as the health status regression loss; Measure the difference between the environmental adaptability index predicted by the model and the theoretical adaptability index calculated according to the environmental parameter records through the mean squared error loss function, as the environmental adaptability regression loss; Dynamically adjust the weight coefficients of each loss term according to the historical gradient magnitudes of the growth stage classification loss, health status regression loss, and environmental adaptability regression loss, so that the optimization speeds of each task during training are balanced; Update the parameters of the initial multi-task deep learning model through the backpropagation algorithm until the joint loss function converges to a preset threshold.
6. The method according to claim 5, wherein The dynamically adjusting the weight coefficients of each loss term includes: In each training iteration, calculate the gradient norms of the growth stage classification loss, health status regression loss, and environmental adaptability regression loss with respect to the model parameters respectively; Determine the current adjustment factor of each loss term according to the ratio of the gradient norm corresponding to each loss term to its historical average value; If the current adjustment factor of any loss term is greater than the preset upper bound threshold, reduce the weight coefficient of the any loss term; if the current adjustment factor is less than the preset lower bound threshold, increase the weight coefficient of the any loss term; Recalculate the weighted sum of the joint loss function according to the adjusted weight coefficients and perform the next training iteration.
7. The method according to claim 4, wherein The method further includes a process of performing data augmentation on the training standardized crop image set, including: Randomly apply rotation transformation, mirror flipping transformation, and scale scaling transformation to each image in the training standardized crop image set to generate enhanced image samples; Perform color jitter processing on the local regions in the enhanced image samples to simulate the crop phenotype changes under different lighting conditions, and obtain jitter enhanced image samples; Combine the jitter enhanced image samples with the original training standardized crop image set to generate an extended training data set; During training, dynamically adjust the batch sampling weights according to the number distribution of samples in different growth stages in the extended training data set, so as to make the model learn the balance of each growth stage category.
8. The method according to claim 7, wherein The dynamically adjusting the batch sampling weights includes: Count the number of samples corresponding to each growth stage category in the extended training data set and calculate the proportion of the number of samples in each growth stage category; Determine the initial sampling weights of each growth stage category according to the reciprocal of the proportion of the number of samples and perform normalization processing on the initial sampling weights; In each training iteration, dynamically correct the initial sampling weights according to the classification accuracy of the model for each growth stage category in the current training stage; If the classification accuracy of any growth stage category is lower than the preset threshold, increase the sampling weight corresponding to the any growth stage category; otherwise, maintain or reduce the sampling weight; Extract batch training samples from the extended training data set according to the corrected sampling weights.
9. A computer system, characterized in that, Includes: One or more processors; Memory; One or more computer programs; Wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Data acquisition and analysis method for AI identification of green plant growth situation
CN117036088A
Intelligent visual identification method for mangrove forest growth condition
CN117456369A