Rice growth state perception and farming decision method based on multi-modal semantic guidance

By using a multi-modal semantic-guided multi-task recognition model, the problem of rice growth status monitoring relying on human experience has been solved. This model enables accurate identification of rice growth status and agricultural decision-making, improving monitoring efficiency and recognition accuracy. It is suitable for intelligent agricultural management in smart agriculture.

CN122135137APending Publication Date: 2026-06-02CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-11
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, the growth status of rice relies on human experience to make judgments, which results in low monitoring efficiency, fragmented information, and difficulty in directly guiding agricultural decisions. Furthermore, multi-task recognition models lack accuracy and generalization ability in complex field environments.

Method used

A multimodal semantic-guided approach is adopted, which uses techniques such as shared feature extraction, multi-scale feature fusion, task-aware feature selection and recombination, and task-specific weighted loss to construct a multi-task recognition model, thereby achieving collaborative recognition of multidimensional growth states of rice and generating agricultural decision-making strategies.

Benefits of technology

It enables precise and collaborative identification of rice growth status and intelligent agricultural decision-making, improving monitoring efficiency and the level of automation in agricultural management, reducing labor costs, and enhancing identification accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135137A_ABST
    Figure CN122135137A_ABST
Patent Text Reader

Abstract

This invention relates to a method for rice growth status perception and agricultural decision-making based on multimodal semantic guidance, belonging to the fields of smart agriculture, computer vision, and artificial intelligence. The method includes acquiring rice canopy images; constructing a dataset; utilizing a small number of samples labeled by agricultural experts and their corresponding descriptive semantic information; introducing a multimodal semantic-guided auxiliary labeling mechanism to generate growth status labels for the remaining samples, thereby reducing data labeling costs and improving labeling consistency; based on the labeled dataset, constructing a multi-task recognition model to collaboratively identify multi-dimensional growth states of rice, such as phenological stage, nutritional status, disease type, disease impact degree, and population growth, within a unified feature space. Based on the multi-task recognition results, rice growth status analysis information is generated, forming agricultural guidance strategies such as fertilization management and pest and disease control, which are then distributed to field execution equipment, realizing an intelligent management closed loop combining rice growth status perception and agricultural decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of smart agriculture, computer vision and artificial intelligence technology. It relates to a method for rice growth status perception and agricultural decision-making based on computer vision, multimodal semantic guidance and multi-task deep learning. It is applicable to the collaborative perception and analysis of multi-dimensional growth status of rice, such as phenological stage, nutritional status, disease type, degree of disease impact and population growth, and generates fertilization management, pest and disease control and field management decision-making strategies based on the perception results. Background Technology

[0002] Throughout the entire growth cycle of rice, information on phenological progression, nutritional status, disease occurrence, and population growth comprehensively reflects the physiological health and production potential of the rice. If nutrient supply to rice is unbalanced or diseases are not identified and treated in a timely manner, it can easily lead to problems such as abnormal tillering, reduced grain number per panicle, or insufficient grain filling, thus adversely affecting rice yield and quality. Therefore, comprehensive, accurate, and timely monitoring and evaluation of rice growth status is a crucial foundation for achieving scientific water and fertilizer management, pest and disease control, and refined field management.

[0003] In traditional rice production management, rice growth status relies primarily on manual field inspections and experience-based judgment. Agricultural technicians or farmers typically assess rice growth by visually observing changes in leaf color, plant height differences, growth progress, and the appearance of disease spots. This method has significant limitations: firstly, manual inspections are inefficient and cannot meet the needs of large-scale rice fields for real-time, continuous monitoring; secondly, judgments are highly dependent on personal experience, lacking unified and objective quantitative standards, which hinders the formation of traceable and comparable growth status data; furthermore, manual field inspections are labor-intensive and easily affected by weather, water levels, and terrain conditions, making it difficult to support the requirements of modern agriculture for large-scale and refined management.

[0004] With the development of smart agriculture and precision agriculture technologies, automatic monitoring of rice growth status based on drones, fixed monitoring equipment, and intelligent vision systems has gradually become a research hotspot. Utilizing computer vision and deep learning technologies to analyze rice images has yielded certain research results in areas such as disease identification and lodging detection. However, addressing the practical needs of multi-factor and multi-task collaborative identification in rice still faces numerous technical challenges, mainly in the following aspects: (1) Difficulty and high cost of data annotation: Deep learning models usually rely on large-scale, high-quality labeled data for training. This requires professional personnel to annotate the data, which is time-consuming, labor-intensive, and costly. This leads to a scarcity of high-quality labeled data, which limits the promotion and application of deep models in real-world scenarios.

[0005] (2) The growth status is complex and the differences between categories are subtle: different phenological stages and nutritional status of rice show continuous changes in appearance. The differences in leaf color, plant type and texture are not significantly discrete. Especially in adjacent growth stages or under mild nutritional stress, traditional methods based on single color or simple texture characteristics are difficult to distinguish accurately.

[0006] (3) The natural environment is complex and the interference factors are prominent: the lighting conditions in the rice planting environment are significantly affected by the weather, and there are problems such as dense shading between rice plants, water surface reflection and background clutter, which leads to large fluctuations in image quality and puts high demands on the robustness and feature expression ability of the model.

[0007] (4) Uneven distribution of samples between tasks: In the actual collected data, the number of samples in different growth states or disease categories varies significantly. Some disease or extreme nutritional states have a low proportion of samples, which can easily lead to the model training process being biased towards the majority of categories, thereby reducing the recognition accuracy of a few key states.

[0008] (5) The model's generalization ability and agricultural application requirements coexist: Rice growth status monitoring not only requires high identification accuracy, but also good generalization ability and computational efficiency to adapt to field edge devices or agricultural Internet of Things systems, so as to achieve continuous operation and real-time feedback.

[0009] Existing studies mostly employ single-task identification models, analyzing rice diseases, nutrition, or growth stages independently. This makes it difficult to comprehensively reflect the overall growth status of rice and directly support agronomic decisions such as fertilization management and pest and disease control. Therefore, there is an urgent need to propose an intelligent analysis method capable of collaboratively identifying multiple growth states of rice in complex field environments, and further supporting agricultural guidance and closed-loop regulation. Summary of the Invention

[0010] In view of this, the purpose of this invention is to overcome the shortcomings of existing technologies, such as reliance on manual experience to judge rice growth status, low monitoring efficiency, fragmented information, and difficulty in directly guiding agricultural decision-making. This invention provides a method for rice growth status perception and agricultural decision-making based on multimodal semantic guidance. This invention enables accurate and coordinated identification of multidimensional growth information, including phenological stage, nutritional status, disease type, disease impact degree, and population growth, using only a small number of high-quality labeled samples, combined with multimodal semantic guidance for rice growth status auxiliary labeling. Based on the identification results, it generates scientific and executable fertilization management and pest and disease control agronomical control strategies, thereby achieving intelligent monitoring of rice growth status and closed-loop management of agricultural decision-making.

[0011] To achieve the above objectives, the present invention provides the following technical solution: A method for rice growth status perception and agricultural decision-making based on multimodal semantic guidance includes the following steps: S1: Data Acquisition and Preprocessing: Acquire images of the rice canopy in the paddy field, perform quality screening on the images (remove blurry, occluded or invalid images), and perform cropping, size normalization and pixel standardization on the screened images to form a sample set of rice images to be labeled. S2: Constructing a multi-task training dataset for rice: Select a small number of representative samples from the set of rice image samples to be labeled, have agricultural experts manually label them and generate semantic constraint information corresponding to the labels to form highly reliable labeled samples; extract image feature vectors, and under the guidance of semantic constraints, use a multimodal semantic reasoning model to learn expert discrimination logic to generate auxiliary labels for the remaining samples, thus completing the construction of the target dataset; S3: Multi-task recognition model construction, training and prediction: A multi-task recognition model is trained using the target dataset. The multi-task recognition model includes a shared feature extraction module, a multi-scale feature fusion module, a task-aware feature selection and recombination module, a task feature adaptation module, a task-specific weighted loss module, and a task prediction branch module. The model is then used to predict the multi-task growth status of rice images.

[0012] Based on deep convolutional neural networks, a multi-task recognition model that combines shared feature extraction and task-aware collaborative learning is constructed to jointly model the multidimensional growth states of rice, specifically including: (1) Construct a shared feature extraction module, and encode the input rice image with multi-scale features through multi-level convolution and residual structure to extract shared feature representations that take into account both local texture information and high-level semantic information; (2) Introduce a multi-scale feature fusion module to perform top-down and horizontal fusion of features at different levels, thereby enhancing the model’s ability to perceive the structural features of rice at different spatial scales. (3) Construct a task-aware feature selection and reorganization module. Based on the different requirements of spatial resolution and semantic level for different recognition tasks, select and reorganize feature representations related to the corresponding tasks from multi-scale features; (4) Design a task feature adaptation module. Based on the discrimination requirements of different recognition tasks, perform differentiated modeling and enhancement of color attribute features and texture structure features to obtain adaptation features that match each task. (5) Construct a task prediction branch module, input the adaptation features into the corresponding task prediction branches respectively, and output the prediction results of each recognition task through feature fusion, attention enhancement and classification mapping operations; (6) Task-specific weighted loss module: Introduces a joint loss function of task weighting and category weighting. Based on the importance of different tasks and the distribution of category samples, the loss of each task is weighted and fused to balance the multi-task training process and improve the overall recognition performance.

[0013] Subsequently, the multi-task recognition model was jointly trained and its parameters were optimized using the constructed target dataset to obtain the optimal model weights.

[0014] S4: Agricultural Information Generation: Generate rice growth status analysis information based on multi-task recognition results, generate comprehensive analysis information on the current growth status of rice, and further generate agronomic control instructions such as fertilization management instructions and pest and disease control instructions that match the growth status. S5: Command Issuance and Execution: The agronomic control commands are sent to the paddy field management execution unit through a standardized communication protocol. The standardized communication protocol includes, but is not limited to, MQTT protocol, Modbus protocol or application programming interface (API). The paddy field management execution unit includes agricultural execution equipment such as intelligent spraying equipment, integrated water and fertilizer management system or agricultural Internet of Things control terminal, so as to realize automated or semi-automated control of paddy fields. S6: Feedback and Adaptive Adjustment: Receives the operation completion status information fed back by the paddy field management execution unit and / or the field environment data collected by the environmental sensors, updates the rice growth parameters or model input data according to the feedback information, and uses it to guide the identification and control decisions of the next cycle, thereby forming an adaptive closed-loop control of rice agricultural management.

[0015] Further, in step S2, constructing the rice multi-task training dataset specifically includes: selecting a small number of representative samples from the set of rice image samples to be labeled, and having agricultural experts manually label the selected samples, while generating descriptive semantic information corresponding to the labeling results. The descriptive semantic information is used to characterize the observable features and discrimination criteria of rice growth status, forming highly reliable labeled samples; extracting visual feature vectors from the highly reliable labeled samples and the remaining rice image samples to be labeled, and under the guidance of semantic constraints, using a multimodal large-scale language model to learn expert discrimination logic based on the visual feature vectors and semantic information, generating auxiliary labels for the images to be labeled; performing expert sampling and correction on the auxiliary labels, generating multi-task label codes, and establishing associations between the multi-task label codes and the corresponding rice image samples to construct the target dataset for multi-task model training.

[0016] Furthermore, in step S3, the shared feature extraction module includes an initial feature encoding module and a multi-level feature extraction module; the initial feature encoding module is as follows:

[0017] in, Indicates the input image The initial feature map is obtained after processing by the initial feature encoding module. This represents the feature mapping function corresponding to the initial feature encoding module. This represents the input image of rice. and These represent two consecutive levels of convolutional feature transformation operations, which include convolution operations and normalization processing, without changing the spatial resolution of the feature map. Through the initial feature encoding module, features are extracted from the input rice image without spatial downsampling, thereby enhancing the ability to express fine-grained spatial information such as leaf edges, texture structure, and color distribution. The multi-level feature extraction module includes multiple sequentially connected feature extraction stages, the first of which... Each feature extraction stage is represented as:

[0018] in, This represents the input feature map from the previous stage. This indicates a feature channel transformation operation. This indicates a spatial downsampling operation. This represents the residual feature extraction operation. By introducing spatial downsampling and residual feature extraction structures at different feature extraction stages, the spatial resolution of the feature map is progressively reduced, while simultaneously enhancing the stability and nonlinear representation capability of the feature representation, thereby obtaining multi-level feature representations with different spatial scales and semantic levels. In the multi-level feature extraction module, the outputs of at least three different feature extraction stages are selected as intermediate feature maps, which are represented as follows:

[0019] in, , , These represent different feature extraction stage numbers, corresponding to feature representations with different spatial resolutions and semantic levels.

[0020] Further, in step S3, the multi-scale feature fusion module is used to perform step-by-step fusion of intermediate feature maps with different spatial resolutions output by the shared feature extraction module. It includes a lateral feature alignment module and a top-down feature fusion module. The lateral feature alignment module is used to perform channel alignment processing on intermediate feature maps of different scales, as shown below:

[0021] in, Indicates the first i Intermediate feature map of layer The lateral alignment feature map obtained after channel alignment processing. , , These represent intermediate feature maps at different spatial resolutions output by the shared feature extraction module. Indicates adoption The channel mapping operation of convolution is used to unify the number of channels in feature maps of different scales to a preset dimension; The top-down feature fusion module is used to upsample high-level semantic features step by step and fuse them with low-level features. The fusion process is represented as follows:

[0022]

[0023]

[0024] in, , , This represents a multi-scale fused feature map that has a unified semantic representation but different spatial resolutions. This indicates an upsampling operation based on bilinear interpolation, used for spatial scale alignment of high-level feature maps; " indicates an element-wise feature fusion operation; This indicates that after fusion, feature smoothing and semantic enhancement are performed through convolutional operations. The multi-scale feature fusion module merges intermediate feature maps with different spatial resolutions into multiple multi-scale fused feature maps with unified semantic representation but different spatial resolutions. , , It preserves high-level semantic information and low-level spatial details, providing multi-scale feature support for feature selection and task modeling for different recognition tasks.

[0025] Furthermore, in step S3, the task-aware feature selection and reorganization module is used to fuse feature maps from multiple scales according to the differentiated requirements of different recognition tasks for spatial resolution, semantic level, and feature structure. , , The corresponding features are selected, and the selected features are scale-aligned and reorganized to generate task features suitable for each task; among them, the task features for the phenological period identification task are... Represented as:

[0026] in, This represents a feature concatenation operation along the channel dimension. Indicates mesoscale features Upsampling to high-resolution features Operations within the same spatial dimension; The task characteristics of the disease identification task and the impact assessment task are represented as follows: ,in, This describes the task characteristics of the disease identification task. This indicates the task characteristics of the impact assessment task. Indicates low-scale features; Task characteristics of nutrient status identification task Represented as: ; Task characteristics of population growth assessment task Represented as: .

[0027] The task-aware feature selection and recombination module selects and recombines fused feature maps of different spatial scales according to task requirements, realizing feature customization for each recognition task. It retains high-resolution spatial detail information while taking into account high semantic level global information, providing input for subsequent task-specific adaptation and prediction.

[0028] Furthermore, in step S3, the task feature adaptation module includes a color feature adaptation module and a texture feature adaptation module; the color feature adaptation module is as follows:

[0029] in, This represents the task features output by the task-aware feature selection and reorganization module. and These represent the convolution transformation operation, Presentation layer normalization operation; This indicates a channel attention enhancement operation, outputting color features. ; The texture feature adaptation module is:

[0030] in, Represents texture features, This represents a depthwise convolution operation. This indicates a collaborative attention operation. This indicates the convolutional channel alignment operation. The outputs of the color feature adaptation module and the texture feature adaptation module constitute the adaptation features for the corresponding task, which are used by the subsequent task prediction branch module for classification.

[0031] Furthermore, in step S3, the task-specific weighted loss module is used to weight the losses of each task according to the differences in category distribution and training difficulty of each recognition task, so as to achieve collaborative optimization of multiple tasks; the task-specific weighted loss module includes: (1) Class weight calculation: For each category of samples in each task, calculate the effective class weight based on the sample distribution. :

[0032] in, Indicate category The number of samples, For smoothing coefficients, These are the normalized class weights; (2) Category-weighted loss: cross-entropy loss Multiply by the class weights to obtain the class-weighted loss. :

[0033] (3) Task-weighted processing: Multiply the category-weighted loss by the corresponding task weight. Receive task-weighted loss :

[0034] in, For task indexing, Task weights are used to adjust the importance or difficulty of task training. (4) Multi-task loss aggregation: sum the task-weighted losses of all tasks to obtain the total loss. :

[0035] in, This represents the set of all training tasks.

[0036] The total loss is used to guide the training of multi-task models, thereby achieving collaborative optimization of each task while taking into account the difficulty of different tasks and class imbalance.

[0037] Furthermore, in step S3, the task prediction branch module is used to receive the task features output by the task feature adaptation module, and according to the discrimination requirements of different recognition tasks, adopt differentiated feature enhancement and convergence structures to output the prediction results of the corresponding tasks. Among them, for the first For each recognition task, the corresponding task prediction branch is represented as follows:

[0038] in, Indicates the first The prediction results for each recognition task Indicates the first The adaptation features corresponding to each task; These represent task-related feature enhancement operators used for spatial, directional, or contextual modeling of input features. This represents a task-related feature aggregation operator used to perform global or local saliency aggregation on enhanced features; This represents a classification mapping operator used to map convergent features to the corresponding category prediction results for a task.

[0039] Among them, feature enhancement operators are used for different task types. and feature convergence operator Satisfy one or a combination of the following constraints: (1) Tasks during the phenological period: This includes pointwise convolution and depthwise separable convolution to enhance the response to small-scale structural features; This includes a combination of local saliency pooling and global average pooling, used to simultaneously preserve key regional features such as the ear and overall contextual information; (2) Task on the degree of disease and its impact: This includes context-aware attention mechanisms and channel attention mechanisms, used to enhance the ability to model the texture structure of lesion regions and their contextual relationships; This includes global feature aggregation operations; (3) Nutritional status task: This includes feature fusion and channel mapping operations for color attributes; This includes global average pooling operations; (4) Group growth task: This includes directional convolution along the horizontal and vertical directions, used to model the row and column structure features of rice planting; This includes global feature aggregation operations.

[0040] The beneficial effects of this invention are as follows: By constructing an intelligent recognition model based on shared feature representations and task collaborative learning, this invention achieves comprehensive perception and precise control of the multidimensional growth state of rice, significantly improving the intelligence and precision of rice production management. Compared with the prior art, this invention has at least the following significant effects: (1) Advantages of data construction that balances labeling cost and labeling quality: In the process of dataset construction, this invention introduces descriptive semantic constraint information provided by agricultural experts and combines it with a large-scale language model for semantic reasoning of growth status. This expands expert experience from discrete labels into transferable semantic information, effectively improving the consistency and rationality of multi-task labels. By guiding the language model to perform auxiliary labeling of unlabeled samples through semantic constraints, the cost of manual labeling is reduced while ensuring that the labeling logic conforms to agronomic cognition. In addition, the combination of expert sampling and correction mechanisms improves the reliability of training data, providing high-quality data support for the stable training and improved recognition accuracy of the multi-task growth status recognition model.

[0041] (2) Multidimensional growth status collaborative identification, significantly improving overall perception capability: This invention uses a multi-task joint modeling method to simultaneously identify various key growth information of rice under the same model framework, such as phenological stage, nutritional status, disease type, degree of disease impact and population growth, avoiding the problem of different growth indicators relying on multiple independent models and information fragmentation in the prior art, and realizing the overall and systematic perception of rice growth status.

[0042] (3) Significantly improved recognition accuracy and robustness in complex field environments: This invention addresses the problems of drastic changes in light, severe plant shading, and complex backgrounds in rice paddy environments. By using a shared feature expression and task-aware feature selection mechanism, it improves the model's ability to express fine-grained features and multi-scale structures, thereby maintaining high recognition accuracy and stability even in complex natural environments.

[0043] (4) Effectively alleviate class imbalance problem and improve the ability to identify small sample states: In the model training process, the present invention introduces a task-specific weighted optimization mechanism, which comprehensively considers the differences in different recognition tasks and their class distributions, so that the model can pay more attention to the growth state class with fewer samples or higher discrimination difficulty during the training process, effectively alleviate the impact of class imbalance on recognition performance and improve the overall recognition effect.

[0044] (5) Closed-loop agricultural regulation capability from perception to decision-making: This invention not only realizes the automatic identification of rice growth status, but also generates corresponding agricultural guidance information such as water and fertilizer management and pest and disease control based on the identification results. It also links with field execution equipment through communication interface to build a closed-loop agricultural management process of "perception-analysis-decision-execution-feedback", which significantly improves the automation and precision of rice field management.

[0045] (6) Reduce labor costs and improve the efficiency of large-scale planting management: This invention reduces the reliance on manual field inspection and experience judgment, and can quickly and continuously obtain information on the growth status of rice in a large area of ​​paddy fields and generate agricultural decisions, significantly reducing labor input costs. It is particularly suitable for large-scale and intensive rice production scenarios.

[0046] (7) Significant economic benefits and promotion value: By improving the timeliness of rice growth status monitoring and the scientific nature of decision-making, this invention helps to optimize water and fertilizer input and pest and disease control measures, reduce resource waste, and improve rice yield and quality, with significant economic benefits and good prospects for promotion and application.

[0047] In summary, this invention is significantly superior to existing technologies in terms of growth status perception accuracy, intelligent agricultural decision-making, management efficiency, and economic benefits, and can provide effective technical support for smart rice cultivation and modern agricultural management.

[0048] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 The overall flowchart of the rice growth status perception and agricultural decision-making method based on multimodal semantic guidance; Figure 2 This is a diagram of a rice growth status perception model. Detailed Implementation

[0050] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0051] Please see Figures 1-2 This invention provides a method for rice growth status perception and agricultural decision-making based on multimodal semantic guidance, specifically including the following steps: S1: Data Acquisition and Preprocessing A representative rice-growing field in a certain district of a city was selected as the data collection example area. The collection time covered a complete rice growth cycle, and images were collected using a high-resolution RGB camera.

[0052] 1) Acquisition conditions: The time period covers 10:00 AM to 4:00 PM, including images under various natural lighting conditions such as front lighting, side lighting, backlighting, sunny days, and cloudy / rainy days. The acquisition device is located approximately 60cm to 140cm away from the top canopy, and the shooting angle is approximately 60° to 80° horizontally.

[0053] 2) Coverage of the acquired targets: The acquired images cover multiple key growth stages of rice, including: (1) Phenological stages: greening stage, tillering stage, jointing stage, booting stage, heading stage, flowering stage, grain filling stage, yellow ripening stage and maturity stage; (2) Nutritional status: Nitrogen, phosphorus and potassium are three nutrient statuses, each including three levels: deficiency, normal and excess; (3) Disease types: including normal state as well as rice blast, sheath blight, rice false smut, rice leaf roller, panicle rot, rice planthopper, golden apple snail, bacterial leaf blight and bacterial leaf streak; (4) Degree of impact: Used to characterize the degree of impact of diseases on rice growth, divided into none, mild and severe; (5) Evaluation of population growth: used to reflect the overall growth status of rice, divided into weak, average and vigorous.

[0054] 3) Data preprocessing: A total of 28,747 original images were collected. Images with severe blurring, overexposure or underexposure were removed. The remaining images were uniformly cropped and sized, resulting in 25,384 high-quality images, which constitute the rice image sample set to be labeled.

[0055] S2: Constructing a multi-task training dataset for rice 1) Multi-task labeling and expert annotation: In this embodiment, to achieve unified expression and efficient parsing of multiple growth state information of rice, a multi-task labeling method based on fixed character positions is adopted to realize a one-to-one correspondence between rice image samples and multi-task growth state labels. The specific steps are as follows: (1) Encoding rules: Each rice image sample corresponds to a unique image ID and a multi-task label string. The multi-task label string is fixed at 7 characters, with each character corresponding to a specific recognition task and arranged in a preset order. The multi-task label string is bound to the image ID to realize the association storage and parsing of image samples and multi-task labels. For specific encoding methods, please refer to Table 1.

[0056] (2) File naming example: ID_00049;G111F20.jpg, where the first part ID_00049 is a unique image number; the second part G111F20 is a multi-task tag string with the following semantics: a. Phenological period (G): Grouting period; b. Nitrogen status (1): Nitrogen levels are normal; c. Phosphorus status (1): Normal phosphorus content; d. Potassium status (1): Potassium levels are normal; e. Disease type (F): Rice blast; f. Degree of impact of disease (2): Severe; g. Population growth (0): Weak.

[0057] Each category is labeled with approximately 100 representative, highly reliable images by agricultural experts, and descriptive semantic information corresponding to the labeling results is provided for each image (for example, the semantic description of the grain-filling stage is: "rice grains are milky white to pale yellow, the grains are full, the plant height is basically fixed, tillering has stopped, and some leaves have begun to show natural senescence"). The standardized expression of multi-task growth status information is achieved through multi-task label strings, while providing a semantic constraint basis for subsequent semantically guided auxiliary labeling.

[0058] 2) Multimodal semantic-guided auxiliary annotation: Visual feature vectors are extracted from the initial high-confidence labeled samples, and the visual feature vectors and corresponding descriptive semantic information are used as semantic inputs. A multimodal semantic reasoning model is introduced to learn the growth status discrimination logic of agricultural experts. Based on this, corresponding multi-task auxiliary annotation labels are generated for the remaining unlabeled rice image samples. The process is used to infer the label generation logic and does not involve direct visual recognition of the original images.

[0059] 3) Expert verification and multi-task label encoding generation: Expert sampling and correction are carried out on the auxiliary annotation results to ensure the consistency and accuracy of the annotation results; the verified annotation results are encoded into multi-task label strings with fixed character positions in a preset order, and an association storage relationship is established with the corresponding rice image samples to form a complete and parsable multi-task annotation dataset.

[0060] 4) Target dataset partitioning: The dataset with completed multi-task annotation and established label association is divided into training set, validation set and test set according to a preset ratio. In this embodiment, the ratio is 7:2:1, which is used for subsequent training, parameter tuning and performance evaluation of the multi-task recognition model.

[0061] S3: Multi-task recognition model construction, training, and prediction: In this embodiment, the target dataset processed by S2 is input into the multi-task growth state recognition model to achieve joint recognition of multiple growth states of rice. The multi-task recognition model operates sequentially in the order of feature extraction, feature fusion, task feature generation, and task prediction, and its specific processing procedure is as follows.

[0062] 1) Shared feature extraction process: (1) Prepare input: Input the prepared rice image Input to the shared feature extraction module, image size is .

[0063] (2) Initial low-level feature encoding: First, the input image is encoded. conduct Convolution + normalization + activation (GELU) outputs a 32-channel feature map; then, repeat... Convolution, normalization, and activation operations are used to further encode low-level features, resulting in the STEM output feature map. The purpose of this step is to extract and preserve low-level texture and color features to capture fine-grained information such as leaves, spikelets, lesions, and leaf veins.

[0064] (3) Low-resolution feature extraction stage: First, using Average pooling (stride=2) downsamples the Stem output; secondly, through... Convolution, normalization, and activation are applied to project the downsampling results into channels, outputting 64-channel features. Then, three residual blocks are applied consecutively to increase non-linear expressive power and receptive field while maintaining resolution. Finally, a low-resolution feature map is output. The resolution is half that of the input, and it mainly contains detailed texture information.

[0065] (4) Low-to-mid-level feature extraction stage: First, for conduct Average pooling downsampling (stride=2) results in an output spatial resolution that is 1 / 4 of the input; secondly, through... Convolution, normalization, and activation are applied to project the channel data, outputting a 128-channel feature map. Then, three residual modules are applied consecutively to enhance the feature representation, finally outputting the mid-to-low-level feature maps. It integrates texture and structural information.

[0066] (5) Mid-to-high level feature extraction stage: First, for conduct Average pooling downsampling (stride=2) results in an output spatial resolution that is 1 / 8 of the input; secondly, through... Convolution, normalization, and activation are used to project the channels, outputting a 256-channel feature map. Finally, five residual modules are applied consecutively to enhance the nonlinear expressive power and stability of the features.

[0067] (6) High-level semantic feature extraction stage: First, the output of the mid-to-high-level feature extraction stage is processed. Average pooling downsampling (stride=2) results in an output spatial resolution that is 1 / 16th of the input; secondly, through... Convolution, normalization, and activation are applied to project the channel data, outputting a 384-channel feature map. Then, six residual modules are applied consecutively to extract high-level semantic information, forming a robust global feature representation. Finally, a high-level feature map is output. It contains rich semantic information.

[0068] 2) Multi-scale feature fusion: First, the shared features are output... Use convolution to perform a unified mapping of channel numbers, forming , , Secondly, After upsampling and Add them together and perform convolution smoothing to obtain... Then, After upsampling and Add them together and then perform convolution to smooth them out, resulting in... Final output Multi-scale fused feature maps.

[0069] 3) Task-aware feature selection and reorganization: The multi-scale fused feature map is reorganized according to task requirements: The task features for phenological period identification are represented as follows:

[0070] in, This represents a feature concatenation operation along the channel dimension. Indicates mesoscale features Upsampling to high-resolution features Operations within the same spatial dimension; the task characteristics of the disease identification task and the impact assessment task are represented as follows:

[0071] The task features for nutrient status identification are represented as follows:

[0072] The task characteristics of the growth assessment task are represented as follows:

[0073] The above feature allocation method is determined based on the different needs of different tasks for spatial resolution and semantic abstraction level, thereby reducing feature interference between multiple tasks.

[0074] 4) Task Feature Adaptation Module: The color feature adaptation module focuses on enhancing the expression of color attributes related to leaf color, ear color, and nutrient status; the texture feature adaptation module focuses on enhancing lesion morphology, leaf texture, and local structural features. The selected and recombined task-aware features are then assigned to matching features for each task according to requirements.

[0075] 5) Task Branch Prediction Module: Receives task features output by the task feature adaptation module, and according to the discrimination requirements of different recognition tasks, adopts differentiated feature enhancement and convergence structures to output the prediction results of the corresponding tasks. Among them, for the first For each recognition task, the corresponding task prediction branch is represented as follows:

[0076] in, Indicates the first The adaptation features corresponding to each task; These represent task-related feature enhancement operators used for spatial, directional, or contextual modeling of input features. This represents a task-related feature aggregation operator used to perform global or local saliency aggregation on enhanced features; This represents a classification mapping operator used to map convergent features to the corresponding category prediction results for a task.

[0077] Among them, feature enhancement operators are used for different task types. and feature convergence operator Satisfy one or a combination of the following constraints: (1) Tasks during the phenological period: This includes pointwise convolution and depthwise separable convolution to enhance the response to small-scale structural features; This includes a combination of local saliency pooling and global average pooling, used to simultaneously preserve key regional features such as the ear and overall contextual information; (2) Task on the degree of disease and its impact: This includes context-aware attention mechanisms and channel attention mechanisms, used to enhance the ability to model the texture structure of lesion regions and their contextual relationships; This includes global feature aggregation operations; (3) Nutritional status task: This includes feature fusion and channel mapping operations for color attributes; This includes global average pooling operations; (4) Group growth task: This includes directional convolution along the horizontal and vertical directions, used to model the row and column structure features of rice planting; This includes global feature aggregation operations.

[0078] 6) Improved Loss Function: This module weights the losses for each recognition task based on differences in category distribution and training difficulty, enabling collaborative optimization across multiple tasks. The task-specific weighted loss module includes the following steps: (1) Class weight calculation: For each category of samples in each task, calculate the effective class weight based on the sample distribution. :

[0079] in, Indicate category The number of samples, For smoothing coefficients, These are the normalized class weights; (2) Class-weighted loss: The cross-entropy loss is multiplied by the class weights to obtain the class-weighted loss:

[0080] (3) Task-weighted processing: Multiply the category-weighted loss by the corresponding task weight. The task-weighted loss is obtained as follows:

[0081] For task indexing, Task weights are used to adjust the importance or difficulty of task training.

[0082] (4) Multi-task loss aggregation: Sum the task-weighted losses of all tasks to obtain the total loss:

[0083] in, This represents the set of all training tasks; the total loss is used to guide the training of multi-task models, thereby achieving collaborative optimization of each task while taking into account the difficulty of different tasks and class imbalance.

[0084] 7) Training environment and hyperparameter settings (1) Training Environment and Hyperparameters: In this embodiment, model training is implemented based on the TensorFlow 2.16.0 deep learning framework, and the training hardware environment is an NVIDIA GeForce RTX4090 graphics processor. The key hyperparameter settings are as follows: input image size 384×512; batch size 32; SGD optimizer, initial learning rate 0.01, momentum 0.937, weight decay 5e-4; cosine annealing strategy for learning rate scheduling; total training epochs 100, and early stopping strategy enabled. Through the above hyperparameter configuration, the model's generalization ability is improved while ensuring the model's convergence stability.

[0085] (2) Multi-task model training process: During the training phase, the multi-task growth state recognition model constructed in S3 is combined with the target dataset constructed in S2 for training. Specifically, in each training round, the model is input into the training set samples in batches, and simultaneously outputs the prediction results of multiple tasks such as phenological stage, nutrient status, disease type, degree of impact, and population growth assessment through forward propagation; subsequently, the multi-task joint loss is calculated according to the task-specific weighted loss function corresponding to each task, and the model parameters are updated through backpropagation. After each training round is completed, the model performance is evaluated once using the validation set, and the loss value and prediction accuracy of each task are recorded to dynamically monitor the training status and convergence of the model.

[0086] (3) Model Evaluation Method: After training, the model parameters with the best overall performance on the validation set are selected as the final model, and the model is evaluated on the independent test set. In this embodiment, for the multi-task classification problem, the classification accuracy, macro-average accuracy, and overall average performance index of each recognition task on the test set are statistically analyzed to comprehensively measure the model's predictive ability on different growth state recognition tasks. Experimental results show that the multi-task growth state recognition model can achieve joint recognition of rice phenological stage, nutritional status, disease type, degree of impact, and population growth under a single model structure. Under the premise of keeping the model parameter scale controllable, it achieves high recognition accuracy and good generalization performance, verifying the effectiveness and practicality of the method of this invention in rice agricultural guidance scenarios.

[0087] S4: Agricultural Information Generation In this embodiment, based on the prediction results of seven tasks output by the multi-task growth state recognition model, the current growth state of rice is comprehensively analyzed, and corresponding agricultural guidance information is generated. The generation process is as follows: Let the prediction result of the rice image be:

[0088] These correspond to seven categories of identification tasks: phenological stage, nitrogen status, phosphorus status, potassium status, disease type, degree of disease impact, and population growth.

[0089] The rules for generating fertilizer management are as follows:

[0090] in, This is a phenological stage-specific regulatory factor used to control whether fertilization is suitable for the current stage and the upper limit of fertilization intensity. This is the result of phenological period prediction; The nitrogen, phosphorus, and potassium nutrient regulation coefficients, These are the predicted levels for nitrogen, phosphorus, and potassium (0: deficiency, 1: normal, 2: excess), used to calculate whether to increase or decrease the amount of fertilizer applied ("+" and "-" indicate that it is recommended to increase or decrease the amount of fertilizer applied).

[0091] The rules for disease prevention and control are as follows:

[0092] in, The results are the predicted disease type, the degree of disease impact (level 0: none, 1: mild, 2: severe), and the population growth (level 0: weak, 1: normal, 2: vigorous).

[0093] S5: Issuance and Execution of Agronomic Instructions In this embodiment, the system will construct corresponding agricultural control instructions based on the agricultural guidance information generated in step S3, and send them to the paddy field management execution unit through a standardized communication interface. For example, when nutrient deficiency is detected, the system variable fertilizer applicator sends a JSON format instruction: {{"type":"n_management","rate":"20%"},{"type":"p_management","rate":"0%"},{"type":"k_management","rate":"10%"}}; Simultaneously, a task is sent to the unmanned sprayer: {"task":"disease","target":"apply_pesticide"}. Agricultural control commands are encapsulated in a structured data format and pushed to the field control system or agricultural IoT platform via a RESTful API communication protocol.

[0094] S6: Feedback reception and adaptive adjustment After the agricultural control commands are executed, the system receives operational feedback information from the paddy field management execution unit and / or field status data collected by environmental sensors. Feedback information includes, but is not limited to, operational completion status, actual execution parameters, and data related to the rice growth environment. Based on this feedback, the system adaptively adjusts the identification and control processes for subsequent cycles, including updating rice growth parameters, adjusting agricultural decision-making strategies, or optimizing model input data, thereby guiding the identification of growth status and agricultural management decisions for the next cycle. Through this approach, a closed-loop rice agricultural management system centered on multi-task growth status identification is constructed, enabling continuous iteration and optimization of rice growth monitoring, decision control, and feedback correction.

[0095] Thus, this invention completes a closed-loop process for a rice growth status perception and agricultural decision-making method based on multimodal semantic guidance. Those skilled in the art can calibrate parameters such as thresholds according to the above description and specific field conditions to realize the application of this invention.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for rice growth status perception and agricultural decision-making based on multimodal semantic guidance, characterized in that, Specifically, the following steps are included: S1: Data acquisition and preprocessing: Acquire images of the rice canopy in the paddy field, screen the images for quality, and then crop, normalize the size, and standardize the pixels of the screened images to form a sample set of rice images to be labeled. S2: Constructing a multi-task training dataset for rice: Select representative samples from the set of rice image samples to be labeled, have agricultural experts manually label them and generate semantic constraint information corresponding to the labels to form highly reliable labeled samples; extract image feature vectors, and under the guidance of semantic constraints, use a multimodal semantic reasoning model to learn expert discrimination logic to generate auxiliary labels for the remaining samples, thus completing the construction of the target dataset; S3: Multi-task recognition model construction, training and prediction: A multi-task recognition model is trained using the target dataset. The multi-task recognition model includes a shared feature extraction module, a multi-scale feature fusion module, a task-aware feature selection and recombination module, a task feature adaptation module, a task-specific weighted loss module, and a task prediction branch module. The model is then used to predict the multi-task growth status of rice images.

2. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S2, constructing a rice multi-task training dataset specifically includes: selecting representative samples from the set of rice image samples to be labeled, and having agricultural experts manually label the selected samples, while generating descriptive semantic information corresponding to the labeling results. The descriptive semantic information is used to characterize the observable features and discrimination criteria of rice growth status, forming highly reliable labeled samples; extracting visual feature vectors from the highly reliable labeled samples and the remaining rice image samples to be labeled, and under the guidance of semantic constraints, using a multimodal large-scale language model to learn expert discrimination logic based on the visual feature vectors and semantic information to generate auxiliary labels for the images to be labeled; performing expert sampling and correction on the auxiliary labels to generate multi-task label codes, and establishing an association between the multi-task label codes and the corresponding rice image samples to construct the target dataset for multi-task model training.

3. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the shared feature extraction module includes an initial feature encoding module and a multi-level feature extraction module; the initial feature encoding module is: in, Indicates the input image The initial feature map is obtained after processing by the initial feature encoding module. This represents the feature mapping function corresponding to the initial feature encoding module. This represents the input image of rice. and These represent two consecutive levels of convolutional feature transformation operations, which include convolution operations and normalization processing, without changing the spatial resolution of the feature map. The multi-level feature extraction module includes multiple sequentially connected feature extraction stages, the first of which... Each feature extraction stage is represented as: in, This represents the input feature map from the previous stage. This indicates a feature channel transformation operation. This indicates a spatial downsampling operation. This represents the residual feature extraction operation; in the multi-level feature extraction module, the outputs of at least three different feature extraction stages are selected as intermediate feature maps, which are represented as follows: in, , , These represent different feature extraction stage numbers, corresponding to feature representations with different spatial resolutions and semantic levels.

4. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the multi-scale feature fusion module is used to perform step-by-step fusion of intermediate feature maps with different spatial resolutions output by the shared feature extraction module. It includes a lateral feature alignment module and a top-down feature fusion module. The lateral feature alignment module is used to perform channel alignment processing on intermediate feature maps of different scales, as shown below: in, Indicates the first i Intermediate feature map of layer The lateral alignment feature map obtained after channel alignment processing. , , These represent intermediate feature maps at different spatial resolutions output by the shared feature extraction module. Indicates adoption The channel mapping operation of convolution is used to unify the number of channels in feature maps of different scales to a preset dimension; The top-down feature fusion module is used to upsample high-level semantic features step by step and fuse them with low-level features. The fusion process is represented as follows: in, , , This represents a multi-scale fused feature map that has a unified semantic representation but different spatial resolutions. This indicates an upsampling operation based on bilinear interpolation, used for spatial scale alignment of high-level feature maps; " indicates an element-wise feature fusion operation; This indicates that after fusion, feature smoothing and semantic enhancement are performed through convolution operations.

5. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the task-aware feature selection and reorganization module is used to fuse feature maps from multiple scales according to the different requirements of different recognition tasks for spatial resolution, semantic level and feature structure. , , The corresponding features are selected, and the selected features are scale-aligned and reorganized to generate task features suitable for each task; among them, the task features for the phenological period identification task are... Represented as: in, This represents a feature concatenation operation along the channel dimension. Indicates mesoscale features Upsampling to high-resolution features Operations within the same spatial dimension; The task characteristics of the disease identification task and the impact assessment task are represented as follows: ,in, This describes the task characteristics of the disease identification task. This indicates the task characteristics of the impact assessment task. Indicates low-scale features; Task characteristics of nutrient status identification task Represented as: ; Task characteristics of population growth assessment task Represented as: .

6. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the task feature adaptation module includes a color feature adaptation module and a texture feature adaptation module; the color feature adaptation module is as follows: in, This represents the task features output by the task-aware feature selection and reorganization module. and These represent the convolution transformation operation, Presentation layer normalization operation; This indicates a channel attention enhancement operation, outputting color features. ; The texture feature adaptation module is: in, Represents texture features, This represents a depthwise convolution operation. This indicates a collaborative attention operation. This indicates the convolution channel alignment operation.

7. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the task-specific weighted loss module is used to weight the loss of each task according to the differences in category distribution and training difficulty of each recognition task, so as to achieve collaborative optimization of multiple tasks. The task-specific weighted loss module includes: (1) Class weight calculation: For each category of samples in each task, calculate the effective class weight based on the sample distribution. : in, Indicate category The number of samples, For smoothing coefficients, These are the normalized class weights; (2) Category-weighted loss: cross-entropy loss Multiply by the class weights to obtain the class-weighted loss. : (3) Task-weighted processing: Multiply the category-weighted loss by the corresponding task weight. Receive task-weighted loss : in, For task indexing, Task weights are used to adjust the importance or difficulty of task training. (4) Multi-task loss aggregation: sum the task-weighted losses of all tasks to obtain the total loss. : in, This represents the set of all training tasks.

8. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, In S3, the task prediction branch module is used to receive the task features output by the task feature adaptation module, and according to the discrimination requirements of different recognition tasks, adopt differentiated feature enhancement and convergence structures to output the prediction results of the corresponding tasks. Among them, for the first For each recognition task, the corresponding task prediction branch is represented as follows: in, Indicates the first The prediction results for each recognition task Indicates the first The adaptation features corresponding to each task; These represent task-related feature enhancement operators used for spatial, directional, or contextual modeling of input features. This represents a task-related feature aggregation operator used to perform global or local saliency aggregation on enhanced features; This represents a classification mapping operator used to map convergent features to the corresponding category prediction results for the task. Among them, feature enhancement operators are used for different task types. and feature convergence operator Satisfy one or a combination of the following constraints: (1) Tasks during the phenological period: This includes pointwise convolution and depthwise separable convolution to enhance the response to small-scale structural features; This includes a combination of local saliency pooling and global average pooling, used to simultaneously preserve key regional features such as the ear and overall contextual information; (2) Task on the degree of disease and its impact: This includes context-aware attention mechanisms and channel attention mechanisms, used to enhance the ability to model the texture structure of lesion regions and their contextual relationships; This includes global feature aggregation operations; (3) Nutritional status task: This includes feature fusion and channel mapping operations for color attributes; This includes global average pooling operations; (4) Group growth task: This includes directional convolution along the horizontal and vertical directions, used to model the row and column structure features of rice planting; This includes global feature aggregation operations.

9. The method for sensing rice growth status and making agricultural decisions according to claim 1, characterized in that, It also includes the following steps: S4: Agricultural Information Generation: Generate rice growth status analysis information based on multi-task recognition results, generate comprehensive analysis information on the current growth status of rice, and further generate agronomic control instructions that match the growth status. S5: Command Issuance and Execution: Agronomic control commands are sent to the paddy field management execution unit through a standardized communication protocol to achieve automated or semi-automated control of the paddy field; S6: Feedback and Adaptive Adjustment: Receives operation completion status information from the paddy field management execution unit and field environmental data collected by environmental sensors. Updates rice growth parameters or model input data based on the feedback information to guide identification and control decisions in the next cycle, thereby forming an adaptive closed-loop control for rice agricultural management.

10. The method for sensing rice growth status and making agricultural decisions according to claim 9, characterized in that, In step S5, the standardized communication protocol includes, but is not limited to, MQTT protocol, Modbus protocol or application programming interface (API), and the paddy field management execution unit includes intelligent spraying equipment, integrated water and fertilizer management system or agricultural Internet of Things (IoT) control terminal.