Agricultural monitoring system and method based on intelligent helmet

Through the integrated wearable terminal and multimodal fusion module of smart helmet, the inaccurate data acquisition and unstable communication problems of agricultural monitoring systems in complex environments are solved, real-time operation guidance and automated processing are realized, and the flexibility and accuracy of operations are improved.

CN120240756APending Publication Date: 2025-07-04BEIJING ZHUOYUAN ZHILIAN TECHNOLOGY CO LTD +1
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510672021.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing agricultural monitoring systems have inaccurate data collection, unstable communication and lagging operation guidance responses in complex environments, making it difficult to meet the needs of flexibility and real-time.

Method used

The agricultural monitoring system based on smart helmets is adopted, and the wearable terminal, active sampling control module, image recognition processing module, multi-modal fusion module and recognition inference module are integrated. Through image acquisition, environmental data fusion and dynamic trend recognition, accurate recognition and real-time feedback are achieved.

Benefits of technology

Stable data collection and communication are achieved in complex environments, improving the response speed and accuracy of operation guidance, and improving the automated processing capabilities of agricultural operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120240756A_ABST
    Figure CN120240756A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent agriculture and intelligent construction, and discloses an agricultural monitoring system and method based on an intelligent helmet, and the system comprises a wearable terminal, an active sampling control module, an image recognition processing module, a multi-modal fusion module, a recognition reasoning module, and a result output module. The method comprises the steps of collecting a field crop image and environment data; calculating an uncertainty index of an identification result according to the current image; if the uncertainty is high, adjusting the acquisition angle of the camera and re-sampling; constructing the image and the agricultural environment data into a multi-modal graph structure and carrying out graph neural network processing; performing dynamic trend identification on a fusion result, and outputting a crop state; and outputting an identification result through a helmet display or linking an agricultural control system to execute related operations. According to the invention, the capabilities of identification display, voice communication, operation guidance and system integration in a complex environment are improved, and the problems of poor real-time performance and weak interaction of an existing scheme are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of smart agriculture and intelligent construction, and particularly to an agricultural monitoring system and method based on an intelligent helmet. Background Art

[0002] Under the current development background of smart agriculture and intelligent construction, data collection and operation management are gradually evolving towards informatization and intelligence. More and more solutions introduce modules such as image recognition, environmental sensing, and remote control, aiming to achieve real-time monitoring and auxiliary decision-making of farmland conditions or construction sites. Such systems usually rely on manual inspections combined with mobile terminals, or sense and collect data by deploying drones or fixed monitoring devices, and are centrally processed by a backend server, and finally the results are fed back to the user side. However, in operation areas with strong scene dynamics or high operation intensity, the response delay of this "sensing - uploading - feedback" linear process is obvious.

[0003] In practical applications, traditional handheld terminals or fixed collection devices have shortcomings in terms of flexibility and operation continuity. Operators need to free their hands to operate the devices, which easily affects the coherence of operation actions and increases fatigue. Although drones can cover large areas, it is difficult to accurately focus on a single target point, and the flight environment is limited by factors such as weather and electromagnetic interference. At the same time, most solutions lack the ability of front-end intelligent discrimination, resulting in the collection quality highly depending on backend algorithm compensation, and the recognition accuracy fluctuates greatly. Especially in low-light, strong wind or complex terrain conditions, data consistency is difficult to guarantee.

[0004] In addition, the instant feedback and remote interaction of on-site information in existing solutions are still insufficient. Once network fluctuations or environmental noise increase, the voice communication quality deteriorates, and the efficiency of remote guidance is greatly reduced. Although some systems have basic communication functions, they lack noise reduction mechanisms and edge preprocessing capabilities, and the accuracy and stability of on-site feedback are difficult to meet the requirements of fine operations. Task scheduling and operation guidance also mostly rely on personnel experience and manual matching, and the service consistency and operation standardization are insufficient.

[0005] Therefore, the present invention proposes an agricultural monitoring system and method based on an intelligent helmet to solve the deficiencies of the prior art. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides an agricultural monitoring system and method based on an intelligent helmet, which solves the problems of inaccurate data collection, unstable communication, and lagging response of operation guidance in complex environments.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: An agricultural monitoring system based on an intelligent helmet, comprising: A wearable terminal is used to collect crop images and environmental information at an agricultural operation site and provide them to an active sampling control module and an image recognition and processing module; An active sampling control module is used to determine whether to perform angle adjustment based on the uncertainty index output by the image recognition and processing module for the initial image, and schedule the camera to collect supplementary images; An image recognition and processing module is used to perform status analysis on the images collected by the wearable terminal and transmit the recognition results to a multi-modal fusion module; A multi-modal fusion module is used to receive the image features output by the image recognition and processing module and perform feature fusion processing in combination with the environmental data from the wearable terminal; A recognition and inference module is used to identify the disease category of the fusion features output by the multi-modal fusion module and form a final determination result; A result output module is used to receive the recognition results of the recognition and inference module and display them through an AR interface or send the results to a background agricultural management platform through a communication unit.

[0008] Preferably, the active sampling control module includes: An image evaluation unit is used to calculate the uncertainty measure of the recognition result based on the initially collected image; An angle scheduling unit is used to control the sampling angle of the camera based on the uncertainty judgment result and re-collect images.

[0009] Preferably, the image evaluation unit is used to calculate the information entropy index, and the calculation method of the information entropy is: For the crop category prediction probability distribution p1, p2,..., p C , the information entropy is defined as: Among them, H represents the information entropy; C represents the number of categories, and p i represents the prediction probability of the i-th category.

[0010] Preferably, the image recognition and processing module includes: An image encoding unit is used to map the crop image into a compressed feature representation; A recognition output unit is used to generate the crop status category prediction probability distribution according to the compressed features.

[0011] Preferably, the image encoding unit uses variational information compression to generate the image latent variable z, where the latent variable conforms to a normal distribution and is trained by maximizing the following objective function, and the loss function used is: Among them, represents the total loss function value; denotes the expectation of the negative log-likelihood of the label y with respect to the conditional probability model based on the latent variable z; p(y|z) is the crop status prediction distribution, q(z|x) is the image encoding distribution, p(z) is the prior distribution, β is the weight parameter, D KL denotes the Kullback-Leibler divergence; z represents the latent encoding representation of the image features; y represents the true status label of the crop or the recognized target category.

[0012] Preferably, the multimodal fusion module includes: a node construction unit for constructing a graph structure including agricultural data nodes of images, temperature, humidity, and soil pH; a feature propagation unit for applying graph convolution processing to the graph structure to extract fused features.

[0013] Preferably, the feature propagation unit performs the following graph convolution processing on the graph structure node features, and its forward propagation formula is as follows: where H (l) represents the node feature matrix of the l-th layer, W (l) is the learnable weight matrix, σ is the activation function, is the normalized adjacency matrix, satisfying

[0014] Preferably, the recognition and inference module includes: a status update unit for performing time series inference and update on the crop recognition category; a trend judgment unit for outputting the final recognition result.

[0015] Preferably, the status update unit is used to construct a dynamic update model of the recognition preference variable φ(t), and its update method is: where λ is the update rate coefficient, S(t) is the recognition activation intensity, and η(t) is the Gaussian perturbation term.

[0016] The present invention also provides an agricultural monitoring method based on an intelligent helmet, including the following steps: S1. Collect field crop images and environmental data; S2. Calculate the uncertainty index of the recognition result according to the current image; S3. If the uncertainty is high, adjust the camera acquisition angle to resample; S4. Construct the image and agricultural environment data into a multimodal graph structure and perform graph neural network processing; S5. Perform dynamic trend recognition on the fusion result and output the crop status; S6. Output the recognition result through the head-mounted display or link the agricultural control system to perform relevant operations.

[0017] The present invention provides an agricultural monitoring system and method based on an intelligent helmet, having the following beneficial effects: 1. The present invention adopts an integrated design of an AR optical engine free-folding structure and a high-brightness display screen, achieving the technical effect of stably displaying recognition results even in complex environments such as strong light and low visibility. Compared with the fixed-view or low-brightness display solutions in the prior art, this structure effectively avoids the problem of line-of-sight occlusion, improving the operation experience of on-site personnel and the information acquisition efficiency.

[0018] 2. The present invention introduces an industrial-grade voice noise reduction model and the hardware arrangement of a continuous microphone array, making voice communication still clearly available in a noisy environment, and enabling real-time remote expert voice guidance. In traditional solutions, the communication quality drops significantly under wind noise and mechanical noise. The present invention avoids such distortion interference and improves the collaborative response speed.

[0019] 3. By integrating a task recognition system based on a knowledge graph and an operation process guidance mechanism, the present invention can automatically match the recognition result with the corresponding operation steps, achieving the goal of accurately pushing agricultural instructions. Compared with the method of manually flipping through operation procedures or manually matching operation suggestions, the present invention greatly reduces the usage threshold and the risk of misoperation.

[0020] 4. The present invention constructs a multi-functional platform architecture integrating image acquisition, environmental perception, intelligent recognition, and result push, significantly improving the automated processing ability of agricultural operations. Traditional solutions often have fragmented modules and rely on manual transfer, resulting in large data response delays and poor real-time performance. The present invention well solves this kind of system integration problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is the system architecture diagram of the present invention; Figure 2 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] Please refer to the attached Figure 1 , an embodiment of the present invention provides an agricultural monitoring system based on an intelligent helmet, including: A wearable terminal for collecting crop images and environmental information at an agricultural operation site and providing them to an active sampling control module and an image recognition and processing module; In the agricultural monitoring system provided by the present invention, the wearable terminal, as a core front-end sensing component, is deployed on the operator's head, and the image and environment acquisition components are embedded in the form of an intelligent helmet, providing an original data basis with time and space tags for subsequent judgment logic. It forms a tightly coupled structure with the active sampling control module and the image recognition and processing module, and needs to achieve continuous, stable, and low-latency data transmission and feedback response in complex field scenarios.

[0024] Generally, the terminal includes structures such as an image acquisition component, an environmental sensor component, a positioning module, an edge computing chip, a data buffer, a wireless communication interface, and a power supply control module.

[0025] In some embodiments, the image acquisition component uses a high-sensitivity CMOS image sensor with a photosensitive size of not less than 1 / 2.8 inches and an effective pixel count of approximately 2 million. The lens uses an F2.0 large aperture design to adapt to the drastic changes in field light intensity. The camera is fixed in front of the helmet, slightly below the line of sight level. The lens bracket has a mechanically adjustable structure, with a maximum support for a ±20° pitch range, and the drive method is micro servo linkage control.

[0026] In this embodiment, the image signal is first sampled in the YUV format, and then compressed into a JPEG image stream and input into the main control chip processing queue; at the same time, an image identification code and an acquisition timestamp are generated. Each frame of image I t The corresponding data structure includes: Image matrix: where h, w, and c respectively represent the height, width, and number of channels of the image; Position identifier: P t =(lat t , lon t ), which is the longitude and latitude collected by the GNSS module; Timestamp: Indicates the image acquisition time; Environmental data vector: For various environmental parameters, the definitions are as follows: e1 is the light intensity (unit: lux); e2 is the air temperature (unit: °C); e3 is the air humidity (unit: %); e4 is the soil humidity (unit: %); e5 is the soil pH value (unit: dimensionless); e6 is the atmospheric pressure (unit: hPa); e7 is the wind speed (unit: m / s); n represents the total number of environmental parameters collected; In a possible implementation, to improve the effectiveness of images under different lighting conditions, an image adaptive exposure control logic is deployed in the terminal. The exposure time T e is dynamically adjusted according to the central brightness mean and the edge brightness deviation. The control function is in the following form: T e = T0 + α·(L c - L ref ) + β·(L e - L c ); Where: T e is the actual exposure time, in milliseconds; T0 is the default base exposure time; L c is the brightness mean of the central region of the image; L e is the brightness mean of the edge of the image; L ref is the preset target brightness threshold; α, β are exposure gain control coefficients.

[0027] The above formula adjusts the exposure by the weighted difference method to ensure that the key crop feature regions in the central region are always within the effective gray-scale dynamic range, improving the clarity of micro-features such as disease spots and insect holes.

[0028] In some embodiments, the wearable terminal integrates an IMU module and a GNSS module. Through the time synchronization mechanism Δt ≤ 100 ms, the consistency of the image frame with the positioning and attitude information is ensured. The terminal attitude angle (roll, pitch, yaw) data is used in the subsequent camera view angle fine-tuning compensation mechanism to provide basic support for the active sampling control.

[0029] As an option, the data caching strategy in the terminal adopts a double-buffer structure, including a short-term cache pool and a persistent cache pool. After the image acquisition is completed, it is first stored in the short-term cache pool. If the frame image is determined by the active sampling control module to be a low-entropy image (i.e., the information entropy index H < θ1, where θ1 is the set threshold), it is marked and added to the persistent cache pool; otherwise, it is cleared and released to reduce redundant storage.

[0030] The calculation method of the information entropy is as follows: Where: H is the predicted information entropy of the image classification result; C is the total number of crop recognition categories; p i is the predicted probability that the crop image belongs to the i-th category, satisfying and 0 < p i < 1.

[0031] The image information entropy, as a measure of the image effectiveness, is used in active sampling to determine whether the current image contains sufficient discriminative features. If the information entropy is within the interval [θ1, θ2], it is represented as a "pending supplementary image", and the angle adjustment process will be triggered.

[0032] The environmental perception module in the terminal updates data every 2 seconds and synchronizes it to the main control cache queue. The data is marked with both the sampling timestamp and the GPS location. The main control unit can dynamically adjust the image acquisition frequency according to the operator's traveling speed, GPS trajectory density, and crop arrangement rules. For example, if the field density increases and the speed decreases, the system will increase the acquisition frequency from 1 frame every 10 seconds to 1 frame every 5 seconds.

[0033] In some implementation manners, the image acquisition component also supports the image pre-classification function based on the edge AI model. The image generates intermediate predictions through a lightweight convolutional network and is attached with soft labels where s i is the logit value of the i-th class. This prediction is only used for sampling control judgment and does not affect the main recognition path.

[0034] The terminal communication part supports Wi-Fi, Bluetooth, and low-power LoRa communication protocols, and adaptively switches the upload method according to the signal strength. When the cache overflows or there is no connection to the background for a long time, the data can be saved in the local encryption chip, and the maximum support for persistent data protection is 256MB.

[0035] Finally, in the field deployment test, the wearable terminal can achieve a continuous operation time of not less than 4 hours, the image acquisition hit rate is increased to more than 84%, and combined with the active sampling mechanism, the coverage rate of lesion features is increased by about 32.4% compared with the traditional fixed-angle acquisition.

[0036] The active sampling control module is used to judge whether to perform angle adjustment based on the uncertainty index output by the image recognition processing module for the initial image, and schedule the camera to collect supplementary images; After the system is initialized and the crop plot configuration is loaded, the wearable terminal starts to automatically perform image acquisition operations. The collected images first enter the image recognition processing module for forward recognition, and this module outputs a set of prediction probability vectors of crop status categories. The active sampling control module is then started based on this to quantify and judge the uncertainty of the prediction result.

[0037] Generally, the data interaction between the active sampling control module and the image recognition processing module is realized through a local high-speed cache, and the information transmission delay is controlled within 5ms to avoid processing blockage.

[0038] In some embodiments, the image evaluation unit not only performs information entropy calculation, but also combines pixel statistical feature analysis to judge whether the image contains strong structural crop feature regions, such as texture-dense regions like leaf veins and spot edges. If the image features are sparse, the system will increase the uncertainty penalty factor of this image.

[0039] In this embodiment, the image evaluation unit first calculates the image prediction information entropy according to the predicted probability distribution P = {p1, p2, …, p C}: where H is the prediction information entropy; C is the total number of crop disease categories; p i is the probability that the image is recognized as the i-th category; log(·) is the logarithmic function with the natural logarithm as the base; all p i satisfy and 0 < p i < 1.

[0040] To enhance the stable judgment ability of the classification confidence, the system introduces an entropy normalization mechanism: where H norm ∈(0, 1) is the normalized entropy value, and its upper limit is the maximum entropy in the uniform distribution.

[0041] Based on the normalized entropy value, the system sets two judgment thresholds θ low and θ high , satisfying 0 < θ low < θ high < 1, and makes the following judgments: If H norm < θ low : It is determined as an image with insufficient information; If H norm > θ high : It is determined as an image with a fuzzy distribution; If H norm ∈[θ low , θ high : The image is acceptable.

[0042] As an option, the system also introduces a confidence peak index to assist in decision-making: p max = max(p1, p2, …, p C ); If both p max < γ and H norm > θ low , the system determines that the current image has information but lacks obvious tendency, triggers the boundary image recognition fallback strategy, and sends the image into the angle supplementary acquisition process.

[0043] In this embodiment, the angle scheduling unit has a four-degree-of-freedom angle configuration strategy, including: Forward / backward tilt adjustment: ±α ° , generally controlled within ±15°; Horizontal offset: ±β °, for avoiding direct sunlight or occlusion; Height micro - lifting mechanism: The central axis of the camera is lifted approximately 5 - 10 cm through an electric lifting structure to adapt to different plant heights; Viewpoint selection polling: Preset an angle set {θ1, θ2, …, θ K}, and sample round by round to obtain the optimal frame.

[0044] In a possible implementation, the angle scheduling gives priority to the center - edge contrast of the image brightness distribution. Define the average central brightness as L c , and the average brightness of the edge region as L e , and the system constructs the following ratio function: where: η is the brightness deviation coefficient; L c , L e are the average central and edge brightnesses respectively; ∈ is a small constant to prevent division by zero, with a value of approximately 1; if η > δ, it indicates that the brightness distribution is severely uneven, and it is recommended to adjust the shooting angle towards the direction of uniform illumination.

[0045] In some embodiments, to avoid frequent invalid sampling, the system adds a sampling penalty factor mechanism. Whenever the angle scheduling is triggered once, the system includes the current viewpoint in the nearest viewpoint list V past , if the difference between the new viewpoint and any angle in V past is less than ξ degrees, the system automatically skips the current scheduling to prevent angle backtracking.

[0046] This embodiment also designs an image memory cache mechanism. The system retains the most recent N frames of images and their evaluation results, and uses the dynamic optimization method to select the frame with the lowest entropy value as the candidate input image for the next - step recognition. This mechanism is enabled in scenarios where the image information is insufficient for a long time or the viewpoint switching is ineffective.

[0047] In addition, the system supports custom configuration of values such as θ low , θ high , γ, δ, etc. for each plot through the configuration parameter interface. For example, in the disease - outbreak area, appropriately increase γ and decrease θ low to improve the sampling sensitivity of the system to early disease symptoms.

[0048] The image recognition processing module is used to perform state analysis on the images collected by the wearable terminal and transmit the recognition results to the multi - modal fusion module; The image recognition and processing module in the system of the present invention is immediately started after the active sampling control module determines that the image is analyzable, and undertakes the tasks of structural encoding, latent variable extraction and state category prediction of crop images. The module is implemented through a two-stage processing structure: the first stage is the image encoding unit, which is used to map the original image into potential compressed features; the second stage is the recognition output unit, which is used to infer the state of the latent variable and generate the classification probability result. The recognition result is further input into the multi-modal fusion module to participate in the comprehensive judgment of diseases.

[0049] Generally, this module runs in a lightweight structure on an embedded computing platform, taking into account both recognition efficiency and model expression ability. When computing resources permit, a deep network structure can also be deployed to improve robustness.

[0050] In this embodiment, the image encoding unit adopts a variational autoencoder structure (Variational Autoencoder, VAE) to compress the redundant features of the image and extract potential key patterns. The input image is defined as: where h represents the number of pixels in the height of the image, w is the number of pixels in the width of the image, and c is the number of channels, and the typical value is 3 (three-channel RGB image).

[0051] The encoder consists of a series of convolutional layers and normalization layers, and outputs the mean vector of the latent variable distribution and the logarithmic variance vector where d represents the size of the latent variable dimension, which is a constant set during model design, such as 64, 128 or 256.

[0052] During the encoding process, latent variable sampling adopts the reparameterization technique to support gradient propagation, and its formula is as follows: where: is the finally sampled latent variable representation; mean vector; standard deviation vector, calculated as is the perturbation variable sampled from the standard normal distribution; ⊙: is the Hadamard multiplication on the corresponding dimension, that is, the element-wise multiplication operation; represents the standard normal distribution with a mean of 0 and a covariance of the identity matrix.

[0053] In a possible implementation, the latent variable z is passed into the recognition output unit, and the recognition output unit includes two fully connected layers and a Softmax normalization layer, which are used to output the crop state prediction probability vector: where p i: The probability value that the image is recognized as the i-th class state; C: The total number of classes, including state types such as "healthy", "early disease spots", "wilt symptoms", "leaf pests", etc.; all probability values satisfy and 0 < p i < 1.

[0054] The recognition result is not only used for disease type determination, but also serves as the basis for information entropy calculation in the active sampling control module, participating in the image confidence analysis and re-sampling judgment process.

[0055] In this embodiment, the training of the entire image recognition processing module is carried out by jointly optimizing the objective function, and the loss function used is: where is the total loss function during the training process; is the expectation of the negative log-likelihood of the model prediction label y based on the sampled latent variable z, used to measure the prediction accuracy; p(y|z): represents the conditional probability distribution, that is, the probability of predicting the label y given the latent variable z; q(z|x): represents the posterior distribution generated by the image encoder; p(z): is the prior distribution, usually set to the standard normal distribution D KL (q(z|x)||p(z)): is the Kullback-Leibler divergence, used to measure the difference between the encoder distribution and the prior distribution, expressed as follows: β: is the weighting coefficient of the KL divergence, controlling the trade-off between the model information compression degree and the reconstruction ability.

[0056] As an option, to enhance the model's recognition ability for crop abnormal areas, the system introduces an attention mechanism module in the encoding network. It includes the combination of channel attention (SE-block) and spatial attention (CBAM structure) to highlight the areas with significant regional differences and prominent texture density in the image, such as the edges of disease spots and the curled leaf tips.

[0057] In some embodiments, the image recognition processing module also supports the image multi-scale resampling mechanism, generating different resolution copies of the input image, such as 128×128, 256×256, 512×512. After each copy extracts features through an independent encoding path, they are fused to enhance the robust extraction ability for disease patterns with different far-near ratios.

[0058] As a supplement, in a multi-target image scenario, the image recognition module can also call an external object detection network (such as YOLO or SSD) to first segment the crop area, then independently encode and discriminate each area, and finally fuse the recognition results of each target to form a comprehensive output. This structure shows more stable recognition effect in the dense planting area.

[0059] A multi-modal fusion module, which is used to receive the image features output by the image recognition processing module and perform feature fusion processing in combination with the environmental data from the wearable terminal; In the implementation system of the present invention, the crop image latent variable output by the image recognition processing module is one of the important inputs for downstream inference. However, if only relying on visual features, the true growth state of the crop may be insufficiently characterized. Therefore, the multi-modal fusion module is designed to receive image features and combine multi-source environmental parameters collected by the wearable terminal, and complete cross-modal information fusion and expression by constructing a graph structure and applying a graph neural network propagation mechanism.

[0060] Generally, the multi-modal fusion module includes two core components, a node construction unit and a feature propagation unit, which are responsible for data structured modeling and feature fusion tasks respectively.

[0061] Node construction unit. In this embodiment, the core task of the node construction unit is to uniformly represent multi-source observation data that is time-synchronized or spatially adjacent as graph nodes, and construct a graph structure using the adjacency structure. Each node represents an agricultural observation sample, and its feature dimension includes image features and multiple environmental sensing parameters. The specific composition is as follows: Image feature vector Compressed latent variable output from the image recognition processing module; Temperature feature Air temperature around the crop collected by the wearable terminal, in degrees Celsius; Humidity feature Relative air humidity percentage, ranging from 0% to 100%; Soil pH feature Collected soil pH value, generally ranging from 3 to 10; Timestamp feature Indicates the data collection time, in Unix timestamp or days; Location feature Used to record the spatial distribution of observation points, such as GPS coordinates or regional grid ID.

[0062] The above features can be concatenated to form a multi-modal node vector where d = d z +1+1+1+1+2 = d z +6.

[0063] In a typical implementation, the connection relationship between nodes is represented by constructing an adjacency matrix The construction rules include but are not limited to the following: Temporal connection: If the time difference between the acquisitions of two nodes is less than the preset threshold Δτ, an edge is established; Spatial connection: If the spatial distance between nodes ||l i -l j ||2 ≤ δ, an edge is established; Content connection: If the cosine similarity of the image features of nodes is higher than the threshold θ, an edge is established.

[0064] To ensure the computational stability of the graph neural network, the system further performs symmetric normalization on the adjacency matrix, which is defined as follows: Where: The original adjacency matrix; The degree matrix, which is a diagonal matrix, and the elements are defined as D ii = The normalized adjacency matrix, which avoids numerical inflation during the feature propagation process.

[0065] The feature propagation unit, after completing the construction of the graph structure, uses the graph convolution method to propagate and fuse the information between nodes to obtain a multi-modal representation rich in neighbor context.

[0066] In this embodiment, the graph convolution network (GCN) method proposed by Kipf et al. is adopted, and its forward propagation formula is as follows: Where, The node feature matrix of the l-th layer, where each row represents a node feature; The node features of the next layer; The weight matrix of the l-th layer, which performs feature dimension transformation; σ(·): The activation function, generally ReLU, is defined as σ(x) = max(0, x); The normalized adjacency matrix, which is used for feature weighted propagation in the above formula.

[0067] Generally, the model uses a 2- to 3-layer graph convolution network to implement feature propagation, and a normalization module such as BatchNorm or LayerNorm is connected after the final output layer to improve the convergence effect.

[0068] As an option, the attention mechanism can also be used to replace the static adjacency matrix to model the asymmetric dependence relationship between multi-modal features. For example, the graph attention network (GAT) mechanism is adopted, and the attention weight between nodes is defined as follows: e ij = LeakyReLU(aT [Wh i ||Wh j ); Wherein: represents the input feature of the i,j-th node; feature mapping weight matrix; attention parameter vector; ||: vector concatenation operation; set of adjacent nodes of the i-th node; α ij : normalized attention weight between nodes.

[0069] The attention mechanism is particularly suitable for scenarios with strong node heterogeneity and can automatically assign importance weights to different modal features.

[0070] The node output feature of the final multi-modal graph structure can be used as a comprehensive representation vector of the crop state for downstream modules such as trend inference modules, classifier modules, etc. Common usage methods include: Classification of the current state of the crop (such as healthy, early disease, severe disease); Detection of abnormal change points; Cross-time and space crop similarity comparison; Time series trend prediction modeling.

[0071] In some embodiments, a global graph embedding representation g = Readout(H (L) ) can also be constructed based on the fused node representation, such as average pooling: used to describe the health status of the overall crop area.

[0072] Recognition and inference module, used to recognize the disease category of the fused features output by the multi-modal fusion module and form a final determination result; After the agricultural monitoring system described in the present invention completes primary perception processing such as image recognition processing and multi-modal feature fusion, it is still necessary to further establish a module with memory and dynamic inference capabilities to adapt to problems such as time series fluctuations, environmental interference, and recognition uncertainty in agricultural field recognition. Therefore, a recognition and inference module is provided in the system structure, and its function is to perform dynamic recognition processing on the fused features output by the previous module (especially the multi-modal fusion module) and make a final stable crop disease category judgment based on the time evolution trend.

[0073] The recognition and inference module includes two core units: a state update unit and a trend judgment unit. The former is used to construct the time evolution process of the recognition preference variable, and the latter is used to make a decision based on the state trend.

[0074] State update unit. Generally, the state update unit introduces a vector variable that evolves over time used to represent the system's current preference degree for each recognition category, where C is the total number of all recognition categories, such as a total of 4 to 8 categories including "health", "fungal disease", "pest", "physiological nutrient deficiency", etc.

[0075] The evolution of this variable follows the following dynamic update formula: where represents the current recognition preference vector, and the i-th element φ i (t) represents the system's preference degree for the i-th state, and its value range is [0, 1], but it does not need to satisfy the sum of 1; represents the category activation intensity output by the multimodal fusion module at the current moment, usually the predicted probability distribution after Softmax normalization, satisfying each dimension State update rate coefficient, which controls the influence amplitude of the current recognition result on the preference variable. When it is larger, the response is more sensitive, and the numerical range is generally set at is the state perturbation term, which models the unmodeled dynamics or perceptual error fluctuations, and is set as a zero-mean Gaussian process, satisfying: where is the perturbation covariance matrix, usually set as the diagonal matrix Σ = σ 2 I, where is the perturbation standard deviation coefficient, usually taking values in the range of 0.01 - 0.05.

[0076] As a numerical implementation method, in an embedded or real-time system, this state evolution process can be discretized into a difference expression form: φ(t + Δt) = φ(t) + λ · (S(t) - φ(t)) · Δt + η(t); where: Δt: is the time step, generally taking the sampling period value, such as Δt = 1 second; φ(t + Δt): is the updated state at the next time step; η(t): is the perturbation at the current step, which can be randomly sampled from a Gaussian distribution.

[0077] To improve the update stability, in some embodiments, a sliding window average strategy is adopted to smooth S(t): where k is the window size, usually set to 3 - 5.

[0078] Trend judgment unit. The trend judgment unit is responsible for outputting the final recognition category after the recognition preference variable φ(t) accumulates and converges, so as to avoid misrecognition caused by short-term mutations.

[0079] In a basic implementation, the determination formula for the final recognition result y * (t) is: That is, the category index i corresponding to the maximum value in the current preference vector is selected as the recognition result.

[0080] To improve robustness and prevent misjudgment when the confidence level is too low, the system introduces a confidence threshold function: If it satisfies: Conf(t) ≥ θ; Then the category is output as a valid recognition, otherwise "to be determined" or the previous recognition result is maintained. The threshold parameter θ ∈ (0.5, 0.9) can be dynamically adjusted according to the risk tolerance.

[0081] Specifically, to determine the trend continuity of the judgment result, the system can also add a stability judgment strategy: when the consistent result y * (t) = y * (t - 1) =... = y * (t - m + 1) is output for m consecutive time steps (such as m = 3), then the recognition value is output to prevent incorrect switching caused by momentary disturbances.

[0082] In a possible implementation, the trend judgment module can also introduce an output buffer mechanism to apply a delay filtering strategy to the category result, such as exponential weighted average classification confidence: where α ∈ (0, 1) is the smoothing coefficient, usually set to 0.3 - 0.7. The final category depends on the maximum component of the smoothed preference vector : The above mechanism enables the system to not only have the ability to suppress short-term prediction anomalies, but also make a stable response to the progressive evolution process of the crop growth state.

[0083] The result output module is used to receive the recognition result of the recognition and inference module and display it through the AR interface or send the result to the background agricultural management platform through the communication unit; The result output module described in the present invention is a key functional module at the processing end in the intelligent agricultural monitoring system. Its main task is to complete the display presentation and communication output of the result after the recognition and inference module outputs the final crop state recognition result, and construct a visible interaction channel that can be perceived by the user and a remote data channel that can be recorded by the background.

[0084] In the system operation process, the status results output by the recognition and reasoning module are represented in structured data, including the recognition category, confidence level, image number, spatial position information, and timestamp; the result output module receives this structured data and processes it according to the mode set by the user, and visual output and remote upload can be executed independently or synchronously.

[0085] The graphics rendering subunit. Generally, the graphics rendering subunit is used to receive the recognition results and generate a visual annotation layer for the augmented reality interface (AR interface), and the display content includes recognition labels, confidence scores, category colors, highlight prompts, etc. Specifically, the structured information output by the recognition and reasoning module is: where: * (t) ∈ {1, 2,..., C}: the recognition category output by the current recognition and reasoning module; The recognition confidence level, which is the maximum value in the category preference variables; The current image frame or node number; The recognition timestamp, with the unit of seconds; The spatial position corresponding to the recognition point, which can be GPS coordinates or local plane coordinates; The three-dimensional vector position of the target point in the camera coordinate system, which is used for AR space anchoring.

[0086] The above data is parsed into a display layer information structure: where: is the status name corresponding to the y * class, such as "leaf spot disease", "water deficiency", etc.; Score: used to display the confidence score value or grade color scale; Pos: mapped as the spatial anchor point coordinates in the augmented reality view; Time: displayed in the form of a floating time bar or used for judging the expiration of recognition.

[0087] In a possible implementation, this module calls the augmented reality rendering interface (such as ARCore or ARKit) to bind the recognition results to the spatial position under the image acquisition perspective, realizing the real-time binding between the information and the real crop image.

[0088] As a supplement, this module also includes a perspective adaptation mechanism and occlusion detection logic to prevent the loss of positioning information when the annotation floats or is blocked by the crop main body; the system supports anchor point locking and dynamic adjustment, and calculates the projection matrix based on the current pose of the head-mounted terminal to correct the graphic projection deviation in real time.

[0089] A communication transmission subunit, which is used to push the structured recognition result to the agricultural management platform and supports mechanisms such as periodic uploading, instant pushing, and abnormal resending.

[0090] Before being sent, the recognition result is encapsulated into a standardized data message: Among them, DeviceID: the unique identifier of the current terminal device; y * (t): the index of the recognized disease category; The corresponding confidence value; τ(t): timestamp, in seconds; l(t): spatial location (GPS longitude and latitude or grid code); Sensor(t): a triple containing the temperature T(t), humidity H(t), and pH value pH(t) at this moment FrameRef: the original image frame number or the pointer to the encoded feature index.

[0091] The above message can be converted into a JSON or ProtocolBuffer encoding structure, and the sending protocol supports communication methods such as MQTT, HTTP(S), TCP, or LoRa to adapt to different channel environments and data rate requirements.

[0092] In some embodiments, the result output module supports an output policy table configured by the user. For example: Results with a confidence level lower than the set threshold θ o ∈(0.3, 0.6) are not displayed; If the status is continuously abnormal, a "trend warning" module will pop up on the display interface; Supports filtering and displaying by category. For example, users can choose to view only the results related to fungal diseases; All result outputs are added with a local timestamp and a unique index number to support subsequent traceability and statistics.

[0093] The present invention also provides an agricultural monitoring method based on an intelligent helmet, including the following steps: S1. Collect images of field crops and environmental data; Generally, the wearable terminal obtains the image information of the crop from the current perspective through the integrated image acquisition component, and at the same time calls multiple environmental sensors to synchronously collect key agricultural parameters such as temperature, humidity, and soil pH value to form structured original observation data.

[0094] S2. Calculate the uncertainty index of the recognition result based on the current image; After receiving the collected image, the system initially recognizes the crop status through a preset model, and calculates the recognition uncertainty index under this image in combination with the feature distribution or the consistency of the classification output to judge the discriminant reliability of the current sample.

[0095] S3. If the uncertainty is high, adjust the camera acquisition angle and resample; As a strategy, if the uncertainty exceeds the set threshold, the system will control the camera to perform a slight angular rotation or position movement, and re-acquire image data from a new observation perspective to improve the stability and determination accuracy of subsequent recognition.

[0096] S4. Construct the image and agricultural environment data into a multi-modal graph structure and perform graph neural network processing; The image features and multi-dimensional environmental parameters are jointly constructed into the nodes of the graph structure, and the edges are connected according to the temporal or spatial association between the nodes; this multi-modal graph structure is input into the graph neural network model for fusion calculation to extract the comprehensive representation features.

[0097] S5. Perform dynamic trend recognition on the fusion result and output the crop status; The system performs time series analysis based on the output result of the graph neural network combined with the historical status information to determine whether the crop is in a healthy, early disease or other state, and outputs the current credible crop recognition category.

[0098] S6. Output the recognition result through the head-mounted display or link the agricultural control system to perform relevant operations; The final recognition result will be displayed in real time on the AR interface of the head-mounted terminal for farmers to refer to for decision-making; at the same time, the system can also transmit the recognition result to the agricultural background platform to trigger control logics such as irrigation and fertilization to achieve automated agricultural linkage.

[0099] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An agricultural monitoring system based on an intelligent helmet, characterized in that, Including: A wearable terminal, which is used to collect crop images and environmental information at the agricultural operation site and provide them to the active sampling control module and the image recognition and processing module; An active sampling control module, which is used to judge whether to perform angle adjustment based on the uncertainty index output by the image recognition and processing module for the initial image, and schedule the camera to collect supplementary images; An image recognition and processing module, which is used to analyze the state of the images collected by the wearable terminal and transmit the recognition results to the multi-modal fusion module; A multi-modal fusion module, which is used to receive the image features output by the image recognition and processing module and perform feature fusion processing in combination with the environmental data from the wearable terminal; A recognition and inference module, which is used to identify the disease category of the fusion features output by the multi-modal fusion module and form a final judgment result; A result output module, which is used to receive the recognition results of the recognition and inference module and display them through the AR interface or send the results to the background agricultural management platform through the communication unit.

2. The agricultural monitoring system based on an intelligent helmet according to claim 1, wherein The active sampling control module includes: An image evaluation unit, which is used to calculate the uncertainty measure of the recognition result based on the initially collected image; An angle scheduling unit, which is used to control the sampling angle of the camera based on the uncertainty judgment result and re-collect images.

3. The agricultural monitoring system based on an intelligent helmet according to claim 2, characterized in that, The image evaluation unit is used to calculate the information entropy index, and the calculation method of the information entropy is as follows: For the predicted probability distributions p1, p2, …, p of crop categories C , the information entropy is defined as: Among them, H represents the information entropy; C represents the number of categories, and p i represents the predicted probability of the i-th category.

4. The agricultural monitoring system based on an intelligent helmet according to claim 1, wherein, The image recognition and processing module includes: An image encoding unit, which is used to map the crop image into a compressed feature representation; A recognition output unit, which is used to generate the prediction probability distribution of the crop state category according to the compressed features.

5. The agricultural monitoring system based on an intelligent helmet according to claim 4, wherein, The image encoding unit uses the variational information compression method to generate the image latent variable z, where the latent variable conforms to the normal distribution and is trained by maximizing the following objective function, and the loss function used is: Among them, represents the total loss function value; represents the expectation of the negative log-likelihood of the conditional probability model of the label y based on the latent variable z; p(y|z) is the crop state prediction distribution, q(z|x) is the image encoding distribution, p(z) is the prior distribution, β is the weight parameter, D KL represents the Kullback-Leibler divergence; z represents the latent coding representation of the image features; y represents the true state label of the crop or the recognized target category.

6. The agricultural monitoring system based on an intelligent helmet according to claim 1, wherein, The multi-modal fusion module includes: A node construction unit, which is used to construct a graph structure containing agricultural data nodes such as images, temperature, humidity, and soil pH; A feature propagation unit, which is used to apply graph convolution processing to the graph structure to extract fusion features.

7. The agricultural monitoring system based on an intelligent helmet according to claim 6, wherein, The feature propagation unit performs the following graph convolution processing on the node features of the graph structure, and its forward propagation formula is as follows: Among them, H (l) represents the node feature matrix of the l-th layer, W (l) is a learnable weight matrix, σ is an activation function, is a normalized adjacency matrix, satisfying 8. The agricultural monitoring system based on an intelligent helmet according to claim 1, characterized in that, The recognition and inference module includes: A state update unit, which is used to perform time series inference and update on the crop recognition category; A trend judgment unit, which is used to output the final recognition result.

9. The agricultural monitoring system based on an intelligent helmet according to claim 8, wherein, The state update unit is used to construct a dynamic update model of the recognition preference variable φ(t), and its update method is: where λ is the update rate coefficient, S(t) is the recognition activation intensity, and η(t) is the Gaussian perturbation term.

10. A method for agricultural monitoring based on an intelligent helmet, applied to an agricultural monitoring system based on an intelligent helmet according to any one of claims 1-9, characterized in that, Including the following steps: S1. Collect field crop images and environmental data; S2. Calculate the uncertainty index of the recognition result according to the current image; S3. If the uncertainty is high, adjust the camera sampling angle to re-sample; S4. Construct the image and agricultural environment data into a multi-modal graph structure and perform graph neural network processing; S5. Perform dynamic trend recognition on the fusion result and output the crop state; S6. Output the recognition result through the helmet display or link the agricultural control system to perform relevant operations.

Citation Information

Cited By

  • Crop insect pest intelligent identification and early warning system based on multi-mode large language model

    CN120931968A

  • A crop pest intelligent identification and early warning system based on a multi-modal large language model

    CN120931968B

  • Crop phenotype identification method and system based on computer vision

    CN121280437A

  • Tunnel disease evolution knowledge graph updating method based on graph memory

    CN122470784A

  • A tunnel disease evolution knowledge graph updating method based on graph memory

    CN122470784B