Grain bin pest status monitoring method, apparatus, system, electronic device, and medium

By combining multimodal data fusion and multi-level text encoding with visible light and near-infrared images and text description information, the problems of low efficiency and poor accuracy in monitoring the status of pests in grain warehouses have been solved, and end-to-end monitoring of pests and their life status has been achieved with accurate identification.

CN122090387BActive Publication Date: 2026-08-04CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
Filing Date
2026-04-24
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies for monitoring the status of pests in grain warehouses suffer from low efficiency, poor accuracy, and difficulty in adapting to complex real-world scenarios. In particular, image monitoring methods are costly and lack representativeness, while non-image monitoring methods cannot directly determine the life status of pests and have poor recognition capabilities.

Method used

A multimodal data fusion method is adopted, which combines visible light images and near-infrared images with multi-level text description information. Through image modality fusion module, multi-level text encoding module, cross-modal feature alignment module and feature classification module, end-to-end pest status monitoring is achieved.

Benefits of technology

It improves the accuracy and robustness of monitoring the status of pests in grain warehouses, enables precise identification of pests and their life status, overcomes the shortcomings of insufficient information from a single data source, and enhances the flexibility and accuracy of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090387B_ABST
    Figure CN122090387B_ABST
Patent Text Reader

Abstract

The application provides a kind of granary pest state monitoring method, device, system, electronic equipment and medium, it is related to artificial intelligence technical field, comprising: obtaining multi-modal data;Multi-modal data includes the visible light image and near infrared image of the granary to be monitored and text description information;Text description information describes the storage environment information of the granary to be monitored and the pest information in the granary to be monitored based on multiple levels;The multi-modal data is input into the pest state monitoring model, and the pest state monitoring model is aligned and identified based on the multi-modal data, and the pest state monitoring result of the granary to be monitored is output.The method and device provided by the application significantly improve the monitoring accuracy and robustness of the model in complex and variable real granary environment through multi-modal data identification;Through multiple levels of image features and text features alignment, the flexibility and accuracy of pest state monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device, system, electronic device, and medium for monitoring the status of pests in grain storage. Background Technology

[0002] Current technical methods for monitoring the mortality status of grain storage pests mainly fall into two categories: traditional monitoring and intelligent monitoring. Traditional monitoring primarily relies on manual inspection, physical or chemical attraction devices combined with laboratory verification, which is inefficient and difficult to implement continuously. Intelligent monitoring includes image monitoring methods and non-image monitoring methods. Image monitoring methods can be further divided into methods that combine automatic sampling equipment with visual recognition, methods that use fixed cameras or drones to capture images and then analyze them using algorithms, and methods that fuse visible light and infrared imaging. Non-image monitoring methods mainly include sound signal analysis and sensor networks monitoring environmental parameters to indirectly infer the pest status.

[0003] However, these technologies still have many drawbacks. Traditional monitoring methods are inefficient, lack real-time performance, and rely on manual labor. Image monitoring methods involve costly sampling equipment that may damage crops, have insufficient representativeness, and have limited image quality and coverage from fixed cameras. Drone applications face challenges such as high cost and technical barriers, interference with pest activity, and insufficient data fusion during imaging. At the algorithm level, they generally rely on small models with weak generalization capabilities, utilize limited data, have poor ability to identify unknown pest species, and most methods are only validated in laboratory environments, making them difficult to handle complex real-world scenarios. Non-image monitoring methods cannot directly and accurately determine the life status of pests, nor can they distinguish species or locate individuals.

[0004] Therefore, how to accurately identify pests and their life state in grain warehouses has become a technical problem that the industry urgently needs to solve. Summary of the Invention

[0005] This application provides a method, device, system, electronic device, and medium for monitoring the status of pests in grain warehouses, which solves the technical problem of how to accurately identify pests and their life status in grain warehouses.

[0006] This application provides a method for monitoring the status of pests in grain storage facilities, including:

[0007] Acquire multimodal data; the multimodal data includes visible light images and near-infrared images of the grain warehouse to be monitored, as well as text description information; the text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels; the multiple levels include scene level, instance level and attribute level;

[0008] The multimodal data is input into the pest status monitoring model, which performs feature alignment and identification on the multimodal data and outputs the pest status monitoring results for the grain warehouse to be monitored.

[0009] In some embodiments, the pest status monitoring model includes an image modality fusion module, a multi-level text encoding module, a cross-modal feature alignment module, and a feature classification module;

[0010] The image modality fusion module is used to extract features and fuse modalities between the visible light image and the near-infrared image to obtain fused visual features;

[0011] The multi-level text encoding module is used to encode the text description information to obtain text features at each level;

[0012] The cross-modal feature alignment module is used to perform cross-modal alignment between the fused visual features and the text features at each level to obtain the alignment features at each level.

[0013] The feature classification module is used to identify the alignment features at each level to obtain the pest status monitoring results.

[0014] In some embodiments, the method further includes:

[0015] The visible light image and the near-infrared image are fused using an image modality fusion module, and the fused image features are enhanced using a spatial attention mechanism to obtain the fused visual features.

[0016] In some embodiments, the step of fusing features from the visible light image and the near-infrared image, and then performing feature enhancement on the fused image features based on a spatial attention mechanism to obtain the fused visual features, includes:

[0017] Feature extraction is performed on the visible light image and the near-infrared image respectively to obtain the first visible light image feature and the first near-infrared image feature;

[0018] The first fused image features are enhanced by a spatial attention mechanism to obtain the first enhanced image features; the first fused image features are obtained by matrix point addition of the first visible light image features and the first near-infrared image features.

[0019] The first enhanced image feature is matrix-added with the first visible light image feature and the first near-infrared image feature respectively to obtain the first visible light fused image feature and the first near-infrared fused image feature.

[0020] The first visible light fused image features and the first near-infrared fused image features are stitched together and feature enhancement based on spatial attention mechanism to obtain the second enhanced image features;

[0021] The second enhanced image feature is matrix-added with the first visible light fused image feature and the first near-infrared fused image feature respectively to obtain the second visible light fused image feature and the second near-infrared fused image feature.

[0022] The fused visual features are obtained by matrix point addition of the second visible light fused image features and the second near-infrared fused image features.

[0023] In some embodiments, the scene-level text description information includes crop background information and real-time environmental data in the grain warehouse to be monitored; the real-time environmental data includes at least one of real-time temperature information, real-time humidity information, and real-time gas monitoring data.

[0024] The text description information at the instance level includes pest species information;

[0025] The textual description information at the attribute level includes pest biological characteristics information.

[0026] In some embodiments, the method further includes:

[0027] Semantic parsing is performed on the text description information at the multiple levels to obtain the triples corresponding to each level;

[0028] A dynamic knowledge graph is constructed based on the triples corresponding to each level.

[0029] In some embodiments, the method further includes:

[0030] The multi-level text encoding module performs vectorized representation and hierarchical layering of entity nodes and relation edges in the dynamic knowledge graph corresponding to the text description information to obtain scene subgraph, instance subgraph and attribute subgraph.

[0031] Feature extraction is performed on the scene sub-image to obtain scene text features, and the scene text features are processed to obtain scene scaling coefficient and scene offset coefficient;

[0032] Feature extraction is performed on the attribute subgraph to obtain initial attribute features, and dynamic feature modulation is performed on the initial attribute features based on the scene scaling coefficient and the scene offset coefficient to obtain attribute text features;

[0033] Feature extraction is performed on the instance subgraph to obtain initial instance features, and the initial instance features and the initial attribute features are processed based on a hierarchical attention mechanism to obtain instance text features.

[0034] In some embodiments, the method further includes:

[0035] The cross-modal feature alignment module globally aligns the fused visual features with the scene-level text features to obtain scene-aligned visual features.

[0036] Multi-scale feature extraction is performed on the scene alignment visual features to obtain visual multi-scale features;

[0037] The visual multi-scale features are aligned with the instance text features at the instance level to obtain instance-aligned multi-scale features.

[0038] Region detection is performed on the instance-aligned multi-scale features to obtain the target detection box and the image region features corresponding to the target detection box;

[0039] The image region features corresponding to the target detection box are locally aligned with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features.

[0040] In some embodiments, the step of globally aligning the fused visual features with the scene text features at the scene level to obtain scene-aligned visual features includes:

[0041] The fused visual features are processed using a global attention mechanism to obtain a scene query vector;

[0042] The scene text features are processed using a common attention mechanism to obtain scene key vectors and scene value vectors;

[0043] The scene query vector, scene key vector, and scene value vector are processed based on a cross-attention mechanism to obtain global alignment features;

[0044] The global alignment features are processed using a feedforward neural network to obtain the scene alignment visual features.

[0045] In some embodiments, aligning the visual multi-scale features with the instance text features at the instance level to obtain instance-aligned multi-scale features includes:

[0046] The visual multi-scale features are processed based on a deformable attention mechanism to obtain an instance query vector;

[0047] The instance text features are processed using a common attention mechanism to obtain instance key vectors and instance value vectors;

[0048] The instance query vector, instance key vector, and instance value vector are processed based on the cross-attention mechanism to obtain region alignment features;

[0049] The region alignment features are processed using a feedforward neural network to obtain the instance alignment multi-scale features.

[0050] In some embodiments, the step of locally aligning the image region features corresponding to the target detection box with the attribute text features of the attribute level to obtain attribute-aligned multi-scale features includes:

[0051] The image region features corresponding to the target detection box are processed based on a multi-head attention mechanism to obtain an attribute query vector;

[0052] The attribute text features are processed using a common attention mechanism to obtain attribute key vectors and attribute value vectors;

[0053] The attribute query vector, attribute key vector, and attribute value vector are processed based on a cross-attention mechanism to obtain local alignment features;

[0054] The local alignment features are processed using a feedforward neural network to obtain the attribute alignment multi-scale features.

[0055] In some embodiments, the training loss of the pest status monitoring model is determined based on at least one of pest classification loss, pest detection loss, pest detection box matching loss, and model knowledge preservation loss.

[0056] The pest classification loss is determined based on the difference between the predicted value and the actual value of the pest status monitoring results corresponding to the multimodal data of the sample.

[0057] The pest detection loss is determined based on the difference between the predicted location and the actual location of the target detection box corresponding to the multimodal data of the sample.

[0058] The pest detection box matching loss is determined based on the matching degree between the predicted position and the actual position of the target detection box corresponding to the multimodal data of the sample.

[0059] The model knowledge retention loss is determined based on the information divergence between the pest state monitoring model trained in the current iteration and the pest state monitoring model trained in the previous iteration.

[0060] This application provides a grain storage pest status monitoring device, comprising:

[0061] The data acquisition module is used to acquire multimodal data, which includes visible light and near-infrared images of the grain warehouse to be monitored, as well as text description information. The text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels, including scene level, instance level and attribute level.

[0062] The status monitoring module is used to input the multimodal data into the pest status monitoring model, and the pest status monitoring model performs feature alignment and identification on the multimodal data, and outputs the pest status monitoring results of the grain warehouse to be monitored.

[0063] This application provides a grain storage pest status monitoring system, including a data acquisition device, a data transmission device, and a data processing device;

[0064] The data acquisition device is installed in the grain warehouse to be monitored and is used to acquire visible light images, near-infrared images and real-time environmental data of the grain warehouse to be monitored.

[0065] The data transmission device is communicatively connected to the data acquisition device and is used to send the visible light image, the near-infrared image and the real-time environmental data to the data processing device.

[0066] The data processing device is communicatively connected to the data transmission device and is used to execute the grain warehouse pest status monitoring method, processing the visible light image, the near-infrared image and the real-time environmental data to obtain the pest status monitoring results of the grain warehouse to be monitored.

[0067] In some embodiments, the data acquisition device includes a circular track deployed on the top of the grain silo to be monitored, and a data acquisition unit that slides along the circular track;

[0068] The data acquisition unit is equipped with an image sensor, a temperature sensor, a humidity sensor, and a gas sensor;

[0069] The image sensor is used to acquire visible light images and near-infrared images of the surface of the grain silo to be monitored;

[0070] The temperature sensor is used to collect real-time temperature information in the grain warehouse to be monitored.

[0071] The humidity sensor is used to collect real-time humidity information in the grain warehouse to be monitored.

[0072] The gas sensor is used to collect real-time gas monitoring data in the grain silo to be monitored.

[0073] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the grain storage pest status monitoring method.

[0074] This application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the grain storage pest status monitoring method.

[0075] This application provides a computer program product, including a computer program that, when executed by a processor, implements the grain storage pest status monitoring method.

[0076] The grain storage pest status monitoring method, device, system, electronic equipment, and medium provided in this application organically combine multimodal data from different sources and of different types, input them into the constructed pest status monitoring model, and obtain pest status monitoring results, realizing end-to-end pest detection and pest status identification. Through multimodal data identification, the shortcomings of insufficient information from a single data source are overcome, significantly improving the monitoring accuracy and robustness of the model in complex and variable real grain storage environments. By aligning image features and text features at multiple levels, the flexibility and accuracy of pest status monitoring are improved. Ultimately, accurate identification of pests and their life status in grain storage is achieved. Attached Figure Description

[0077] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0078] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0079] Figure 1 This is a flowchart illustrating the grain storage pest status monitoring method provided in this application.

[0080] Figure 2 This is a schematic diagram of the pest status monitoring model provided in this application.

[0081] Figure 3 This is a schematic diagram of the image modality fusion module provided in this application.

[0082] Figure 4 This is a schematic diagram of the structure of the multi-level text encoding module provided in this application.

[0083] Figure 5 This is a schematic diagram of the cross-modal feature alignment module provided in this application.

[0084] Figure 6 This is a schematic diagram of the structure of the global alignment module provided in this application.

[0085] Figure 7 This is a schematic diagram of the structure of the region alignment module provided in this application.

[0086] Figure 8 This is a schematic diagram of the structure of the local alignment module provided in this application.

[0087] Figure 9 This is a schematic diagram of the grain storage pest status monitoring device provided in this application.

[0088] Figure 10 This is a schematic diagram of the grain storage pest status monitoring system provided in this application.

[0089] Figure 11 This is the overall architecture diagram of the grain storage pest status monitoring system provided in this application.

[0090] Figure 12 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0091] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0092] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps, units, or modules is not necessarily limited to those explicitly listed, but may include other steps, units, or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0093] Grain storage pest monitoring refers to the use of specific technologies to monitor pests stored in grain warehouses, determine whether they are dead, and assess the mortality rate. This monitoring typically utilizes intelligent devices, such as image recognition technology and sensors, to acquire real-time information on pest activity and uses data analysis to determine their survival status. Its significance lies in ensuring the safety and quality of stored grain.

[0094] To address the shortcomings of existing technologies in monitoring the status of pests in grain warehouses, such as low efficiency, poor accuracy, and inability to adapt to complex real-world scenarios. Figure 1 This is a flowchart illustrating the grain storage pest status monitoring method provided in this application, as shown below. Figure 1 As shown, the method includes steps 110 and 120.

[0095] Step 110: Acquire multimodal data; the multimodal data includes visible light and near-infrared images of the grain warehouse to be monitored, as well as text description information; the text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels.

[0096] Specifically, the execution entity of the grain storage pest status monitoring method provided in this application embodiment is a grain storage pest status monitoring device or system. This device can be implemented in software, such as a grain storage pest status monitoring program running on a server; or it can be implemented in hardware, such as a processor, mobile terminal, computer, or server that executes the grain storage pest status monitoring method.

[0097] The grain silo to be monitored is a warehouse used for storing grain crops. In one specific embodiment, the grain silo to be monitored is a circular silo used for storing common crops such as corn and wheat. The grain stored in the grain silo to be monitored may contain pests, and it is necessary to monitor the life status of these pests.

[0098] The multimodal data in this application embodiment includes, but is not limited to, visible light and near-infrared images of the grain silo to be monitored, as well as textual descriptive information. These data collectively constitute a comprehensive description of the grain silo environment and the state of pests within it.

[0099] Visible light images, also referred to as red-green-blue (RGB) images or color images in this embodiment, are primarily used to capture high-resolution visual details such as the morphology, color, and texture of pests. For example, visible light images can clearly show the surface gloss of pests, the spots and textures on their elytra, and the shape of their limbs, which are important bases for determining the species of pests and for preliminary observation of their activity.

[0100] Near-infrared (NIR) images are primarily used to reflect the biochemical characteristics of insect body surfaces, such as changes in moisture content and protein decomposition. Live and dead insects exhibit significant differences in reflectance and absorption characteristics under the NIR spectrum due to cessation of metabolism, water loss, and protein denaturation. Therefore, NIR images can provide intrinsic biological information directly related to the state of life that is not available in visible light images, and they are more robust to changes in light intensity.

[0101] The text description information is used to describe the storage environment and pest information of the grain silos under monitoring from multiple levels. Storage environment information describes the crop background information and real-time environmental data in the grain silos. Pest information includes information related to the species and biological characteristics of the pests.

[0102] Textual description information is a type of structured knowledge introduced in this application embodiment, used to provide semantic guidance and background constraints for the model's analysis process, and to associate semantic information with image content. This hierarchical textual information helps the model to make logical inferences and judgments from the macro environment to micro features, improving the flexibility and accuracy of detection. In this application embodiment, the multiple levels may specifically include scene level, instance level, and attribute level.

[0103] Scene-level text descriptions describe the macro-environmental background of the grain storage facility to be monitored. This information defines the overall context in which the monitoring task takes place. For example, scene-level text could be "a corn storage facility with a temperature of 28 degrees Celsius and a relative humidity of 75%." This information helps the model adaptively adjust its criteria for judging pest characteristics based on different environmental conditions (such as temperature, humidity, crop type, and whether it has undergone chemical treatment).

[0104] The instance-level text description information specifies the particular category of the target pest that needs to be monitored. For example, the instance-level text could be "corn weevil" or "grain borer." This level of information allows the model to focus on specific types of pests and utilize relevant knowledge for identification.

[0105] The attribute-level text description information is used to detail the specific biological or physical characteristics that the target pest should possess in different states (especially the living or dead states). For example, for the "dead" state of the corn weevil, the attribute-level text can be described as "the body surface loses its luster and darkens in color," "the antennae are stiff and extended," and "the body is shriveled and the abdomen is sunken," etc. This level of information provides a refined and verifiable basis for the model to make final state confirmation.

[0106] The acquisition of the aforementioned multimodal data can be achieved in various ways. In one specific embodiment, it can be collected through automated monitoring equipment deployed in the grain silo. For example, an integrated device including a visible light camera, a near-infrared camera, and environmental sensors such as temperature, humidity, and gas concentration can periodically scan the surface of the grain silo to collect images and environmental data. Textual description information can be pre-collected and organized from agricultural knowledge bases, pest control manuals, academic papers, and other sources to construct a structured knowledge base. In one specific embodiment, textual knowledge about common stored grain pests can be obtained based on multi-source data such as books, web scraping, and the grain silo's internal knowledge base, encompassing pest body length, habits, mortality status, and near-infrared observation range.

[0107] Step 120: Input the multimodal data into the pest status monitoring model. The pest status monitoring model will perform feature alignment and identification on the multimodal data and output the pest status monitoring results of the grain warehouse to be monitored.

[0108] The pest status monitoring model includes an image modality fusion module, a multi-level text encoding module, a cross-modal feature alignment module, and a feature classification module.

[0109] The image modality fusion module is used to extract features and fuse modalities from visible light and near-infrared images to obtain fused visual features; the multi-level text encoding module is used to encode text description information to obtain text features at each level; the cross-modal feature alignment module is used to perform cross-modal alignment between the fused visual features and the text features at each level to obtain aligned features at each level; and the feature classification module is used to identify the aligned features at each level to obtain pest status monitoring results.

[0110] Specifically, after acquiring the aforementioned multimodal data, it is fed as input to a pre-trained pest status monitoring model. After processing and analysis, the model outputs monitoring results on the current pest status within the grain warehouse. The pest status monitoring results may include, but are not limited to: pest species identification, their location in the image (e.g., represented by the coordinates of the target detection box (boundary box), and their life status determination (e.g., "surviving," "dead," "dying," etc.).

[0111] To achieve the above functions, the pest status monitoring model proposed in this application is structurally designed to include several core modules that work together. Figure 2 This is a schematic diagram of the structure of the pest status monitoring model provided in this application, as shown below. Figure 2As shown, the pest status monitoring model 200 may include an image modality fusion module 210, a multi-level text encoding module 220, a cross-modal feature alignment module 230, and a feature classification module 240. The image modality fusion module extracts and fuses features from visible light and near-infrared images to obtain fused visual features; the multi-level text encoding module encodes textual description information to obtain textual features at each level; the cross-modal feature alignment module aligns the fused visual features with the textual features at each level to obtain aligned features at each level; and the feature classification module identifies the aligned features at each level to obtain the pest status monitoring results.

[0112] The image modality fusion module processes input visible light and near-infrared images separately and effectively fuses the visual information from these two different modalities. This module first extracts features from both images to obtain their respective deep feature representations. Then, it combines these features using a specific fusion strategy to form a unified and more information-rich fused visual feature. The goal is to ensure that the fused feature combines the high-resolution morphological information of the visible light image with the biological state information of the near-infrared image, thus providing a higher-quality visual foundation for subsequent recognition tasks. This module can be implemented using various deep learning models such as Convolutional Neural Networks (CNNs) and Visual Transformers.

[0113] The multi-level text encoding module processes input text descriptions at multiple levels, converting human-readable natural language text into machine-processable numerical vector representations—i.e., text features at each level. Specifically, it generates corresponding scene text features, instance text features, and attribute text features for scene-level, instance-level, and attribute-level text descriptions, respectively. This module enables the model to understand the prior knowledge and semantic information contained in the text. This module can be implemented using pre-trained language models such as Recurrent Neural Networks (RNNs) and Bidirectional Encoder Representations from Transformers (BERT) models, or techniques such as Graph Neural Networks (GNNs).

[0114] The function of the cross-modal feature alignment module is to establish a semantic connection between visual information and textual knowledge. It receives fused visual features from the image modality fusion module and textual features from various levels of the multi-level text encoding module, and matches and interacts with them through an alignment mechanism. This process can be understood as using textual features as a "guide" or "query" to guide the model to focus on and understand relevant regions and patterns within the fused visual features. For example, scene textual features are used to calibrate the model's interpretation of overall visual features, instance textual features help the model locate a specific pest in an image, and attribute textual features guide the model to examine the pest's subtle features to determine its lifespan. After alignment, the module outputs aligned features at each level; these features are visual features that have been associated or calibrated with textual knowledge and contain richer semantic information. This module can be implemented using techniques such as attention mechanisms.

[0115] The function of the feature classification module is to make a final judgment and output the result based on the features generated by the preceding modules and after sufficient information fusion and alignment. It receives the final aligned features from the cross-modal feature alignment module and identifies the pest category and life status through one or more classifiers. Ultimately, the pest status monitoring result output by this module is the goal to be achieved by the method provided in this application embodiment.

[0116] The grain storage pest status monitoring method provided in this application organically combines multimodal data from different sources and of different types, inputs them into the constructed pest status monitoring model, and obtains pest status monitoring results, realizing end-to-end pest detection and pest status identification. By recognizing multimodal data, it overcomes the shortcomings of insufficient information from a single data source, significantly improving the monitoring accuracy and robustness of the model in complex and variable real grain storage environments. By aligning image features and text features at multiple levels, it improves the flexibility and accuracy of pest status monitoring. Ultimately, it achieves accurate identification of pests and their life status in grain storage.

[0117] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.

[0118] The related technologies have failed to effectively integrate the color and structural information of visible light and near-infrared images when processing them; they lack in-depth analysis of the correlation between the two; and they are insufficient in extracting pest features under complex conditions such as changes in lighting and environmental interference.

[0119] Therefore, this application proposes an innovative fusion scheme, which is specifically optimized for grain warehouse pest detection scenarios. By conducting in-depth research on the complementarity and correlation between visible light images and near-infrared images under different lighting and environmental conditions, the modal fusion strategy is further optimized to more effectively extract and utilize information from the two modal images, thereby improving the accuracy of pest detection and death status identification.

[0120] In some embodiments, the method further includes:

[0121] The image modality fusion module performs feature fusion on visible light and near-infrared images, and then enhances the fused image features based on a spatial attention mechanism to obtain fused visual features.

[0122] Specifically, considering the complexity and subtle differences in insect textures, this application embodiment first performs feature fusion on visible light images and near-infrared images, and then inputs the fused image features into a specific feature enhancement module (referred to as SS block in this application embodiment) to perform feature enhancement based on a spatial attention mechanism to obtain fused visual features.

[0123] In one specific embodiment, feature fusion of visible light and near-infrared images can be achieved by stitching together the feature maps obtained from the two images after passing through a preliminary feature extraction network (such as a convolutional neural network) along the channel dimension. In another specific embodiment, element-wise mathematical operations can be performed, such as element-wise addition, multiplication, or averaging, to achieve deeper levels of interaction.

[0124] Spatial attention mechanisms enable models to automatically and selectively focus on the most informative regions of image features while ignoring unimportant background areas. In one specific implementation, this mechanism generates a spatial weight mask for the fused image feature map. This mask has the same size as the spatial dimension of the feature map, and the weight value at each location represents the importance of that spatial location. Regions that may contain pests will have higher weight values, while regions identified as background will have lower weight values. Subsequently, this learned spatial weight mask is multiplied by the original fused image features to achieve feature enhancement. This process is equivalent to performing a weighting operation on the feature map: amplifying the feature responses of important regions (pests) while suppressing the feature responses of unimportant regions (background).

[0125] The grain storage pest status monitoring method provided in this application embodiment enhances the image features after fusing visible light and near-infrared images based on the spatial attention mechanism, which can better capture the minute texture features on the surface of pests and significantly improve the accuracy of pest identification.

[0126] In some embodiments, feature fusion is performed on visible light images and near-infrared images, and feature enhancement based on a spatial attention mechanism is applied to the fused image features to obtain fused visual features, including:

[0127] Feature extraction is performed on the visible light image and the near-infrared image respectively to obtain the first visible light image feature and the first near-infrared image feature;

[0128] The first fused image features are enhanced by a spatial attention mechanism to obtain the first enhanced image features; the first fused image features are obtained by matrix point addition of the first visible light image features and the first near-infrared image features.

[0129] The first enhanced image feature is matrix-added with the first visible light image feature and the first near-infrared image feature respectively to obtain the first visible light fused image feature and the first near-infrared fused image feature.

[0130] The first visible light fused image features and the first near-infrared fused image features are stitched together and feature enhancement based on a spatial attention mechanism are performed to obtain the second enhanced image features;

[0131] The second enhanced image feature is matrix-added with the first visible light fused image feature and the first near-infrared fused image feature respectively to obtain the second visible light fused image feature and the second near-infrared fused image feature.

[0132] The fused visual features are obtained by matrix point addition of the second visible light fused image features and the second near-infrared fused image features.

[0133] Specifically, Figure 3 This is a schematic diagram of the image modality fusion module provided in this application, as shown below. Figure 3 As shown, the image modality fusion module includes Convolutional blocks ( conv) Convolutional blocks ( The modules include conv, concat (feature connection module), and SS block (feature enhancement module).

[0134] Features of the initial near-infrared image can be extracted in advance using near-infrared visual coding networks and visible light visual coding networks, respectively. Features of the initial visible light image The images are then fused in the image modality fusion module to generate fused visual features. .

[0135] Near-infrared visual coding networks can employ the Shifted Window Transformer (Swin Transformer) model. Visible light visual coding networks can employ the Grounding DINO encoder.

[0136] Based on the acquired initial visible light image features Features of the initial near-infrared image You can enter them separately. Preprocessing is performed in the convolutional block to adjust the number of channels and extract features, resulting in the first visible light image features and the first near-infrared image features.

[0137] Regarding global texture and color, a first fused image feature can be obtained by matrix-adding the first visible light image feature and the first near-infrared image feature. This fused image feature is then input into a feature enhancement module (SS block) for feature enhancement based on a spatial attention mechanism, resulting in a first enhanced image feature. This enhanced image feature is then matrix-added with both the first visible light image feature and the first near-infrared image feature to obtain the first visible light fused image feature and the first near-infrared fused image feature. Through this process, the resulting image features fuse global information and texture and color feature information.

[0138] In terms of local structure, the first visible light fused image features and the first near-infrared fused image features can be stitched together using a feature concat module, and then input sequentially. Convolutional blocks A convolutional block and a feature enhancement module (SS block) are used to obtain the second enhanced image feature. This second enhanced image feature is then matrix-added with the first visible light fused image feature and the first near-infrared fused image feature, respectively, to obtain the second visible light fused image feature and the second near-infrared fused image feature. Finally, the second visible light fused image feature and the second near-infrared fused image feature are matrix-added to obtain the fused visual feature.

[0139] In grain storage environments with varying lighting conditions (such as shadows and strong light), near-infrared images exhibit stronger resistance to lighting interference compared to visible light images. By employing matrix point addition operations, features from visible light and near-infrared images are deeply fused, fully leveraging the advantages of near-infrared images under low light or shadow conditions while preserving the rich color details of visible light images. This fusion method not only enhances the model's robustness to lighting changes but also improves its ability to identify pest features.

[0140] This application's embodiments utilize a specially designed fusion network structure to deeply analyze the complementarity of visible light images and near-infrared images under different environmental conditions. For example, regarding the reflectivity of insect body surfaces, visible light images excel at capturing color differences, while near-infrared images better reflect the texture and structural features of insect bodies. Convolutional blocks can include sequentially connected Depth-separable convolutional blocks ( The DWconv module, Layer Normalization (LN) module, and Parametric Rectified Linear Unit (PReLU) activation function module. Depthwise separable convolutional blocks are used for local feature extraction; layer normalization modules stabilize the training process and prevent gradient explosion or vanishing; activation function modules provide non-linear activation, increasing the model's expressive power. Depthwise separable convolutional blocks and layer normalization modules further optimize the feature extraction process, resulting in more compact and discriminative feature vectors. After insect pests die, their surface texture changes significantly, exhibiting characteristics such as dull color and loose structure. Image fusion methods in related technologies may not accurately distinguish between live and dead pests. By introducing an activation function module, the network's sensitivity to low-contrast features is enhanced, enabling better identification of subtle texture changes in dead insects.

[0141] The grain storage pest status monitoring method provided in this application embodiment can achieve in-depth mining and fusion of visible light image and near-infrared image information through multiple feature fusion and feature enhancement. It considers both the fusion of global information and the preservation and enhancement of local details, so that the final output fused visual features can accurately represent multiple attributes of pests such as color, shape, texture and biological state, providing extremely high-quality feature input for subsequent accurate identification tasks.

[0142] In some embodiments, feature enhancement based on a spatial attention mechanism is performed on the fused image features, including:

[0143] The feature enhancement module performs feature enhancement on the fused image features based on a spatial attention mechanism.

[0144] The feature enhancement module includes a channel attention submodule, a feature map splitting submodule, and a normalization submodule connected in sequence.

[0145] The channel attention submodule is used to process image features based on the channel attention mechanism and adjust the weights of feature channels;

[0146] The feature map splitting submodule is used to segment the image features after feature channel weight adjustment into multiple sub-feature maps;

[0147] The normalization submodule is used to normalize each sub-feature map.

[0148] Specifically, the feature enhancement module (SS block) in this embodiment incorporates a spatial attention mechanism, which can dynamically adjust the importance weights of visible light images and near-infrared images in different regions to ensure that key features are effectively preserved.

[0149] The feature enhancement module consists of a channel attention submodule (SE-Net), a feature map splitting submodule (Split), and a normalization submodule (Softmax), which are connected in sequence.

[0150] The channel attention submodule processes image features according to the channel attention mechanism, adaptively adjusting channel weights to enhance important information features. The feature map splitting submodule divides the feature map output by the channel attention submodule into multiple sub-feature maps. The normalization submodule normalizes the segmented sub-feature maps to ensure a reasonable weight distribution among them.

[0151] The grain storage pest status monitoring method provided in this application focuses on important features through channel attention, and then performs feature map splitting and normalization, enabling the model to dynamically weight the feature map with very fine granularity, thereby greatly improving the model's ability to extract and enhance the features of key targets (such as small pests) from complex backgrounds.

[0152] Existing pest status identification methods in related technologies often rely on image or sensor data, lacking a deep understanding of environmental context and biological knowledge, leading to misjudgments in complex scenarios. Furthermore, text feature extraction methods in related technologies are typically based on static knowledge bases or rule matching, making it difficult to handle newly emerging pest species, abnormal mortality patterns, or dynamically changing storage environments. Therefore, this application proposes a multi-level text encoding network for identifying the mortality status of stored grain pests, aiming to solve core problems in related technologies such as shallow semantic modeling and static knowledge updates. By constructing a dynamically evolving lightweight knowledge graph as a semantic skeleton and designing a graph neural network architecture that integrates hierarchical attention and dynamic feature modulation mechanisms, deep modeling and structured decoupling output of three levels of semantic information—"scene—instance—attribute" are achieved.

[0153] In some embodiments, the scene-level text description information includes crop background information and real-time environmental data in the grain warehouse to be monitored; the real-time environmental data includes at least one of real-time temperature information, real-time humidity information, and real-time gas monitoring data.

[0154] The text description information at the instance level includes information about the pest species;

[0155] The textual description information at the attribute level includes information on the biological characteristics of the pests.

[0156] Specifically, to support three-level semantic modeling of "scene-instance-attribute," this application's embodiments design a structured and scalable multi-source text data system, which includes: scene-level text description information (describing various grain storage environment information); instance-level text description information (listing various pest species that may appear in this environment); and attribute-level text description information (encompassing the biological characteristics of pests). This detailed data structure not only helps to accurately identify and distinguish the survival status of pests but also improves the accuracy and reliability of pest monitoring in grain warehouses.

[0157] Crop background information is used to clarify the context for pest monitoring. Different types of grains (such as corn, wheat, rice, and soybeans) vary significantly in color, shape, size, and texture, and these differences directly affect the visibility of pests in images. For example, finding a brownish-yellow corn weevil among yellow corn kernels is much more difficult than finding it among white flour. By providing the model with textual information such as "crop background: corn," the model can pre-adjust its visual recognition strategy, thereby more effectively locating pests in specific contexts.

[0158] Real-time environmental data describes the microclimate conditions inside the grain silo, which directly affect the physiological activities and life state of pests. Real-time environmental data can include real-time temperature information, real-time humidity information, and real-time gas monitoring data.

[0159] Real-time temperature information refers to the actual temperature inside the grain silo, such as "Temperature: 10 degrees Celsius." Temperature is a key factor affecting the metabolic rate of pests. For example, in low-temperature environments, many pests may enter a state of inactivity, feigning death, or diapause, and their visual characteristics (such as immobility) are very similar to those of actual death. Real-time temperature information from the grain silo can be obtained and input into a model for processing. Once the model recognizes the current low-temperature environment, it will be more accurate in assessing the state of the pests.

[0160] Real-time humidity information refers to the real-time humidity inside the grain warehouse, such as "relative humidity: 65%". Humidity also affects the survival and activity of pests. Both excessively high and low humidity can cause stress to pests.

[0161] Real-time gas monitoring data refers to monitoring information such as the concentration of various gases in the grain warehouse, for example, "phosphine concentration: 1 / 50,000,000 (ppm)". The composition of gases, especially the concentration of fumigants (such as phosphine), is direct evidence to determine whether pests have died due to chemical treatment. At the same time, the respiration of pests causes local changes in carbon dioxide concentration, which can also serve as an indirect indicator of their life activities.

[0162] The instance-level text description information defines the specific target category that the model needs to identify, which can include pest species information. This information provides the model with an identity label for the pest to be identified. For example, "Pest Name: Corn Weevil". It can also contain richer biological classification information, such as "Classification: Coleoptera, Weevil Family". This transforms the model's identification task from an open-ended "identify pests" problem into a more targeted "identify corn weevil" problem, helping to improve the accuracy and specificity of the identification.

[0163] The attribute-level text description information is used to describe the fine-grained biological characteristics exhibited by instances (pests) in different states, and may include pest biological characteristic information. Pest biological characteristic information is the microscopic basis for distinguishing different pest species and judging their life state, such as body length, color, and living characteristics.

[0164] The grain storage pest status monitoring method provided in this application improves the flexibility and accuracy of pest status monitoring by structuring text description information into three logical levels: scene, instance, and attribute. This allows the model to utilize the semantic information of the text description and the correlation between image content.

[0165] In some embodiments, the method further includes:

[0166] Semantic parsing is performed on text description information at multiple levels to obtain triples corresponding to each level;

[0167] A dynamic knowledge graph is constructed based on the triples corresponding to each level.

[0168] Specifically, semantic parsing can be performed on text description information at multiple levels, including named entity recognition and relation extraction, to convert text description information at each level into structured triples (head entity, relation, tail entity) corresponding to each level, such as (maize weevil, has, body length 2–3 mm), (high temperature, affects, mobility), (non-motor + stiffness, manifested as, death), etc.

[0169] The initial knowledge graph is obtained by combining the triples corresponding to each level. Subsequent feedback text from administrators (such as "Found a black insect corpse, grayish-brown in appearance, motionless") can be introduced as a dynamic data source to determine if it is a new entity or relation. If the entity or relation does not exist, an embedding is generated through a zero-shot or few-shot learning mechanism; if it is a combination of known entities, a new inference path is established, and the new triples are incrementally inserted into the initial knowledge graph, triggering a graph structure update to obtain a dynamic knowledge graph. A lightweight fine-tuning method is selected each time parameters are updated.

[0170] The grain storage pest status monitoring method provided in this application introduces semantic parsing and dynamic knowledge graph construction, realizing the transformation process from raw text to structured knowledge graph. This enables subsequent multi-level text encoding modules to efficiently encode and reason about this knowledge. Moreover, through a dynamic update mechanism, the entire system is endowed with the ability to continuously learn and evolve, improving the flexibility and accuracy of pest status monitoring.

[0171] Based on the dynamic knowledge graph constructed above, this application embodiment designs a multi-level text encoding network to learn text information features at different levels, and realizes the extraction of structured "scene-instance-attribute" three-level semantic features from the dynamic knowledge graph to adapt to text-image alignment at different scales.

[0172] In some embodiments, the method further includes:

[0173] By using a multi-level text encoding module, entity nodes and relation edges in the dynamic knowledge graph corresponding to the text description information are vectorized and layered to obtain scene subgraph, instance subgraph and attribute subgraph;

[0174] Feature extraction is performed on the scene sub-image to obtain scene text features, and the scene text features are processed to obtain scene scaling factor and scene offset factor;

[0175] Feature extraction is performed on the attribute subgraph to obtain initial attribute features, and dynamic feature modulation is performed on the initial attribute features based on the scene scaling factor and scene offset factor to obtain attribute text features;

[0176] Feature extraction is performed on the instance subgraph to obtain initial instance features. Then, the initial instance features and initial attribute features are processed based on a hierarchical attention mechanism to obtain instance text features.

[0177] Specifically, Figure 4 This is a structural diagram of the multi-level text encoding module provided in this application, as shown below. Figure 4 As shown, the multi-level text encoding module may include a subgraph acquisition module, a scene text feature extraction module, an instance text feature extraction module, and an attribute text feature extraction module.

[0178] The subgraph acquisition module can be implemented using a pre-trained language model (such as BERT). Entity nodes (e.g., "corn weevil," "body zombie") and relation edges (e.g., "have," "behave as") in the dynamic knowledge graph are encoded into low-dimensional semantic vectors by the subgraph acquisition module, generating initial embedding vectors. For newly added dynamic entities (e.g., "black insect corpse"), semantic similarity matching or zero-shot generation are used to map them to the same vector space, ensuring representation consistency during knowledge graph expansion.

[0179] The dynamic knowledge graph contains three types of entity nodes: scene nodes (S), instance nodes (I), and attribute nodes (A). The edge types in the dynamic knowledge graph include S→I (e.g., "corn barn → exists → corn weevil," representing the collinear relationship between the environment and the pest), I→A (e.g., "corn weevil → has → stiffness," representing the attribution relationship between the pest and its characteristics), and A→A (e.g., "stiffness + no movement → inference → death," representing the logical combination relationship between death features). The subgraph acquisition module hierarchically stratifies these vectorized representations of nodes and edges, obtaining scene subgraphs, instance subgraphs, and attribute subgraphs.

[0180] The scene text feature extraction module can be implemented using a graph neural network to extract features from scene subgraphs and obtain scene text features. .

[0181] The attribute text feature extraction module can be implemented using a graph neural network to extract features from the attribute subgraph and obtain initial attribute features. .

[0182] In real-world grain storage management scenarios, the same morphological or behavioral characteristics of pests can have drastically different biological meanings under different environmental conditions. For example, in low-temperature environments, the stagnation of pest activity may be a sign of "feigned death" rather than actual death.

[0183] To enable attribute text features to adapt to the environment, this application proposes a dynamic feature modulation mechanism based on scene context. Specifically, two parallel multi-layer perceptrons (MLPs) can be added to the scene text feature extraction module to process the scene text features and obtain scene scaling coefficients. and scene offset coefficient .

[0184] The attribute text feature extraction module performs dynamic feature modulation on the initial attribute features based on these two coefficients to obtain the attribute text features. This can be expressed as a formula:

[0185] .

[0186] in, This indicates the XOR operation.

[0187] To achieve accurate identification of pest survival status, it is necessary to effectively aggregate multidimensional features from the attribute level at the instance level. However, the contribution of different attributes to status judgment varies significantly under specific environmental and biological contexts. For example, "rigor mortis" in a low-temperature storage environment may only be a manifestation of metabolic inhibition, and its weight as a mortality criterion should be appropriately reduced; while under normal or high-temperature conditions, "rigor mortis" accompanied by "lack of movement" more strongly points to actual mortality and should be given higher attention.

[0188] To enable the model to have dynamic semantic focusing capabilities, this application's embodiments design a hierarchical attention mechanism to adaptively allocate weights when aggregating attribute features at instance nodes. Specifically, the instance text feature extraction module extracts features from the instance subgraph to obtain initial instance features. And based on the hierarchical attention mechanism, the initial instance features are... and initial attribute features Processing is performed to obtain instance text features. .

[0189] Given instance nodes in a dynamic knowledge graph and its associated set of attribute nodes , This represents the number of attribute nodes.

[0190] It can be based on the initial instance characteristics of the node. The attention network is used to calculate and initialize attribute features. Similarity score of attribute features before dynamic feature modulation :

[0191] .

[0192] in, and For learnable parameter matrix, For attention context vectors, This is the transpose operator. The label of the attribute node.

[0193] The attention weights are then obtained by normalization using the Softmax function. :

[0194] .

[0195] Ultimately, the comprehensive characteristics of instance nodes Obtained by weighted summation, where The attribute text features modulated by the scene context:

[0196] .

[0197] The grain storage pest status monitoring method provided in this application, through knowledge graph layering, graph neural network encoding, dynamic feature modulation, and hierarchical attention mechanisms, enables the model to dynamically identify key criteria based on the current context and suppress interfering features, thereby improving the robustness of discrimination in complex scenarios. The model not only extracts semantic information from each level but also cleverly captures and utilizes complex dependencies between different levels (such as the influence of the scene on attributes and the definition of attributes on instances), thus providing extremely rich, accurate, and highly context-aware text features for subsequent cross-modal alignment. The model achieves a shift from "passively receiving features" to "actively selecting evidence," significantly enhancing the fine-grained understanding and interpretable reasoning ability of pest mortality states.

[0198] By aligning visual features with textual descriptions, the model can not only identify elements in an image but also understand the underlying context and semantic relationships. Effectively fusing multi-scale image data with multi-level textual information can significantly improve the model's understanding and accuracy when handling different image-text alignment tasks. Therefore, this application proposes a global alignment module (scene environment adaptation), a regional alignment module (instance localization), and a local alignment module (state category verification) based on global, regional, and local image-text features. This achieves cross-modal image-text feature alignment and improves the accuracy of pest mortality identification in complex scenarios.

[0199] In some embodiments, the method further includes:

[0200] The cross-modal feature alignment module globally aligns the fused visual features with the scene-level text features to obtain scene-aligned visual features.

[0201] Multi-scale feature extraction is performed on the scene alignment visual features to obtain visual multi-scale features;

[0202] By aligning the visual multi-scale features with the instance-level instance text features, we obtain instance-aligned multi-scale features.

[0203] Region detection is performed on instance-aligned multi-scale features to obtain the target detection box and the image region features corresponding to the target detection box;

[0204] The image region features corresponding to the target detection box are locally aligned with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features.

[0205] Specifically, Figure 5 This is a schematic diagram of the cross-modal feature alignment module provided in this application, as shown below. Figure 5As shown, the cross-modal feature alignment module can include a global alignment module, a multi-scale feature extraction module, a region alignment module, a region detection module, and a local alignment module.

[0206] The global alignment module performs image and text feature alignment at the scene level, which is used to globally align the fused visual features with the scene text features at the scene level to obtain scene-aligned visual features.

[0207] The multi-scale feature extraction module can be implemented using a Feature Pyramid Network (FPN) to extract multi-scale features from scene-aligned visual features, thereby obtaining visual multi-scale features.

[0208] The region alignment module performs image-text feature alignment at the instance level, which is used to align visual multi-scale features with instance text features at the instance level to obtain instance-aligned multi-scale features.

[0209] The region detection module can be implemented using a region detection model (such as the prediction head in the RetinaNet object detection network) to perform region detection on instance-aligned multi-scale features, thereby obtaining the object detection box (bbox) and the corresponding image region of interest (ROI) features.

[0210] The local alignment module performs image-text feature alignment at the attribute level, which is used to locally align the image region features corresponding to the target detection box with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features.

[0211] The grain storage pest status monitoring method provided in this application greatly improves the accuracy of pest status monitoring and enhances the interpretability of the entire model through a three-level progressive cross-modal image and text feature alignment of global alignment (scene adaptation), regional alignment (instance location), and local alignment (attribute verification).

[0212] In some embodiments, the fused visual features are globally aligned with scene text features at the scene level to obtain scene-aligned visual features, including:

[0213] The fused visual features are processed using a global attention mechanism to obtain a scene query vector;

[0214] The scene text features are processed based on the ordinary attention mechanism to obtain scene key vectors and scene value vectors;

[0215] The scene query vector, scene key vector, and scene value vector are processed based on the cross-attention mechanism to obtain global alignment features;

[0216] The scene alignment visual features are obtained by processing the global alignment features using a feedforward neural network.

[0217] Specifically, Figure 6 This is a schematic diagram of the structure of the global alignment module provided in this application, as shown below. Figure 6 As shown, the global alignment module can include a global attention module, a regular attention module, a cross attention module, and a feedforward neural network module.

[0218] fusion of visual features The input is a global attention module, which focuses on global image information features. Based on the global attention mechanism, the fused visual features are processed to obtain the scene query vector (Query, Q). The global attention mechanism can fully capture global contextual information.

[0219] Scene text features The input is taken into a standard attention module, which focuses on the internal information features of the text. Based on the standard attention mechanism, the scene text features are processed to obtain the scene key vector (Key, K) and scene value vector (Value, V). The standard attention mechanism is the most basic attention mechanism, capable of dynamically focusing on different parts of the input.

[0220] The scene query vector, scene key vector, and scene value vector are input into the cross-attention module. The cross-attention module processes the scene query vector, scene key vector, and scene value vector based on the cross-attention mechanism to obtain global alignment features.

[0221] The global alignment features are input into the Feed Forward Network (FFN) module to obtain the output scene alignment visual features.

[0222] The grain storage pest status monitoring method provided in this application can better understand the content of the entire scene by outputting scene-aligned visual features, thus improving the accuracy of pest mortality identification in complex scenes.

[0223] In some embodiments, the visual multi-scale features are region-aligned with instance-level instance text features to obtain instance-aligned multi-scale features, including:

[0224] Visual multi-scale features are processed based on a deformable attention mechanism to obtain instance query vectors;

[0225] The instance text features are processed using a common attention mechanism to obtain instance key vectors and instance value vectors;

[0226] The instance query vector, instance key vector, and instance value vector are processed based on the cross-attention mechanism to obtain region alignment features;

[0227] The region alignment features are processed using a feedforward neural network to obtain instance alignment multi-scale features.

[0228] Specifically, Figure 7 This is a schematic diagram of the structure of the region alignment module provided in this application, as shown below. Figure 7 As shown, the region alignment module includes a deformable attention module, a regular attention module, a cross attention module, and a feedforward neural network module.

[0229] Visual multiscale features The input is a deformable attention module, which processes visual multi-scale features based on the deformable attention mechanism to obtain the instance query vector. The deformable attention mechanism no longer focuses on all spatial locations, but instead predicts a small number of learnable offset coordinates for each query point, focusing only on features around these offset points. This significantly reduces computational cost and allows the model to more flexibly focus on irregular regions that may contain important information.

[0230] Instance text features The input is a standard attention module, which focuses on the internal information features of the text. Based on the standard attention mechanism, the instance text features are processed to obtain instance key vectors and instance value vectors.

[0231] The instance query vector, instance key vector, and instance value vector are input into the cross-attention module. The cross-attention module processes the instance query vector, instance key vector, and instance value vector based on the cross-attention mechanism to obtain the region alignment features.

[0232] The region alignment features are input into the feedforward neural network module to obtain the output instance alignment multi-scale features.

[0233] The grain storage pest status monitoring method provided in this application can better identify specific objects in the image by outputting instance-aligned multi-scale features, thus improving the accuracy of pest mortality identification in complex scenarios.

[0234] In some embodiments, the image region features corresponding to the target detection box are locally aligned with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features, including:

[0235] The image region features corresponding to the target detection box are processed based on the multi-head attention mechanism to obtain the attribute query vector;

[0236] The attribute text features are processed based on the ordinary attention mechanism to obtain attribute key vectors and attribute value vectors;

[0237] The attribute query vector, attribute key vector, and attribute value vector are processed based on the cross-attention mechanism to obtain local alignment features;

[0238] By processing local alignment features using a feedforward neural network, attribute alignment multi-scale features are obtained.

[0239] Specifically, Figure 8 This is a schematic diagram of the structure of the local alignment module provided in this application, as shown below. Figure 8 As shown, the local alignment module includes a multi-head attention module, a regular attention module, a cross-attention module, and a feedforward neural network module.

[0240] Image region features The input is processed by a multi-head attention module, which uses a multi-head attention mechanism to process image region features and obtain an attribute query vector. The multi-head attention mechanism linearly projects the query (Q), key (K), and value (V) onto multiple different subspaces (i.e., multiple "heads"), allowing each "head" to independently learn to focus on different aspects of the input information.

[0241] attribute text features The input is a standard attention module, which focuses on the internal information features of the text. Based on the standard attention mechanism, the attribute text features are processed to obtain attribute key vectors and attribute value vectors.

[0242] The attribute query vector, attribute key vector, and attribute value vector are input into the cross-attention module. The cross-attention module processes the attribute query vector, attribute key vector, and attribute value vector based on the cross-attention mechanism to obtain local alignment features.

[0243] The local alignment features are input into the feedforward neural network module to obtain the output attribute-aligned multi-scale features.

[0244] The grain storage pest status monitoring method provided in this application can better understand the specific attributes of objects in the image and identify the status category by finally outputting attribute-aligned multi-scale features, thereby improving the accuracy of pest mortality identification in complex scenarios.

[0245] In some embodiments, the training loss of the pest status monitoring model is determined based on at least one of pest classification loss, pest detection loss, pest detection box matching loss, and model knowledge preservation loss.

[0246] Among them, the pest classification loss is determined based on the difference between the predicted value and the actual value of the pest status monitoring results corresponding to the multimodal data of the sample;

[0247] The pest detection loss is determined based on the difference between the predicted and actual locations of the target detection boxes corresponding to the pests in the multimodal data of the samples.

[0248] The pest detection box matching loss is determined based on the matching degree between the predicted location and the actual location of the target detection box corresponding to the multimodal data of the sample.

[0249] The model knowledge retention loss is determined based on the information divergence between the pest state monitoring model trained in the current iteration and the pest state monitoring model trained in the previous iteration.

[0250] Specifically, pest classification losses To measure the difference between predicted and true values ​​(actual labels), cross-entropy loss is used, which can be expressed as:

[0251] .

[0252] In the formula, The number of multimodal data points in the sample; The true values ​​of pest status monitoring results corresponding to the multimodal data of the samples; These are the predicted values ​​of pest status monitoring results corresponding to the multimodal data of the samples.

[0253] Image annotation of pests in multimodal data samples includes bounding boxes. Pest categories, bounding boxes, and mortality status can be annotated using methods such as manual annotation and large-scale model pre-inference.

[0254] Pest detection losses The difference between the predicted location and the actual location of the target detection box (bounding box) containing the pest can be represented as:

[0255] .

[0256] In the formula, IOU is the standard intersection-union ratio; It is the Euclidean distance between the center points of the two object detection boxes; It is the minimum value of the diagonal length of the two object detection boxes, used for standardization. ; and It is used to adjust the balancing weights of the aspect ratio terms; It is the difference in aspect ratio of the target detection bounding box; and These are the predicted width of the target detection box and the actual width of the target detection box, respectively.

[0257] Pest detection box matching loss The formula used to measure the matching degree between the predicted and actual locations of object detection boxes is:

[0258] .

[0259] In the formula, This represents the optimal match obtained using the Hungarian algorithm. and These are the target detection boxes in the optimal match.

[0260] Model knowledge retention loss The formula used to measure the degree of knowledge retention of a pest status monitoring model during iterative training can be expressed as follows:

[0261] .

[0262] In the formula, This is the pest status monitoring model trained in the previous iteration; This is the pest status monitoring model trained in the current iteration; This is known as information divergence (Kullback-Leibler Divergence), also called KL divergence.

[0263] Finally, the training loss of the pest status monitoring model can be expressed as:

[0264] .

[0265] in, , , , The calculation weights are the corresponding values ​​for each type of loss.

[0266] The grain storage pest status monitoring method provided in this application constructs the training loss of the pest status monitoring model through pest classification loss, pest detection loss, pest detection box matching loss and model knowledge preservation loss, thereby improving the high accuracy and robustness of the model in pest status monitoring in complex scenarios.

[0267] The apparatus provided in the embodiments of this application is described below. The apparatus described below can be referred to in correspondence with the method described above.

[0268] Figure 9 This is a schematic diagram of the grain warehouse pest status monitoring device provided in this application, as shown below. Figure 9 As shown, the device includes:

[0269] The data acquisition module 910 is used to acquire multimodal data, including visible light and near-infrared images of the grain warehouse to be monitored, as well as text description information. The text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels.

[0270] The status monitoring module 920 is used to input multimodal data into the pest status monitoring model, which then performs feature alignment and identification on the multimodal data and outputs the pest status monitoring results for the grain warehouse to be monitored.

[0271] The grain storage pest status monitoring device provided in this application organically combines multimodal data from different sources and of different types, inputs them into the constructed pest status monitoring model, and obtains pest status monitoring results, realizing end-to-end pest detection and pest status identification. Through multimodal data identification, it overcomes the shortcomings of insufficient information from a single data source, significantly improving the monitoring accuracy and robustness of the model in complex and variable real grain storage environments. By aligning image features and text features at multiple levels, it improves the flexibility and accuracy of pest status monitoring. Ultimately, it achieves accurate identification of pests and their life status in grain storage.

[0272] Figure 10 This is a schematic diagram of the grain storage pest status monitoring system provided in this application, as shown below. Figure 10 As shown, the grain storage pest status monitoring system 1000 includes a data acquisition device 1010, a data transmission device 1020, and a data processing device 1030.

[0273] The data acquisition device is installed in the grain warehouse to be monitored to acquire visible light images, near-infrared images and real-time environmental data of the grain warehouse;

[0274] The data transmission device is communicatively connected to the data acquisition device and is used to send visible light images, near-infrared images and real-time environmental data to the data processing device.

[0275] The data processing device is communicatively connected to the data transmission device and is used to execute the grain warehouse pest status monitoring method in the above embodiments. It processes visible light images, near-infrared images and real-time environmental data to obtain the pest status monitoring results of the grain warehouse to be monitored.

[0276] Specifically, the data transmission device can be a 5G mobile communication base station. Given that grain silos are generally enclosed spaces, a stable data transmission device is needed to ensure that image data, sensor data, and location data can be transmitted to the processing center in real time. A 5G communication base station can provide high-speed data transmission.

[0277] The data processing device can be a cloud server. This server includes services for detecting stored grain pests and identifying pest mortality status. It is responsible for receiving all data transmitted from 5G communication base stations and automatically detecting pests and identifying their mortality status.

[0278] The grain storage pest status monitoring system provided in this application organically combines multimodal data from different sources and of different types, inputs them into the constructed pest status monitoring model, and obtains pest status monitoring results, realizing end-to-end pest detection and pest status identification. Through multimodal data identification, it overcomes the shortcomings of insufficient information from a single data source, significantly improving the monitoring accuracy and robustness of the model in complex and variable real grain storage environments. By aligning image features and text features at multiple levels, it improves the flexibility and accuracy of pest status monitoring. Ultimately, it achieves accurate identification of pests and their life status in grain storage.

[0279] In some embodiments, the data acquisition device includes a circular track deployed on the top of the grain silo to be monitored, and a data acquisition unit that slides along the circular track;

[0280] The data acquisition unit is equipped with an image sensor, a temperature sensor, a humidity sensor, and a gas sensor;

[0281] Image sensors are used to acquire visible light and near-infrared images of the surface of the grain silo to be monitored;

[0282] Temperature sensors are used to collect real-time temperature information in the grain silos to be monitored;

[0283] Humidity sensors are used to collect real-time humidity information in the grain silos to be monitored;

[0284] Gas sensors are used to collect real-time gas monitoring data in the grain silos to be monitored.

[0285] Specifically, Figure 11 This is an overall architecture diagram of the grain storage pest status monitoring system provided in this application, as shown below. Figure 11 As shown, the grain silo to be monitored can be a circular silo. The data acquisition device includes a circular track deployed on the top of the grain silo to be monitored, and a data acquisition unit that slides along the circular track.

[0286] A vertical support is mounted on the top of the circular grain silo, with at least one annular, slidable, fixed mounting track (circular track) connected to the end of the support. The data acquisition unit is set on the circular track and can move inside the grain silo to ensure that grain in all areas can be photographed and sensed.

[0287] The data acquisition unit is equipped with an image sensor, a temperature sensor, a humidity sensor, and a gas sensor.

[0288] Image sensors are used to acquire visible light and near-infrared images of the surface of the grain silo to be monitored.

[0289] Temperature sensors are used to collect real-time temperature information in the grain silos under monitoring. Humidity sensors are used to collect real-time humidity information in the grain silos under monitoring. Changes in temperature and humidity affect the metabolism and survival rate of pests.

[0290] Gas sensors are used to collect real-time gas monitoring data in the grain silos being monitored. Specific gases released during pest activity or decomposition can help determine whether pests are dead.

[0291] The entire system operation process is as follows:

[0292] Step 1: The data acquisition unit, set up on the circular track, starts from the starting position to capture visible light and near-infrared images of the shallow grain surface and collects sensor data.

[0293] Step 2: The data acquisition unit moves from the starting point to the ending position by means of preset stop points or time control, and takes pictures and collects sensor data at each stop point during the movement, while storing the position data to ensure that the data collection of the corresponding area is complete.

[0294] Step 3: The data acquisition devices on other tracks inside the grain warehouse repeat steps 1 and 2 until all data inside the grain warehouse has been collected.

[0295] Step 4: During the data collection process, the data is uploaded to the data processing device (cloud server) via the data transmission device (5G communication base station).

[0296] Step 5: The cloud server calls the deployed grain pest identification service to identify the uploaded data frame by frame in order to detect the types and mortality status of grain pests in the area.

[0297] The grain storage pest status monitoring system provided in this application has the following technical effects:

[0298] (1) A multimodal grain storage pest mortality status monitoring network model is provided for wide-area scenarios to realize end-to-end pest detection and pest mortality status identification. By integrating multimodal data, hierarchical query and other technologies, it solves the problems of difficulty in identifying grain storage pest mortality status, lack of support for identification in complex scenarios, and lack of support for identification of unknown pest categories (open set detection).

[0299] (2) A method for monitoring pests on a circular grain storage track is provided, which can be used for real-time image monitoring of pests in the entire grain storage area. The low-altitude track setting method is not only cost-controllable and high-definition monitoring effect, but also non-contact and low-noise monitoring, which solves the problems of high monitoring cost, poor monitoring effect and poor real-time performance of existing monitoring. At the same time, it is equipped with multi-source data monitoring to solve the problem of difficulty in real-time high-definition non-contact pest monitoring in circular wide-area grain storage.

[0300] (3) A multi-level text encoding network based on dynamic knowledge graph is provided. By constructing a lightweight knowledge graph that can evolve dynamically as a semantic skeleton, and designing a graph neural network architecture that integrates hierarchical attention and dynamic feature modulation mechanism, the deep modeling and structural decoupling output of the three-level semantic information of "scene-instance-attribute" is realized, which helps to realize image-text alignment in complex scenes (multiple crops, environmental backgrounds) and solve the problem of accurate identification of pest deaths in complex scenes.

[0301] (4) A three-level cross-modal alignment architecture is provided, including a global alignment architecture (scene environment adaptation), a regional alignment architecture (instance localization), and a local alignment architecture (state category verification), to achieve cross-modal dynamic routing. This architecture realizes a refined detection process from scene understanding → instance localization → attribute verification through progressive three-level cross-modal interaction, meeting the multi-granularity requirements of warehouse pest mortality detection, while adapting to pest detection in different complex background environments, and solving the problem of difficult accurate identification of pest mortality in complex scenarios.

[0302] (5) A fusion architecture for visible light images and near-infrared images is provided, which performs color and structure fusion for the two types of imaging, deeply explores the correlation between the two, and solves the problem of inaccurate identification of pest death in multi-source fusion.

[0303] Figure 12 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 12 As shown, the electronic device may include: a processor 1210, a communications interface 1220, a memory 1230, and a communications bus 1240, wherein the processor 1210, the communications interface 1220, and the memory 1230 communicate with each other via the communications bus 1240. The processor 1210 can call logical commands in the memory 1230 to execute the methods described in the above embodiments, for example:

[0304] Acquire multimodal data; the multimodal data includes visible light and near-infrared images of the grain warehouse to be monitored, as well as text description information; the text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels; input the multimodal data into the pest status monitoring model, which performs feature alignment and recognition on the multimodal data and outputs the pest status monitoring results of the grain warehouse to be monitored.

[0305] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0306] The processor in the electronic device provided in this application embodiment can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effect, which will not be repeated here.

[0307] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.

[0308] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.

[0309] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0310] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0311] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0312] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method of monitoring the status of a grain store for pests, characterised by, include: Acquire multimodal data; the multimodal data includes visible light and near-infrared images of the grain warehouse to be monitored, as well as textual description information; The text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels; the multiple levels include scene level, instance level and attribute level; The multimodal data is input into the pest status monitoring model, which performs feature alignment and identification on the multimodal data and outputs the pest status monitoring results of the grain warehouse to be monitored. The pest status monitoring model includes an image modality fusion module, a multi-level text encoding module, a cross-modal feature alignment module, and a feature classification module. The image modality fusion module is used to extract features and fuse modalities between the visible light image and the near-infrared image to obtain fused visual features; The multi-level text encoding module is used to encode the text description information to obtain text features at each level; The cross-modal feature alignment module is used to perform cross-modal alignment between the fused visual features and the text features at each level to obtain the alignment features at each level. The feature classification module is used to identify the alignment features at each level to obtain the pest status monitoring results. The scene-level text description information includes crop background information and real-time environmental data in the grain warehouse to be monitored; the real-time environmental data includes at least one of real-time temperature information, real-time humidity information, and real-time gas monitoring data. The text description information at the instance level includes pest species information; The text description information at the attribute level includes pest biological characteristics information; The method further includes: The cross-modal feature alignment module globally aligns the fused visual features with the scene-level text features to obtain scene-aligned visual features. Multi-scale feature extraction is performed on the scene alignment visual features to obtain visual multi-scale features; The visual multi-scale features are aligned with the instance text features at the instance level to obtain instance-aligned multi-scale features. Region detection is performed on the instance-aligned multi-scale features to obtain the target detection box and the image region features corresponding to the target detection box; The image region features corresponding to the target detection box are locally aligned with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features.

2. The method of claim 1, wherein, The method further includes: The visible light image and the near-infrared image are fused using an image modality fusion module, and the fused image features are enhanced using a spatial attention mechanism to obtain the fused visual features.

3. The method of claim 2, wherein, The process of fusing features from the visible light image and the near-infrared image, and then enhancing the fused image features based on a spatial attention mechanism to obtain the fused visual features, includes: Feature extraction is performed on the visible light image and the near-infrared image respectively to obtain the first visible light image feature and the first near-infrared image feature; The first fused image features are enhanced by a spatial attention mechanism to obtain the first enhanced image features; the first fused image features are obtained by matrix point addition of the first visible light image features and the first near-infrared image features. The first enhanced image feature is matrix-added with the first visible light image feature and the first near-infrared image feature respectively to obtain the first visible light fused image feature and the first near-infrared fused image feature. The first visible light fused image features and the first near-infrared fused image features are stitched together and feature enhancement based on spatial attention mechanism to obtain the second enhanced image features; The second enhanced image feature is matrix-added with the first visible light fused image feature and the first near-infrared fused image feature respectively to obtain the second visible light fused image feature and the second near-infrared fused image feature. The fused visual features are obtained by matrix point addition of the second visible light fused image features and the second near-infrared fused image features.

4. The method of claim 1, wherein, The method further includes: Semantic parsing is performed on the text description information at the multiple levels to obtain the triples corresponding to each level; A dynamic knowledge graph is constructed based on the triples corresponding to each level.

5. The method of claim 1, wherein, The method further includes: The multi-level text encoding module performs vectorized representation and hierarchical layering of entity nodes and relation edges in the dynamic knowledge graph corresponding to the text description information to obtain scene subgraph, instance subgraph and attribute subgraph. Feature extraction is performed on the scene sub-image to obtain scene text features, and the scene text features are processed to obtain scene scaling coefficient and scene offset coefficient; Feature extraction is performed on the attribute subgraph to obtain initial attribute features, and dynamic feature modulation is performed on the initial attribute features based on the scene scaling coefficient and the scene offset coefficient to obtain attribute text features; Feature extraction is performed on the instance subgraph to obtain initial instance features, and the initial instance features and the initial attribute features are processed based on a hierarchical attention mechanism to obtain instance text features.

6. The method of claim 1, wherein, The step of globally aligning the fused visual features with the scene-level text features to obtain scene-aligned visual features includes: The fused visual features are processed using a global attention mechanism to obtain a scene query vector; The scene text features are processed using a common attention mechanism to obtain scene key vectors and scene value vectors; The scene query vector, scene key vector, and scene value vector are processed based on a cross-attention mechanism to obtain global alignment features; The global alignment features are processed using a feedforward neural network to obtain the scene alignment visual features.

7. The method of claim 1, wherein, The step of aligning the visual multi-scale features with the instance text features at the instance level to obtain instance-aligned multi-scale features includes: The visual multi-scale features are processed based on a deformable attention mechanism to obtain an instance query vector; The instance text features are processed using a common attention mechanism to obtain instance key vectors and instance value vectors; The instance query vector, instance key vector, and instance value vector are processed based on the cross-attention mechanism to obtain region alignment features; The region alignment features are processed using a feedforward neural network to obtain the instance alignment multi-scale features.

8. The method of claim 1, wherein, The step of locally aligning the image region features corresponding to the target detection box with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features includes: The image region features corresponding to the target detection box are processed based on a multi-head attention mechanism to obtain an attribute query vector; The attribute text features are processed using a common attention mechanism to obtain attribute key vectors and attribute value vectors; The attribute query vector, attribute key vector, and attribute value vector are processed based on a cross-attention mechanism to obtain local alignment features; The local alignment features are processed using a feedforward neural network to obtain the attribute alignment multi-scale features.

9. The method of monitoring the status of grain pests according to any one of claims 1 to 8, characterized in that, The training loss of the pest status monitoring model is determined based on at least one of pest classification loss, pest detection loss, pest detection box matching loss, and model knowledge preservation loss. The pest classification loss is determined based on the difference between the predicted value and the actual value of the pest status monitoring results corresponding to the multimodal data of the sample. The pest detection loss is determined based on the difference between the predicted location and the actual location of the target detection box corresponding to the multimodal data of the sample. The pest detection box matching loss is determined based on the matching degree between the predicted position and the actual position of the target detection box corresponding to the multimodal data of the sample. The model knowledge retention loss is determined based on the information divergence between the pest state monitoring model trained in the current iteration and the pest state monitoring model trained in the previous iteration.

10. A grain bin pest status monitoring device, characterized by, include: The data acquisition module is used to acquire multimodal data, which includes visible light and near-infrared images of the grain warehouse to be monitored, as well as text description information. The text description information describes the storage environment information and pest information of the grain warehouse to be monitored based on multiple levels; the multiple levels include scene level, instance level and attribute level; The status monitoring module is used to input the multimodal data into the pest status monitoring model, and the pest status monitoring model performs feature alignment and recognition on the multimodal data, and outputs the pest status monitoring results of the grain warehouse to be monitored. The pest status monitoring model includes an image modality fusion module, a multi-level text encoding module, a cross-modal feature alignment module, and a feature classification module. The image modality fusion module is used to extract features and fuse modalities between the visible light image and the near-infrared image to obtain fused visual features; The multi-level text encoding module is used to encode the text description information to obtain text features at each level; The cross-modal feature alignment module is used to perform cross-modal alignment between the fused visual features and the text features at each level to obtain the alignment features at each level. The feature classification module is used to identify the alignment features at each level to obtain the pest status monitoring results. The scene-level text description information includes crop background information and real-time environmental data in the grain warehouse to be monitored; the real-time environmental data includes at least one of real-time temperature information, real-time humidity information, and real-time gas monitoring data. The text description information at the instance level includes pest species information; The text description information at the attribute level includes pest biological characteristics information; The cross-modal feature alignment module globally aligns the fused visual features with the scene-level text features to obtain scene-aligned visual features. Multi-scale feature extraction is performed on the scene alignment visual features to obtain visual multi-scale features; The visual multi-scale features are aligned with the instance text features at the instance level to obtain instance-aligned multi-scale features. Region detection is performed on the instance-aligned multi-scale features to obtain the target detection box and the image region features corresponding to the target detection box; The image region features corresponding to the target detection box are locally aligned with the attribute text features at the attribute level to obtain attribute-aligned multi-scale features.

11. A grain bin pest status monitoring system, comprising: Includes data acquisition devices, data transmission devices, and data processing devices; The data acquisition device is installed in the grain warehouse to be monitored and is used to acquire visible light images, near-infrared images and real-time environmental data of the grain warehouse to be monitored. The data transmission device is communicatively connected to the data acquisition device and is used to send the visible light image, the near-infrared image and the real-time environmental data to the data processing device. The data processing device is communicatively connected to the data transmission device and is used to execute the grain warehouse pest status monitoring method according to any one of claims 1 to 9, processing the visible light image, the near-infrared image and the real-time environmental data to obtain the pest status monitoring results of the grain warehouse to be monitored.

12. The grain bin pest status monitoring system of claim 11, wherein, The data acquisition device includes a ring track deployed on the top of the grain silo to be monitored, and a data acquisition unit that slides along the ring track; The data acquisition unit is equipped with an image sensor, a temperature sensor, a humidity sensor, and a gas sensor; The image sensor is used to acquire visible light images and near-infrared images of the surface of the grain silo to be monitored; The temperature sensor is used to collect real-time temperature information in the grain warehouse to be monitored. The humidity sensor is used to collect real-time humidity information in the grain warehouse to be monitored. The gas sensor is used to collect real-time gas monitoring data in the grain silo to be monitored.

13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the grain storage pest status monitoring method according to any one of claims 1 to 9. 14.A non-transitory computer-readable storage medium having stored thereon a computer program. When the computer program is executed by the processor, it implements the grain storage pest status monitoring method according to any one of claims 1 to 9.

15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the grain storage pest status monitoring method according to any one of claims 1 to 9.