A cross-modal representation learning and retrieval method and system for grain production

By combining a bidirectional guided fusion network of images and text with a global semantic guidance method and the optimal transmission theory, the problem of cross-modal feature alignment between images and text in agricultural data was solved, enabling efficient image and text data retrieval and intelligent management in the grain production process.

CN120705355BActive Publication Date: 2025-12-02AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511195014.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-02
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies lack effective cross-modal data fusion methods in the agricultural field, making it difficult to achieve fine-grained semantic alignment and efficient retrieval between images and text, resulting in poor application of cross-modal data in food production.

Method used

A bidirectional guided fusion network for image and text is used for multi-granular semantic alignment. Combining global semantic guidance and optimal transmission theory, a cross-modal representation learning and retrieval system is constructed. Through feature extraction and fusion of image-text pairs and video-text pairs, fast and accurate matching and retrieval of images and text are achieved.

Benefits of technology

It improves the accuracy and retrieval efficiency of semantic matching between images and text, enhances the accuracy of cross-modal feature fusion and the intelligent application capabilities of agricultural data, and supports tasks such as intelligent identification of pests and diseases and dynamic monitoring of seedling conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705355B_ABST
    Figure CN120705355B_ABST
Patent Text Reader

Abstract

This application discloses a cross-modal representation learning and retrieval method and system for grain production, relating to the field of agricultural informatization. The method includes: performing multi-granularity semantic alignment on image-text pairs in the grain production process based on a bidirectional guided image-text fusion network to obtain semantically segmented images; performing image spatial decoupling and temporal enhancement on video-text pairs in the grain production process based on global semantic guidance to obtain structured semantic image features; constructing a text feature library and an image feature library; determining a transmission plan matrix based on the modality of the data to be retrieved; generating query features for the data to be retrieved based on the transmission plan matrix; and outputting text query results or image query results using a similarity measurement method based on the query features, text feature library, and image feature library. This application enables deep fusion of cross-modal features, improves the accuracy of image-text semantic matching, and achieves fast and accurate matching and retrieval between images and text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of agricultural informatization, and in particular to a cross-modal representation learning and retrieval method and system for grain production. Background Technology

[0002] Throughout the entire grain production process, the acquisition and effective utilization of data from multiple sources are crucial for improving agricultural production efficiency and achieving precision management. With the rapid development of remote sensing and information technology, the agricultural sector has accumulated a wealth of heterogeneous, multimodal data resources, such as remote sensing images, crop monitoring images, meteorological records, pest and disease analysis reports, and expert advice. Image data provides rich spatial visual information, while text data offers detailed background descriptions and contextual semantics; the two are highly complementary. Fully utilizing this multimodal data can help achieve intelligent perception and precise decision-making throughout the entire agricultural production process.

[0003] However, traditional agricultural data processing methods mostly focus on single-modal processing and lack modeling of the deep semantic relationships between images and text, resulting in poor performance in the analysis and retrieval of cross-modal data. Especially in the field of food production, data from different sources have significant differences in scale. Remote sensing images have spatiotemporal continuity, while text data is diverse in expression. Existing data-driven machine learning algorithms struggle to establish effective unified semantic representations across multi-scale and multi-modal data, further limiting the intelligent application of agricultural big data.

[0004] Furthermore, a "heterogeneous gap" exists between agricultural images and text, meaning that differences in modality distribution and feature representation prevent direct alignment of cross-modal features, reducing retrieval efficiency and understanding accuracy. Related technologies for the fusion and utilization of image and text information mainly focus on the following methods: one type uses images as the primary focus and text as supplementary explanation, performing only simple keyword matching; another type uses deep neural networks to extract features from images and text separately, achieving alignment through attention mechanisms or semantic matching algorithms.

[0005] The above methods have made some progress in general image and text retrieval and image annotation tasks, but their applicability in agricultural scenarios has the following shortcomings: (1) Difficulty in granular alignment: Agricultural images often contain multiple complex targets (such as crops, lesions, and agricultural machinery), while text descriptions may involve multiple objects and spatiotemporal backgrounds. Existing methods mostly remain at the overall level of matching, making it difficult to achieve fine-grained semantic alignment. (2) Lack of a unified cross-modal expression mechanism: Agricultural data often comes from diverse sources and has different organizational forms. There is a lack of effective mechanisms to uniformly model and share the representation of information from different modalities, which limits the integration and utilization of data and the development of downstream retrieval and analysis tasks. (3) Low retrieval efficiency: Faced with large-scale agricultural multimodal data, existing retrieval methods are difficult to balance matching efficiency and accuracy, making it difficult to meet the needs of intelligent agricultural applications for response speed and information accuracy.

[0006] Therefore, there is an urgent need to propose a cross-modal feature alignment and fusion method suitable for agricultural graphic data scenarios, to achieve deep alignment of images and text at the semantic level, construct a unified feature space, and improve the retrieval efficiency and understanding ability of large-scale graphic data. Summary of the Invention

[0007] The purpose of this application is to provide a cross-modal representation learning and retrieval method and system for grain production, which can effectively capture fine-grained semantic relationships between different modal data such as images, videos and texts in the grain production process, realize deep fusion of cross-modal features, improve the accuracy of image and text semantic matching, and achieve fast and accurate matching and retrieval between images and texts.

[0008] To achieve the above objectives, this application provides the following solution:

[0009] Firstly, this application provides a cross-modal representation learning and retrieval method for grain production, including:

[0010] Based on a pre-trained bidirectional guided fusion network, multi-granular semantic alignment is performed on image-text pairs in the grain production process to obtain semantically segmented images; the bidirectional guided fusion network includes a feature extraction module, a bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module;

[0011] Based on global semantic guidance, image spatial decoupling and temporal enhancement are performed on video-text pairs in the grain production process to obtain structured semantic image features; the videos in the video-text pairs are remote sensing image sequences with temporal order.

[0012] Based on the image-text pairs, the video-text pairs, the semantic segmentation images, and the structured semantic image features, a text feature library and an image feature library are constructed.

[0013] The transmission plan matrix is ​​determined based on the modality of the data to be retrieved, and the query features of the data to be retrieved are generated based on the transmission plan matrix. Based on the query features of the data to be retrieved, the text feature library, and the image feature library, the text query results or image query results are output using a similarity measurement method.

[0014] Secondly, this application provides a cross-modal representation learning and retrieval system for grain production, including:

[0015] The image-text learning module is used to perform multi-granular semantic alignment of image-text pairs in the grain production process based on a pre-trained bidirectional guided fusion network to obtain semantically segmented images; the bidirectional guided fusion network includes a feature extraction module, a bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module;

[0016] The video-text learning module is used to perform image spatial decoupling and temporal enhancement on video-text pairs in the grain production process based on global semantic guidance, so as to obtain structured semantic image features; the videos in the video-text pairs are remote sensing image sequences with temporal order.

[0017] The feature library construction module is used to construct a text feature library and an image feature library based on the image text pairs, the video text pairs, the semantic segmentation images, and the structured semantic image features;

[0018] The cross-modal retrieval module is used to determine the transmission plan matrix based on the modality of the data to be retrieved, generate query features of the data to be retrieved based on the transmission plan matrix, and output text query results or image query results using a similarity measurement method based on the query features of the data to be retrieved, the text feature library, and the image feature library.

[0019] According to the specific embodiments provided in this application, this application has the following technical effects:

[0020] This application provides a cross-modal representation learning and retrieval method and system for grain production. By designing a bidirectional guided fusion network, it achieves multi-granular semantic alignment between images and text, effectively enhancing the consistent expression of linguistic and visual information in the feature space and improving the accuracy of cross-modal feature fusion. Based on global semantic guidance, it performs image spatial decoupling and temporal enhancement on video-text pairs in the grain production process, fully exploring the spatiotemporal feature correlations in remote sensing images, and improving the modeling and generalization capabilities of key agricultural semantics. Furthermore, it introduces optimal transport theory to optimize feature mapping, achieving fast and accurate matching and retrieval between images and text. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a cross-modal representation learning and retrieval method for grain production provided in an embodiment of this application.

[0023] Figure 2 This is a schematic diagram of the overall structure of a bidirectional text-image guidance fusion network in one embodiment of this application.

[0024] Figure 3 This is a schematic diagram of the process of global semantic guidance for spatial decoupling and temporal enhancement of remote sensing images in one embodiment of this application.

[0025] Figure 4 This is a schematic diagram of the framework for optimal transmission of cross-modal unified representation learning and efficient bidirectional image and text retrieval in grain production, as shown in one embodiment of this application.

[0026] Figure 5 This is a schematic diagram of the functional modules of a cross-modal representation learning and retrieval system for grain production, provided as an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] This application aims to address the "heterogeneous gap" in feature representation of massive amounts of multimodal data, including remote sensing images, near-ground images, and agricultural texts generated throughout the entire grain production process. It develops an end-to-end deployable, strongly correlated, and semantically consistent image-text retrieval method from three dimensions: cross-modal deep fusion modeling, spatiotemporal enhancement understanding of temporal images, and unified semantic representation of heterogeneous data. This aims to improve the fusion representation capability and intelligent retrieval efficiency of cross-modal data. To achieve the above objectives, this application employs single-image feature extraction based on convolutional neural networks, fusion of linguistic, visual, and spatial coordinate features based on bidirectional guided image-text fusion networks, global guided spatial decoupling and temporal enhancement of images, and minimizing the transmission cost of original and unified representations to reduce semantic loss during compression of unified representations. From the perspective of data-driven learning and generalization, this approach achieves unified and correlated representation and efficient retrieval of cross-modal content.

[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] In one exemplary embodiment, such as Figure 1 As shown, a cross-modal representation learning and retrieval method for grain production is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the cross-modal representation learning and retrieval method for grain production includes the following steps 1 to 4.

[0031] Step 1 constructs a multi-granularity semantic alignment method based on a bidirectional guided fusion network of images and text. This addresses the shortcomings of existing cross-modal interaction methods, which often rely on manually designed interaction modes, lack flexibility, and struggle to achieve efficient guidance and deep fusion of image and text information, resulting in low cross-modal understanding accuracy and limited feature representation capabilities. The method achieves multi-granularity semantic alignment between images and text through a bidirectional guided fusion network. Specifically, it involves: first, using a convolutional neural network to extract visual features from the input image; then, using a recurrent neural network to encode the natural language text and obtain its contextual features. Simultaneously, spatial coordinate features are generated for each image location to enhance the image's positional information representation capability. In the feature fusion stage, contextual features, visual features, and spatial coordinate features are cascaded and integrated, and cross-modal feature inference is performed through a bidirectional guided fusion module. In this module, visual information, guided by text, strengthens its semantic expressive power, while text, guided by semantic cues from the image, further refines its description, thereby enhancing semantic consistency and feature complementarity between images and text. This multi-granularity semantic alignment mechanism effectively captures the fine-grained and global semantic correspondence between images and text. Finally, based on the fully convolutional network, the fused features are classified into target regions, and the results are upsampled by the deconvolution module to obtain pixel-level region segmentation results in the image corresponding to the text description content. This can significantly improve the understanding and matching accuracy of cross-modal information between images and text in the grain production process, and promote efficient linkage and joint analysis between image and text data.

[0032] Step 2 constructs a remote sensing image spatial decoupling and temporal enhancement method based on global semantic guidance to improve the expressive power of key regions and the perception accuracy of spatiotemporal changes. The specific method includes: First, in the spatial feature extraction stage, a global semantic guidance mechanism is introduced. By integrating global semantic vectors extracted from multiple temporal remote sensing images, the spatial feature representation of the current frame remote sensing image is guided by saliency and decoupled structurally, dynamically distinguishing between high and low semantic saliency regions. This operation helps highlight regional features in remote sensing images that are highly relevant to agricultural task objectives, while suppressing redundant background information and enhancing the discriminativeness and interpretability of the image's semantic structure. Second, in the temporal modeling stage, a temporal enhancement mechanism is designed, employing differentiated modeling strategies for spatial regions with different semantic saliency. For high saliency regions, their prior advantages in the semantic space are utilized, combined with semantic comparison with low saliency regions, to further enhance region discrimination capabilities; for low saliency regions, an inter-frame accumulation strategy within the image sequence is adopted to capture fine-grained temporal change trends, enhancing their dynamic representation capabilities and avoiding the loss of important temporal information. Finally, through a synergistic mechanism of spatial decoupling and temporal enhancement, we achieve high-expressive modeling of key regions in agricultural remote sensing images and deep information mining in the temporal dimension. This significantly improves the model's ability to perceive and express spatial heterogeneity and temporal evolution in agricultural scenarios, providing a solid spatiotemporal semantic foundation for cross-modal representation learning and further assisting in the accurate execution of image-text semantic alignment and downstream intelligent tasks.

[0033] Step 4 constructs a cross-modal unified representation learning and efficient bidirectional image-text retrieval mechanism for grain production based on optimal transmission. This addresses the issue of multi-source heterogeneous modal data involved in grain production, including remote sensing images, ground monitoring images, and agricultural text records. These data differ significantly in representation, structural dimension, and semantic abstraction level, leading to a "semantic gap" between different modalities and hindering semantic alignment and joint modeling. The goal is to achieve fusion representation and efficient matching of multimodal data in grain production within a unified semantic space. Specific methods include: First, the representation process of heterogeneous modal data is uniformly modeled as an optimal mapping problem from the original modal feature space to a shared unified semantic space. To achieve effective alignment between modalities, optimal transmission theory is introduced, and a transmission cost function for inter-modal alignment is constructed. By minimizing the transmission distance between the original visual / text features and the target unified representation, semantic loss during feature compression and fusion is effectively reduced, improving the consistency and discriminativeness of multimodal semantic expression. Secondly, during model training and inference, a fast-converging and structurally well-defined joint optimization strategy is designed to achieve collaborative learning of the modality mapping matrix and the representation decoder, enhancing the model's adaptability and generalization ability to complex and diverse agricultural scenarios. This mechanism can dynamically adjust the semantic weights between different modalities to achieve robust representations in diverse scenarios. Finally, a unified representation mechanism supports bidirectional image-text retrieval, enabling bidirectional matching queries of "image to text" and "text to image" within a shared semantic space. This allows for efficient retrieval of image-text pairs semantically relevant to agricultural tasks from large-scale agricultural multimodal data, providing accurate and reliable information support for tasks such as intelligent identification of pests and diseases, dynamic monitoring of seedling conditions, and automatic understanding of agricultural behaviors, significantly improving the level of integrated management and intelligent services for cross-modal agricultural data.

[0034] The details of each step are described below.

[0035] Step 1: Based on a pre-trained bidirectional guided image-text fusion network, multi-granularity semantic alignment is performed on image-text pairs in the grain production process to obtain semantically segmented images. The bidirectional guided image-text fusion network includes a feature extraction module, a bidirectional guided image-text fusion module, a multi-scale fusion module, and a semantic segmentation module.

[0036] Specifically, the processing procedure of the bidirectional text-image guidance fusion network is as follows: Figure 2 As shown. Step 1 includes steps 11 to 14.

[0037] Step 11: The feature extraction module performs feature extraction and feature concatenation on the image and text in the image-text pair to obtain the contextual features of each word in the text and the multimodal hybrid features of each image region. Different image regions correspond to different layers of the image.

[0038] In a specific application example, the feature extraction module includes a DeepLab ResNet-101 v2 network, a relative spatial coding network, a long short-term memory network, and a feature concatenation network. Step 11 includes steps 111 to 114.

[0039] Step 111: The image is processed using a pre-trained DeepLab ResNet-101 v2 network to extract features, resulting in a visual feature vector for each image region. The visual feature vectors of different image regions correspond to the backbone outputs of the last three layers of the DeepLab ResNet-101 v2 network.

[0040] Specifically, the images are first uniformly scaled and zero-padded to a size of 320×320 to ensure consistency of network input. Then, the images are fed into a pre-trained DeepLab ResNet-101 v2 network to extract its multi-scale visual features. The outputs of the last three backbone layers of this network (Res3, Res4, Res5) are selected as multi-granular visual semantic representations to capture the local texture and global semantics of the images.

[0041] Res3 extracts mid-level semantic features, mainly including local texture information and edge structures, and has a strong ability to preserve details. Its output feature map has a dimension of 40×40×512, i.e., a spatial resolution of 40×40 and a channel count of [missing information]. The visual feature vector at each spatial location is 512. The dimension is 512.

[0042] Res4 extracts mid-to-high-level semantic features, capturing more abstract regional semantics. Its output feature map has a dimension of 20×20×1024, reducing the spatial resolution by half and the number of channels... The value is 1024. The visual feature vector for each location. The dimension is 1024.

[0043] Res5 extracts the deepest semantic features, containing the richest global semantic information, but has the lowest spatial resolution. Its output feature map has a dimension of 10×10×2048 and a number of channels... The value is 2048. The visual feature vector for each location. The dimension is 2048.

[0044] Visual features are complete 3D tensors output by a layer in the DeepLab ResNet-101 v2 network, with dimensions H×W×C (height×width×number of channels). Visual feature vectors, on the other hand, refer to the feature representation of a spatial location (pixel) in this feature map with respect to the number of channels, used to describe the semantic information of the corresponding region in the image, denoted as [image vector]. ;in, That is, the firsti Visual feature vectors of image regions.

[0045] Step 112: Use a relative spatial coding network to generate positional features for each region of the image, thereby obtaining the spatial coordinate features of each image region.

[0046] To enhance the spatial location awareness of target regions in images, this application uses a relative spatial coding method to generate location features for each region, typically represented by an 8-dimensional vector. , .in, For the first i Spatial coordinate characteristics of an image region.

[0047] Step 113: Extract the text using a Long Short-Term Memory (LSTM) network. Each word Contextual features .in, This represents the total number of words in the text (such as harvester, in the field, harvesting wheat). For the first T The context features of the last word constitute the context feature vector, and the final hidden state of the LSTM is output. To represent the meaning of the entire sentence.

[0048] Step 114: For any image region, a feature concatenation network is used to concatenate the visual feature vector of the image region, the spatial coordinate features of the image region, and the contextual features of the last word in the text to obtain the multimodal hybrid features of the image region. ;in, For the first i Multimodal hybrid features of an image region.

[0049] Step 12: Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, perform bidirectional guided fusion of text and image through the bidirectional guided fusion module to obtain the enhanced features of each image region.

[0050] In a specific application example, the bidirectional text-image guidance fusion module includes a visually guided language attention mechanism and a language-guided visual attention mechanism. The visually guided language attention mechanism uses the contextual features of each word in the text as guidance to measure the semantic contribution of each word to each image region. The language-guided visual attention mechanism uses the linguistic contextual features of each image region as guidance to calculate the spatial semantic dependencies between image regions. Step 12 includes steps 121 and 122.

[0051] Step 121: Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, a visually guided language attention mechanism is used to determine the language contextual features of each image region.

[0052] Specifically, visual features are used to guide text context modeling, and a weighted attention distribution is learned to measure the first... t The word corresponds to the first i Semantic contribution of each image region: ; ;in, For the first t The word corresponds to the first i The semantic contribution of each image region For the first i The linguistic context features of each image region are visually guided adaptive language representations. The learnable parameters for convolution operations are primarily aimed at reducing the dimensionality of the hybrid features input to the visually guided language attention module, thereby reducing the number of parameters and improving inference speed.

[0053] Step 122: Based on the linguistic context features and multimodal hybrid features of each image region, a language-guided visual attention mechanism is used to determine the enhancement features of each image region.

[0054] Specifically, the following formula is used to calculate the first... i The linguistic context features of the first image region affect the first j The importance of the linguistic context features of each image region: ; ;in, For the first j Multimodal hybrid features of an image region The learnable parameters for the convolution operation primarily aim to reduce the dimensionality of the hybrid features input to the language-guided visual attention module, thereby decreasing the number of parameters and improving inference speed. and These are learnable parameters used to... Dimensionality reduction is performed to decrease the number of parameters in the language-guided visual attention module. For the first i The linguistic context features of the first image region affect the first j The initial importance of the linguistic context features of an image region. For the first i The linguistic context features of the first image region affect the first j The importance of the linguistic context features of each image region The higher the value, the more significant the [value]. iThe image region and the first j The stronger the correlation between image regions.

[0055] The following formula is used to obtain the first... i Enhancement features for each image region : ;in, B To exclude the first i Image regions outside of the image regions, i.e., if i =3, then j =4 or 5, if i =4, then j =3 or 5, if i =5, then j =3 or 4, , , All parameters are learnable, and the main goal is to upscale the updated visual features and concatenate them with adaptive language features to supplement the contextual language features. The language-guided visual attention mechanism can enhance the structured contextual relationships between image regions, enabling adaptive guidance of language semantics for image region perception.

[0056] Step 13: The enhanced features of each image region are fused at multiple scales using the multi-scale fusion module to obtain a fused feature map.

[0057] In a specific application example, the multi-scale fusion module includes a void space pyramid pooling submodule and a bidirectional gating fusion submodule. Step 13 includes steps 131 to 132.

[0058] Step 131: The Atrous Spatial Pyramid Pooling (ASPP) submodule is used to perform multi-scale semantic modeling on the enhanced features of each image region, extracting contextual information under different receptive fields to obtain multimodal features at different levels. ;in, For the first i The multimodal features of each image region correspond to the feature maps output by the Res3, Res4, and Res5 branches of the DeepLab ResNet-101 v2 network, respectively.

[0059] Step 132: Using the bidirectional gated fusion submodule, bidirectional gated fusion (BDGF) is performed on the multimodal features at different levels through top-down and bottom-up paths respectively to obtain the fused feature map.

[0060] The specific fusion operation is as follows:

[0061] ;

[0062] ;

[0063] ;

[0064] ;

[0065] ;

[0066] in, for and The fused feature map for and The fused feature map This is the fused feature map obtained from the top-down path. This is the fused feature map obtained from the bottom-up path. For the final fused feature map, and These are gating functions applied in both upward and downward directions after concatenating feature maps from different layers along the channel dimension. They consist of 3×3 convolutions and a sigmoid activation function, used to regulate the information flow and fusion between multiple feature layers. and These are feature maps from different layers in the network, such as the outputs of layers Res3 and Res4. It is and The features from the two layers are spliced ​​together along the channel dimension to merge the information. To extract fused local features by performing a convolution operation on the concatenated feature maps, Sig applies the Sigmoid function to the convolution result, outputting a gating coefficient (typically between [0,1]) that represents the "pass rate" or "importance" of the feature flow. This refers to the bottom-up direction of information flow. D This refers to the top-down direction of information flow. This is element-wise multiplication.

[0067] Step 14: Perform pixel-by-pixel semantic classification on the fused feature map using the semantic segmentation module to obtain a semantic segmentation image.

[0068] Specifically, a fully convolutional integral network is used to fuse the feature maps. A pixel-by-pixel semantic classification operation is performed to predict image regions that are semantically consistent with the input text, resulting in a classification feature map. A deconvolution module is then used to upsample the classification feature map, generating a semantic segmentation image with the same resolution as the original image. This ultimately achieves accurate segmentation and annotation of semantic target regions in grain production images.

[0069] Step 1 uses key steps such as bidirectional image-text guidance, attention mechanism modeling, and multi-scale fusion to bridge the semantic gap between images and text, effectively improving cross-modal alignment and retrieval accuracy, and providing core support for intelligent perception and decision-making in grain production scenarios.

[0070] Step 2 involves performing image spatial decoupling and temporal enhancement on video-text pairs related to the grain production process based on global semantic guidance, resulting in structured semantic image features. The videos in the video-text pairs are temporally sequenced remote sensing image sequences.

[0071] Specifically, such as Figure 3 As shown, step 2 includes steps 21 to 26.

[0072] Step 21: Sample the video in the video text pair to obtain multiple frames of remote sensing images, and perform global semantic extraction and spatial feature extraction on the multiple frames of remote sensing images to obtain a global semantic vector and a spatial feature representation of each frame of remote sensing image.

[0073] In a specific application example, step 21 includes steps 211 to 213.

[0074] Step 211: A restricted random sampling strategy is used to extract multiple frames of remote sensing images from a continuous sequence of remote sensing images, represented as follows: ,in, A This refers to the number of frames in the remote sensing image (e.g., the number of days or months of remote sensing observation). The height of the image. The width of the image. The number of channels in the image (e.g., 3 for RGB).

[0075] Step 212: Perform temporal average pooling and global average pooling operations on multiple frames of remote sensing images to extract global semantic vectors. : ;in, For the first Frame remote sensing image. The average of all pixels in time and space is used as an abstract global semantics that spans time and space, representing the core semantics of the image throughout the time series and guiding subsequent local feature selection.

[0076] Step 213: Use a pre-trained deep convolutional neural network (such as ResNet-50) to extract features from each frame of the remote sensing image, and obtain the spatial feature representation of each frame of the remote sensing image. ;in, For the first Spatial feature representation of frame remote sensing images .

[0077] Step 22: For any frame of remote sensing image, based on the global semantic vector, decouple the spatial feature representation of the remote sensing image to obtain the high semantic saliency region features and low semantic saliency region features of the remote sensing image.

[0078] In a specific application example, step 22 includes steps 221 to 224.

[0079] Step 221, the global semantic vector Perform a linear transformation and expand it into a feature tensor with the same spatial size as the remote sensing image. It is used to guide attention learning after being spliced ​​with local spatial features.

[0080] Step 222, convert the feature tensor Spatial feature representation of the remote sensing image By cascading, fused features are obtained. .in This indicates that two tensors are concatenated along the channel dimension to perform feature concatenation.

[0081] Step 223: Based on the fusion features, a saliency response map is generated using a convolutional layer and activation function.

[0082] Specifically, the fusion feature input consists of two... A guided correlation estimation module consisting of convolution, batch normalization, and ReLU is used, and then a significant response map is generated through the Sigmoid activation function. ;in, Representing the A saliency response map of a frame-by-frame remote sensing image, where each spatial location represents its semantic importance. This represents the Sigmoid activation function, ranging from (0,1), used to generate a saliency weight map. For two The learnable weights that make up the convolutional layers are typically lightweight convolutional or attention networks with learnable parameters.

[0083] Step 224: Perform saliency-guided decoupling on the spatial feature representation of the remote sensing image based on the saliency response map to obtain the high semantic saliency region features and low semantic saliency region features of the remote sensing image.

[0084] Specifically, using the formula Get the first Highly semantically saliency region features of frame remote sensing images Using formula Get the first Low semantic saliency region features of frame remote sensing images .in, This represents an element-wise multiplication operation used to weight remote sensing images based on saliency response maps. This operation enables dynamic enhancement of regions in remote sensing images that are significantly relevant to agricultural tasks. High semantic saliency region features represent regions in the image that are more strongly relevant to agricultural tasks, while low semantic saliency region features represent redundant or background regions in the image.

[0085] Step 23: Based on the high semantic saliency region features and low semantic saliency region features of multiple frames of remote sensing images, perform temporal enhancement modeling on each frame of remote sensing images to obtain the forward saliency region features, backward saliency region features, forward cumulative low saliency features, and backward cumulative low saliency features of each frame of remote sensing images.

[0086] In a specific application example, step 23 includes steps 231 to 236.

[0087] Step 231: Generate initial cumulative low-saliency features based on the low semantic saliency region features of multiple frames of remote sensing images: ;in, This is an initial cumulative low significance feature.

[0088] Step 232, using forward-enhancing memory units to process the first... High semantic saliency region features of frame remote sensing images and the first Differential modeling is performed using forward cumulative low-significance features of the first frame of remote sensing images to extract the first... Forward semantic change information of frame remote sensing images: ;in, For the first Forward semantic change information of frame remote sensing images For the first Low saliency features of forward accumulation in frame-by-frame remote sensing images. For the first Highly semantically saliency region features of frame remote sensing images and This represents two distinct 1×1 convolutional layers.

[0089] Step 233: Employ channel attention mechanism to extract the first... The response weights of the forward semantic change information of the frame remote sensing image are used to determine the response weights for the first frame. Time-series contrast enhancement was performed on the highly semantically salient region features of the frame remote sensing image to obtain the first... Forward saliency region features of frame remote sensing images.

[0090] Specifically, regarding the first Frame remote sensing images, first of all Perform spatial average pooling, then use the formula Extract response weights; where, For the first Response weights of frame-by-frame remote sensing images for Information for spatial average pooling. These are the parameter weights for the linear transformation. Then, the formula is used. Get the first Contrast enhancement features of frame remote sensing images Then, based on the contrast enhancement features, the forward saliency region features are obtained. .

[0091] Step 234, using backward enhancement memory units to process the first... High semantic saliency region features of frame remote sensing images and the first Differential modeling is performed using backward cumulative low-significance features of the frame remote sensing image to extract the first... Backward semantic change information of frame remote sensing images.

[0092] Step 235: Employ channel attention mechanism to extract the first... The response weights of the backward semantic change information of the frame remote sensing image, and the response weights are used to adjust the first frame. Time-series contrast enhancement was performed on the highly semantically salient region features of the frame remote sensing image to obtain the first... Backward saliency region features of frame remote sensing images.

[0093] The processing procedures for steps 234 and 235 are similar to those for steps 232 and 233, and will not be repeated here. The final result is the... Forward saliency region features of frame remote sensing images are , No. The backward saliency region features of a frame remote sensing image are .

[0094] Step 236, according to the first Cumulative low significance features of frame remote sensing images and the first Low semantic saliency region features of frame remote sensing images to determine the first Forward cumulative low significance features and backward cumulative low significance features of frame remote sensing images : .in, It is a residual block.

[0095] in, , A The number of frames in the remote sensing image. Time Forward cumulative low saliency features of frame remote sensing images and the first The backward cumulative low significance features of the frame remote sensing image are all the initial cumulative low significance features.

[0096] Step 24: Obtain the semantic enhancement feature representation of each frame of remote sensing image based on the forward salient region features and the backward salient region features of each frame of remote sensing image.

[0097] Specifically, and The features are concatenated after global average pooling and then passed through a fully connected layer to generate the final semantically enhanced feature representation: , ;in, For the first Semantic enhancement feature representation of frame remote sensing images, These are learnable parameters for global average pooling.

[0098] Step 25: Based on the forward cumulative low saliency features of the last frame of remote sensing image and backward cumulative low significance characteristics The final cumulative feature representation is obtained as follows: ;in, For the final cumulative feature representation, These are learnable parameters for linear optimization.

[0099] Step 26: The semantic enhancement feature representation of each frame of remote sensing image is fused with the final accumulated feature representation to obtain structured semantic image features.

[0100] Specifically, the first The semantically enhanced feature representation of a frame of remote sensing image is fused with the final cumulative feature representation to obtain structured semantic image features for agricultural production scenarios. This feature can be aligned with semantic representations in the text modality for cross-modal contrastive learning or retrieval tasks.

[0101] In another exemplary embodiment, step 2 further includes step 25, which trains the parameters involved in steps 21 to 26 based on the loss function.

[0102] In a specific application example, step 25 includes steps 251 to 254.

[0103] S251: Regarding and Feature supervision is performed using image-level and video-level online instance matching loss functions. Validation loss is used to supervise the similarity between remote sensing image sequences, measuring the error in determining whether image-image (or video) pairs belong to the same class. The validation loss function is expressed as: ;in, To verify the loss function value, To measure the similarity judgment error between image-image (or video) pairs to determine whether they belong to the same class, For the first n The feature vectors of the images / videos to be matched No. n Feature vectors of a comparison image / video Let be the similarity function, representing the similarity score between features. It is typically a cosine similarity or a sigmoid mapping of embedding space distance. Represents the tag value, if and If they belong to the same category, then =1, otherwise 0.

[0104] S252: Yes Frame-level online instance matching (OIM) loss is employed to enhance region discrimination. The frame-level OIM loss function is applied to the semantic feature-supervised learning of each frame in the image sequence, and is expressed as: ;in, The frame-level OIM loss function value. M The number of remote sensing image sequences. R This represents the total number of categories (e.g., the number of different objects / regions / crop species in the training set). Indicates the first In the remote sensing image sequence, the th Semantic enhancement features of frame images (from highly saliency regions). Indicates the first r The weight vector corresponding to the category is used to represent the category in the embedding space, and its value is updated online based on the image features during training. If the first The first of the remote sensing image sequences Frame belongs to the r If the category (such as an agricultural plot or semantic category) is specified, the value is 1; otherwise, it is 0.

[0105] S253: Yes Sequence-level OIM loss is employed to capture fine-grained semantic evolution trends. Remote sensing sequence-level OIM loss is used for global aggregation representation supervision of low-saliency region features at the end of the sequence (e.g., the last frame), expressed as: ;in, The value of the sequence-level OIM loss function. If the first The remote sensing image sequence belongs to the _th _th r If the class is specified, the value is 1; otherwise, it is 0.

[0106] S254: Combining the above three losses, construct the total loss function for the joint training objective function: ;in, This is the total loss function value. These are all weight hyperparameters, which control the proportion of influence of different loss terms on the final objective function. They need to be tuned experimentally.

[0107] Step 3: Construct a text feature library and an image feature library based on the image text pairs, the video text pairs, the semantic segmentation images, and the structured semantic image features.

[0108] Step 4: Determine the transmission plan matrix based on the modality of the data to be retrieved, and generate query features of the data to be retrieved based on the transmission plan matrix. Based on the query features of the data to be retrieved, the text feature library, and the image feature library, output the text query results or image query results using a similarity measurement method.

[0109] To achieve unified semantic representation and efficient bidirectional retrieval of multi-source heterogeneous remote sensing images, monitoring images, and agricultural text records in grain production scenarios, this application proposes a cross-modal unified representation learning and efficient bidirectional image-text retrieval mechanism for grain production based on optimal transmission in step 4. The overall framework is as follows: Figure 4 As shown.

[0110] In a specific application example, step 4 includes steps 41 to 44.

[0111] Step 41: Construct the optimal transport mapping. Specifically, step 41 includes steps 411 and 412.

[0112] Step 411, assuming there is Modality (e.g., the number of modalities from different data sources such as remote sensing images, monitoring images, and text), the first The discrete distribution of class modes in the feature space is as follows ,in The distribution vector of the unified semantic space is denoted as , The corresponding number The feature point transfer cost matrix of the modality Represented as: ;in, Indicates the first In the class modality, the first The nth feature point is mapped to the unified space. The cost incurred for each feature point For the first The number of feature points of a modality (i.e., the number of support points discretely distributed in the original feature space for that modality). The distribution vector of the original mode. For the first In the class modality, the first The distribution of feature points The distribution vector of the unified space has elements that sum to 1. Q To unify the number of feature points in the semantic space (i.e., the number of support points in the final shared space).

[0113] Step 412, for each type of mode, the first The class mode can be solved by the following optimal transport, and the result after optimal transport is the first... Modal optimal transmission plan matrix Represented as: ;in, Let be the transmission plan matrix, representing the first... The mapping weights from feature points of a modality to feature points in a unified space satisfy the feasible set of the transmission plan. , and They represent lengths of Q and A vector of all 1s. The constraints are given. The resulting joint distribution is... That is, it can be characterized by Migrate to The optimal joint distribution probability. Unlike traditional fusion mechanisms, The first one can be clearly given The relationships between class modalities and unified representation dimensions are established, and the results are supported by clear theoretical guarantees.

[0114] Step 42: Align and fuse the unified representation with the modal representation. Specifically, step 42 includes steps 421 and 422.

[0115] Step 421, using the optimal transmission plan matrix for each mode. Generate the projection representation of this mode in a unified space. : ;in For the first The original eigenvector matrix of the modality, total There are feature points, each of which is . d Dimensional vector.

[0116] Step 422: Weight all modal projection representations. Weighted fusion, representing the first The contribution ratio of each modality to the final unified representation satisfies non-negativity and sums to 1, thus yielding the unified representation: , ;in, The weighted fusion result representing all modal projection representations is the final unified representation matrix.

[0117] Step 43, retrieval model optimization. Specifically, step 43 includes steps 431 and 432.

[0118] Step 431: Calculate retrieval similarity. Define image query features in a unified space. Features of text retrieval database The similarity is expressed as cosine similarity: In a unified space, These represent the unified representation matrices (or vector sets) of the input image and text, respectively. It is an inner product operation used to calculate the cosine similarity numerator. Represents the Frobenius norm or Euclid norm of a vector or matrix, depending on the context, used for normalization. This is the feature similarity function.

[0119] Step 432, Retrieval Objective Function. The overall objective function employs a joint optimal transmission loss and retrieval comparison loss. for: ;in, The set of positive sample pairs contains all "correctly matched" image-text pairs. To retrieve and compare the loss weight coefficients, This is the Sigmoid activation function, used to map similarity to (0,1).

[0120] Step 433, end-to-end optimization. The above loss term is backpropagated and combined with the feature library update mapping network, feature encoder and decoder constructed in step 3 to achieve fast convergence and well-defined joint optimization.

[0121] Step 44: Perform bidirectional image and text search. Specifically, step 44 includes steps 441 and 442.

[0122] Step 441, Image-to-Text Search. Given image query features. Based on similarity in the text feature library Sort the data and output the Z most relevant items after sorting by similarity.

[0123] Step 442, Text-to-Image Search. Given a text query... Sort the image features by similarity in the image feature library and output the top Z most relevant items after sorting by similarity.

[0124] This retrieval process can achieve a response time within seconds in large-scale agricultural multimodal databases and ensure high-precision matching.

[0125] This application addresses the "semantic gap" problem caused by heterogeneous data from multiple sources, multiple modalities, and multiple scales in agricultural production. It constructs a unified cross-modal feature representation and efficient retrieval mechanism, achieving multi-granular semantic alignment between images and text through a bidirectional guided fusion network. This effectively enhances the consistent representation of linguistic and visual information in the feature space, improving the accuracy of cross-modal feature fusion. Combined with a proposed remote sensing image spatial decoupling and temporal enhancement method based on global semantic guidance, it fully explores the spatiotemporal feature correlations in remote sensing images, improving the model's ability to model and generalize key agricultural semantics. Furthermore, it constructs a unified cross-modal representation and efficient retrieval mechanism for heterogeneous data in grain production, introducing optimal transport theory for feature mapping optimization to minimize semantic loss during inter-modal mapping, achieving fast and accurate matching and retrieval between images and text. Overall, this application is widely applicable to typical scenarios such as large-scale agricultural remote sensing monitoring, crop growth identification, and agricultural record analysis, demonstrating significant application value and promising prospects in improving the efficiency of agricultural information utilization and supporting intelligent decision-making and management.

[0126] Based on the same inventive concept, this application also provides a cross-modal representation learning and retrieval system for grain production to implement the methods described above. The solution provided by this system is similar to the implementation scheme described in the above methods. Therefore, the specific limitations of one or more embodiments of the cross-modal representation learning and retrieval system for grain production provided below can be found in the limitations of the cross-modal representation learning and retrieval method for grain production described above, and will not be repeated here.

[0127] In one exemplary embodiment, such as Figure 5 As shown, a cross-modal representation learning and retrieval system for grain production is provided, including: an image text learning module 51, a video text learning module 52, a feature library construction module 53, and a cross-modal retrieval module 54.

[0128] The image-text learning module 51 is used to perform multi-granularity semantic alignment of image-text pairs in the grain production process based on a pre-trained bidirectional guided image-text fusion network, resulting in a semantically segmented image. The bidirectional guided image-text fusion network includes a feature extraction module, a bidirectional guided image-text fusion module, a multi-scale fusion module, and a semantic segmentation module.

[0129] The video-text learning module 52 is used to perform image spatial decoupling and temporal enhancement on video-text pairs in the grain production process based on global semantic guidance, to obtain structured semantic image features. The videos in the video-text pairs are temporally sequenced remote sensing image sequences.

[0130] The feature library construction module 53 is used to construct a text feature library and an image feature library based on the image text pair, the video text pair, the semantic segmentation image, and the structured semantic image features.

[0131] The cross-modal retrieval module 54 is used to determine the transmission plan matrix according to the modality of the data to be retrieved, generate the query features of the data to be retrieved based on the transmission plan matrix, and output the text query results or image query results using a similarity measurement method based on the query features of the data to be retrieved, the text feature library and the image feature library.

[0132] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0133] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0134] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0136] In this application, all actions to acquire signals, information, or data are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0137] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0138] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0139] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0140] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A cross-modal representation learning and retrieval method for grain production, characterized in that, The method includes: Based on a pre-trained bidirectional guided image-text fusion network, multi-granularity semantic alignment of image-text pairs in the grain production process is performed to obtain semantically segmented images. The bidirectional guided image-text fusion network includes a feature extraction module, a bidirectional guided image-text fusion module, a multi-scale fusion module, and a semantic segmentation module. The feature extraction module includes a DeepLab ResNet-101v2 network, a relative spatial coding network, a long short-term memory network, and a feature concatenation network. The bidirectional guided image-text fusion module includes a visually guided language attention mechanism and a language-guided visual attention mechanism. The visually guided language attention mechanism uses the contextual features of each word in the text as guidance to measure the semantic contribution of each word in the text to each image region. The language-guided visual attention mechanism uses the linguistic contextual features of each image region as guidance to calculate the spatial semantic dependencies between image regions. The multi-scale fusion module includes a hollow spatial pyramid pooling submodule and a bidirectional gated fusion submodule. Specifically, based on a pre-trained bidirectional guided image-text fusion network, multi-granular semantic alignment of image-text pairs in the grain production process is performed to obtain semantic segmentation images, including: The feature extraction module performs feature extraction and feature concatenation on the image and text in the image-text pair to obtain the context features of each word in the text and the multimodal hybrid features of each image region; wherein, different image regions correspond to different layers of the image; Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, the image-text bidirectional guided fusion module performs image-text bidirectional guided fusion to obtain the enhanced features of each image region. The enhanced features of each image region are fused at multiple scales using the multi-scale fusion module to obtain a fused feature map. The semantic segmentation module performs pixel-by-pixel semantic classification on the fused feature map to obtain a semantic segmentation image; Based on global semantic guidance, image spatial decoupling and temporal enhancement are performed on video-text pairs in the grain production process to obtain structured semantic image features; the videos in the video-text pairs are remote sensing image sequences with temporal order. Based on the image-text pairs, the video-text pairs, the semantic segmentation images, and the structured semantic image features, a text feature library and an image feature library are constructed. The transmission plan matrix is ​​determined based on the modality of the data to be retrieved, and the query features of the data to be retrieved are generated based on the transmission plan matrix. Based on the query features of the data to be retrieved, the text feature library, and the image feature library, the text query results or image query results are output using a similarity measurement method.

2. The cross-modal representation learning and retrieval method for grain production according to claim 1, characterized in that, The feature extraction module performs feature extraction and feature concatenation on the image and text in the image-text pair to obtain the context features of each word in the text and the multimodal hybrid features of each image region, specifically including: The image is used to extract features using a pre-trained DeepLab ResNet-101v2 network to obtain the visual feature vector of each image region; the visual feature vectors of different image regions correspond to the backbone output of the last three layers in the DeepLab ResNet-101v2 network. A relative spatial coding network is used to generate positional features for each region of the image, thereby obtaining the spatial coordinate features of each image region; A long short-term memory network is used to extract the contextual features of each word in the text; For any image region, a feature concatenation network is used to concatenate the visual feature vector of the image region, the spatial coordinate features of the image region, and the contextual features of the last word in the text to obtain the multimodal hybrid features of the image region.

3. The cross-modal representation learning and retrieval method for grain production according to claim 1, characterized in that, Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, the image-text bidirectional guided fusion module performs image-text bidirectional guided fusion to obtain enhanced features for each image region, specifically including: Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, a visually guided language attention mechanism is used to determine the language contextual features of each image region. Based on the linguistic context features and multimodal hybrid features of each image region, a language-guided visual attention mechanism is used to determine the enhancement features of each image region.

4. The cross-modal representation learning and retrieval method for grain production according to claim 1, characterized in that, The enhanced features of each image region are fused at multiple scales using the multi-scale fusion module to obtain a fused feature map, specifically including: The hollow spatial pyramid pooling submodule is used to perform multi-scale semantic modeling on the enhanced features of each image region, extract contextual information under different receptive fields, and obtain multimodal features at different levels; The bidirectional gating fusion submodule is used to perform bidirectional gating fusion of multimodal features at different levels through top-down and bottom-up paths to obtain a fused feature map.

5. The cross-modal representation learning and retrieval method for grain production according to claim 1, characterized in that, Based on global semantic guidance, image spatial decoupling and temporal enhancement are performed on video-text pairs in the grain production process to obtain structured semantic image features, specifically including: The video in the video text pair is sampled to obtain multiple frames of remote sensing images, and global semantic extraction and spatial feature extraction are performed on the multiple frames of remote sensing images to obtain a global semantic vector and a spatial feature representation of each frame of remote sensing image; For any frame of remote sensing image, the spatial feature representation of the remote sensing image is decoupled based on the global semantic vector to obtain the high semantic saliency region features and low semantic saliency region features of the remote sensing image. Based on the high semantic saliency region features and low semantic saliency region features of multiple frames of remote sensing images, temporal enhancement modeling is performed on each frame of remote sensing images to obtain the forward saliency region features, backward saliency region features, forward cumulative low saliency features, and backward cumulative low saliency features of each frame of remote sensing images. The semantic enhancement feature representation of each frame of remote sensing image is obtained based on the forward salient region features and the backward salient region features of each frame of remote sensing image. Based on the forward cumulative low significance features and backward cumulative low significance features of the last frame of remote sensing image, the final cumulative feature representation is obtained; The semantic enhancement feature representation of each frame of remote sensing image is fused with the final accumulated feature representation to obtain structured semantic image features.

6. The cross-modal representation learning and retrieval method for grain production according to claim 5, characterized in that, Global semantic extraction and spatial feature extraction are performed on multiple frames of remote sensing images to obtain a global semantic vector and a spatial feature representation of each frame of remote sensing image, specifically including: Temporal average pooling and global average pooling operations are performed on multiple frames of remote sensing images to extract global semantic vectors; A pre-trained deep convolutional neural network is used to extract features from each frame of remote sensing image to obtain the spatial feature representation of each frame of remote sensing image.

7. The cross-modal representation learning and retrieval method for grain production according to claim 5, characterized in that, Based on the global semantic vector, the spatial feature representation of the remote sensing image is decoupled to obtain high semantic saliency region features and low semantic saliency region features of the remote sensing image, specifically including: The global semantic vector is linearly transformed and expanded into a feature tensor with the same spatial size as the remote sensing image; The feature tensor is concatenated with the spatial feature representation of the remote sensing image to obtain a fused feature. Based on the fusion features, a salient response map is generated using convolutional layers and activation functions; Based on the saliency response map, the spatial feature representation of the remote sensing image is decoupled by saliency guidance to obtain the high semantic saliency region features and low semantic saliency region features of the remote sensing image.

8. The cross-modal representation learning and retrieval method for grain production according to claim 5, characterized in that, Based on the high semantic saliency region features and low semantic saliency region features of multiple frames of remote sensing images, temporal enhancement modeling is performed on each frame of remote sensing images to obtain the forward saliency region features, backward saliency region features, forward cumulative low saliency features, and backward cumulative low saliency features of each frame of remote sensing images, specifically including: Initial cumulative low-saliency features are generated based on the low semantic saliency region features of multiple frames of remote sensing images; A forward-enhancing memory unit is used to model the difference between the high semantic saliency region features of the a-th frame remote sensing image and the forward cumulative low saliency features of the a-1 frame remote sensing image, and to extract the forward semantic change information of the a-th frame remote sensing image. The channel attention mechanism is used to extract the response weights of the forward semantic change information of the a-th frame remote sensing image, and the high semantic saliency region features of the a-th frame remote sensing image are enhanced by time series comparison based on the response weights to obtain the forward saliency region features of the a-th frame remote sensing image. Backward enhancement memory units are used to model the differences between the high semantic saliency region features of the a-th frame remote sensing image and the backward cumulative low saliency features of the a-1 frame remote sensing image, and to extract the backward semantic change information of the a-th frame remote sensing image. The channel attention mechanism is used to extract the response weights of the backward semantic change information of the a-th frame remote sensing image, and the high semantic saliency region features of the a-th frame remote sensing image are enhanced by time series comparison based on the response weights to obtain the backward saliency region features of the a-th frame remote sensing image. Based on the cumulative low saliency features of the (a-1)th frame remote sensing image and the low semantic saliency region features of the a frame remote sensing image, the forward cumulative low saliency features and backward cumulative low saliency features of the a frame remote sensing image are determined. Where, 0 < a ≤ A, A is the number of frames of the remote sensing images. When a = 1, both the forward cumulative low-significance features and the backward cumulative low-significance features of the (a - 1)-th frame of remote sensing images are the initial cumulative low-significance features.

9. A cross-modal representation learning and retrieval system for grain production, characterized in that, The system applies the cross-modal representation learning and retrieval method for food production according to any one of claims 1-8. The system includes: An image-text learning module, configured to perform multi-granularity semantic alignment on image-text pairs in the food production process based on a pre-trained graph-text bidirectional guidance fusion network to obtain a semantic segmentation image. The graph-text bidirectional guidance fusion network includes a feature extraction module, a graph-text bidirectional guidance fusion module, a multi-scale fusion module, and a semantic segmentation module; A video-text learning module, configured to perform image space decoupling and temporal enhancement on video-text pairs in the food production process based on global semantic guidance to obtain structured semantic image features. The video in the video-text pairs is a sequence of remote sensing images with temporality; A feature library construction module, configured to construct a text feature library and an image feature library according to the image-text pairs, the video-text pairs, the semantic segmentation image, and the structured semantic image features; A cross-modal retrieval module, configured to determine a transmission plan matrix according to the modality of the data to be retrieved, generate a query feature of the data to be retrieved based on the transmission plan matrix, and output a text query result or an image query result by using a similarity measurement method according to the query feature of the data to be retrieved, the text feature library, and the image feature library.

Citation Information

Patent Citations

  • Image text retrieval method and system based on context-guided multi-modal association

    CN116737979A

  • Construction safety risk early warning method and system based on cross-modal visual language retrieval

    CN120146549A