Cross-modal representation learning and retrieval method and system for grain production

Through the bidirectional guided fusion network of images and text and the global semantic guidance method, combined with the optimal transmission theory, the semantic alignment and fusion problems of agricultural cross-modal data were solved, and the fast and accurate matching and retrieval of images and texts were achieved, thereby improving the efficiency of cross-modal data utilization in grain production.

CN120705355AActive Publication Date: 2025-09-26AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI

Patent Information

Application Number
CN202511195014.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-09-26
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies lack effective cross-modal data fusion methods in the agricultural field, making it difficult to achieve fine-grained semantic alignment and efficient retrieval between images and text, resulting in poor application of cross-modal data in food production.

Method used

A bidirectional guided fusion network of images and texts is used for multi-granularity semantic alignment. Combining global semantic guidance and optimal transmission theory, a cross-modal representation learning and retrieval system is constructed. Through feature extraction and fusion of image-text pairs and video-text pairs, fast and accurate matching and retrieval of images and texts are achieved.

Benefits of technology

It improves the semantic matching accuracy and retrieval efficiency between images and texts, enhances the accuracy of cross-modal feature fusion, supports fast and accurate matching and retrieval between images and texts, and improves the response speed and information accuracy of intelligent agricultural applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705355A_ABST
    Figure CN120705355A_ABST
Patent Text Reader

Abstract

The invention discloses a grain production-oriented cross-modal representation learning and retrieval method and system, and relates to the field of agricultural informationization, and the method comprises the steps: carrying out the multi-granularity semantic alignment of an image text pair in a grain production process based on an image-text bidirectional guidance fusion network, and obtaining a semantic segmentation image; performing image space decoupling and time sequence enhancement on a video text pair in a grain production process based on global semantic guidance to obtain a structured semantic image feature; constructing a text feature library and an image feature library; and determining a transmission plan matrix according to the modality of the to-be-retrieved data, generating query features of the to-be-retrieved data based on the transmission plan matrix, and outputting a text query result or an image query result by adopting a similarity measurement method according to the query features of the to-be-retrieved data, the text feature library and the image feature library. According to the method, deep fusion of cross-modal features can be realized, the semantic matching accuracy of the image and the text is improved, and rapid and accurate matching and retrieval between the image and the text are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of agricultural informatization, and in particular to a cross-modal representation learning and retrieval method and system for food production. Background Art

[0002] Throughout the grain production process, the acquisition and effective utilization of multi-source data are crucial for improving agricultural production efficiency and achieving precision management. With the rapid development of remote sensing and information technologies, the agricultural sector has accumulated a vast amount of heterogeneous, multimodal data resources, such as remote sensing imagery, crop monitoring images, meteorological records, pest and disease analysis reports, and expert advice. Image data provides rich spatial visual information, while text data provides detailed background descriptions and contextual semantics. The two are highly complementary. Leveraging this multimodal data will facilitate intelligent perception and precise decision-making across the entire agricultural production process.

[0003] However, traditional agricultural data processing methods mostly focus on single-modality processing and lack the ability to model the deep semantic connections between images and text, resulting in poor results in cross-modal data analysis and retrieval. This is particularly true in the field of grain production, where data from different sources vary significantly in scale. Remote sensing images possess spatiotemporal continuity, while text data has diverse expressions. Existing data-driven machine learning algorithms struggle to establish effective, unified semantic representations across multi-scale and multi-modal data, further limiting the intelligent application of agricultural big data.

[0004] Furthermore, there's a "heterogeneity gap" between agricultural images and text. This is due to differences in distribution and feature expression between modalities, which prevents direct alignment of cross-modal features, reducing retrieval efficiency and comprehension accuracy. Related technologies for the fusion and utilization of image and text information primarily focus on the following approaches: One approach prioritizes images, with text as additional annotation and performing only simple keyword matching; the other employs deep neural networks to extract image and text features separately, achieving alignment through attention mechanisms or semantic matching algorithms.

[0005] The above methods have made some progress in general image-text retrieval, image annotation and other tasks, but their applicability in agricultural scenarios has the following shortcomings: (1) Difficulty in granular alignment: Agricultural images often contain multiple complex targets (such as crops, disease spots, and agricultural machinery), and text descriptions may involve multiple objects and spatiotemporal backgrounds. Existing methods mostly stay at the overall level of matching, making it difficult to achieve fine-grained semantic alignment. (2) Lack of a unified cross-modal expression mechanism: Agricultural data often comes from diverse sources and has different organizational forms. There is a lack of an effective mechanism for unified modeling and shared representation of different modal information, which limits the fusion and utilization of data and the implementation of downstream retrieval and analysis tasks. (3) Low retrieval efficiency: Faced with large-scale agricultural multimodal data, existing retrieval methods are difficult to balance matching efficiency and accuracy, and it is difficult to meet the requirements of agricultural intelligent applications for response speed and information accuracy.

[0006] Therefore, it is urgent to propose a cross-modal feature alignment and fusion method suitable for agricultural image and text data scenarios to achieve deep alignment of images and texts at the semantic level, construct a unified feature space, and improve the retrieval efficiency and comprehension ability of large-scale image and text data. Summary of the Invention

[0007] The purpose of this application is to provide a cross-modal representation learning and retrieval method and system for food production, which can effectively capture the fine-grained semantic relationship between different modal data such as images, videos and texts in the food production process, realize the deep fusion of cross-modal features, improve the accuracy of semantic matching between images and texts, and realize fast and accurate matching and retrieval between images and texts.

[0008] To achieve the above objectives, this application provides the following solutions: In the first aspect, this application provides a cross-modal representation learning and retrieval method for food production, including: Based on a pre-trained image-text bidirectional guided fusion network, multi-granularity semantic alignment is performed on image and text pairs in the grain production process to obtain semantic segmentation images; the image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module; Based on global semantic guidance, the video text pairs in the grain production process are subjected to image space decoupling and temporal enhancement to obtain structured semantic image features; the video in the video text pair is a remote sensing image sequence with temporal sequence; Constructing a text feature library and an image feature library based on the image-text pairs, the video-text pairs, the semantic segmentation images, and the structured semantic image features; A transmission plan matrix is ​​determined according to the modality of the data to be retrieved, and query features of the data to be retrieved are generated based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

[0009] Secondly, this application provides a cross-modal representation learning and retrieval system for food production, including: An image-text learning module is used to perform multi-granular semantic alignment of image-text pairs in the grain production process based on a pre-trained image-text bidirectional guided fusion network to obtain semantically segmented images. The image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module. A video-text learning module is used to perform image space decoupling and temporal enhancement on video-text pairs of the grain production process based on global semantic guidance to obtain structured semantic image features; the video in the video-text pair is a remote sensing image sequence with temporal sequence; A feature library construction module, configured to construct a text feature library and an image feature library based on the image-text pairs, the video-text pairs, the semantically segmented images, and the structured semantic image features; The cross-modal retrieval module is used to determine a transmission plan matrix according to the modality of the data to be retrieved, and generate query features of the data to be retrieved based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

[0010] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a cross-modal representation learning and retrieval method and system for food production. By designing a bidirectional guided fusion network of images and text, multi-granularity semantic alignment between images and text is achieved, which effectively enhances the consistent expression of language information and visual information in the feature space and improves the accuracy of cross-modal feature fusion. Based on global semantic guidance, the video-text pairs in the food production process are decoupled in image space and enhanced in time sequence, which fully explores the spatiotemporal feature correlation in remote sensing images and improves the modeling and generalization capabilities of key agricultural semantics. The optimal transmission theory is further introduced to optimize feature mapping to achieve fast and accurate matching and retrieval between images and text. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0012] Figure 1 A flowchart of a cross-modal representation learning and retrieval method for food production provided in one embodiment of the present application.

[0013] Figure 2 This is a schematic diagram of the overall structure of the image and text bidirectional guidance fusion network in one embodiment of the present application.

[0014] Figure 3 Schematic diagram of the process of spatial decoupling and temporal enhancement of remote sensing images guided by global semantics in one embodiment of the present application.

[0015] Figure 4 Schematic diagram of the framework for optimally transmitted cross-modal unified representation learning and efficient bidirectional retrieval of grain production images and texts in one embodiment of the present application.

[0016] Figure 5 A schematic diagram of the functional modules of a cross-modal representation learning and retrieval system for food production provided in one embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] This application aims to solve the problem of "heterogeneous gap" in feature expression of multimodal data such as massive remote sensing images, near-ground images and agricultural texts generated throughout the entire process of grain production. Starting from the three dimensions of cross-modal deep fusion modeling, spatiotemporal enhancement understanding of time series images and unified semantic expression of heterogeneous data, a set of end-to-end deployable, cross-modal strong correlation and high semantic expression consistency image and text retrieval methods are formed to improve the fusion expression ability and intelligent retrieval efficiency of cross-modal data. To achieve the above purpose, this application uses single image feature extraction based on convolutional neural networks, fusion of language features, visual features and spatial coordinate features based on a bidirectional guided fusion network of images and texts, global guided spatial decoupling and temporal enhancement based on images, minimizing the transmission cost of the original representation and the unified representation to reduce the semantic loss generated by the unified representation during the compression process, and other methods to achieve unified correlation representation and efficient retrieval of cross-modal content from the perspective of data-driven learning and generalization.

[0019] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] In an exemplary embodiment, Figure 1 As shown, a cross-modal representation learning and retrieval method for food production is provided. The method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or can be executed jointly by a terminal and a server. In an embodiment of the present application, the cross-modal representation learning and retrieval method for food production includes the following steps 1 to 4.

[0021] In step 1, a multi-granularity semantic alignment method based on a bidirectional guided fusion network for image and text is constructed. This addresses the problem that existing cross-modal interaction methods generally rely on manually designed interaction patterns, lack flexibility, and struggle to achieve efficient guidance and deep fusion of image and text information, resulting in low cross-modal understanding accuracy and limited feature representation capabilities. The bidirectional guided fusion network for image and text achieves multi-granularity semantic alignment between images and text. The specific method involves first extracting visual features from the input image using a convolutional neural network, encoding the natural language text using a recurrent neural network to obtain contextual features of the text. Simultaneously, spatial coordinate features are generated for each image location to enhance the image's positional information representation capabilities. In the feature fusion stage, contextual, visual, and spatial coordinate features are concatenated and integrated, and cross-modal feature inference is performed using a bidirectional guided fusion module for image and text. In this module, the visual information is guided by the text to enhance its semantic expression, while the text is further accurately described using semantic cues from the image, thereby enhancing semantic consistency and feature complementarity between the image and text. This multi-granularity semantic alignment mechanism effectively captures both fine-grained and global semantic correspondences between image and text. Finally, the fused features are classified into target areas based on a fully convolutional network, and the results are upsampled through a deconvolution module to obtain pixel-level region segmentation results corresponding to the text description content in the image. This can significantly improve the understanding and matching accuracy of cross-modal information between images and texts in the food production process, and promote efficient linkage and joint analysis between image and text data.

[0022] In step 2, a global semantic guidance-based spatial decoupling and temporal enhancement method for remote sensing images is constructed to improve the representation of key regions and the perception of spatiotemporal changes. The specific method includes: First, in the spatial feature extraction stage, a global semantic guidance mechanism is introduced. By integrating global semantic vectors extracted from multiple temporal remote sensing images, the spatial feature representation of the current frame is saliency-guided and structurally decoupled, dynamically distinguishing high-saliency regions from low-saliency regions. This operation helps highlight regional features highly relevant to agricultural mission objectives in remote sensing images while suppressing redundant background information, enhancing the discriminability and interpretability of the image semantic structure. Second, in the temporal modeling stage, a temporal enhancement mechanism is designed, employing differentiated modeling strategies for spatial regions of varying semantic saliency. For high-saliency regions, this strategy leverages their prior advantages in semantic space and combines semantic comparison with low-saliency regions to further enhance regional discrimination. For low-saliency regions, an inter-frame accumulation strategy within the image sequence is employed to capture fine-grained temporal variation trends, enhance their dynamic representation, and avoid the loss of important temporal information. Finally, through the synergistic mechanism of spatial decoupling and temporal enhancement, we can achieve highly expressive modeling of key areas in agricultural remote sensing images and deep information mining in the temporal dimension, significantly improving the model's perception and expression capabilities of spatial heterogeneity and temporal evolution in agricultural scenarios. This provides a solid spatiotemporal semantic foundation for cross-modal representation learning, further assisting in the semantic alignment of images and texts and the precise execution of downstream intelligent tasks.

[0023] Through step 4, a cross-modal unified representation learning and efficient bidirectional retrieval mechanism for grain production images and text based on optimal transfer is constructed. This addresses the problem of multi-source heterogeneous modal data in the grain production process, including remote sensing imagery, ground monitoring imagery, and agricultural text records. These data exhibit significant differences in representation, structural dimensions, and semantic abstraction levels, leading to a "semantic gap" between modalities and making semantic alignment and joint modeling difficult. This approach enables the fused representation and efficient matching of multimodal grain production data in a unified semantic space. The specific method includes: First, the representation process of heterogeneous modal data is unified and modeled as an optimal mapping problem from the original modal feature space to a shared unified semantic space. To achieve effective alignment between modalities, optimal transfer theory is introduced, and a transfer cost function for inter-modal alignment is constructed. By minimizing the transfer distance between the original visual / textual features and the target unified representation, the semantic loss caused by feature compression and fusion is effectively reduced, improving the consistency and discriminability of multimodal semantic representation. Secondly, during model training and inference, a fast-converging and well-structured joint optimization strategy is designed to achieve collaborative learning of the modality mapping matrix and the representation decoder, enhancing the model's adaptability and generalization capabilities for complex and diverse agricultural scenarios. This mechanism can dynamically adjust the semantic weights between different modalities to achieve robust representation in diverse scenarios. Finally, a unified representation mechanism supports bidirectional image-text retrieval processes, enabling "image-to-text" and "text-to-image" bidirectional matching queries in a shared semantic space. This allows for efficient retrieval of image-text pairs semantically related to agricultural tasks within large-scale agricultural multimodal data, providing accurate and reliable information support for tasks such as intelligent identification of pests and diseases, dynamic monitoring of seedling conditions, and automatic understanding of agricultural behavior, significantly improving the integrated management and intelligent service capabilities of agricultural cross-modal data.

[0024] The details of each step are described below.

[0025] Step 1: Perform multi-granular semantic alignment on image and text pairs in the grain production process based on a pre-trained image-text bidirectional guided fusion network to obtain a semantically segmented image. The image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module.

[0026] Specifically, the processing process of the image-text bidirectional guidance fusion network is as follows: Figure 2 Step 1 includes the following steps 11 to 14.

[0027] Step 11: The feature extraction module performs feature extraction and feature concatenation on the image and text in the image-text pair, respectively, to obtain contextual features of each word in the text and multimodal mixed features of each image region, wherein different image regions correspond to different layers of the image.

[0028] In a specific application example, the feature extraction module includes a DeepLab ResNet-101 v2 network, a relative spatial encoding network, a long short-term memory network, and a feature concatenation network. Step 11 includes the following steps 111 to 114.

[0029] Step 111: Use a pre-trained DeepLab ResNet-101 v2 network to perform feature extraction on the image to obtain a visual feature vector for each image region. The visual feature vectors for different image regions correspond to the backbone outputs of the last three layers of the DeepLab ResNet-101 v2 network.

[0030] Specifically, the images were first uniformly scaled and zero-padded to 320×320 to ensure consistency of the network input. Subsequently, the images were fed into a pre-trained DeepLab ResNet-101 v2 network to extract multi-scale visual features. The outputs of the last three backbone layers of the network (Res3, Res4, and Res5) were selected as multi-granular visual semantic representations to capture both the local texture and global semantics of the image.

[0031] Among them, Res3 extracts mid-level semantic features, which mainly include local texture information and edge structure, and has a strong ability to retain details. Its output feature map dimension is 40×40×512, that is, the spatial resolution is 40×40, and the number of channels is is 512, the visual feature vector at each spatial position The dimension is 512.

[0032] Res4 extracts mid- to high-level semantic features and captures more abstract regional semantics. Its output feature map dimension is 20×20×1024, the spatial resolution is reduced by half, and the number of channels is is 1024. The visual feature vector at each position The dimension is 1024.

[0033] Res5 extracts the deepest semantic features and contains the richest global semantic information, but has the lowest spatial resolution. Its output feature map dimension is 10×10×2048, with a channel number of is 2048. The visual feature vector at each position The dimension is 2048.

[0034] The visual feature is a complete three-dimensional tensor output by a layer in the DeepLab ResNet-101 v2 network, with dimensions H×W×C (height×width×number of channels). The visual feature vector refers to the feature representation of a spatial position (pixel) in the feature map in terms of the number of channels, which is used to describe the semantic information of the corresponding area in the image, denoted as ;in, That is thei The visual feature vector of an image region.

[0035] Step 112: Generate position features for each region of the image using a relative spatial encoding network to obtain spatial coordinate features of each image region.

[0036] In order to enhance the spatial position perception ability of the target area of ​​the image, this application uses the relative spatial encoding method to generate position features for each area, which is usually represented by an 8-dimensional vector. , .in, For the i The spatial coordinate features of an image region.

[0037] Step 113: Use a Long Short-Term Memory (LSTM) network to extract the text. Each word in Contextual features .in, is the total number of words in the text (such as harvester, in the field, harvest wheat). For the T The context features of the last word form the context feature vector and output the final hidden state of LSTM To represent the semantics of the entire sentence.

[0038] Step 114: For any image region, a feature stitching network is used to stitch the visual feature vector of the image region, the spatial coordinate features of the image region, and the context features of the last word in the text to obtain a multimodal hybrid feature of the image region: ;in, For the i Multimodal mixed features of image regions.

[0039] Step 12: Based on the contextual features of each word in the text and the multimodal mixed features of each image region, the image-text bidirectional guided fusion module is used to perform image-text bidirectional guided fusion to obtain enhanced features of each image region.

[0040] In a specific application example, the image-text bidirectional guided fusion module includes a visually guided language attention mechanism and a language-guided visual attention mechanism. The visually guided language attention mechanism uses the contextual features of each word in the text as a guide to measure the semantic contribution of each word to each image region. The language-guided visual attention mechanism uses the language contextual features of each image region as a guide to calculate the spatial semantic dependencies between image regions. Step 12 includes the following steps 121 and 122.

[0041] Step 121 : Based on the contextual features of each word in the text and the multimodal mixed features of each image region, a visually guided language attention mechanism is used to determine the language contextual features of each image region.

[0042] Specifically, visual features are used to guide text context modeling, and a weighted attention distribution is learned to measure the t Word pair i Semantic contribution of image regions: ; ;in, For the t Word pair i The semantic contribution of each image region, For the i The language context features of each image region are adaptive language representations under visual guidance. It is a learnable parameter of the convolution operation. Its main goal is to reduce the dimension of the mixed features input to the visually guided language attention module, reduce the number of parameters, and improve the inference speed.

[0043] In step 122 , based on the language context features of each image region and the multimodal mixed features of each image region, a language-guided visual attention mechanism is used to determine the enhanced features of each image region.

[0044] Specifically, the following formula is used to calculate the i The language context features of the image region are j The importance of language context features in each image region: ; ;in, For the j Multimodal mixed features of image regions, The main goal is to reduce the dimension of the mixed features input to the language-guided visual attention module, reduce the number of parameters, and improve the inference speed. and is a learnable parameter used to Perform dimensionality reduction to reduce the number of parameters in the language-guided visual attention module. For the i The language context features of the image region are j The preliminary importance of language context features in each image region, For the i The language context features of the image region are j The importance of language context features in each image region, The higher the value, the iimage regions and j The stronger the correlation between image regions.

[0045] Use the following formula to get the i Enhanced features of image regions : ;in, B To exclude i image area outside the image area, that is, if i =3, then j =4 or 5, if i =4, then j =3 or 5, if i =5, then j =3 or 4, 、 、 These are all learnable parameters. Their primary goal is to upgrade the updated visual features and concatenate them with adaptive language features to complement contextual language features. The language-guided visual attention mechanism can enhance the structured contextual relationships between image regions, enabling the adaptive guidance of language semantics on the perception of image regions.

[0046] Step 13: Perform multi-scale fusion on the enhanced features of each image region through the multi-scale fusion module to obtain a fused feature map.

[0047] In a specific application example, the multi-scale fusion module includes a dilated spatial pyramid pooling submodule and a bidirectional gated fusion submodule. Step 13 includes the following steps 131 to 132.

[0048] Step 131: Use the Atrous Spatial Pyramid Pooling (ASPP) submodule to perform multi-scale semantic modeling on the enhanced features of each image region, extract context information under different receptive fields, and obtain multimodal features at different levels. ;in, For the i The multimodal features of the image regions correspond to the feature maps output by Res3, Res4, and Res5 of the DeepLab ResNet-101 v2 network.

[0049] In step 132 , the bidirectional gated fusion submodule is used to perform bidirectional gated fusion (BDGF) on the multimodal features at different levels through top-down and bottom-up paths to obtain a fused feature map.

[0050] The specific fusion operations are as follows: ; ; ; ; ; in, for and The fused feature map, for and The fused feature map, is the fusion feature map obtained from the top-down path, is the fusion feature map obtained from the bottom-up path, is the final fusion feature map, and These are gating functions for the upward and downward directions after the feature maps of different layers are spliced ​​along the channel dimension. They are composed of 3×3 convolution and Sigmoid activation function, which are used to regulate the information flow and fusion between multiple layers of features. and It is the feature map from different layers in the network, such as the output of Res3, Res4 and other layers, It will and Splicing along the channel dimension, splicing the features of the two layers to merge the information, To perform a convolution operation on the concatenated feature map to extract the fused local features, Sig is to use the Sigmoid function on the convolution result and output the gating coefficient (usually between [0,1]), which indicates the "pass rate" or "importance" of the feature flow. The direction of information flow is from bottom to top. D The direction of information flow is from top to bottom. is element-wise multiplication.

[0051] Step 14: Perform pixel-by-pixel semantic classification on the fused feature map through the semantic segmentation module to obtain a semantic segmentation image.

[0052] Specifically, a full convolution classification network is used to fusion feature maps Perform pixel-by-pixel semantic classification to predict image regions that are semantically consistent with the input text, generating a classification feature map. A deconvolution module is then used to upsample the classification feature map to generate a semantic segmentation image with the same resolution as the original image, ultimately achieving accurate segmentation and labeling of semantically targeted regions in grain production images.

[0053] In step 1, through key steps such as bidirectional guidance of images and text, attention mechanism modeling, and multi-scale fusion, the semantic gap between images and text is bridged, effectively improving the accuracy of cross-modal alignment and retrieval, and providing core support for intelligent perception and decision-making in food production scenarios.

[0054] Step 2: Based on global semantic guidance, the video text pairs of the grain production process are subjected to image space decoupling and temporal enhancement to obtain structured semantic image features. The video in the video text pair is a remote sensing image sequence with temporal sequence.

[0055] Specifically, if Figure 3 As shown, step 2 includes the following steps 21 to 26.

[0056] Step 21: Sample the video in the video-text pair to obtain multiple frames of remote sensing images, and perform global semantic extraction and spatial feature extraction on the multiple frames of remote sensing images to obtain a global semantic vector and a spatial feature representation of each frame of remote sensing image.

[0057] In a specific application example, step 21 includes the following steps 211 to 213.

[0058] Step 211, using restricted random sampling strategy (Restricted Random Sampling) to extract multiple frames of remote sensing images from the continuous remote sensing image sequence, expressed as ,in, A is the number of frames of remote sensing images (e.g. the number of days or months of remote sensing observations), is the height of the image, is the width of the image, is the number of channels of the image (such as 3 for RGB).

[0059] Step 212: Perform temporal average pooling and global average pooling operations on multiple frames of remote sensing images to extract the global semantic vector : ;in, For the Frame remote sensing image. It is obtained by averaging all pixels in time and space, and serves as an abstract global semantics across time and space, representing the core semantics of the image in the entire time series and guiding the subsequent local feature selection.

[0060] Step 213: Use a pre-trained deep convolutional neural network (such as ResNet-50) to extract features from each frame of remote sensing image to obtain a spatial feature representation of each frame of remote sensing image. ;in, For the Spatial feature representation of frame remote sensing images, .

[0061] Step 22: for any frame of remote sensing image, decouple the spatial feature representation of the remote sensing image based on the global semantic vector to obtain high semantic saliency region features and low semantic saliency region features of the remote sensing image.

[0062] In a specific application example, step 22 includes the following steps 221 to 224 .

[0063] Step 221: The global semantic vector Perform linear transformation and expand it into a feature tensor consistent with the spatial size of the remote sensing image , which is used to guide attention learning after being spliced ​​with local spatial features.

[0064] Step 222: transform the feature tensor Spatial feature representation of the remote sensing image Cascade to obtain fusion features .in Indicates concatenating two tensors in the channel dimension for feature concatenation.

[0065] Step 223: Generate a saliency response map using a convolutional layer and an activation function based on the fused features.

[0066] Specifically, the fusion feature input consists of two The guided correlation estimation module composed of convolution, batch normalization and ReLU, and then generates a significant response map through the Sigmoid activation function: ;in, Representative Saliency response map of frame remote sensing image, each spatial position represents its semantic importance, Represents the Sigmoid activation function, which ranges from (0, 1) and is used to generate the saliency weight map. Because of two Convolutional layers consist of learnable weights, typically lightweight convolutional or attention networks with learnable parameters.

[0067] Step 224 : performing saliency-guided decoupling on the spatial feature representation of the remote sensing image according to the saliency response map to obtain high semantically significant region features and low semantically significant region features of the remote sensing image.

[0068] Specifically, the formula Get the first High semantic saliency region features of frame remote sensing images, , using the formula Get the first Low semantic saliency region features of frame remote sensing images, .in, Represents an element-by-element multiplication operation used to weight remote sensing images according to the saliency response map. This operation dynamically enhances regions in remote sensing images that are significantly relevant to agricultural tasks. High semantically salient region features represent regions in the image that are more strongly related to agricultural tasks, while low semantically salient region features represent redundant or background regions in the image.

[0069] Step 23: Based on the high semantic saliency region features and low semantic saliency region features of the multiple frames of remote sensing images, temporal enhancement modeling is performed on each frame of the remote sensing image to obtain the forward saliency region features, backward saliency region features, forward cumulative low saliency features, and backward cumulative low saliency features of each frame of the remote sensing image.

[0070] In a specific application example, step 23 includes the following steps 231 to 236.

[0071] Step 231: Generate initial cumulative low-saliency features based on low semantically significant region features of multiple frames of remote sensing images: ;in, is the initial cumulative low significance feature.

[0072] Step 232, using the forward enhancement memory unit to High semantic saliency region features of frame remote sensing images and the The forward accumulation of low-significance features of the frame remote sensing image is used to perform difference modeling and extract the first Forward semantic change information of frame remote sensing images: ;in, For the Forward semantic change information of frame remote sensing images, For the Forward accumulation of low-significance features of frame remote sensing images, For the High semantic saliency region features of frame remote sensing images, and Represents two different 1×1 convolutional layers.

[0073] Step 233, using the channel attention mechanism, extract the The response weight of the forward semantic change information of the frame remote sensing image is The high semantic saliency region features of the frame remote sensing image are enhanced by time series contrast, and the first Forward salient region features of frame remote sensing images.

[0074] Specifically, for the Frame remote sensing image, first Perform spatial average pooling and then use the formula Extract response weights; where, For the Response weight of frame remote sensing image, for Information for spatial average pooling, is the parameter weight of the linear transformation. Then use the formula Get the first Contrast enhancement features of frame remote sensing images , and then obtain the forward salient region features based on the contrast enhancement features .

[0075] Step 234, using the backward enhancement memory unit to High semantic saliency region features of frame remote sensing images and the The backward accumulation of low-significance features of frame remote sensing images is used for difference modeling to extract the first Backward semantic change information of frame remote sensing images.

[0076] Step 235, using the channel attention mechanism, extract the The response weight of the backward semantic change information of the frame remote sensing image is used, and the first The high semantic saliency region features of the frame remote sensing image are enhanced by time series contrast, and the first Backward salient region features of frame remote sensing images.

[0077] The processing of steps 234 and 235 is similar to that of steps 232 and 233, and will not be repeated here. The forward salient region feature of the frame remote sensing image is , No. The backward salient region feature of the frame remote sensing image is .

[0078] Step 236, according to Cumulative low-saliency features of frame remote sensing images and the The low semantic saliency region features of the frame remote sensing image are used to determine the Forward accumulation of low-significance features and backward accumulation of low-significance features of frame remote sensing images : .in, is the residual block.

[0079] in, , A is the number of frames of remote sensing images, Time Forward accumulation of low-significance features of frame remote sensing images and the first The backward accumulated low-saliency features of the frame remote sensing image are all the initial accumulated low-saliency features.

[0080] Step 24: Obtain semantic enhancement feature representation of each frame of remote sensing image based on the forward salient region features and the backward salient region features of each frame of remote sensing image.

[0081] Specifically, and After global average pooling and concatenation, the final semantic enhancement feature representation is generated through the fully connected layer: , ;in, For the Semantic enhancement feature representation of frame remote sensing images, is the learnable parameter of global average pooling.

[0082] Step 25: Based on the forward accumulated low-significance features of the last frame of remote sensing image and backward accumulation of low-significance features , and the final cumulative feature representation is obtained: ;in, is the final cumulative feature representation, are learnable parameters for linear optimization.

[0083] Step 26: fusing the semantic enhancement feature representation of each frame of remote sensing image with the final accumulated feature representation to obtain structured semantic image features.

[0084] Specifically, the The semantic enhancement feature representation of the frame remote sensing image is fused with the final cumulative feature representation to obtain structured semantic image features for agricultural production scenes. , this feature can be aligned with the semantic representation in the text modality for cross-modal contrastive learning or retrieval tasks.

[0085] In another exemplary embodiment, step 2 further includes step 25 of training the parameters involved in steps 21 to 26 based on the loss function.

[0086] In a specific application example, step 25 includes the following steps 251 to 254 .

[0087] S251: For and , use the image-level online instance matching loss function and the video-level online instance matching loss function for feature supervision, and use the verification loss to supervise the similarity between remote sensing image sequences, which is used to measure the similarity judgment error between image-image (or video) pairs and whether they are from the same class. The verification loss function is expressed as: ;in, To verify the loss function value, To measure the similarity judgment error between image-image (or video) pairs, whether they are from the same class, For the n feature vectors of the image / video to be matched, No. n feature vectors of the contrasted images / videos, is a similarity function that represents the similarity score between features, usually cosine similarity or sigmoid mapping of embedding space distance, Represents the label value, if and Belong to the same category, then =1 if the value is set to true, otherwise 0.

[0088] S252: Yes The frame-level online instance matching (OIM) loss is used to enhance regional discrimination. The frame-level OIM loss function is applied to supervised learning of semantic features of each frame in the image sequence, which is expressed as: ;in, is the frame-level OIM loss function value, M is the number of remote sensing image sequences, R is the total number of categories (e.g. the number of different objects / regions / crop types in the training set), Indicates the In the remote sensing image sequence Semantic enhancement features of frame images (from high saliency regions). Indicates the r The weight vector corresponding to the category is used to represent the category in the embedding space, and its value is updated online according to the image features during training. , if The first remote sensing image sequence Frame belongs to r If the field is a certain type of land (such as an agricultural plot or a semantic category), it is 1, otherwise it is 0.

[0089] S253: Yes Sequence-level OIM loss is used to capture the fine-grained semantic evolution trend. Remote sensing sequence-level OIM loss is used for global aggregation representation supervision of low-significance region features at the end of the sequence (such as the last frame), which is expressed as: ;in, is the sequence-level OIM loss function value, , if The remote sensing image sequence belongs to r class, then it is 1, otherwise it is 0.

[0090] S254: Combining the above three losses, construct the total loss function of the joint training objective function: ;in, is the total loss function value, They are all weight hyperparameters, which control the proportion of the influence of different loss terms on the final objective function and need to be adjusted through experiments.

[0091] Step 3: construct a text feature library and an image feature library based on the image-text pairs, the video-text pairs, the semantic segmentation images and the structured semantic image features.

[0092] Step 4: Determine a transmission plan matrix according to the modality of the data to be retrieved, and generate query features of the data to be retrieved based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

[0093] In order to achieve unified semantic expression and efficient bidirectional retrieval of multi-source heterogeneous remote sensing images, monitoring images and agricultural text records in the food production scenario, this application proposes a cross-modal unified representation learning and efficient food production image-text bidirectional retrieval mechanism based on optimal transmission through step 4. The overall framework is as follows: Figure 4 shown.

[0094] In a specific application example, step 4 includes the following steps 41 to 44.

[0095] Step 41 : Constructing an optimal transmission mapping. Specifically, step 41 includes the following steps 411 and 412 .

[0096] Step 411, assuming there is Type of modality (such as remote sensing images, monitoring images, text and other modalities of different data sources), The discrete distribution of the class mode in the feature space is ,in , the distribution vector of the unified semantic space is recorded as , The corresponding Transfer cost matrix between feature points of similar modalities Expressed as: ;in, Indicates that the In the class mode The feature points are mapped to the unified space The cost of a feature point, For the The number of characteristic points of the mode (that is, the number of support points of the mode in the original feature space). is the distribution vector of the original mode, For the In the class mode The distribution of feature points, is the distribution vector of the unified space, the sum of each element is 1, Q is the number of feature points in the unified semantic space (i.e., the number of support points in the final shared space).

[0097] Step 412: for each type of mode, The mode can be solved by the following optimal transmission. Optimal transmission plan matrix for quasi-modal Expressed as: ;in, is the transmission plan matrix, which represents the The mapping weights of the feature points of the class modality to the feature points of the unified space satisfy the feasible set of the transmission plan. , and Respectively represent the length Q and A vector of all 1s. is the constraint condition. The joint distribution obtained is It can be characterized by Migrate to The optimal joint distribution probability of . Different from the traditional fusion mechanism, Can clearly give the The relationship between class modalities and unified representation of various dimensions is established, and the results obtained have clear theoretical guarantees.

[0098] Step 42: Align and fuse the unified representation with the modal representation. Specifically, step 42 includes the following steps 421 and 422.

[0099] Step 421: The optimal transmission plan matrix of each mode Generate the projection representation of the modality in the unified space : ;in For the The original eigenvector matrix of the class mode, feature points, each point is d dimensional vector.

[0100] Step 422: weight the projection representations of all modalities Weighted fusion, indicating the The contribution ratio of the class mode to the final unified representation satisfies the non-negative and sums to 1, and the unified representation is obtained: , ;in, Represents the weighted fusion result of all modal projection representations, that is, the final unified representation matrix.

[0101] Step 43, retrieval model optimization. Specifically, step 43 includes the following steps 431 and 432.

[0102] Step 431, calculate the retrieval similarity. In the unified space, define the image query features Text retrieval library features The similarity is cosine similarity: In the unified space, A unified representation matrix (or vector set) representing the input image and text respectively. Is the inner product operation, used to calculate the cosine similarity numerator, Represents the Frobenius norm or Euclid norm of a vector or matrix, depending on the context, used for normalization. is the feature similarity function.

[0103] Step 432, search objective function. Use the combined optimal transmission loss and search contrast loss, the overall goal is for: ;in, is a set of positive sample pairs, containing all “correctly matched” image-text pairs, To retrieve the contrast loss weight coefficient, is the Sigmoid activation function, which is used to map the similarity to (0,1).

[0104] Step 433: End-to-end optimization. The above loss term is back-propagated and combined with the feature library constructed in step 3 to update the mapping network, feature encoder, and decoder components to achieve fast convergence and well-structured joint optimization.

[0105] Step 44 , image-text bidirectional search is performed. Specifically, step 44 includes the following steps 441 and 442 .

[0106] Step 441: Image search for text. Given the image query feature , based on the similarity in the text feature library Sorting, output the top Z most relevant items sorted by similarity.

[0107] Step 442: Text search image. Given a text query , sort by similarity on the image feature library, and output the top Z most relevant items sorted by similarity.

[0108] This retrieval process can achieve a response within seconds in large-scale agricultural multimodal databases and ensure high-precision matching.

[0109] This application addresses the "semantic gap" problem of multi-source, multi-modal, and multi-scale heterogeneous data in the agricultural production process, and constructs a unified cross-modal feature expression and efficient retrieval mechanism. By designing a bidirectional guided fusion network for images and text, multi-granular semantic alignment between images and text is achieved, effectively enhancing the consistent expression of language information and visual information in the feature space, and improving the accuracy of cross-modal feature fusion. Combined with the proposed remote sensing image spatial decoupling and temporal enhancement method based on global semantic guidance, the temporal and spatial feature correlations in remote sensing images are fully explored, and the model's modeling and generalization capabilities for key agricultural semantics are improved. In addition, a cross-modal unified representation and efficient retrieval mechanism for heterogeneous data in grain production is constructed, and the optimal transmission theory is introduced for feature mapping optimization to minimize the semantic loss in the inter-modal mapping process, thereby achieving fast and accurate matching and retrieval between images and text. Overall, this application can be widely applied to typical scenarios such as large-scale agricultural remote sensing monitoring, crop growth identification, and agricultural record analysis. It has significant application value and promotion prospects in improving the efficiency of agricultural information utilization and supporting intelligent decision-making and management.

[0110] Based on the same inventive concept, the embodiments of the present application also provide a cross-modal representation learning and retrieval system for food production for implementing the above-mentioned method. The implementation solution provided by this system is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations of one or more embodiments of the cross-modal representation learning and retrieval system for food production provided below can be found in the above-mentioned limitations of the cross-modal representation learning and retrieval method for food production, and will not be repeated here.

[0111] In an exemplary embodiment, Figure 5 As shown, a cross-modal representation learning and retrieval system for food production is provided, including: an image text learning module 51, a video text learning module 52, a feature library construction module 53 and a cross-modal retrieval module 54.

[0112] The image-text learning module 51 is used to perform multi-granular semantic alignment on image-text pairs in the grain production process based on a pre-trained image-text bidirectional guided fusion network to obtain semantically segmented images. The image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module.

[0113] The video text learning module 52 is used to perform image space decoupling and temporal enhancement on the video text pairs of the grain production process based on global semantic guidance to obtain structured semantic image features. The video in the video text pair is a remote sensing image sequence with temporal sequence.

[0114] The feature library construction module 53 is used to construct a text feature library and an image feature library according to the image-text pairs, the video-text pairs, the semantic segmentation images and the structured semantic image features.

[0115] The cross-modal retrieval module 54 is used to determine a transmission plan matrix according to the modality of the data to be retrieved, and generate query features of the data to be retrieved based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

[0116] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0117] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0118] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0120] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0121] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0122] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0123] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0124] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A cross-modal representation learning and retrieval method for food production, characterized by: The method comprises: Based on a pre-trained image-text bidirectional guided fusion network, multi-granularity semantic alignment is performed on image and text pairs in the grain production process to obtain semantic segmentation images; the image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module; Based on global semantic guidance, the video text pairs in the grain production process are subjected to image space decoupling and temporal enhancement to obtain structured semantic image features; the video in the video text pair is a remote sensing image sequence with temporal sequence; Constructing a text feature library and an image feature library based on the image-text pairs, the video-text pairs, the semantic segmentation images, and the structured semantic image features; A transmission plan matrix is ​​determined according to the modality of the data to be retrieved, and query features of the data to be retrieved are generated based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

2. The cross-modal representation learning and retrieval method for food production according to claim 1 is characterized in that: Based on the pre-trained image-text bidirectional guided fusion network, multi-granular semantic alignment of image and text pairs in the grain production process is performed to obtain semantic segmentation images, including: The feature extraction module performs feature extraction and feature concatenation on the image and text in the image-text pair, respectively, to obtain contextual features of each word in the text and multimodal mixed features of each image region; wherein different image regions correspond to different layers of the image; Based on the contextual features of each word in the text and the multimodal mixed features of each image region, the image-text bidirectional guided fusion module is used to perform image-text bidirectional guided fusion to obtain enhanced features of each image region; Performing multi-scale fusion on the enhanced features of each image region through the multi-scale fusion module to obtain a fused feature map; The semantic segmentation module performs pixel-by-pixel semantic classification on the fused feature map to obtain a semantic segmentation image.

3. The cross-modal representation learning and retrieval method for food production according to claim 2 is characterized in that: The feature extraction module includes DeepLab ResNet-101 v2 network, relative spatial encoding network, long short-term memory network and feature splicing network; The feature extraction module performs feature extraction and feature splicing on the image and text in the image-text pair respectively to obtain context features of each word in the text and multimodal mixed features of each image region, specifically including: A pre-trained DeepLab ResNet-101 v2 network is used to perform feature extraction on the image to obtain a visual feature vector for each image region; the visual feature vectors of different image regions correspond to the backbone outputs of the last three layers of the DeepLab ResNet-101 v2 network; Using a relative spatial encoding network to generate position features for each region of the image, thereby obtaining spatial coordinate features of each image region; Using a long short-term memory network to extract contextual features of each word in the text; For any image region, a feature splicing network is used to splice the visual feature vector of the image region, the spatial coordinate features of the image region and the context features of the last word in the text to obtain a multimodal mixed feature of the image region.

4. The cross-modal representation learning and retrieval method for food production according to claim 2, characterized in that: The image-text bidirectional guided fusion module includes a visually guided language attention mechanism and a language-guided visual attention mechanism. The visually guided language attention mechanism uses the contextual features of each word in the text as a guide to measure the semantic contribution of each word in the text to each image region. The language-guided visual attention mechanism uses the language contextual features of each image region as a guide to calculate the spatial semantic dependency between image regions. Based on the contextual features of each word in the text and the multimodal hybrid features of each image region, the image-text bidirectional guided fusion module is used to perform image-text bidirectional guided fusion to obtain enhanced features of each image region, specifically including: Determining the language context features of each image region using a visually guided language attention mechanism based on the context features of each word in the text and the multimodal mixed features of each image region; According to the language context features and multimodal mixed features of each image region, a language-guided visual attention mechanism is adopted to determine the enhanced features of each image region.

5. The cross-modal representation learning and retrieval method for food production according to claim 2, characterized in that: The multi-scale fusion module includes a dilated space pyramid pooling submodule and a bidirectional gated fusion submodule; The multi-scale fusion module performs multi-scale fusion on the enhanced features of each image region to obtain a fused feature map, which specifically includes: The atrous spatial pyramid pooling submodule is used to perform multi-scale semantic modeling on the enhanced features of each image region, extract contextual information under different receptive fields, and obtain multimodal features at different levels; The bidirectional gating fusion submodule is used to perform bidirectional gating fusion on multimodal features at different levels through top-down and bottom-up paths to obtain a fusion feature map.

6. The cross-modal representation learning and retrieval method for food production according to claim 1, characterized in that: Based on global semantic guidance, we perform image space decoupling and temporal enhancement on the video and text pairs in the grain production process to obtain structured semantic image features, including: Sampling the video in the video-text pair to obtain multiple frames of remote sensing images, and performing global semantic extraction and spatial feature extraction on the multiple frames of remote sensing images to obtain a global semantic vector and a spatial feature representation of each frame of remote sensing image; For any frame of remote sensing image, decoupling the spatial feature representation of the remote sensing image based on the global semantic vector to obtain high semantic saliency region features and low semantic saliency region features of the remote sensing image; According to the high semantic saliency region features and low semantic saliency region features of multiple frames of remote sensing images, temporal enhancement modeling is performed on each frame of remote sensing images respectively, and the forward saliency region features, backward saliency region features, forward cumulative low saliency features and backward cumulative low saliency features of each frame of remote sensing images are obtained; The semantic enhancement feature representation of each frame of remote sensing image is obtained based on the forward salient region features and backward salient region features of each frame of remote sensing image. The final cumulative feature representation is obtained based on the forward accumulated low-significance features and the backward accumulated low-significance features of the last frame of remote sensing image; The semantic enhancement feature representation of each frame of remote sensing image is fused with the final accumulated feature representation to obtain structured semantic image features.

7. The cross-modal representation learning and retrieval method for food production according to claim 6, characterized in that: Perform global semantic extraction and spatial feature extraction on multiple frames of remote sensing images to obtain a global semantic vector and spatial feature representation of each frame of remote sensing image, specifically including: Perform temporal average pooling and global average pooling operations on multiple frames of remote sensing images to extract global semantic vectors; A pre-trained deep convolutional neural network is used to extract features from each frame of remote sensing image to obtain the spatial feature representation of each frame of remote sensing image.

8. The cross-modal representation learning and retrieval method for food production according to claim 6, characterized in that: Based on the global semantic vector, the spatial feature representation of the remote sensing image is decoupled to obtain features of high semantically significant regions and features of low semantically significant regions of the remote sensing image, specifically including: Performing a linear transformation on the global semantic vector and expanding it into a feature tensor consistent with the spatial size of the remote sensing image; Cascading the feature tensor with the spatial feature representation of the remote sensing image to obtain a fusion feature; According to the fusion features, a convolutional layer and an activation function are used to generate a saliency response map; The spatial feature representation of the remote sensing image is subjected to saliency-guided decoupling according to the saliency response map, so as to obtain high semantic saliency region features and low semantic saliency region features of the remote sensing image.

9. The cross-modal representation learning and retrieval method for food production according to claim 6, characterized in that: According to the high semantic saliency region features and low semantic saliency region features of multiple frames of remote sensing images, temporal enhancement modeling is performed on each frame of remote sensing images respectively, and the forward saliency region features, backward saliency region features, forward cumulative low saliency features and backward cumulative low saliency features of each frame of remote sensing images are obtained, specifically including: Generate initial accumulated low-saliency features based on low semantic saliency region features of multiple frames of remote sensing images; Use forward-enhanced memory units to High semantic saliency region features of frame remote sensing images and the The forward accumulation of low-significance features of the frame remote sensing image is used to perform difference modeling and extract the first Forward semantic change information of frame remote sensing images; Adopt channel attention mechanism to extract The response weight of the forward semantic change information of the frame remote sensing image is The high semantic saliency region features of the frame remote sensing image are enhanced by time series contrast, and the first Forward salient region features of frame remote sensing images; Backward reinforcement memory unit is used to High semantic saliency region features of frame remote sensing images and the The backward accumulation of low-significance features of frame remote sensing images is used for difference modeling to extract the first Backward semantic change information of frame remote sensing images; Adopt channel attention mechanism to extract The response weight of the backward semantic change information of the frame remote sensing image is used, and the first The high semantic saliency region features of the frame remote sensing image are enhanced by time series contrast, and the first Backward salient region features of frame remote sensing images; According to Cumulative low-saliency features of frame remote sensing images and the The low semantic saliency region features of the frame remote sensing image are used to determine the Forward accumulation of low-significance features and backward accumulation of low-significance features of frame remote sensing images; in, , A is the number of frames of remote sensing images, Time Forward accumulation of low-significance features of frame remote sensing images and the first The backward accumulated low-saliency features of the frame remote sensing image are all the initial accumulated low-saliency features.

10. A cross-modal representation learning and retrieval system for food production, characterized by: The system applies the cross-modal representation learning and retrieval method for food production according to any one of claims 1 to 9, and the system includes: An image-text learning module is used to perform multi-granular semantic alignment of image-text pairs in the grain production process based on a pre-trained image-text bidirectional guided fusion network to obtain semantically segmented images. The image-text bidirectional guided fusion network includes a feature extraction module, an image-text bidirectional guided fusion module, a multi-scale fusion module, and a semantic segmentation module. A video-text learning module is used to perform image space decoupling and temporal enhancement on video-text pairs of the grain production process based on global semantic guidance to obtain structured semantic image features; the video in the video-text pair is a remote sensing image sequence with temporal sequence; A feature library construction module, configured to construct a text feature library and an image feature library based on the image-text pairs, the video-text pairs, the semantically segmented images, and the structured semantic image features; The cross-modal retrieval module is used to determine a transmission plan matrix according to the modality of the data to be retrieved, and generate query features of the data to be retrieved based on the transmission plan matrix. According to the query features of the data to be retrieved, the text feature library and the image feature library, a similarity measurement method is used to output text query results or image query results.

Citation Information

Patent Citations

  • Image text retrieval method and system based on context-guided multi-modal association

    CN116737979A

  • Image-text retrieval method and system based on semantic information reasoning and cross-modal interaction

    CN118133839A

  • Construction safety risk early warning method and system based on cross-modal visual language retrieval

    CN120146549A

  • Cross-modal remote sensing image-text retrieval method based on multistage semantic collaborative matching

    CN120336574A

Cited By

  • Data processing method and device, equipment, medium and product

    CN121681872A