A multi-modal image fine-grained segmentation model and an image segmentation method thereof

By using a multimodal image fine-grained segmentation model, combined with semantic hypergraph parsing and matching, the problem of fine-grained segmentation in complex scenes of rapeseed remote sensing images was solved, and high-precision rapeseed morphology recognition was achieved.

CN122289694APending Publication Date: 2026-06-26SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST UNIV
Filing Date
2026-04-22
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify detailed features in rapeseed remote sensing images, especially in complex scenarios. Traditional methods cannot take into account the complementary information from multiple data sources, and deep learning models lack sufficient semantic understanding of complex scenarios, resulting in inaccurate rapeseed trait segmentation.

Method used

We employ a multimodal image fine-grained segmentation model, combined with a semantic hypergraph parser and a semantic hypergraph matcher. Through multi-layer convolutional neural networks, adaptive hypergraph learning, and cross-modal visual association tensor computation, we simulate human-like intuitive reasoning to mine semantic cues and higher-order associations in multimodal data.

Benefits of technology

It achieves high-precision fine-grained segmentation of rapeseed traits in complex scenarios, effectively addresses noise and plant overlap issues, preserves trait details, and improves segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289694A_ABST
    Figure CN122289694A_ABST
Patent Text Reader

Abstract

This invention relates to the field of remote sensing image processing technology, specifically to a fine-grained segmentation model for modal images and its image segmentation method. The model includes a semantic hypergraph parser and a semantic hypergraph matcher. The image segmentation method includes the following steps: S1: Establish a multimodal remote sensing image training dataset; S2: Build a multimodal image segmentation model based on a semantic cognitive hypergraph; S3: Initialize model parameters and train the multimodal image segmentation model with the goal of minimizing the loss function, thus completing the model construction; S4: Input the multimodal remote sensing image to be processed into the fine-grained segmentation model for multimodal remote sensing images and output pixel-level semantic segmentation results. This invention combines a semantic cognitive framework based on human-like intuitive reasoning with hypergraph modeling to simulate the cognitive logic of human-like intuitive reasoning. At the same time, it mines semantic clues and higher-order associations in multimodal data, which can effectively deal with complex problems such as strong noise and plant superposition, while preserving the detailed features of traits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a fine-grained segmentation model for modal images and its image segmentation method. Background Technology

[0002] In recent years, rapeseed has become my country's largest oilseed crop, with a stable annual planting area of ​​over 100 million mu (approximately 6.67 million hectares). Its oil production accounts for nearly 60% of China's total oilseed crop production, making it a core variety for ensuring my country's self-sufficiency in edible vegetable oil. It holds an irreplaceable strategic position in stabilizing agricultural production and maintaining food and oilseed security. Accurately distinguishing key traits such as rapeseed leaves and flowers is a crucial prerequisite for rapeseed yield prediction, ultimately affecting both yield and quality. However, current multimodal rapeseed data collected via UAV remote sensing is highly susceptible to factors such as field climate and planting density: high temperature and humidity increase image noise, dense planting leads to plant overlap, and backlighting or cloudy weather causes blurred boundaries. These issues make it difficult to clearly identify the detailed features of rapeseed traits. Existing technologies, such as traditional image segmentation methods, rely heavily on single-modal features, failing to consider complementary information from multiple data sources; commonly used filtering and denoising techniques sacrifice trait details and struggle to handle boundary confusion caused by plant overlap; and ordinary deep learning segmentation models lack sufficient semantic understanding of complex scenes, failing to accurately capture high-order relationships among rapeseed traits. These shortcomings make it difficult for existing technologies to achieve fine-grained segmentation of rapeseed traits, which severely restricts the accuracy of subsequent growth assessment and management decisions. Summary of the Invention

[0003] The purpose of this invention is to provide a multimodal image fine-grained segmentation model and its image segmentation method. The segmentation model simulates the cognitive logic of human intuitive reasoning, and the segmentation method based on the segmentation model can simultaneously mine semantic clues and higher-order associations of multimodal data. It can effectively deal with complex problems such as strong noise and plant superposition, while preserving the detailed features of traits, providing a reliable technical solution for achieving high-precision fine-grained segmentation of rapeseed traits.

[0004] To achieve the above objectives, the present invention provides a multimodal image fine-grained segmentation model including a semantic hypergraph parser and a semantic hypergraph matcher; Semantic hypergraph parsers, including: The feature extraction module uses a multi-layer convolutional neural network to extract features from single-modal rapeseed remote sensing images, obtaining the original single-modal feature matrix; The node mapping module maps feature representations to a set of K node embeddings through the embedding layer; The hypergraph construction module is used to calculate the feature similarity between any two nodes, and based on this similarity, select nodes with similar features and connect them through hyperedges to form an initial hypergraph; and, An adaptive hypergraph learning module designs an adaptive hypergraph learning algorithm based on an information convergence mechanism.

[0005] A semantic hypergraph matcher comprises: A correlation tensor calculation module is configured to calculate a cross-modal visual correlation tensor according to semantic hypergraph node features of different modalities according to a neural tensor network structure. A multi-modal fusion module is configured to perform adaptive multi-modal fusion representation of semantic hypergraph nodes by using a semantic analysis module constructed based on an autoencoder or a Transformer neural network framework. A dynamic inversion module performs dynamic inversion based on the cognitive representation learned by the model to improve the accuracy of fine-grained segmentation of the rapeseed remote sensing image.

[0006] The application further provides an image segmentation method of the multi-modal image fine-grained segmentation model, comprising the following steps: S1: establishing a multi-modal remote sensing image training data set; S2: building a multi-modal image segmentation model based on a semantic cognitive hypergraph, which comprises a semantic hypergraph parser and a semantic hypergraph matcher, the semantic hypergraph parser is configured to parse input data of each modality into an irregular hypergraph to mine deep semantic clues, and then generate a global semantic representation corresponding to a single modality; the semantic hypergraph matcher is configured to integrate and form a high-order cognitive representation of a complex dynamic scene through the internal semantic clues of different modal semantic hypergraphs; S3: initializing model parameters, training the multi-modal image segmentation model with the objective of minimizing a loss function, and completing the construction of the multi-modal remote sensing image fine-grained segmentation model; S4: inputting a multi-modal remote sensing image to be processed into the multi-modal remote sensing image fine-grained segmentation model, and outputting a pixel-level semantic segmentation result.

[0007] Further, the step S1 specifically comprises collecting rapeseed remote sensing images by using remote sensing technology, preprocessing, and creating a training data set.

[0008] The data is derived from multi-modal growth data of rapeseed collected by an unmanned aerial vehicle remote sensing system, wherein the construction of the training data set is to plant rapeseed samples with known traits in the field, and collect multi-modal remote sensing data of different growth periods by using an unmanned aerial vehicle, as reference data corresponding to standard traits.

[0009] Further, in the step S3, the cross-modal visual correlation tensor expression is: wherein is a tensor operation.

[0010] In the embodiment, It can be the Hadamard accumulation or the external accumulation.

[0011] Furthermore, in step S2, the semantic hypergraph parser integrates the inherent semantic cues of different modal semantic hypergraphs to form a high-order cognitive representation of complex dynamic scenes, specifically including the following steps: Multi-layer convolutional neural networks are used to extract features from single-modal remote sensing images. By alternating the stacking of convolutional layers and pooling layers, the hierarchical features of the image are gradually mined to obtain the original feature matrix of the single-modality image. The feature representation is then mapped to a set of K node embeddings through an embedding layer. ; Gaussian similarity is used as the distance metric to calculate the feature similarity between any two nodes, and n nodes with similar features are selected based on the similarity and connected by hyperedges to form an initial hypergraph. An adaptive hypergraph learning algorithm is designed based on an information aggregation mechanism to dynamically update the hypergraph structure. In the process of aggregating the structural feature information of neighboring nodes, it captures the inherent semantic clues about complex dynamic scenes in single-modal remote sensing images.

[0012] Furthermore, the original feature matrix of the single mode is ,in The total number of pixels in the image. Image height and width respectively For a single modal feature dimension, m represents the m-th modality.

[0013] The mapping process of feature representation to a set of K node embeddings satisfies in, This represents the number of pixels in the cluster. For node embedding, This is a single-modal pixel feature vector, corresponding to the row vectors in the F matrix.

[0014] The information aggregation formula is as follows: This invention presents an image segmentation method based on a multimodal image fine-grained segmentation model. By combining a semantic cognitive framework based on human intuitive reasoning with hypergraph modeling, it simulates the cognitive logic of human intuitive reasoning and simultaneously mines semantic clues and higher-order associations in multimodal data. This method can effectively address complex problems such as strong noise and plant superposition while preserving detailed trait features. By efficiently capturing higher-order relationships in complex dynamic scenes, it provides a reliable technical solution for achieving high-precision fine-grained segmentation of rapeseed traits. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Fig. 1 This is a flowchart illustrating an image segmentation method based on a multimodal image fine-grained segmentation model according to the present invention.

[0017] Fig. 2 This is a schematic diagram of the internal structure of the semantic hypergraph parser of this invention.

[0018] Fig. 3 This is a schematic diagram of the internal structure of the semantic hypergraph matcher of the present invention. Detailed Implementation

[0019] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0020] Please see Figs. 1 to 3 This invention provides a multimodal image fine-grained segmentation model, including a semantic hypergraph parser and a semantic hypergraph matcher; Semantic hypergraph parsers, including: The feature extraction module uses a multi-layer convolutional neural network to extract features from single-modal rapeseed remote sensing images, obtaining the original single-modal feature matrix; The node mapping module maps feature representations to a set of K node embeddings through the embedding layer; The hypergraph construction module is used to calculate the feature similarity between any two nodes, and based on this similarity, select nodes with similar features and connect them through hyperedges to form an initial hypergraph; and, The adaptive hypergraph learning module designs an adaptive hypergraph learning algorithm based on an information convergence mechanism.

[0021] Semantic hypergraph matchers, including: The association tensor calculation module is used to calculate the cross-modal visual association tensor according to the semantic hypergraph node features of different modalities and the neural tensor network structure. The multimodal fusion module is configured to perform adaptive semantic hypergraph node multimodal fusion representation using a semantic parsing module built on an autoencoder or Transformer neural network framework; and, The dynamic inversion module performs dynamic inversion based on the cognitive representation learned by the model, thereby improving the accuracy of fine-grained segmentation of rapeseed remote sensing images.

[0022] The high-order structure design of the semantic hypergraph parser can effectively improve the semantic capture capability of complex rapeseed features without sacrificing model time complexity. The dynamic update mechanism of adaptive hypergraph learning can effectively avoid information loss or semantic conflicts caused by differences in multimodal feature dimensions without adding extra parameters, ensuring model stability while enjoying the fine-grained segmentation accuracy improvement brought by high-order relationship modeling. Meanwhile, the semantic hypergraph matcher adopts a node association strategy based on Gaussian similarity, effectively avoiding the problem of increased model complexity and reduced efficiency in processing actual field data caused by the introduction of an attention layer. The association tensor computation module can construct cross-modal visual association tensors through various operations such as Hadamard product, outer product, or other tensor operations, capturing potentially more complex high-order associations to provide a larger representation learning space.

[0023] This invention also discloses an image segmentation method based on a multimodal image fine-grained segmentation model, comprising the following steps: S1: Establish a multimodal remote sensing image training dataset; S2: Construct a multimodal image segmentation model based on semantic cognitive hypergraphs. This model includes a semantic hypergraph parser and a semantic hypergraph matcher. The semantic hypergraph parser is used to parse the input data of each modality into irregular hypergraphs to mine deep semantic cues, and then initially generate the global semantic representation of the corresponding single modality. The semantic hypergraph matcher is used to integrate the inherent semantic cues of the semantic hypergraphs of different modalities and form a high-order cognitive representation of complex dynamic scenes. S3: Initialize the model parameters and train the multimodal image segmentation model with the goal of minimizing the loss function, thus completing the construction of the fine-grained segmentation model for multimodal remote sensing images; S4: Input the multimodal remote sensing image to be processed into the multimodal remote sensing image fine-grained segmentation model, and output pixel-level semantic segmentation results.

[0024] Step S1 specifically includes acquiring remote sensing images of rapeseed using remote sensing technology, preprocessing them, and creating a training dataset.

[0025] In step S1, remote sensing images of rapeseed are acquired using remote sensing technologies such as drones, preprocessed, and a training dataset is created.

[0026] In step S2, the semantic hypergraph matcher integrates the inherent semantic cues of different modal semantic hypergraphs to form a high-order cognitive representation of complex dynamic scenes. It can select and create rapeseed multimodal remote sensing image datasets with different noise effects, and train rapeseed multimodal remote sensing image segmentation models that specifically process the corresponding noise as described in step S3.

[0027] In step S3, based on the multimodal remote sensing image segmentation model and its training parameters, the model is trained to form a fine-grained segmentation model for rapeseed multimodal remote sensing images with the goal of minimizing the loss function. The optimizer used in step S3 for neural network training is the Adam optimizer. When optimizing network parameters, the Adam optimizer adaptively adjusts the first and second momentum of the gradient, effectively avoiding the model from converging to a local optimum and accelerating the optimization efficiency.

[0028] In step S4, the rapeseed multimodal remote sensing image to be processed is preprocessed and then input into the segmentation model to output fine-grained segmentation results.

[0029] Please see Fig. 2 In one embodiment, the semantic hypergraph parser parses the input data of each modality into an irregular hypergraph to mine deep semantic clues, and then initially generates a global semantic representation corresponding to the single modality. The specific implementation process is as follows: First, the feature extraction module uses a multi-layer convolutional neural network to extract features from single-modal rapeseed remote sensing images. By alternately stacking convolutional and pooling layers, it gradually mines hierarchical features such as local texture, spectral response, and morphological structure of the image to obtain the original single-modal feature matrix. (in The total number of pixels in the image. These are the image height and width, respectively. (This represents the single-modal feature dimension, where m denotes the m-th modality).

[0030] Subsequently, the node mapping module maps this feature representation to a set of K node embeddings through the embedding layer. The mapping process satisfies in The number of pixels in the cluster is represented by this mapping, which clusters pixels with similar characteristics into the same node, preserving semantic consistency while reducing computational complexity.

[0031] Next, the hypergraph construction module uses Gaussian similarity as the distance metric to calculate the feature similarity between any two nodes, and selects n nodes with similar features based on this similarity. These nodes are then connected by hyperedges to form an initial hypergraph. Finally, the adaptive hypergraph learning module designs an adaptive hypergraph learning algorithm based on an information convergence mechanism, specifically according to the following information convergence formula: By dynamically updating the hypergraph structure and continuously strengthening the semantic associations related to rapeseed traits while weakening the false associations caused by noise interference during the process of aggregating the structural feature information of neighboring nodes, the inherent semantic clues about the complex dynamic scene of rapeseed field in single-modal remote sensing images can be accurately captured.

[0032] Please see Fig. 3 In another embodiment, for the purpose of this embodiment, the semantic hypergraph matcher integrates the intrinsic semantic cues of semantic hypergraphs from different modalities to form a high-order cognitive representation of complex dynamic scenes. The association tensor calculation module calculates the cross-modal visual association tensor according to the neural tensor network structure based on the node features of semantic hypergraphs from different modalities. Generally speaking, it can be represented as ,in This allows for Hadamard product, outer product, or other tensor operations. Subsequently, the multimodal fusion module utilizes a semantic parsing module built on an autoencoder or Transformer neural network framework to perform adaptive semantic hypergraph node multimodal fusion representations. This parses the inherent semantic cues and cross-modal visual association information within the remote sensing data, achieving accurate cognition of dynamic and complex scenes within the multimodal remote sensing data. Finally, the dynamic inversion module performs dynamic inversion based on the cognitive representation learned by the model, improving the accuracy of fine-grained segmentation of rapeseed remote sensing images. The initial weights of the model are generated using a Gaussian random function, and the initial parameters of the hypergraph learning module are set to a uniform distribution to ensure the stability of the model's initial state.

[0033] The loss function to be minimized is defined as follows: in, This represents all learnable parameters in the model. and These correspond to the height and width of the input remote sensing image, respectively. This refers to the total number of semantic categories that need to be segmented. This represents the probability predicted by the model. It is the regularization coefficient. Model parameters L2 norm, Defined as an indicator function, this function returns 1 when the true class of the i-th pixel is the j-th class, and 0 otherwise.

[0034] Compared with existing technologies, the fine-grained segmentation method for rapeseed multimodal remote sensing images provided by this invention combines a human-like semantic cognitive framework and a hypergraph structure. It mines single-modal semantic cues in the semantic hypergraph parser and integrates cross-modal association information in the semantic hypergraph matcher. By leveraging the high-order modeling capabilities of the hypergraph, it captures the relationships in complex dynamic scenes, effectively improving the accuracy of fine-grained segmentation of rapeseed traits and providing an efficient technical solution for agricultural remote sensing monitoring.

[0035] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A multimodal image fine-grained segmentation model, characterized in that, This includes a semantic hypergraph parser and a semantic hypergraph matcher; Semantic hypergraph parsers, including: The feature extraction module uses a multi-layer convolutional neural network to extract features from single-modal rapeseed remote sensing images, obtaining the original single-modal feature matrix; The node mapping module maps feature representations to a set of K node embeddings through the embedding layer; The hypergraph construction module is used to calculate the feature similarity between any two nodes, and based on this similarity, select nodes with similar features and connect them through hyperedges to form an initial hypergraph; and, The adaptive hypergraph learning module designs an adaptive hypergraph learning algorithm based on an information convergence mechanism.

2. The multimodal image fine-grained segmentation model as described in claim 1, characterized in that, Semantic hypergraph matchers, including: The association tensor calculation module is used to calculate the cross-modal visual association tensor according to the semantic hypergraph node features of different modalities and the neural tensor network structure. The multimodal fusion module is configured to perform adaptive semantic hypergraph node multimodal fusion representation using a semantic parsing module built on an autoencoder or Transformer neural network framework; and, The dynamic inversion module performs dynamic inversion based on the cognitive representation learned by the model, thereby improving the accuracy of fine-grained segmentation of rapeseed remote sensing images.

3. An image segmentation method based on a multimodal image fine-grained segmentation model according to any one of claims 1 to 2, characterized in that, Includes the following steps: S1: Establish a multimodal remote sensing image training dataset; S2: Construct a multimodal image segmentation model based on semantic cognitive hypergraphs. This model includes a semantic hypergraph parser and a semantic hypergraph matcher. The semantic hypergraph parser is used to parse the input data of each modality into irregular hypergraphs to mine deep semantic cues, and then initially generate the global semantic representation of the corresponding single modality. The semantic hypergraph matcher is used to integrate the inherent semantic cues of the semantic hypergraphs of different modalities and form a high-order cognitive representation of complex dynamic scenes. S3: Initialize the model parameters and train the multimodal image segmentation model with the goal of minimizing the loss function, thus completing the construction of the fine-grained segmentation model for multimodal remote sensing images; S4: Input the multimodal remote sensing image to be processed into the multimodal remote sensing image fine-grained segmentation model, and output pixel-level semantic segmentation results.

4. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 3, characterized in that, Step S1 specifically includes acquiring remote sensing images of rapeseed using remote sensing technology, preprocessing them, and creating a training dataset.

5. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 3, characterized in that, In step S3, the cross-modal visual association tensor expression is: in This refers to tensor operations.

6. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 3, characterized in that, In step S2, the semantic hypergraph parser is used to integrate the inherent semantic cues of different modal semantic hypergraphs and form a high-order cognitive representation of complex dynamic scenes. Specifically, this includes the following steps: Multi-layer convolutional neural networks are used to extract features from single-modal remote sensing images. By alternating the stacking of convolutional layers and pooling layers, the hierarchical features of the image are gradually mined to obtain the original feature matrix of the single-modality image. Subsequently, the feature representation is mapped to a set of K node embeddings through an embedding layer. ; Gaussian similarity is used as the distance metric to calculate the feature similarity between any two nodes, and n nodes with similar features are selected based on the similarity and connected by hyperedges to form an initial hypergraph. An adaptive hypergraph learning algorithm is designed based on an information aggregation mechanism to dynamically update the hypergraph structure. In the process of aggregating the structural feature information of neighboring nodes, it captures the inherent semantic clues about complex dynamic scenes in single-modal remote sensing images.

7. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 6, characterized in that, The original feature matrix of the single-modality is: ,in This represents the total number of pixels in the image. These are the image height and width, respectively. For a single modal feature dimension, m represents the m-th modality.

8. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 6, characterized in that, The mapping process of feature representations to a set of K node embeddings satisfies: in, This represents the number of pixels in the cluster. For node embedding, This is a single-modal pixel feature vector, corresponding to the row vectors in the F matrix.

9. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 6, characterized in that, The information aggregation formula is as follows: in, Refers to any link The super-edge, These are the weight parameters for information aggregation.

10. The image segmentation method of a multimodal image fine-grained segmentation model as described in claim 3, characterized in that, In step S3, the loss function to be minimized is defined as: in, This represents all learnable parameters in the model. and These correspond to the height and width of the input remote sensing image, respectively. This refers to the total number of semantic categories that need to be segmented. This represents the probability predicted by the model. It is the regularization coefficient. Model parameters L2 norm, Defined as an indicator function, this function returns 1 when the true class of the i-th pixel is the j-th class, and 0 otherwise.