RGB-T image multi-modal semantic segmentation method and system
By introducing feature extraction and fusion methods of Mobius transform and hyperbolic spatial embedding in the Transformer layer, the modal difference and edge distortion problems of traditional methods when processing multimodal images are solved, and a higher precision multimodal semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202510661107.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
AI Technical Summary
When the traditional Transformer layer processes multimodal images (such as RGB-T), due to linear transformation based on Euclidean space, it is difficult to effectively handle the modal differences and edge distortion between infrared and visible light images, resulting in inaccurate feature extraction and fusion.
The Mobius transformation is used to replace the linear transformation in the visual Transformer, and the multimodal features are extracted and fusion in hyperbolic space through the Mobius self-attention and cross-attention mechanism.
Effectively cope with modal differences and edge distortion in multimodal images, improve the fusion accuracy and boundary processing capabilities of RGB-T images, and achieve higher precision multimodal semantic segmentation.
Smart Images

Figure CN120182610A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to an RGB-T image multimodal semantic segmentation method and system. Background Art
[0002] In unmanned aerial vehicle and vehicle-mounted systems, images captured by infrared and visible light sensors will produce significant spatial distortions due to different imaging methods (such as focal length, spectral response, sensor type, etc.). Current multimodal image processing (such as RGB-T feature fusion) usually relies on the Transformer layer (especially the Vision Transformer, ViT) in the deep learning framework to extract deep features of images. The ViT structure realizes the extraction of high-level features through the self-attention mechanism. However, due to its linear transformation based on Euclidean space, it is difficult to handle the feature differences and relative distortions of multimodal inputs (such as infrared and visible light), especially in terms of the accuracy of feature fusion and registration.
[0003] Among them, the traditional Transformer layer is based on the attention mechanism of Euclidean space and lacks adaptability to the distortions and modal differences existing in infrared and visible light images. Even though the cross-attention mechanism performs well in multimodal fusion, its linear transformation will still cause feature loss, especially when dealing with features with severe distortions in the edge region. Therefore, the traditional Vision Transformer (ViT) based on the linear self-attention mechanism of Euclidean space cannot effectively handle the problems of feature alignment and fusion difficulties caused by modal differences and distortions between infrared (T) and visible light (RGB) images, has poor adaptability to such multimodal distortions, and is difficult to achieve high-precision multimodal feature fusion. Summary of the Invention
[0004] The present invention aims to solve the problem of inaccurate feature extraction and fusion in multi-modal images (such as RGB-T) due to modal differences and edge distortions. In particular, the traditional attention mechanism of the Transformer layer based on Euclidean space cannot adapt to the distortions and modal differences existing in infrared and visible light images. For this reason, the technical solution of the present invention provides an RGB-T image multi-modal semantic segmentation method and system. This method realizes the extraction and fusion of multi-modal features by using the Möbius transformation to replace the linear transformation in the vision Transformer. The Möbius attention mechanism can perform non-linear mapping on multi-modal features in the hyperbolic space, and can more effectively cope with modal differences and edge distortions. Therefore, the technical solution of the present invention is based on the fusion structure of the Möbius self-attention and cross-attention mechanisms, combined with a dual-branch network and a multi-scale fusion strategy, which can effectively improve the RGB-T image fusion accuracy and boundary processing ability, better solve the spatial heterogeneity and edge distortion problems between RGB-T modalities, and thus achieve higher-precision multi-modal semantic segmentation. In addition, the weights of the Möbius transformation are trainable, which helps to further improve the adaptability and fusion accuracy of the model for multi-modal feature fusion.
[0005] For this reason, the present invention provides the following technical solutions:
[0006] An RGB-T image multi-modal semantic segmentation method, comprising the following steps:
[0007] Multi-modal data preprocessing, obtaining an RGB image and an infrared image and performing preprocessing, the preprocessing at least includes normalization processing;
[0008] Multi-modal feature extraction, using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch to respectively extract RGB image features and infrared image features. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N sequentially connected Möbius self-attention layers, and the Möbius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1;
[0009] Feature fusion, using N fusion layers based on Möbius cross-attention to form a feature fusion layer for fusing RGB image features and infrared image features to obtain multi-scale features;
[0010] Among them, the Möbius self-attention mechanism of the Möbius self-attention layer and the Möbius cross-attention mechanism of the fusion layer based on Möbius cross-attention respectively perform non-linear mapping of features in a non-Euclidean space by using Möbius multiplication and Möbius addition to realize the alignment and fusion of multi-modal features;
[0011] The spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates the segmentation prediction results through the Softmax function;
[0012] Among them, the network is trained using the sampling samples and their labels of RGB images and infrared images in semantic segmentation applications, and then the trained network is used for semantic segmentation prediction.
[0013] Preferably, the other non-Euclidean space is a hyperbolic space, or a spherical space or an elliptical space or a mixed space composed of any combination of a hyperbolic space, a spherical space, and an elliptical space.
[0014] Preferably, if the other non-Euclidean space is a hyperbolic space, the Möbius multiplication is expressed as:
[0015] ;
[0016] The Möbius addition is expressed as:
[0017] ;
[0018] In the formula, , are the Möbius multiplication symbol and the Möbius addition symbol in the hyperbolic space with curvature c, respectively; , , are custom variable symbols used to represent the objects of Möbius multiplication and Möbius addition, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, represents the result of the standard matrix multiplication, is the modulus of the vector, represents the curvature parameter of the hyperbolic space represents inner product.
[0019] Preferably, the data processing process of the Möbius self-attention layer is as follows:
[0020] First, the initial features are input into the Möbius self-attention layer. Among them, the initial features of the first Möbius self-attention layer are the preprocessed RGB image or infrared image;
[0021] Then, through the transformation of the Möbius mapping layer, the non-linear mapping of the features in the other non-Euclidean space is realized to obtain the query vector, the key vector, and the value vector:
[0022] Next, calculate the attention weights in his non-Euclidean space through the multi-head self-attention mechanism, and perform Möbius multiplication with the value vector after non-linear mapping to obtain a new feature representation;
[0023] Next, use a feed-forward network layer fully embedded in his non-Euclidean space to process the new feature representation, and the feed-forward network layer consists of two Möbius linear transformations and non-linear activation;
[0024] Finally, perform residual connection and normalization operations in his non-Euclidean space on the output of the feed-forward network layer;
[0025] Among them, the output features of the previous Möbius self-attention layer serve as the initial features of the next Möbius self-attention layer.
[0026] Preferably, if his non-Euclidean space is a hyperbolic space, the non-linear mapping representation of features in his non-Euclidean space through the Möbius mapping layer transformation is:
[0027] ;
[0028] ;
[0029] ;
[0030] In the formula, , , are the query vector, key vector, and value vector obtained by non-linear mapping of the Möbius self-attention layer i respectively, , , and , , are the trainable weights and biases of the Möbius mapping layer, used to map RGB features to the hyperbolic space; and are the Möbius multiplication and Möbius addition symbols in the Möbius transformation respectively, is the initial feature input to the Möbius self-attention layer i;
[0031] Calculate the attention weights in the hyperbolic space through the multi-head self-attention mechanism , defined as:
[0032] ;
[0033] In the formula, is 's dimension, T is the matrix transpose symbol, and use the attention weight and the value to calculate the new feature representation :
[0034] ;
[0035] The processing process using the feed - forward network layer fully embedded in the hyperbolic space is as follows:
[0036] ;
[0037] ;
[0038] In the formula, and are trainable weight matrices in the hyperbolic space, and are bias vectors, is the result after the Möbius linear transformation and non - linear activation in the first layer of the feed - forward network layer, is the final output after the linear mapping and bias in the second layer;
[0039] The output of the feed - forward layer is obtained by using residual connection and normalization operations in the hyperbolic space , as follows:
[0040] ;
[0041] Among them, the function is:
[0042] ;
[0043] In the formula, z is a self - defined variable symbol.
[0044] Preferably, the processing process of the Möbius cross - attention fusion layer is as follows:
[0045] First, map the RGB image features extracted by the Möbius self - attention layer at the same level on the RGB image feature extraction branch into query vectors; map the infrared image features extracted by the Möbius self - attention layer at the same level on the infrared image feature extraction branch into key vectors and value vectors;
[0046] ;
[0047] ;
[0048] ;
[0049] In the formula, , , They are the query vector, key vector, and value vector corresponding to the i-th Möbius cross-attention fusion layer respectively; , , They are the biases corresponding to the query vector, key vector, and value vector respectively; 、 、 They are the trainable weights of the cross-attention layer corresponding to the query vector, key vector, and value vector respectively;
[0050] Then, based on the query vector, key vector, and value vector, calculate the multi-scale features after fusion of the Möbius cross-attention fusion layer:
[0051] ;
[0052] In the formula, is the multi-scale feature corresponding to the i-th Möbius cross-attention fusion layer, is the key vector 's dimension.
[0053] Preferably, if the non-Euclidean space is a hyperbolic space, when training the network with the sampling samples and their labels of RGB images and infrared images, calculate the alignment error between the RGB image features and the infrared image features through the hyperbolic distance loss function;
[0054] Among them, the alignment error is expressed as:
[0055] ;
[0056] Among them, the formula for the hyperbolic distance is:
[0057] ;
[0058] In the formula, L is the total feature alignment loss value, is the Möbius embedding mapping function in the hyperbolic space, is the embedded feature of the i-th infrared image, is the embedded feature of the i-th visible light image, and That is, the mentioned above, just to distinguish whether it is an infrared image or a visible light image extracted, is the trade-off coefficient in the loss function, is the set of trainable parameters of the entire model, N is the number of samples participating in the loss calculation, and arcosh is the inverse hyperbolic cosine function.
[0059] In addition, the present invention also provides a semantic segmentation system based on the above semantic segmentation method, including:
[0060] A multimodal data preprocessing module, which is used to obtain RGB images and infrared images and perform preprocessing, and the preprocessing at least includes normalization processing;
[0061] A multimodal feature extraction module, which is used to respectively extract RGB image features and infrared image features by using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence, and the Mobius self-attention layers of the two branches correspond one by one, where N is a positive integer greater than 1;
[0062] A feature fusion module, which is used to use N fusion layers based on Mobius cross-attention to form a feature fusion layer for fusing RGB image features and infrared image features to obtain multi-scale features;
[0063] Among them, the Mobius self-attention mechanism of the Mobius self-attention layer and the Mobius cross-attention mechanism of the fusion layer based on Mobius cross-attention respectively perform feature non-linear mapping in non-Euclidean space by using Mobius multiplication and Mobius addition to realize the alignment and fusion of multimodal features;
[0064] A spatial decoder, which aligns the resolutions of multi-scale features at different levels through upsampling operations, and then projects the multi-scale features of N levels after fusion into the category space, and generates a segmentation prediction result through the Softmax function;
[0065] And / or a training module, which is used to perform network training by using the sampling samples and their labels of RGB images and infrared images in semantic segmentation applications, and then use the trained network to perform semantic segmentation prediction.
[0066] The present invention also provides an electronic terminal, including: one or more processors; a memory storing one or more computer programs; wherein, the processor calls the computer program to implement:
[0067] The steps of an RGB-T image multimodal semantic segmentation method.
[0068] The present invention also provides a computer-readable storage medium, storing a computer program, and the computer program is called by the processor to implement:
[0069] The steps of an RGB-T image multimodal semantic segmentation method.
[0070] Advantageous effects
[0071] Compared with the existing methods, the advantages of the present invention are:
[0072] 1. The present invention proposes multi-modal Möbius cross-attention fusion, that is, adopts the Möbius cross-attention mechanism based on hyperbolic space to align and fuse RGB and T (infrared) image features in a non-linear manner. This mechanism realizes the geometric consistency of multi-modal features in hyperbolic space, thus effectively coping with the differences between modalities, achieving precise complementarity and alignment of information, and improving the accuracy and consistency of multi-modal fusion.
[0073] 2. The hyperbolic space decoder structure proposed by the present invention is a hyperbolic space decoder based on Möbius interpolation. This decoder smoothly expands and integrates the fused features of different resolutions in hyperbolic space through multi-scale upsampling and feature fusion, so as to achieve high-resolution and refined segmentation output while maintaining the geometric structure.
[0074] 3. The present invention optimizes the loss function of hyperbolic space, that is, introduces the distance metric of hyperbolic space in loss calculation, and uses the hyperbolic distance in the Poincaré disk model to measure the alignment error between RGB and T features. Compared with the traditional Euclidean distance, the hyperbolic distance is more suitable for describing the non-linear relationship of multi-modal features, improves the accuracy of fused feature alignment, and ensures better alignment effect in multi-modal segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 is a schematic flowchart of the semantic segmentation method provided by an embodiment of the present invention;
[0076] Figure 2 is a schematic diagram of the network architecture provided by an embodiment of the present invention;
[0077] Figure 3 where (a) is the feature representation in Euclidean space, Figure 3 and (b) is the feature representation in hyperbolic space. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] A RGB-T image multi-modal semantic segmentation method provided by the present invention is used for multi-modal image feature extraction, fusion and then realizing semantic segmentation. Among them, the core of the present invention is to introduce the Möbius transformation mechanism to replace the attention mechanism in the vision Transformer to improve the extraction and fusion effect of multi-modal features. Among them, taking infrared (T) and visible light (RGB) images as processing objects, it can be specifically applied to semantic segmentation in fields such as unmanned aerial vehicle / vehicle-mounted systems. Therefore, the technical idea of the semantic segmentation method provided by the present invention is as follows:
[0079] Preprocessing of multi-modal data, obtaining RGB images and infrared images and performing preprocessing, and the preprocessing at least includes normalization processing;
[0080] Multi-modal feature extraction: A dual-branch structure consisting of an RGB image feature extraction branch and an infrared image feature extraction branch is used to extract RGB image features and infrared image features respectively. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence. The Mobius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1.
[0081] Feature fusion: A feature fusion layer is used to fuse the RGB image features and the infrared image features to obtain multi-scale features. This feature fusion layer is composed of N Mobius cross-attention fusion layers.
[0082] Among them, the Mobius self-attention mechanism of the Mobius self-attention layer and the Mobius cross-attention mechanism of the Mobius cross-attention fusion layer both use Mobius multiplication and Mobius addition to perform feature non-linear mapping in non-Euclidean space respectively, so as to achieve the alignment and fusion of multi-modal features.
[0083] The spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates a segmentation prediction result through the Softmax function.
[0084] Among them, non-Euclidean space generally refers to hyperbolic space, spherical space or elliptic space, etc., to achieve the embedding and fusion of multi-modal features. Spherical space performs well in processing data with global consistency or closed structure, and can be used to describe some scenes with less edge distortion. However, compared with hyperbolic space, spherical space has insufficient expressive ability for hierarchical structure data, and there will be a problem of decreasing information density when processing high-dimensional data. Therefore, although spherical and elliptic spaces are used as alternative solutions, the embodiments of the present invention preferably use hyperbolic space. Taking hyperbolic space as an example will be described below. In some other embodiments, a generalized non-Euclidean space model can also be adopted, that is, a model that combines the characteristics of multiple geometric spaces (such as hyperbolic-spherical hybrid space). For example, the edge region can be embedded into the spherical space to retain its local smoothness, while the deep structure can be embedded into the hyperbolic space to capture hierarchical features. This method can utilize the advantages of different geometric spaces for feature embedding and fusion. However, this alternative solution has certain flexibility in theory, but may require higher computational costs.
[0085] The present invention will be further described below in conjunction with embodiments.
[0086] Embodiment 1:
[0087] This embodiment takes the hyperbolic space as an example to illustrate a multi-modal semantic segmentation method for RGB-T images, which includes the following steps:
[0088] S1: Data preprocessing.
[0089] In this embodiment, RGB and infrared images (T images) are obtained, and the RGB images and infrared images are normalized (for example, the pixel values are scaled to the range of [0, 1]), and the image size is adjusted to match the requirements of the network input. This step ensures that the input data has a unified scale and feature distribution through image normalization and denoising, etc., laying a data foundation for subsequent feature extraction and fusion. Among them, in practical applications, RGB and infrared images (T images) are generally collected by multi-modal sensors of drones or vehicle-mounted systems.
[0090] S2: Multi-modal feature extraction.
[0091] The present invention uses a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch to extract RGB image features and infrared image features respectively, that is, to embed and map image features in the hyperbolic space.
[0092] A: RGB feature extraction branch.
[0093] As Figure 2 shown, there are 4 Mobius self-attention layers connected in sequence on the RGB feature extraction branch of this embodiment, that is, the RGB image is feature-extracted through the Mobius self-attention layer. The specific process is as follows:
[0094] 2-1: Input the initial features of the RGB image. is the image feature input to the i-th Mobius self-attention layer, where i represents the Mobius self-attention layer level of the RGB feature extraction branch, and i = 0, 1, 2, 3, 4. When i = 0, = , represents the normalized RGB image (the preprocessed RGB image), norm represents the normalization operation, represents the RGB image, and the output of the previous Mobius self-attention layer is used as the input of the next Mobius self-attention layer.
[0095] 2-2: Implement non-linear mapping of features through the Mobius mapping layer transformation, specifically: ;
[0096] Among them, , , They are the query vector, key vector, and value vector obtained by the nonlinear mapping of the Möbius self-attention layer i on the RGB feature extraction branch, respectively; , , are all trainable weights of , , are all bias parameters that can be adaptively adjusted during training to map the RGB features to the hyperbolic space, where q, k, and v correspond to the query vector, key vector, and value vector respectively; and are the Möbius multiplication and Möbius addition in the Möbius transformation. In this embodiment, through the nonlinear mapping characteristics in the hyperbolic space, the RGB features can adapt to edge distortion during mapping and better align with the infrared features.
[0097] Among them, the Möbius matrix multiplication is defined as:
[0098] ;
[0099] In the formula, , and x are all user-defined variables used to demonstrate the Möbius matrix multiplication; is the hyperbolic tangent function, is the inverse hyperbolic tangent function, represents the result of the standard matrix multiplication, and are the norms of the vectors respectively, represents the curvature parameter of the hyperbolic space, and 2 is taken in this embodiment. The Möbius multiplication operation in this embodiment ensures geometric consistency within the hyperbolic space.
[0100] The Möbius matrix addition is defined as:
[0101] ;
[0102] In the formula, are all user-defined variables used to demonstrate the Möbius matrix addition. represents the inner product. This addition ensures the geometric consistency of the bias.
[0103] 2 - 3: Calculate the attention weights in the hyperbolic space through the multi-head self-attention mechanism , which is defined as:
[0104] ;
[0105] Among them, is the dimension of. The dot product operation within this hyperbolic space ensures the nonlinear interaction between features. Then, the attention weights are used Sum value Calculate a new feature representation :
[0106] ;
[0107] 2 - 4: After the self - attention module, a feed - forward network layer that fully embeds in the hyperbolic space is used to further process the features. The feed - forward network layer consists of two layers of Möbius linear transformations and non - linear activations, defined as:
[0108] ;
[0109] ;
[0110] Among them, and are trainable weight matrices in the hyperbolic space, and are bias vectors. To preserve the continuity and information between layers within the hyperbolic space, residual connections and normalization operations in the hyperbolic space are adopted for the output of the feed - forward network layer.
[0111] ;
[0112] Among them, is defined as:
[0113] ;
[0114] In the formula, is a custom variable used to demonstrate function.
[0115] So far, each part of the operation of this feature extraction layer is completed in the hyperbolic space, thus ensuring geometric consistency. The finally obtained can be used as the input feature for the next Möbius self - attention layer i + 1, enabling the RGB features to maintain their structural information and geometric relationships within the hyperbolic space under the step - by - step processing of the network hierarchy. This way enables the network to better capture the differential features between modalities in the non - Euclidean space, thus adapting to the multi - modal fusion task under complex scenarios.
[0116] B: Feature extraction branch of the infrared image (T branch).
[0117] Such as Figure 2As shown in the figure, the feature extraction branch of the infrared image in this embodiment is provided with four sequentially connected Mobius self-attention layers. In the T branch, the initial features of the infrared image are extracted at four levels. Each level extracts features through the self-attention module in the hyperbolic space, similar to the processing process of the RGB branch, but without sharing weights. That is, the attention mapping parameters of the infrared branch and the RGB branch (including the weight matrices and biases of Q / K / V) are independently initialized and optimized separately during the training process to better adapt to the differences in spatial distribution and texture features between the two modalities. Therefore, the implementation process is not described in detail.
[0118] The difference lies in: the initial input features of the infrared image , that is, the normalized T image . Since the infrared image has only one channel, in order to be symmetric with the three-channel structure of the RGB image, it is repeated three times to obtain a three-channel input.
[0119] Specifically, the feature extraction steps of the T branch at each level are consistent with those of the RGB branch, including the Mobius self-attention mechanism, multi-head attention layer, feed-forward network layer, and residual connection in the hyperbolic space. The T image features generated at each level are represented as and are used to perform multi-modal feature fusion with the corresponding RGB level features in the subsequent stage. In the T branch, the infrared features are embedded into the hyperbolic space through the Mobius transformation to maintain the structural consistency with the RGB features. The hyperbolic mapping of the infrared features makes them in the same geometric space as the RGB features, which is more convenient for subsequent feature fusion.
[0120] S3: Feature fusion layer, that is, multiple Mobius cross-attention fusion layers are set, which correspond one-to-one with the Mobius self-attention layers of the RGB image feature extraction branch and the infrared image feature extraction branch. The Mobius cross-attention mechanism is used to align and fuse the RGB and T features, realizing information interaction and unified representation between modalities, so as to provide high-quality fusion features for further segmentation tasks.
[0121] As Figure 2 shown, this embodiment is provided with four fusion layers of Mobius cross-attention. For each fusion layer i, the inputs are the features and from the RGB and T branches respectively, and these features have been extracted in the hyperbolic space.
[0122] First, the RGB features are mapped into query vectors, and the T features are mapped into key vectors and value vectors:
[0123] ;
[0124] ;
[0125] ;
[0126] In the formula, , , are the query vector, key vector, and value vector corresponding to the fusion layer i, respectively; , , are the biases corresponding to the query vector, key vector, and value vector, respectively; , , are the trainable weights of the Möbius cross-attention layer corresponding to the query vector, key vector, and value vector, respectively, and are used for Möbius transformation in the hyperbolic space.
[0127] Then, calculate the fused features at each level in the hyperbolic space:
[0128] ;
[0129] In the formula, is the fused multi-scale feature corresponding to the fusion layer i, is the dimension of the key vector .
[0130] S4: Hyperbolic space decoder. By performing an upsampling operation to align the resolutions of the multi-scale features at different levels, and then fusing the multi-scale features of N levels and projecting them into the category space, the segmentation prediction result is generated through the Softmax function. The design purpose of the hyperbolic space decoder is to fuse the multi-scale features of four levels and generate pixel-level segmentation predictions while maintaining geometric consistency.
[0131] For the features at each level, first embed them into the same hyperbolic space dimension through a Möbius linear transformation. For the features at a lower resolution level, align them with the highest resolution level through an upsampling operation in the hyperbolic space. Specifically, the upsampling operation uses Möbius interpolation in the hyperbolic space to smoothly expand the low-resolution features to the target resolution while preserving the geometric properties of the hyperbolic space:
[0132] ;
[0133] ;
[0134] In the formula, represents the fused feature, represents the fused feature after upsampling, is the low-level feature to be upsampled, and Both represent the same function operation. Generally refers to low-resolution input features. α is an interpolation factor for controlling the upsampling multiple, which is used to control the scale and position of upsampling. For example represents 2x upsampling. The function maps the features in the hyperbolic space to the Euclidean space for interpolation. This can smoothly expand the low-resolution features in the Euclidean space.
[0135] In this embodiment, the setting of the interpolation factor α mainly refers to the hierarchical relationship in the network structure of the present invention (see Figure 2 ). In the feature fusion stage, the resolution of deeper layers (such as , etc.) is lower than the output of the shallow layer (such as ). Therefore, the feature maps are aligned to a unified spatial resolution through Möbius upsampling. Specifically, the value of α corresponds to the actual required spatial scale multiple. For example: if corresponds to 1 / 8 resolution, and the target layer is 1 / 1, then α = 8; if is 1 / 4, α = 4.
[0136] The features of each layer after upsampling are fused through Möbius addition to generate a comprehensive feature representation:
[0137] ;
[0138] Finally, the fused features are projected into the class space through a linear mapping, and the segmentation prediction is generated through the Softmax function:
[0139] ;
[0140] In the formula, , are the weight and bias respectively.
[0141] It should be understood that according to the foregoing implementation process, after collecting semantic segmentation samples (corresponding to RGB images and infrared images) and setting labels, network training is performed to optimize each weight value. Since the network training process is a conventional technical means in the art, it will not be described and limited in detail herein.
[0142] Among them, in order to improve the network performance, the present invention optimizes the loss function. Specifically: the alignment error between RGB and T features is calculated through the hyperbolic distance loss function, and the weights of the Möbius transformation are optimized during the training process so that the RGB and T features can be accurately aligned after fusion. In the Poincaré disk model, the hyperbolic distance is used to measure the similarity between feature points, and is defined as follows:
[0143] ;
[0144] Among them, p and q represent the RGB and T features in the hyperbolic space.
[0145] ;
[0146] In the formula, L is the total feature alignment loss value, is the Möbius embedding mapping function in the hyperbolic space, is the embedded feature of the i-th infrared image, is the embedded feature of the i-th visible light image, is the trade-off coefficient in the loss function, is the set of model trainable parameters, and N is the number of samples participating in the loss calculation.
[0147] In the embodiment of the present invention, by optimizing the loss function and the weight parameters of the training model, the RGB and T features are matched in the hyperbolic space. And the optimized loss function can effectively handle the modality difference and spatial distortion problems through the hyperbolic distance metric, and can further ensure the alignment accuracy of the RGB and T features.
[0148] It should be understood that in the above embodiment, the hyperbolic space is taken as an example. Therefore, all calculations are based on the construction of the hyperbolic space (negative curvature), and the core calculations involved include Möbius mapping, residual connection, upsampling interpolation, and feature alignment loss function, etc. These operations rely on the geometric properties of the hyperbolic space. In other feasible embodiments, if other types of spaces are used, corresponding adjustments are required, especially reflected in:
[0149] 1) Distance calculation function: If other spaces (such as spherical space) are used, the "arcosh-type hyperbolic distance function" should be replaced with the distance metric function of the corresponding space, such as spherical arc length, spherical inner product cosine distance, etc.;
[0150] 2) Möbius addition and multiplication: In this embodiment, Möbius operations are defined in the hyperbolic space. If other geometric spaces are used, the relevant operations need to be replaced with their equivalent mapping mechanisms in the spherical or elliptical space;
[0151] 3) Normalization and interpolation operations: Such as Norm(c) and Möbius interpolation, etc., also rely on the hyperbolic geometry model. Other spaces should introduce equivalent normalization methods and spatial interpolation strategies;
[0152] 4) Feedforward layer structure: Currently, linear transformation and activation methods embedded in the hyperbolic space are adopted. If other geometric spaces are used, an embedding function can be redefined based on the target space.
[0153] In summary, in this embodiment, the traditional vision Transformer layer is replaced with a feature extraction and fusion method based on Möbius transformation and hyperbolic space embedding. On the one hand, Möbius self-attention and cross-attention mechanisms are proposed. That is, in the feature extraction stage, RGB and T (infrared) image features are extracted in the hyperbolic space through multiple Möbius self-attention layers. Compared with the traditional self-attention mechanism based on Euclidean space, it cannot well handle the non-linear feature differences between infrared and visible light images. The Möbius self-attention mechanism set in the embodiment of the present invention generates features that are geometrically more in line with the non-linear feature changes in the actual physical space through Möbius matrix multiplication and addition in the hyperbolic space, thus effectively coping with modality differences and edge distortions. In the feature fusion stage, the Möbius cross-attention mechanism is used for alignment and fusion of multi-modal features. The cross-attention mechanism defined in the hyperbolic space can ensure the geometric consistency of RGB and T features, making the features between modalities match more closely in the hyperbolic space and improving the multi-modal fusion accuracy. On the other hand, this embodiment also realizes feature decoding in the hyperbolic space, that is, a decoder embedded in the hyperbolic space is proposed, and multi-scale upsampling of features is realized through Möbius interpolation. During the decoding process, by upsampling and fusing features of different resolutions, the spatial resolution can be effectively restored in the hyperbolic space while maintaining the consistency of feature geometric relationships. This design helps to generate more accurate segmentation results in the non-Euclidean space. On the third hand, this embodiment optimizes the loss function in the hyperbolic space, that is, a distance metric in the hyperbolic space is introduced into the loss function, and the hyperbolic distance in the Poincaré disk model is used to calculate the alignment error between RGB and T features. Compared with the traditional Euclidean distance, the distance in the hyperbolic space is more suitable for describing the non-linear similarity between modalities, and can effectively handle modality differences and spatial distortion problems during the optimization process, improving the alignment accuracy.
[0154] Experimental verification:
[0155] To verify the effectiveness and accuracy of the technical solution of this application, the present invention conducts experiments on the standard dataset MFNet for multi-modal semantic segmentation, and evaluates the performance of this technical solution in the multi-class object segmentation task. The MFNet dataset provides RGB and T (infrared) image pairs and their corresponding segmentation labels, covering multiple categories such as vehicles, pedestrians, and bicycles. The scenes are complex and diverse, with significant modality differences and blurred object boundaries. This technical solution uses a two-branch deep learning network, combined with the Möbius transformation mechanism, to efficiently extract and fuse RGB and T features in the hyperbolic space, aiming to improve the accuracy of multi-modal segmentation.
[0156] In this example, the MFNet dataset is first preprocessed, including normalizing and adjusting the resolution of RGB images and T images to meet the network input requirements. The preprocessed RGB and T images are respectively subjected to feature extraction through the Möbius self-attention module of the dual-branch network. Among them, the RGB branch focuses on extracting texture features in the scene, such as vehicle boundaries and road structures; the T branch highlights the target thermal distribution at night or under low-light conditions through infrared features. The features of the two modalities are aligned and fused in the hyperbolic space through the Möbius cross-attention mechanism, thus generating high-quality feature representations required for the semantic segmentation task.
[0157] Figure 3 (a) and Figure 3 (b) of this example further illustrate the mapping in the hyperbolic space, taking the feature fusion stage and using the Möbius cross-attention mechanism to align and fuse RGB and T features as an example. , , are all obtained by operating on the original features through hyperbolic space mapping. Figure 3 (a) and Figure 3 (b) regard Linear RGB and Hyperbolic RGB in the linear mapping in Euclidean space and the feature alignment process in hyperbolic space as , regard Linear RGB and Hyperbolic RGB as and . It can be seen that:
[0158] ;
[0159] The consistency information contained depends on , , the distance. The closer the distance, the higher the consistency. At the input stage, RGB and T features are respectively represented by light gray long arrows and dark gray long arrows. In Euclidean space, the features after linear mapping (short arrows) maintain the original direction, and the differences between modalities are not effectively reduced. In hyperbolic space, using Möbius mapping, the T feature (dark gray short arrow) moves closer to the RGB feature (light gray short arrow) direction, and the angle between modalities is significantly reduced, achieving more efficient feature alignment.
[0160] Figure 3 (b) The black arc represents the adjustment path (geodesic) of the T feature, showing the geometric change of the feature from the initial position to the aligned position. This property of the hyperbolic space enables more sufficient information interaction between modalities and provides higher-quality fused features for downstream segmentation tasks. 。Compared with the linear mapping in the traditional Euclidean space, the alignment effect in the hyperbolic space is more significant.
[0161] In the decoding stage, the fused multi-scale features are upsampled layer by layer and integrated with multi-modal features through a decoder embedded in the hyperbolic space, and finally a pixel-level segmentation result is generated. The experimental results in Table 1 show that the technical solution of this application achieves an mIoU (mean intersection over union) of 58.70% on the MFNet dataset, which is better than traditional methods, especially outstanding in target categories such as cars, pedestrians, and bicycles. For example, the IoU of vehicle segmentation reaches 89.29%, and the IoU of pedestrian segmentation reaches 75.20%, showing significant improvements compared with existing methods respectively. The visualization of the segmentation results further verifies the ability of this solution to accurately capture the boundaries and details of target regions in complex scenes.
[0162] In addition, the technical solution of this application also performs well in small target categories (such as speed bumps), achieving an IoU of 60.37%. This indicates that through the non-linear mapping in the hyperbolic space, this solution effectively alleviates the limitations of RGB and infrared images in modal differences and boundary blurring, and significantly improves the overall accuracy and robustness of multi-modal segmentation.
[0163] The experimental results fully prove the effectiveness and accuracy of the technical solution of this application in the multi-modal semantic segmentation task. Through the deep fusion of RGB and T features, this solution not only solves the core problem of modal alignment, but also shows significant advantages in dealing with small targets and complex boundaries, providing an innovative and practical solution for the multi-modal segmentation task.
[0164] Table 1 Partial segmentation results on the MFNet RGB-T dataset
[0165] Example 2:
[0166] The present invention also provides a semantic segmentation system based on the above semantic segmentation method, including: a multi-modal data preprocessing module, a multi-modal feature extraction module, a feature fusion module, a spatial decoder, and a training module.
[0167] The multi-modal data preprocessing module is used to obtain RGB images and infrared images and perform preprocessing, and the preprocessing at least includes normalization processing; the multi-modal feature extraction module is used to extract RGB image features and infrared image features respectively by using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence, and the Mobius self-attention layers of the two branches correspond one by one, where N is a positive integer greater than 1; the feature fusion module is used to fuse the RGB image features and the infrared image features by using a feature fusion layer composed of N Mobius cross-attention fusion layers to obtain multi-scale features; the spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, and then projects the multi-scale features of N levels after fusion into the category space, and generates a segmentation prediction result through the Softmax function; the training module is used to perform network training by using the sampled samples and their labels of RGB images and infrared images in semantic segmentation applications, and then perform semantic segmentation prediction by using the trained network.
[0168] In some embodiments, if the system loads a pre-trained network, it is also feasible not to include the training module, that is, to perform actual semantic segmentation by using the multi-modal data preprocessing module, the multi-modal feature extraction module, the feature fusion module, the spatial decoder, and the pre-trained network.
[0169] It should also be understood that for the specific implementation processes of each module, please refer to the above method content, which will not be elaborated herein. The above division of functional modules is only for illustration. In some embodiments, some functional modules can be combined, some functional modules can be split, and each functional module can be implemented in software, hardware, or a combination of software and hardware. Among them, the software and hardware devices include, but are not limited to, general computer devices, programmable gate arrays, digital signal processors, microprocessors, and their corresponding programming or burning software.
[0170] Embodiment 3:
[0171] The present invention also provides an electronic terminal, including: one or more processors; a memory storing one or more computer programs; wherein, the processor calls the computer program to implement: the steps of an RGB-T image multi-modal semantic segmentation method.
[0172] Among them, specifically execute:
[0173] Multi-modal data preprocessing, obtaining RGB images and infrared images and performing preprocessing, and the preprocessing at least includes normalization processing;
[0174] Multi-modal feature extraction: A dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch is used to extract RGB image features and infrared image features respectively. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence. The Mobius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1.
[0175] Feature fusion: An N-layer feature fusion layer based on Mobius cross-attention is used to fuse RGB image features and infrared image features to obtain multi-scale features.
[0176] The spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates a segmentation prediction result through the Softmax function.
[0177] In some embodiments, the network model can be pre-trained and can be called or loaded. In some embodiments, it is also necessary to implement network training using the sampled samples and their labels of RGB images and infrared images in semantic segmentation applications, and then use the trained network for semantic segmentation prediction.
[0178] For the specific implementation process of each step, please refer to the description of the foregoing method.
[0179] It should be understood that in the embodiments of the present invention, the so-called processor may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0180] The present invention also provides a computer-readable storage medium storing a computer program, and the computer program is called by a processor to implement the steps of a method for multi-modal semantic segmentation of RGB-T images.
[0181] Among them, specifically execute:
[0182] Multi-modal data preprocessing, obtaining RGB images and infrared images and performing preprocessing, where the preprocessing at least includes normalization processing;
[0183] Multi-modal feature extraction, using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch to extract RGB image features and infrared image features respectively. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence. The Mobius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1;
[0184] Feature fusion, using N fusion layers based on Mobius cross-attention to form a feature fusion layer for fusing RGB image features and infrared image features to obtain multi-scale features;
[0185] The spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates a segmentation prediction result through the Softmax function.
[0186] In some embodiments, the network model can be pre-trained and can be called or loaded. In some embodiments, it is also necessary to implement network training using the sampled samples and their labels of RGB images and infrared images in semantic segmentation applications, and then use the trained network for semantic segmentation prediction.
[0187] For the specific implementation process of each step, please refer to the description of the foregoing method.
[0188] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the readable storage medium can also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0189] Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned readable storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0190] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes. The present application is a device that, according to the flowchart and / or block diagram of the method, device (system), and computer program product of the embodiments of the present application, is used to generate instructions executed by a processor to implement the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.
[0191] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art according to the technical solution of the present invention, whether modified or replaced, as long as they do not depart from the purpose and scope of the present invention, also belong to the protection scope of the present invention.
Claims
1. A method for multi-modal semantic segmentation of RGB-T images, characterized in that: It includes the following steps: Multi-modal data preprocessing, obtaining RGB images and infrared images and performing preprocessing, where the preprocessing at least includes normalization processing; Multi-modal feature extraction, using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch to extract RGB image features and infrared image features respectively. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Mobius self-attention layers connected in sequence. The Mobius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1; Feature fusion, using N fusion layers based on Mobius cross-attention to form a feature fusion layer for fusing RGB image features and infrared image features to obtain multi-scale features; Among them, the Mobius self-attention mechanism of the Mobius self-attention layer and the Mobius cross-attention mechanism of the fusion layer based on Mobius cross-attention both use Mobius multiplication and Mobius addition to perform feature non-linear mapping in non-Euclidean space respectively, realizing the alignment and fusion of multi-modal features; The spatial decoder aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates a segmentation prediction result through the Softmax function; Among them, the sampling samples and their labels of RGB images and infrared images in semantic segmentation applications are used for network training, and then the trained network is used for semantic segmentation prediction.
2. The method according to claim 1, characterized in that: The non-Euclidean space is a hyperbolic space, or a spherical space or an elliptical space or a mixed space composed of any combination of hyperbolic space, spherical space, and elliptical space.
3. The method according to claim 2, characterized in that: If the non-Euclidean space is a hyperbolic space, the Mobius multiplication is expressed as: ; The Mobius addition is expressed as: ; In the formula, , are the Möbius multiplication symbol and the Möbius addition symbol in the hyperbolic space with curvature c, respectively; , , are custom variable symbols used to represent the objects of Möbius multiplication and Möbius addition, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, represents the result of standard matrix multiplication, is the modulus of the vector, represents the curvature parameter of the hyperbolic space, represents the inner product of.
4. The method according to claim 1, characterized in that: The data processing process of the Mobius self-attention layer is as follows: First, input the initial features into the Mobius self-attention layer. Among them, the initial features of the first Mobius self-attention layer are the preprocessed RGB images or infrared images; Then, through the Mobius mapping layer transformation, non-linear mapping of the features in non-Euclidean space is realized to obtain query vectors, key vectors, and value vectors: Next, calculate the attention weights in non-Euclidean space through the multi-head self-attention mechanism, and perform Mobius multiplication calculation with the value vectors after non-linear mapping to obtain a new feature representation; Next, use a feed-forward network layer fully embedded in non-Euclidean space to process the new feature representation. The feed-forward network layer consists of two Mobius linear transformations and non-linear activation; Finally, perform residual connection and normalization operations in non-Euclidean space on the output of the feed-forward network layer; Among them, the output features of the previous Mobius self-attention layer are used as the initial features of the next Mobius self-attention layer.
5. The method according to claim 4, characterized in that: If the non-Euclidean space is a hyperbolic space, the non-linear mapping of the features in non-Euclidean space through the Mobius mapping layer transformation is expressed as: ; ; ; Wherein, , , are the query vector, key vector, and value vector obtained by the non-linear mapping of the Möbius self-attention layer i, respectively, , , and , , are the trainable weights and biases of the Möbius mapping layer, used to map RGB features to the hyperbolic space; and are the Möbius multiplication and Möbius addition symbols in the Möbius transformation, respectively, is the initial feature input to the Möbius self-attention layer i; Calculate the attention weights in hyperbolic space through the multi-head self-attention mechanism , defined as: ; In the formula, is 's dimension, T is the matrix transpose symbol, and the attention weight and the value are used to calculate the new feature representation : ; The processing process of using a feed-forward network layer fully embedded in the hyperbolic space is: ; ; wherein, and are trainable weight matrices in hyperbolic space, and are bias vectors, is the result after the Möbius linear transformation and non-linear activation of the first layer in the feed-forward network layer, is the final output after the linear mapping and bias of the second layer; The output of the feedforward layer is obtained by using residual connections and normalization operations in the hyperbolic space to get , as follows: ; Among them, The function is: ; In the formula, z is a custom variable symbol.
6. The method according to claim 1, characterized in that: The processing process of the Mobius cross-attention fusion layer is as follows: First, map the RGB image features extracted by the Möbius self-attention layer at the same level on the RGB image feature extraction branch into query vectors; map the infrared image features extracted by the Möbius self-attention layer at the same level on the infrared image feature extraction branch into key vectors and value vectors; ; ; ; Wherein, , , are the query vector, key vector, and value vector corresponding to the i-th Möbius cross-attention fusion layer, respectively; , , are the biases corresponding to the query vector, key vector, and value vector, respectively; , , are the trainable weights of the cross-attention layers corresponding to the query vector, key vector, and value vector, respectively; Then, calculate the multi-scale features after fusion by the Möbius cross-attention fusion layer based on the query vectors, key vectors, and value vectors: ; In the formula, is the multi-scale feature corresponding to the i-th Mobius cross-attention fusion layer, is the key vector is the dimension.
7. The method according to claim 1, characterized in that: If the non-Euclidean space is a hyperbolic space, when training the network with the sampled samples and their labels of the RGB image and the infrared image, calculate the alignment error between the RGB image features and the infrared image features through the hyperbolic distance loss function; Among them, the alignment error is expressed as: ; Among them, the formula for the hyperbolic distance is: ; where \(L\) is the total feature alignment loss value, is the Möbius embedding mapping function in the hyperbolic space, is the embedded feature of the \(i\)-th infrared image, is the embedded feature of the \(i\)-th visible light image, is the trade-off coefficient in the loss function, is the set of trainable parameters, \(N\) is the number of samples participating in the loss calculation, and arcosh is the inverse hyperbolic cosine function.
8. A semantic segmentation system based on the method according to any one of claims 1 - 7, characterized in that: Including: A multi-modal data preprocessing module, used to obtain the RGB image and the infrared image and perform preprocessing, and the preprocessing at least includes normalization processing; A multi-modal feature extraction module, used to respectively extract the RGB image features and the infrared image features by using a dual-branch structure composed of an RGB image feature extraction branch and an infrared image feature extraction branch. Both the RGB image feature extraction branch and the infrared image feature extraction branch are provided with N Möbius self-attention layers connected in sequence, and the Möbius self-attention layers of the two branches correspond one by one, and N is a positive integer greater than 1; A feature fusion module, used to use N fusion layers based on Möbius cross-attention to form a feature fusion layer to fuse the RGB image features and the infrared image features to obtain multi-scale features; Among them, the Möbius self-attention mechanism of the Möbius self-attention layer and the Möbius cross-attention mechanism of the fusion layer based on Möbius cross-attention respectively perform non-linear mapping of features in the non-Euclidean space by using Möbius multiplication and Möbius addition to realize the alignment and fusion of multi-modal features; A spatial decoder, which aligns the resolutions of multi-scale features at different levels through upsampling operations, then fuses the multi-scale features of N levels and projects them into the category space, and generates a segmentation prediction result through the Softmax function; And / or a training module, used to train the network with the sampled samples and their labels of the RGB image and the infrared image in the semantic segmentation application, and then use the trained network to perform semantic segmentation prediction.
9. An electronic terminal, characterized in that: Including: One or more processors; A memory storing one or more computer programs; Among them, the processor calls the computer program to implement: The steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and the computer program is called by the processor to implement: The steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Cross-modal pedestrian re-identification method based on bimodal alignment
CN117935307A
Multi-modal image semantic segmentation method and system based on cross-level guide fusion
CN118864866A
Cited By
Chronic patient screening and personalized processing method and system based on large model
CN121237407A