Image semantic segmentation algorithm based on multi-modal feature fusion and application

The multimodal feature fusion method using federated learning solves the problem of feature extraction and fusion in low-latency and high-precision scenarios in existing image semantic segmentation algorithms, achieving more efficient image semantic segmentation results.

CN117152431BActive Publication Date: 2025-10-21NINGBO UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311049781.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-10-21
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

Existing image semantic segmentation algorithms perform poorly in real-world applications requiring low latency and high accuracy, making it difficult to effectively extract and fuse detailed features from images and text.

Method used

A federated learning-based multimodal feature fusion method is adopted. Single-modal feature extraction is completed independently through different extraction modules, and feature fusion is performed after the verification task is completed in the fusion module. Finally, the ASPP module is used for image semantic segmentation prediction.

Benefits of technology

It improves the detailed feature extraction capability of the image feature extraction layer, obtains more image and text information, and achieves better image semantic segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004403864570000052
    Figure BDA0004403864570000052
Patent Text Reader

Abstract

The application relates to the technical field of image semantic segmentation, in particular to an image semantic segmentation algorithm based on multi-modal feature fusion and application, which comprises a data set management module, an extraction module, a fusion module and a segmentation module, and specifically comprises the following steps: a single modal feature extraction task is independently completed by different extraction modules in a cooperative mode of a federal learning concept; the completion of the single modal feature extraction task is checked by the fusion module; after it is determined that the single modal feature extraction task is completed, the extracted single modal features are fused to obtain multi-modal fusion features; and finally, the multi-modal fusion features are taken as input, and an ASPP module is used to realize image semantic segmentation prediction. The application can effectively improve the ability of an image feature extraction layer to extract detailed features, more image detailed information can be obtained by extracting and fusing image and text features, and a better image semantic segmentation effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image semantic segmentation, and specifically to an image semantic segmentation algorithm based on multimodal feature fusion and its application. Background Art

[0002] Semantic segmentation is a key task in computer vision. Its primary goal is to identify the semantic category of each pixel in an image. As the foundation of intelligent environmental perception, semantic segmentation has been widely used in various artificial intelligence scenarios in recent years, such as human-computer interaction systems, autonomous driving, and navigation systems.

[0003] For real-world applications, low latency and high accuracy are often required, which makes the task of image segmentation more challenging.

[0004] Although the current deep learning-based image semantic segmentation algorithm has improved the image segmentation effect, due to the complexity of real-world application scenarios, existing semantic segmentation methods still face many challenges. Summary of the Invention

[0005] In order to improve the ability of the image feature extraction layer to extract detailed features and improve the image segmentation effect, the present invention provides an image semantic segmentation algorithm based on multimodal feature fusion, which can be applied in human-computer interaction systems, autonomous driving and navigation systems.

[0006] In order to achieve the above-mentioned purpose of the invention, the following technical solutions are provided:

[0007] An image semantic segmentation algorithm based on multimodal feature fusion has an architecture comprising: a dataset management module, an extraction module, a fusion module, and a segmentation module. Specifically, the algorithm utilizes the collaborative approach of the federated learning concept to delegate the extraction task of a single modality feature to different extraction modules for independent completion, and utilizes the fusion module to check the completion status of the single modality feature extraction task. After determining that all single modality feature extraction tasks have been completed, the extracted single modality features are fused to obtain multimodal fusion features. Finally, the multimodal fusion features are used as input, and the ASPP module is used to implement image semantic segmentation prediction.

[0008] Preferably, the specific steps of the image semantic segmentation algorithm are as follows:

[0009] Step 1: The extraction module E downloads the initial task model ITM from the dataset management module M;

[0010] Step 2: The extraction module E uses the embedded feature extraction network model to complete the extraction task specified in the initial task model ITM, extracts the single modal features, and transmits the single modal features to the fusion module F;

[0011] Step 3: The fusion module F fuses the single modal features to obtain multimodal fusion features, and transmits the multimodal fusion features to the segmentation module D for image semantic segmentation prediction.

[0012] Preferably, the dataset management module M stores an initial task model ITM, and the initial task model ITM includes an image semantic segmentation task Tiss and an interaction module IM subordinate to the image semantic segmentation task Tiss.

[0013] Preferably, the image semantic segmentation task Tiss includes: an image P waiting for semantic segmentation, a text T-I describing the target I to be segmented in the image P, a text T-II describing the target II to be segmented in the image P, and a text T-III describing the target III to be segmented in the image P.

[0014] Preferably, the extraction module E includes an extraction module EP, an extraction module ET-I, an extraction module ET-II, and an extraction module ET-III, which download the initial task model ITM from the data set management module M respectively.

[0015] Preferably, the extraction module EP obtains the task share t0, the task dispatch number Tdn, and the task record number Trn from the initial task model ITM, uses the Darknet53 network as the image feature extraction layer, and extracts the image feature f from the image P. P ;

[0016] Based on this, the extraction module EP performs the following operations:

[0017] (1) Calculate the task execution number Ten belonging to the task share t0 = h(f P ) Tdn ;

[0018] (2) The task execution number Ten and the image feature f P Transmitted to the fusion module F.

[0019] Preferably, the extraction module ET-Ⅰ obtains the task share t1, the task dispatch number Tdn-Ⅰ, and the task record number Trn-Ⅰ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target to be segmented I from the text T-Ⅰ. T-Ⅰ ;

[0020] Based on this, the extraction module ET-I performs the following operations:

[0021] (1) Calculate the task execution number Ten-I belonging to task share t1 = h(f T-Ⅰ ) Tdn-Ⅰ ;

[0022] (2) The task execution number Ten-Ⅰ and the text feature f T-Ⅰ Transmit to fusion module F;

[0023] The extraction module ET-Ⅱ combines the task execution number Ten-Ⅱ and the text feature f T-Ⅱ Transmitted to the fusion module F, the extraction module ET-Ⅲ takes the task execution number Ten-Ⅲ and the text feature f T-Ⅲ Transmitted to the fusion module F.

[0024] Preferably, the fusion module F confirms whether the single modal feature extraction task is completed by calculating the task completion parameters Ptcc-I and Ptcc-II and verifying whether the equation Ptcc-I=Ptcc-II holds.

[0025] Preferably, the fusion module F extracts the image features f P With the text feature f T-Ⅰ , text features f T-Ⅱ , text features f T-Ⅲ The fusion is performed to obtain the feature vector MFf-E of the multimodal fusion feature MFf, which is as follows:

[0026] MFf-E=F(f P-E W P , f T-Ⅰ W T-Ⅰ , f T-Ⅱ W T-Ⅱ , f T-Ⅲ W T-Ⅲ );

[0027] f P-E is the image feature f P The eigenvector of

[0028] F(·) represents the ReLU function;

[0029] W P is the image feature f P The projection weight matrix of

[0030] W T-Ⅰ is the text feature f T-Ⅰ The projection weight matrix of

[0031] W T-Ⅱ is the text feature f T-Ⅱ The projection weight matrix of

[0032] W T-Ⅲ is the text feature f T-Ⅲ The projection weight matrix of .

[0033] The image semantic segmentation algorithm based on multimodal feature fusion can be applied in human-computer interaction systems, automatic driving and navigation systems.

[0034] Compared with the prior art, the present invention has the following beneficial technical effects:

[0035] By utilizing the collaborative approach of the federated learning concept, the task of extracting single modal features is delivered to different extraction modules for independent completion, and the fusion module is used to check the completion status of the single modal feature extraction task. After determining that all single modal feature extraction tasks are completed, the extracted single modal features are fused to obtain multimodal fusion features. Finally, with the multimodal fusion features as input, the ASPP module is used to realize image semantic segmentation prediction. This framework design can effectively improve the ability of the image feature extraction layer to extract detailed features. Extracting features from images and texts and fusing them can also obtain more image detail information, achieving better image semantic segmentation effects. DETAILED DESCRIPTION

[0036] Example 1:

[0037] Image semantic segmentation algorithm based on multimodal feature fusion, its specific architecture includes: dataset management module M, extraction module E, fusion module F, segmentation module D;

[0038] The extraction module E extracts the features of the image and text in the dataset management module M. The features of the two are fused through the fusion module F. The obtained multimodal features are used for segmentation prediction using the segmentation module D;

[0039] The specific process of the above image semantic segmentation algorithm is as follows:

[0040] The extraction module E downloads the initial task model ITM from the data set management module M;

[0041] The extraction module E uses the embedded feature extraction network model to complete the extraction task specified in the initial task model ITM, extracts the single modal features, and transmits the single modal features to the fusion module F;

[0042] The fusion module F fuses the single modal features to obtain multimodal fusion features, and transmits the multimodal fusion features to the segmentation module D. The segmentation module D takes the multimodal fusion features as input and uses the ASPP module to realize image semantic segmentation prediction.

[0043] Example 2:

[0044] The dataset management module M stores the initial task model ITM, which includes the image semantic segmentation task Tiss and the interaction module IM belonging to the image semantic segmentation task Tiss;

[0045] The image semantic segmentation task Tiss includes: an image P waiting for semantic segmentation, a text T-I describing the target I to be segmented in the image P, a text T-II describing the target II to be segmented in the image P, and a text T-III describing the target III to be segmented in the image P;

[0046] The interaction module IM includes the following interaction rules:

[0047] Interaction rule 1: Let p be a large prime number, G1 and G2 be two multiplicative cyclic groups with the same order p, and g be the generator of G1;

[0048] Based on this: First, define a bilinear map e: G1×G1→G2. For any g,h∈G1, there exists a,b∈Z p , so that e(g a , h b ) = e(g, h) ab ;

[0049] Then define a vector hash function Among them, n numbers g1, g2, ..., g n ∈G1, vector v=(v1, v2, ..., v n )∈Z n p , for any two messages m1, m2 and two real numbers w1, w2, such that

[0050] Finally, define a hash function h(m)∈G2, where m represents the message;

[0051] Interaction rule 2: The interaction module IM divides the image semantic segmentation task Tiss into four execution shares, which are t0, t0, t0, and t3 respectively.

[0052] Based on interaction rules 1 and 2, interaction rule 3 is defined as follows:

[0053] Select the task dispatch number Tdn∈Z belonging to the task share t0 p * , and calculate the task record number Trn=g belonging to the task share t0 accordingly Tdn modp;

[0054] Select the task dispatch number Tdn-I∈Z belonging to the task share t1 p *, and calculate the task record number Trn-I=g belonging to task share t1 based on this Tdn -Imodp;

[0055] Select the task dispatch number Tdn-Ⅱ∈Z belonging to the task share t2 p * , and calculate the task record number Trn-Ⅱ=g belonging to task share t2 based on this Tdn-Ⅱ modp;

[0056] Select the task dispatch number Tdn-Ⅲ∈Z belonging to the task share t3 p * , and calculate the task record number Trn-Ⅲ=g belonging to task share t3 based on this Tdn-Ⅲ modp.

[0057] Example 3:

[0058] The extraction module E includes the extraction module EP, the extraction module ET-I, the extraction module ET-II, and the extraction module ET-III, which download the initial task model ITM from the data set management module M respectively;

[0059] The extraction module EP obtains the task share t0, task dispatch number Tdn, and task record number Trn from the initial task model ITM, uses the Darknet53 network as the image feature extraction layer, and extracts the image feature f from the image P. P ;

[0060] Based on this, the extraction module EP performs the following operations:

[0061] (1) Calculate the task execution number Ten belonging to the task share t0 = h(f P ) Tdn ;

[0062] (2) The task execution number Ten and the image feature f P Transmit to fusion module F;

[0063] The extraction module ET-Ⅰ obtains the task share t1, task dispatch number Tdn-Ⅰ, and task record number Trn-Ⅰ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target to be segmented Ⅰ from the text T-Ⅰ. T-Ⅰ ;

[0064] Based on this, the extraction module ET-I performs the following operations:

[0065] (1) Calculate the task execution number Ten-I belonging to task share t1 = h(f T-Ⅰ ) Tdn-Ⅰ;

[0066] (2) The task execution number Ten-Ⅰ and the text feature f T-Ⅰ Transmit to fusion module F;

[0067] The extraction module ET-Ⅱ obtains the task share t2, task dispatch number Tdn-Ⅱ, and task record number Trn-Ⅱ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target to be segmented Ⅱ from the text T-Ⅱ T-Ⅱ ;

[0068] Based on this, the extraction module ET-II performs the following operations:

[0069] (1) Calculate the task execution number Ten-II belonging to task share t2 = h(f T-Ⅱ ) Tdn-Ⅱ ;

[0070] (2) The task execution number Ten-Ⅱ and the text feature f T-Ⅱ Transmit to fusion module F;

[0071] The extraction module ET-Ⅲ obtains the task share t3, task dispatch number Tdn-Ⅲ, and task record number Trn-Ⅲ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target to be segmented Ⅲ from the text T-Ⅲ. T-Ⅲ ;

[0072] Based on this, the extraction module ET-III performs the following operations:

[0073] (1) Calculate the task execution number Ten-III belonging to task share t3 = h(f T-Ⅲ ) Tdn-Ⅲ ;

[0074] (2) The task execution number Ten-Ⅲ and the text feature f T-Ⅲ Transmitted to the fusion module F.

[0075] Example 4:

[0076] Fusion Module F:

[0077] On the one hand, the initial task model ITM is downloaded from the data set management module M, and the task record number Trn belonging to the task share t0, the task record number Trn-Ⅰ belonging to the task share t1, the task record number Trn-Ⅱ belonging to the task share t2, and the task record number Trn-Ⅲ belonging to the task share t3 are obtained from the initial task model ITM;

[0078] On the other hand, the task execution number Ten and image feature f transmitted by the extraction module EP are receivedP , extract the task execution number Ten-Ⅰ and text features f transmitted by module ET-Ⅰ T-Ⅰ , extract the task execution number Ten-Ⅱ and text features f transmitted by module ET-Ⅱ T-Ⅱ , extract the task execution number Ten-Ⅲ and text features f transmitted by module ET-Ⅲ T-Ⅲ ;

[0079] Based on this, the fusion module F performs the following operations:

[0080] (1) Calculate the task completion parameter Ptcc-I = e(Trn, h(f P ))×e(Trn-Ⅰ,h(f T-Ⅰ ))×e(Trn-Ⅱ,h(f T-Ⅱ ))×e(Trn-Ⅲ, h(Trn-Ⅲ));

[0081] (2) Calculate the task completion parameter Ptcc-II = e(g, Ten × Ten-I × Ten-II × Ten-III);

[0082] (3) If the equation Ptcc-Ⅰ=Ptcc-Ⅱ holds, it proves that the extraction task of a single modal feature is completed and the next step is started; otherwise, the next step cannot be executed;

[0083] (4) The extracted image features f P With the text feature f T-Ⅰ , text features f T-Ⅱ , text features f T-Ⅲ The fusion is performed to obtain the feature vector MFf-E of the multimodal fusion feature MFf, which is as follows:

[0084] MFf-E=F(f P-E W P , f T-Ⅰ W T-Ⅰ , f T-Ⅱ W T-Ⅱ , f T-Ⅲ W T-Ⅲ );

[0085] f P-E is the image feature f P The eigenvector of

[0086] F(·) represents the ReLU function;

[0087] W P is the image feature f P The projection weight matrix of

[0088] W T-Ⅰ is the text feature fT-Ⅰ The projection weight matrix of

[0089] W T-Ⅱ is the text feature f T-Ⅱ The projection weight matrix of

[0090] W T-Ⅲ is the text feature f T-Ⅲ The projection weight matrix of

[0091] The fusion module F transmits the feature vector MFf-E of the multimodal fusion feature MFf to the segmentation module D. The segmentation module D takes the multimodal fusion feature MFf as input and uses the ASPP module to perform image semantic segmentation prediction.

[0092] The Darknet53 network, GRU network, and ASPP module used in the above embodiments are existing technologies.

Claims

1. Image semantic segmentation algorithm based on multimodal feature fusion, characterized by: The architecture of the image semantic segmentation algorithm includes: a dataset management module, an extraction module, a fusion module, and a segmentation module. Specifically, the algorithm utilizes the collaborative approach of federated learning to assign the extraction task of a single modality feature to different extraction modules for independent completion. The fusion module is used to check the completion of the single modality feature extraction task. After determining that all single modality feature extraction tasks have been completed, the extracted single modality features are fused to obtain multimodal fusion features. Finally, the multimodal fusion features are used as input to implement image semantic segmentation prediction using the ASPP module. The specific steps of the image semantic segmentation algorithm are as follows: Step 1: The extraction module E downloads the initial task model ITM from the dataset management module M; Step 2: The extraction module E uses the embedded feature extraction network model to complete the extraction task specified in the initial task model ITM, extracts the single modal features, and transmits the single modal features to the fusion module F; Step 3: The fusion module F fuses the single modal features to obtain multimodal fusion features, and transmits the multimodal fusion features to the segmentation module D for image semantic segmentation prediction. The specific implementation process is as follows: The fusion module F performs the following procedures: Procedure 1: Download the initial task model ITM from the dataset management module M, and obtain the task record number Trn belonging to task share t0, the task record number Trn-I belonging to task share t1, the task record number Trn-II belonging to task share t2, and the task record number Trn-III belonging to task share t3 from the initial task model ITM; Procedure 2: Receive the task execution number Ten and image feature f transmitted by the extraction module EP P , extract the task execution number Ten-Ⅰ and text features f transmitted by module ET-Ⅰ T-Ⅰ , extract the task execution number Ten-Ⅱ and text features f transmitted by module ET-Ⅱ T-Ⅱ , extract the task execution number Ten-Ⅲ and text features f transmitted by module ET-Ⅲ T-Ⅲ ; Based on this, the fusion module F performs the following operations: (1) Calculate the task completion parameter Ptcc-Ⅰ=e(Trn,h(f P ))×e(Trn-Ⅰ,h(f T-Ⅰ ))×e(Trn-Ⅱ,h(f T-Ⅱ ))×e(Trn-Ⅲ, h(Trn-Ⅲ)); Where h(·)∈G1 is a hash function; e represents the bilinear mapping function: G1×G1→G2; G1 and G2 are two multiplicative cyclic groups of the same order p; The generator of G1 is g; p is a large prime number; (2) Calculate the task completion parameter Ptcc-Ⅱ=e(g, Ten×Ten-Ⅰ×Ten-Ⅱ×Ten-Ⅲ); (3) If the equation Ptcc-Ⅰ=Ptcc-Ⅱ holds, it proves that the extraction task of a single modal feature is completed and the next step is started; otherwise, the next step cannot be executed; (4) The extracted image features f P With the text feature f T-Ⅰ , text features f T-Ⅱ , text features f T-Ⅲ The fusion is performed to obtain the feature vector MFf-E of the multimodal fusion feature MFf, which is as follows: MFf-E=F(f P-E W P ,f T-Ⅰ W T-Ⅰ ,f T-Ⅱ W T-Ⅱ ,f T-Ⅲ W T-Ⅲ ); f P-E is the image feature f P The eigenvector of F(·) represents the ReLU function; W P is the image feature f P The projection weight matrix of W T-Ⅰ is the text feature f T-Ⅰ The projection weight matrix of W T-Ⅱ is the text feature f T-Ⅱ The projection weight matrix of W T-Ⅲ is the text feature f T-Ⅲ The projection weight matrix of The fusion module F transmits the feature vector MFf-E of the multimodal fusion feature MFf to the segmentation module D. The segmentation module D takes the multimodal fusion feature MFf as input and uses the ASPP module to perform image semantic segmentation prediction.

2. The image semantic segmentation algorithm based on multimodal feature fusion according to claim 1, characterized in that: The dataset management module M stores an initial task model ITM, which includes an image semantic segmentation task Tiss and an interaction module IM belonging to the image semantic segmentation task Tiss.

3. The image semantic segmentation algorithm based on multimodal feature fusion according to claim 2, characterized in that: The image semantic segmentation task Tiss includes: an image P waiting for semantic segmentation, a text T-I describing an object I to be segmented in the image P, a text T-II describing an object II to be segmented in the image P, and a text T-III describing an object III to be segmented in the image P.

4. The image semantic segmentation algorithm based on multimodal feature fusion according to claim 3 is characterized in that: The extraction module E includes an extraction module EP, an extraction module ET-I, an extraction module ET-II, and an extraction module ET-III, which download the initial task model ITM from the data set management module M respectively.

5. The image semantic segmentation algorithm based on multimodal feature fusion according to claim 4 is characterized in that: The extraction module EP obtains the task share t0, task dispatch number Tdn, and task record number Trn from the initial task model ITM, uses the Darknet53 network as the image feature extraction layer, and extracts the image feature f from the image P. P ; Based on this, the extraction module EP performs the following operations: (1) Calculate the task execution number Ten belonging to the task share t0 = h(f P ) Tdn ; (2) The task execution number Ten and the image feature f P Transmitted to the fusion module F.

6. The image semantic segmentation algorithm based on multimodal feature fusion according to claim 4, characterized in that: The extraction module ET-Ⅰ obtains the task share t1, task dispatch number Tdn-Ⅰ, and task record number Trn-Ⅰ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target to be segmented Ⅰ from the text T-Ⅰ. T-Ⅰ ; Based on this, the extraction module ET-I performs the following operations: (1) Calculate the task execution number Ten-I belonging to task share t1 = h(f T-Ⅰ ) Tdn-Ⅰ ; (2) The task execution number Ten-Ⅰ and the text feature f T-Ⅰ Transmit to fusion module F; The extraction module ET-Ⅱ obtains the task share t2, task dispatch number Tdn-Ⅱ, and task record number Trn-Ⅱ from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target II to be segmented from the text T-Ⅱ. T-Ⅱ ; Based on this, the extraction module ET-II performs the following operations: (1) Calculate the task execution number Ten-Ⅱ belonging to task share t2 = h(f T-Ⅱ ) Tdn-Ⅱ ; (2) The task execution number Ten-Ⅱ and the text feature f T-Ⅱ Transmit to fusion module F; The extraction module ET-III obtains the task share t3, the task dispatch number Tdn-III, and the task record number Trn-III from the initial task model ITM, uses the GRU network as the text feature extraction layer, and extracts the text feature f of the target III to be segmented from the text T-III. T-Ⅲ ; Based on this, the extraction module ET-III performs the following operations: (1) Calculate the task execution number Ten-Ⅲ=h(f T-Ⅲ ) Tdn-Ⅲ ; (2) The task execution number Ten-Ⅲ and the text feature f T-Ⅲ Transmitted to the fusion module F.

7. The image semantic segmentation algorithm based on multimodal feature fusion according to any one of claims 1 to 6, characterized in that: Image semantic segmentation algorithms based on multimodal feature fusion can be applied in human-computer interaction systems, autonomous driving and navigation systems.

Citation Information

Patent Citations

  • Multi-modal federated learning privacy protection method and system

    CN115859367A

  • Data augmentation method and system based on multimodal image federated segmentation

    CN116580188A