A reference image segmentation method and system based on multi-level feature fusion
Through the multi-level feature fusion method, lightweight fusion blocks and single-shot attention mechanisms are used to solve the problem of heavy hardware resource burden in the prior art, and efficient image and text information fusion and segmentation are achieved, reducing computing overhead.
Patent Information
- Application Number
- CN202411248576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-09-06
AI Technical Summary
The existing multi-scale feature fusion method has a large calculation overhead in reference image segmentation, heavy hardware resource burden, and it is difficult to efficiently process cross-modal tasks of image and text information.
Using a method based on multi-level feature fusion, text information is encoded into text features, the original image is encoded into image feature maps of multiple different scales, and feature fusion is performed through lightweight fusion blocks and a single-shot attention mechanism, combining pixel decoding and Query decoding to generate a segmented map of the target information.
The amount of parameters in the feature fusion and decoding stages is reduced, the burden on hardware resources is reduced, and the efficient fusion and segmentation effect of image and text information is ensured.
Smart Images

Figure CN119049058B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and more particularly, to a reference image segmentation method and system based on multi-level feature fusion. Background Art
[0002] Reference image segmentation is a cross-modal task that involves processing two sensory information of vision and text. Given a picture and a reference text pointing to an object in the picture, it processes the picture and the reference text to locate and segment the target object under the indication of the reference text, and generates a segmentation map of the area where the target object is located. This process involves pixel-level reasoning of two modalities of image and text. To obtain an ideal effect, it is necessary to fully align and deeply fuse image and text information.
[0003] Existing multi-scale feature fusion methods often use the attention mechanism. However, the attention mechanism itself has a certain computational cost, and the model has a high complexity in the feature fusion stage, imposing a large burden on hardware resources. Summary of the Invention
[0004] To solve the above at least one defect, the present application proposes a reference image segmentation method and system based on multi-level feature fusion to solve the problem of the large burden on hardware resources caused by multi-scale feature fusion.
[0005] According to the first aspect of the present application, there is provided a reference image segmentation method based on multi-level feature fusion, including:
[0006] Encoding text information into text features, and encoding an original image with target information into multiple image feature maps of different scales;
[0007] Fusing each scale of image feature map with text features respectively to generate multiple image feature maps of different scales that fuse text information;
[0008] Performing pixel decoding on the multiple scales of image feature maps that fuse text information to obtain mask features and multiple enhanced image feature maps of different scales;
[0009] Performing Query decoding on the multiple enhanced image feature maps of different scales by using the mask features and the text features to obtain a segmentation map containing target information.
[0010] Preferably, in the process of fusion, each of the multiple scales of image features and the text features undergoes an attention mechanism operation once. After the result obtained is element-wise multiplied with the projection result of the image features, the result after the element-wise multiplication is projected to obtain multiple scales of image feature maps that fuse text information;
[0011] The primary attention mechanism operation includes performing matrix multiplication on image feature maps of multiple different scales and the text features, normalizing the obtained result, and then performing matrix multiplication on the normalized result and the text features.
[0012] Preferably, the fusion process is specifically represented by the following formula:
[0013] Q q =V i W Q1 ,K k =FW K1 ,V v =FW V1
[0014]
[0015] Among them, the subscript i represents the i-th scale, V i represents the image feature map of the i-th scale, F represents the text feature, W Q1 、W Q2 、W K1 、W V1 、W F are all linear transformation matrices; Q q represents the result after multiplying the image feature map of the i-th scale by W Q1 ; K k represents the result after multiplying the text feature by W K1 ; V v represents the result after multiplying the text feature by W V1 ; ⊙ represents element-wise multiplication; Softmax represents the normalization function; d k represents the dimension of the key vector, Fuse i (V i ,F) represents the image feature map of the i-th scale fused with text information, and n represents the total number of scales.
[0016] Based on the fusion process, each stage of image feature encoding is independent of the fusion operation and is not restricted by the completion of the feature fusion of the previous scale, ensuring the flexibility of feature encoding.
[0017] Preferably, encoding the text information into text features specifically includes:
[0018] Presetting text information and converting the text information into a feature sequence;
[0019] Encoding the feature sequence to obtain text features,
[0020] and / or,
[0021] Encoding the original image with target information into image feature maps of multiple different scales specifically includes:
[0022] Encoding the original image with target information, and extracting a feature map at each encoding stage to obtain image feature maps of multiple different scales.
[0023] In an optional solution, a tokenizer of a text encoder is used to convert the text information into a feature sequence, and specifically, Bert-base can be selected as the text encoder;
[0024] Using an image encoder to encode the original image with target information, and different mainstream image encoders can be selected as the visual backbone network for the image encoder, such as Resnet50, Swin-Transformer-Tiny, and
[0025] Swin-Transformer-Base, etc.
[0026] Preferably, the pixel decoding adopts a multi-scale deformable attention module and a feed-forward neural network, and decodes the image feature maps of multiple scales fused with the text information based on the attention module and the feed-forward neural network to perform interaction between feature maps of different scales, obtaining a mask feature and enhanced image feature maps of multiple different scales.
[0027] Preferably, the Query decoding adopts a learnable Object Query decoder, and the Object Query decoder includes multiple decoding layers. Each decoding layer includes a text cross-attention module, an image cross-attention module, and a feed-forward neural network FFN. The Query decoding process includes:
[0028] Assuming there are N decoding layers, initializing the Query to represent the target information and using it as the output to enter the first decoding layer;
[0029] For the first to Nth decoding layers, receiving the output of the previous decoding layer and interacting with the text feature through the text cross-attention module to generate a Query carrying the text information. The Query carrying the text information is processed by dot product segmentation with the mask feature to obtain an attention mask. The attention mask, the enhanced image feature map, and the Query carrying the text information are jointly interacted through the image cross-attention module, and the result after interaction is enhanced by the feed-forward neural network FFN in terms of non-linear expression ability and used as the output;
[0030] For the Nth decoding layer, the result after interaction through the image cross-attention module is enhanced with non-linear expression ability by the feed-forward neural network FFN, and then processed by dot product segmentation with the mask feature to obtain a segmentation map containing target information.
[0031] In an alternative solution, to avoid the unstable Query-Target bipartite matching problem that may be caused by multiple Object Queries, the Query decoder uses a single learnable Object Query, which can alternately interact with image and text information, and sequentially query multi-scale image information and text information through multiple layers of Query decoder layers to update itself.
[0032] Before the Query carrying text information enters the image cross-attention module, it first undergoes dot product segmentation processing with the mask feature to obtain a stage output. The stage output undergoes downsampling and binarization operations to obtain a binarization operation result. The weighted masking operation is performed on the binarization operation result, and the result after weighted masking is used as the attention mask of the image cross-attention mechanism. The attention mask acts as the region of interest, excluding the regions of no interest to reduce noise interference;
[0033] The Query carrying text information interacts with the enhanced image feature map and the attention mask through the image cross-attention module together. The result after interaction is used as the output to enter the next decoding layer, and / or, the result after interaction undergoes dot product segmentation processing with the mask feature to obtain a segmentation map containing target information.
[0034] Optionally, the Query decoding uses the weighted sum of the dice loss function and the cross-entropy loss function ce as the final loss function. The calculation formula of the final loss function is expressed as:
[0035]
[0036] where N is the number of layers of the Query decoder, pred j is the stage output, j represents the jth decoding layer, target is the segmentation map containing target information, λ dice is the weight of the dice loss function, λ ce is the weight of the cross-entropy loss function ce, dice_loss and ce_loss are the weighted results of the dice loss function and the cross-entropy loss function ce respectively, and loss is the final loss function.
[0037] According to the second aspect of the present application, a reference image segmentation system based on multi-level feature fusion is provided, including an encoding module, a pixel decoding module, a Query decoding module, and a fusion module;
[0038] The encoding module includes a text encoding module and an image encoding module. The text encoding module is used to encode text information into text features, and the image encoding module is used to encode the original image with target information into multiple image feature maps of different scales;
[0039] The fusion module is used to fuse the image feature maps of each scale with the text features respectively to generate multiple image feature maps of different scales that fuse text information;
[0040] The pixel decoding module is used to perform pixel decoding on the image feature maps of multiple scales that fuse text information to obtain mask features and multiple enhanced image feature maps of different scales;
[0041] The Query decoding module is used to perform Query decoding on the multiple enhanced image feature maps of different scales by using the mask features and the text features to obtain a segmentation map containing target information.
[0042] According to the third aspect of the present application, an electronic device is provided, including:
[0043] A memory for storing one or more computer programs;
[0044] A processor, when the one or more computer programs are executed by the processor, implementing a reference image segmentation method based on multi-level feature fusion according to the first aspect above.
[0045] According to the fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement a reference image segmentation method based on multi-level feature fusion according to the first aspect above when executed.
[0046] Based on any of the above aspects, a reference image segmentation method and system based on multi-level feature fusion provided by the embodiments of the present application adopt a fusion strategy based on lightweight fusion blocks to reduce the number of parameters in the stage of fusing text and image features and the decoding stage, thereby reducing the burden on hardware resources. In addition, the fusion block is based on a single attention mechanism, which ensures the fusion level of the bimodal while having a small number of parameters and simple operations. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 Flowchart of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0049] Figure 2 Overall framework diagram of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0050] Figure 3 Schematic diagram of the fusion process of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0051] Figure 4 Visualization diagram of the feature attention map of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0052] Figure 5 Main structure diagram of the pixel decoder and Query decoder of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0053] Figure 6 Schematic diagram of the functional modules of a reference image segmentation system based on multi-level feature fusion provided in this embodiment.
[0054] Figure 7 Schematic diagram of the structure of the electronic device provided in this embodiment.
[0055] Figure 8 Schematic application scenario diagram of a reference image segmentation method based on multi-level feature fusion provided in this embodiment.
[0056] Figure 9 Data table of the number of parameters of the system, fusion block, and pixel decoder when different image encoders are used in this embodiment. Icons: 100, server; 200, terminal; 501, Text Cross-Attention Module Attn t ; 502, Image Cross-Attention Module Attn v ; 503, Feed-Forward Neural Network FFN; 504, Object Query; 601, Encoding Module; 602, Fusion Module; 603, Pixel Decoding Module; 604, Query Decoding Module; 710, Electronic Device; 711, Memory; 712, Processor; 713, Communication Module; 714, Input / Output Interface; 715, Bus. Detailed implementation manners
[0057] The accompanying drawings of this application are only for illustrative purposes and should not be construed as a limitation to this application. To better illustrate the following embodiments, some components in the drawings will be omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0058] In order to enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0059] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0060] Existing multi-scale feature fusion methods often adopt the attention mechanism, and the attention mechanism itself has a certain computational overhead, and the complexity of the model is relatively high in the feature fusion stage, which places a large burden on hardware resources.
[0061] On the one hand, a reference image segmentation method based on multi-level feature fusion provided by an embodiment of this application, as Figure 1 shown, the method at least includes,
[0062] S1. Encode the text information into text features, and encode the original image with target information into multiple image feature maps of different scales;
[0063] S2. Respectively fuse each scale of the image feature map with the text features to generate multiple image feature maps of different scales that fuse the text information;
[0064] S3. Perform pixel decoding on the multiple-scale image feature maps that fuse the text information to obtain mask features and multiple image feature maps of different scales that are enhanced;
[0065] S4. Use the mask feature and the text feature to perform Query decoding on the enhanced image feature maps of multiple different scales to obtain a segmentation map containing target information.
[0066] The following combines specific embodiments and Figure 2 , and details the technical solutions provided in this application.
[0067] S1. Encode text information into text features, and encode the original image with target information into image feature maps of multiple different scales.
[0068] In this embodiment, the text encoder G t is used to encode text information into text feature F, and the image encoder G v is used to encode the original image with target information into multiple different scale image feature maps V, where p represents the total number of scales.
[0069] Optionally, select Bert-base as the text encoder G in the embodiments of this application t . In this embodiment, the text encoding process may include
[0070] Preset text information T, and use the tokenizer of the text encoder G t to convert it into a token sequence e, and then send the token sequence e into the text encoder G t to obtain text feature F. The encoding process can be expressed as:
[0071]
[0072] where L is the sequence length, and C l represents the text feature dimension number, represents the parameter quantity of the text feature.
[0073] Optionally, the image encoder G v can adopt different mainstream image encoders as the visual backbone network, such as Res-50, Swin-Tiny, and Swin-Base, etc.
[0074] Exemplarily, in this embodiment, the image encoding process includes
[0075] For a three-channel picture with height and width of H and W respectively use the image encoder G v to extract its 4-scale image feature maps, and the process can be expressed as:
[0076]
[0077] where Among them, H i and W i respectively represent the height and width of the image feature map at the i-th scale. represents the number of dimensions of the image features in the i-th image channel. i represents the i-th scale. The heights and widths of the feature maps at the four scales are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the height and width of the original image respectively. Each scale of feature map is extracted from one encoding stage. V i represents the image feature map at the i-th scale. V represents the set of feature maps. I i represents the image feature map at the i-th size.
[0078] S2. The image feature maps at each scale are respectively fused with the text features to generate multiple different-scale image feature maps fused with text information.
[0079] In this embodiment, a lightweight fusion block based on a single attention mechanism is adopted to perform an attention mechanism operation on the multiple different-scale image feature maps V i respectively with the text feature F. Among them, the image feature acts as the query, and the text feature acts as the key and value. The specific process includes:
[0080] The multiple different-scale image feature maps V i are respectively multiplied by the text feature F in matrix multiplication, and the obtained result is normalized. The normalized result is multiplied by the text feature F in matrix multiplication, and the obtained result is respectively multiplied element-wise by the corresponding image feature map V i projection result. The result after element-wise multiplication is projected to obtain multiple different-scale image feature maps V i ′ ,
[0081] The main function of the fusion is to incorporate image features at different semantic levels and scale sizes into text information, so that image-text features at different levels can all perceive the text information and align the image semantics with the text semantics.
[0082] Specifically, as Figure 3 shown, the fusion process can be specifically represented by the following formula:
[0083] Q q =V i W Q1 , K k =FW K1 , V v =FW V1
[0084]
[0085] Among them, the subscript i represents the i-th scale, V′ represents the image feature map fused with text features, and V i represents the image feature map of the i-th scale, F represents text features, and W Q1 、W Q2 、W K1 、W V1 、W F are all linear transformation matrices; Q q represents the result after multiplying the image feature map of the i-th scale by W Q1 ; K k represents the result after multiplying the text features by W K1 ; V v represents the result after multiplying the text feature F by W V1 ; ⊙ represents element-wise multiplication; Softmax represents the normalization function; d k represents the dimension of the key vector, and Fuse i (V i , F) represents the image feature map of the i-th scale fused with text information.
[0086] Exemplarily, Figure 4 represents the visualization of the feature attention map within the fusion block during the fusion process. Above the text is the original image. (a) to (d) are the attention maps of the input image features and the word "wheelchair" in the input text, and (e) is the final mask of the model for the target image.
[0087] It can be understood that (a) indicates that at the initial stage of the fusion of the image and text, the fusion block pays less attention to the lower semantic levels in the image features, mainly focusing on information such as the texture, contour, and key points of the image, which is similar to the nature of traditional convolutional neural networks. The evolution process of the attention maps from (a) to (d) shows that as the image encoding stage progresses, the degree of fusion of image-text features gradually deepens, and the attention area of the fusion block gradually changes from global texture contour information to the position information of the local reference object referred to by the text content.
[0088] Based on the above fusion process, each stage of the image feature encoding is independent of the fusion operation and is not restricted by whether the feature fusion of the previous scale is completed, ensuring the flexibility of feature encoding.
[0089] S3. The image feature maps of multiple scales fused with text information are pixel-decoded to obtain mask features and multiple enhanced image feature maps of different scales.
[0090] In an optional implementation manner, Multi-Scale Deformable with a multi-scale deformable attention mechanism can be used as the pixel decoder G p , which is used to model the interaction between feature maps of different scales. The pixel decoder Gp comprises multiple pixel decoding layers, each pixel decoding layer comprising a multi-scale deformable attention module and a feed-forward neural network with two layers of non-linear activation functions, the pixel decoder G p decodes the image feature maps of multiple scales fused with text information to obtain a mask feature and multiple enhanced image feature maps of different scales, the pixel decoder G p The decoding process of can be expressed as:
[0091]
[0092] where, Fv i represents the enhanced image feature map of the i-th scale, M is the mask feature,, V i ′ is the image feature map of the i-th scale fused with text features. represents the number of parameters of the enhanced image feature map, represents the number of parameters of the mask feature, C d represents the number of dimensions, H i 、W i are the height and width of the i-th enhanced image feature map Fv i respectively, i represents the i-th scale, H m 、W m are the height and width of the m-th mask feature M respectively, m represents the m-th scale.
[0093] Exemplarily, when the original image is a three-channel picture I with height and width of H and W respectively, the image encoder G v is used to extract its image feature maps of 4 scales, and after fusion, the image feature maps of 4 scales fused with text information are obtained. After pixel decoding, 3 scales of enhanced image feature maps Fv i are obtained, and the height and width of the 3 scales of enhanced image feature maps Fv i are 1 / 32, 1 / 16, and 1 / 8 of the height H and width W of the corresponding original picture I in sequence.
[0094] In this embodiment, Swin-Base, Swin-Tiny, and Res-50 are respectively used as the visual backbone network in the image feature extraction stage, and the effects of the fusion block and the pixel decoder on the model parameters and running speed are measured. The results are as Figure 9 shown,
[0095] When Swin-Base, Swin-Tiny, and Res-50 are respectively used as the feature extraction backbone network, the number of parameters is approximately 211 million, 142 million, and 160 million in sequence.
[0096] It is understandable that in Swin-base and Swin-tiny, the number of parameters of the fusion block accounts for no more than 5% of the total number of parameters, while in Res-50, the number of parameters of the fusion block accounts for nearly 17%. This is because the feature dimension of the Res-50 backbone network is larger than that of the other two backbone networks.
[0097] In the three networks, the proportion of the number of parameters of the pixel decoder does not exceed 2%. Considering the impact of the pixel decoder and the fusion block on the segmentation performance, these two modules only bring a small parameter burden and have considerable parameter efficiency.
[0098] S4. Use the masked feature and the text feature to perform Query decoding on the enhanced image feature maps of multiple different scales to obtain a segmentation map containing target information.
[0099] Preferably, the Query decoding uses a learnable Object Query decoder, which is used to combine image information and text information and generate information representing the prototype of the target object.
[0100] In an alternative solution, in order to avoid the unstable Query-Target bipartite matching problem that may be caused by multiple Object Queries, the Query decoder uses a single learnable Object Query 504, which alternately interacts with image and text information, and alternately queries multi-scale image information and text information through multiple layers of Query decoder layers in turn to update itself.
[0101] In this embodiment, as Figure 5 shown, the Object Query decoder includes N decoding layers, and each decoding layer includes a text cross-attention module Attn t 501, an image cross-attention module Attn v 502 and a feed-forward neural network FFN 503. Among them, FFN 503 is composed of two linear layers with residual connections and non-linear activation functions.
[0102] The Object Query 504 represents the prototype of the target object with a certain feature and is responsible for detecting / segmenting such objects in the image during the decoding process.
[0103] Exemplarily, the Query decoding process includes,
[0104] Randomly initialize a single learnable Object Query 504 using a standard normal distribution to represent the target object, denoted as Q0, where, Represents the number of parameters of the Object Query, C q Represents the dimension of the Object Query, and the Q0 enters the first decoding layer as the output;
[0105] For the m-th decoding layer (0 < m < N), the text cross-attention module Attn in the decoding layer t 501 receives the output Q0 of the previous decoding layer as the query of the attention mechanism, and uses F as the key and value. Through the text cross-attention module Attn t 501 performs interaction to generate a Query carrying text information, denoted as The process can be expressed as:
[0106]
[0107] Among them, Q represents the query of the attention mechanism, K represents the key of the attention mechanism, Val represents the value of the attention mechanism, and C l represents the number of text feature dimensions, are all learnable weight matrices in the attention mechanism.
[0108] The Before entering the image cross-attention module Attn v 502, it is first dot-multiplied and segmented with the mask feature M to obtain the stage output pred m , in this embodiment, a simple dot-multiplication segmentation head is introduced to perform the segmentation map generation operation. The segmentation head consists of a 3-layer linear feed-forward network MLP with a non-linear activation function. The generation process of the stage output pred m can be expressed as:
[0109]
[0110] The stage output pred m Undergoes downsampling and binarization operations (the threshold value can be set to 0.5) to obtain M m , performs a weighted masking operation on M m , to obtain Finally, is used as the attention mask of the image cross-attention mechanism. For the i-th layer of the image attention cross-module, the input enhanced image feature map Fv i and This masked attention mechanism can be expressed as:
[0111]
[0112] Among them, MaskedAttn vm represents the attention mask of the m-th layer, C l represents the number of dimensions of the text feature,
[0113] are all learnable weight matrices in the attention mechanism.
[0114] It can be understood that the attention mask acts as a soft region of interest, which can exclude the uninterested regions in advance to reduce the interference of noise. Among them, is obtained through operations such as binarization, downsampling, and tensor flattening of the stage output, The value at the (x, y) position can be calculated by the following formula:
[0115]
[0116] interacts with the mask feature M, and the processed result interacts with the enhanced image feature map Fv i successively passes through the image cross-attention module Attn v 502 and the feed-forward neural network FFN503 to interact and obtain Q m , the Q m serves as the output of the m-th decoding layer and enters the (m + 1)-th decoding layer;
[0117] For the N-th decoding layer, it receives the output Q of the (N - 1)-th decoding layer N-1 , the Q N-1 interacts with the text feature through the text cross-attention module Attn t 501 to generate a Query carrying text information. The Query carrying text information interacts with the mask feature M, and the processed result interacts with the enhanced image feature map and the Query carrying text information through the image cross-attention module Attn v 502 to interact. The result after interaction enhances the non-linear expression ability through the feed-forward neural network FFN, and then performs dot product segmentation processing with the mask feature M to obtain a segmentation map containing target information.
[0118] In this embodiment, the Query decoder uses the weighted sum of the dice loss function and the cross-entropy loss function ce as the final loss function. The calculation formula of the final loss function is expressed as:
[0119]
[0120] Among them, N is the number of layers of the Query decoder, pred jis the stage output of the j-th decoding layer, where j represents the j-th decoding layer, target is the segmentation map containing target information, and λ dice is the weight of the dice loss function, λ ce is the weight of the cross-entropy loss function ce. dice_loss and ce_loss are the weighted results of the dice loss function and the cross-entropy loss function ce respectively, and loss is the final loss function.
[0121] On the other hand, this embodiment provides a reference image segmentation system based on multi-level feature fusion, as Figure 6 shown, including
[0122] An encoding module 601, including a text encoding module and an image encoding module. The text encoding module is used to encode text information into text features, and the image encoding module is used to encode the original image with target information into multiple image feature maps of different scales;
[0123] A fusion module 602 is used to fuse the image feature maps of each scale with the text features respectively to generate multiple image feature maps of different scales that have fused text information;
[0124] A pixel decoding module 603 is used to perform pixel decoding on the image feature maps of multiple scales that have fused text information to obtain mask features and multiple enhanced image feature maps of different scales;
[0125] A Query decoding module 604 is used to perform Query decoding on the enhanced image feature maps of multiple different scales by using the mask features and the text features to obtain a segmentation map containing target information.
[0126] It can be understood that the above system embodiment and the above method embodiment can correspond to each other. Similar descriptions of the system embodiment can refer to the method embodiment. To avoid repetition, it will not be elaborated here. A reference image segmentation system based on multi-level feature fusion provided by an embodiment of the present application can execute a reference image segmentation method based on multi-level feature fusion provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. The functional modules of the reference image segmentation system based on multi-level feature fusion can be implemented in the form of hardware, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules.
[0127] Optionally, the software module can be located in a random access memory, a read-only memory, a programmable read-only memory, a flash memory, an electrically erasable programmable memory, a register, and other storage media are all possible. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiment.
[0128] The present application embodiment provides an electronic device 710, whose structure is as follows: Figure 7 The electronic device 710 may be Figure 8 The server 100 or the terminal 200 is shown.
[0129] like Figure 7 As shown, the electronic device 710 includes a memory 711, a processor 712, a communication module 713 and an input / output interface 714, etc. Optionally, the memory 711, the processor 712, the communication module 713 and the input / output interface 714 can be connected and communicated through a bus 715.
[0130] The memory 711 is used to store one or more computer programs and transfer the code of the computer program to the processor 712; when the one or more computer programs are executed by the processor 711, a reference image segmentation method based on multi-level feature fusion in an embodiment of the present application is implemented.
[0131] Optionally, the electronic device 710 can be connected to a network via a communication module 713 to communicate with other devices, such as a terminal or a server, through the network to achieve data interaction. The electronic device 710 can be various forms of digital computers, such as desktop computers, servers, workstations, mainframe computers or other types of computers. The electronic device 710 can also be various forms of mobile terminals, such as smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.) and other similar mobile terminals.
[0132] Optionally, the electronic device 710 can be connected to a required input / output device, such as a keyboard, a display device, etc., through the input / output interface 714. The electronic device 710 itself can have a display device, and can also be connected to other external display devices through the input / output interface 714. Optionally, a storage device, such as a hard disk, can be connected through the input / output interface 714, so that data in the electronic device 710 can be stored in the storage device, or data in the storage device can be read, and data in the storage device can also be stored in the memory 711.
[0133] It is understandable that the input / output interface 714 may be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 714 may be a component of the electronic device 710 or an external device connected to the electronic device 710 when needed.
[0134] Optionally, the memory 711 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory or the like, and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory or the like.
[0135] Optionally, the computer program stored in the processor 711 may be divided into one or more modules. The one or more modules are stored in the memory 711 and executed by the processor 712 to complete the method provided by the present embodiment. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 710.
[0136] Optionally, the processor 712 may be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the processor 712 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, and may also be any suitable controller, microcontroller, processor, etc. The processor 712 executes each method and process of the present embodiment. Exemplarily, it is a method for reference image segmentation based on multi-level feature fusion according to an embodiment of the present application.
[0137] Optionally, the bus 715 may include a path for transmitting information. The bus 715 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus or the like. According to different functions, the bus 715 may be divided into an address bus, a data bus, a control bus, etc.
[0138] In an alternative implementation, an embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to execute the method of the above method embodiment. A part or all of the computer program may be loaded and / or installed on the memory 711 of the electronic device 710. When the computer program is executed by the processor 712, one or more steps of a method for reference image segmentation based on multi-level feature fusion according to an embodiment of the present application may be executed.
[0139] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.
[0140] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A reference image segmentation method based on multi-level feature fusion, characterized in that Including, encoding text information into text features, and encoding an original image with target information into multiple image feature maps of different scales; fusing the image feature maps of each scale with the text features respectively to generate multiple image feature maps of different scales fused with text information; using a lightweight fusion block based on a single attention mechanism to perform attention mechanism operations on the multiple image feature maps of different scales and the text features respectively; based on the fusion process, each stage of image feature encoding is independent of the fusion operation and is not restricted by the completion of the feature fusion of the previous scale; performing pixel decoding on the multiple-scale image feature maps fused with text information to obtain mask features and multiple enhanced image feature maps of different scales; the pixel decoding adopts a multi-scale deformable attention module and a feed-forward neural network; using the mask features and the text features to perform Query decoding on the multiple enhanced image feature maps of different scales to obtain a segmentation map containing target information; the Query decoding adopts a learnable Object Query decoder, and the Object Query decoder includes multiple decoding layers, each decoding layer includes a text cross-attention module, an image cross-attention module and a feed-forward neural network FFN, and the Query decoding process includes, assuming there are N decoding layers, initializing a Query to represent target information and entering the first decoding layer as an output; for the first to Nth decoding layers, receiving the output of the previous decoding layer and interacting with the text features through the text cross-attention module to generate a Query carrying text information, performing dot product segmentation processing on the Query carrying text information and the mask features to obtain an attention mask, and jointly interacting the attention mask, the enhanced image feature map and the Query carrying text information through the image cross-attention module, and enhancing the non-linear expression ability of the result after interaction through the feed-forward neural network FFN as an output; for the Nth decoding layer, enhancing the non-linear expression ability of the result after interaction through the image cross-attention module through the feed-forward neural network FFN, and then performing dot product segmentation processing with the mask features to obtain a segmentation map containing target information.
2. The reference image segmentation method based on multi-level feature fusion according to claim 1, wherein In the fusion process, the multiple image features of different scales and the text features respectively undergo an attention mechanism operation, and the result is element-wise multiplied with the image feature projection result, and the result after the element-wise multiplication is projected to obtain multiple-scale image feature maps fused with text information.
3. The method for segmenting a reference image based on multi-level feature fusion according to claim 2, wherein The fusion process is specifically represented by the following formula: Among them, the subscript represents the th scale, represents the image feature map of the th scale, represents the text feature, , , , , are all linear transformation matrices; represents the result after multiplying the image feature map of the th scale by ; represents the result after multiplying the text feature by ; represents the result after multiplying the text feature by ; represents element-wise multiplication; represents the normalization function; represents the dimension of the key vector, represents the image feature map of the th scale that has incorporated text information, represents the total number of scales.
4. A method for segmenting a reference image based on multi-level feature fusion according to claim 1, characterized in that The encoding of text information into text features specifically includes: presetting text information and converting the text information into a feature sequence; encoding the feature sequence to obtain text features; and / or, The encoding of the original image with target information into multiple image feature maps of different scales specifically includes: encoding the original image with target information, and extracting a feature map once in each encoding stage to obtain multiple image feature maps of different scales.
5. A reference image segmentation method based on multi-level feature fusion according to claim 1, characterized in that Decode the image feature maps of multiple scales fused with text information based on the attention module and the feed-forward neural network to perform interaction between feature maps of different scales, and obtain the mask feature and the enhanced image feature maps of the multiple different scales.
6. A method for segmenting a reference image based on multi-level feature fusion according to claim 1, characterized in that, The Query decoding uses the weighted sum of the loss function and the cross-entropy loss function as the final loss function, and the calculation formula of the final loss function is expressed as: Among them, is the number of layers of the Query decoder, is the stage output, represents the decoding layer, is the segmented graph containing target information, is the weight of the loss function, is the weight of the cross-entropy loss function and and are respectively the weighted results of the loss function and the cross-entropy loss function and is the final loss function.
7. A reference image segmentation system based on multi-level feature fusion, characterized in that, It includes an encoding module, a pixel decoding module, a Query decoding module, and a fusion module; The encoding module includes a text encoding module and an image encoding module. The text encoding module is used to encode text information into text features, and the image encoding module is used to encode the original image with target information into image feature maps of multiple different scales; The fusion module is used to fuse the image feature maps of each scale with the text features respectively to generate image feature maps of multiple different scales fused with text information; a lightweight fusion block based on a single-shot attention mechanism is adopted to perform attention mechanism operations on the image feature maps of the multiple different scales and the text features respectively; based on the fusion process, each stage of image feature encoding is independent of the fusion operation and is not restricted by the completion of the feature fusion of the previous scale; The pixel decoding module is used to perform pixel decoding on the image feature maps of multiple scales fused with text information to obtain the mask feature and enhanced image feature maps of multiple different scales; the pixel decoding adopts a multi-scale deformable attention module and a feed-forward neural network; The Query decoding module is used to perform Query decoding on the enhanced image feature maps of the multiple different scales by using the mask feature and the text features to obtain a segmentation map containing target information; The Query decoding adopts a learnable Object Query decoder. The Object Query decoder includes multiple decoding layers. Each decoding layer includes a text cross-attention module, an image cross-attention module, and a feed-forward neural network FFN. The Query decoding process includes: Assume that there are N decoding layers, initialize the Query, which is used to represent the target information and enters the first decoding layer as the output; For the first to Nth decoding layers, receive the output of the previous decoding layer and interact with the text features through the text cross-attention module to generate a Query carrying text information. The Query carrying text information is subjected to dot product segmentation processing with the mask feature to obtain an attention mask. The attention mask, the enhanced image feature map, and the Query carrying text information are jointly interacted through the image cross-attention module, and the result after interaction is enhanced by the feed-forward neural network FFN to improve the non-linear expression ability and used as the output; For the Nth decoding layer, the result after interaction through the image cross-attention module is enhanced by the feed-forward neural network FFN to improve the non-linear expression ability, and then subjected to dot product segmentation processing with the mask feature to obtain a segmentation map containing target information.
8. An electronic device, characterized in that, It includes: A memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements a reference image segmentation method based on multi-level feature fusion according to any one of claims 1-6.
9. A computer-readable storage medium storing computer instructions for causing a processor to execute a method for segmenting a reference image based on multi-level feature fusion as described in any one of claims 1-6.
Citation Information
Patent Citations
Image semantic segmentation method based on large convolution kernel backbone network
CN116612283A
Creference image segmentation model training method and reference image segmentation method
CN116993976A
Context-aware reference image segmentation method, system and device and storage medium
CN117078942A