Multi-modal large model image segmentation method and device based on hierarchical lexical representation
Through the multimodal large model image segmentation method of hierarchical word element representation and causal attention mechanism, the shortcomings of traditional methods in complex scene understanding and model training are solved, precise segmentation of natural language description goals is achieved, and the performance of multimodal image segmentation is improved.
Patent Information
- Application Number
- CN202510297247.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing multimodal image segmentation method performs poorly when dealing with complex scenes and detailed features, lacking hierarchical learning and causal attention, resulting in insufficient accuracy and generalization of segmentation results.
The multimodal large model image segmentation method based on hierarchical word element representation is adopted, and the mask image is encoded into a one-dimensional mask word element sequence through a mask marker. The causal attention mechanism is used to generate mask word elements, and a three-stage training strategy and a hierarchical mask loss function are combined to achieve a gradual generation from shape prototype to local details.
显著提升了多模态图像分割的性能,能够更准确地理解自然语言描述的目标对象并进行精准分割,提高了模型在复杂场景下的理解和分割能力。
Smart Images

Figure CN120298683A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and specifically relates to a multi-modal large model image segmentation method and device based on hierarchical token representation. Background Art
[0002] Existing multi-modal image segmentation methods have obvious deficiencies. Traditional systems perform poorly in dealing with complex scenes and detailed features, and it is difficult to accurately understand and locate the target objects described in natural language.
[0003] In addition, there are bottlenecks in model training in the existing technology. Most systems lack hierarchical learning of mask representations and fail to effectively integrate the semantic understanding ability of large-scale pre-trained models, resulting in insufficient accuracy and generalization of segmentation results.
[0004] There are technical shortcomings in mask generation in existing systems. Lack of a sequential generation mechanism based on causal attention, it is difficult to achieve progressive refinement of mask features, affecting the fineness of segmentation effects. Solving these problems is of great significance for improving the performance of multi-modal image segmentation. Summary of the Invention
[0005] Aiming at the problems in the existing technology, this application provides a multi-modal large model image segmentation method and device based on hierarchical token representation, which can effectively solve the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improve the performance of multi-modal image segmentation.
[0006] To solve at least one of the above problems, this application provides the following technical solutions:
[0007] In a first aspect, this application provides a multi-modal large model image segmentation method based on hierarchical token representation, including:
[0008] Encoding an input mask image into a one-dimensional mask token sequence through a mask tokenizer, where the mask tokenizer generates a current mask token using a causal attention mechanism according to the mask blocks corresponding to the mask image and the generated mask tokens. The front tokens of the mask token sequence represent the position and shape prototype of the target area, and the rear tokens represent the local detailed features of the target area. Each mask token in the mask token sequence is conditionally generated based on its previous mask token. Inputting the mask token sequence into a vector quantization layer for token quantization to obtain mask token codes, and a mask decoder reconstructs the mask image according to the mask token codes;
[0009] The first - stage training of the mask tokenizer is carried out through a mask reconstruction task, the second - stage training is carried out using a large - scale multimodal model, the mask token code is augmented into the vocabulary of the large - scale multimodal model, the mask token is integrated into the large - scale multimodal model based on low - resolution image data, joint training is carried out according to a segmentation dataset and a general dataset, a hierarchical mask loss function is used to supervise the training of the mask token sequence at different hierarchical levels, and the third - stage fine - tuning training is carried out using high - resolution image data to obtain a mask semantic segmentation model;
[0010] Receive the to - be - processed image and the corresponding target object description in the referential expression segmentation request, input the to - be - processed image and the target object description into the mask semantic segmentation model, the mask semantic segmentation model generates the corresponding mask token sequence based on the target object description, the mask decoder generates the corresponding segmentation mask image according to the mask token sequence, and the mask semantic segmentation model identifies the target area referred to by the target object description in the to - be - processed image according to the segmentation mask image.
[0011] Further, it also includes: performing block processing on the input mask image to obtain a mask image block sequence, inputting the mask image block sequence and a number of fixed hidden vectors into the encoder network of the mask tokenizer for joint self - attention calculation, and after the above - mentioned encoding process, the target mask token is obtained for the hidden vectors.
[0012] The encoder network adopts conditional autoregressive modeling. Each attention layer is bidirectional attention between mask picture blocks and unidirectional attention between hidden vectors. Each hidden vector is calculated conditional on the mask picture block features and all the previous hidden vectors, which ensures that the hidden vectors can express the mask image from coarse - grained to fine - grained, so that the hidden vectors are encoded into the accurate one - dimensional mask token sequence.
[0013] Input the conditionally generated mask token sequence into the vector quantization layer. The vector quantization layer maintains a preset number of codebook vectors. By calculating the Euclidean distance between each mask token in the mask token sequence and the codebook vectors, the codebook vector with the smallest distance is selected as the corresponding mask token code. The mask decoder uses a Transformer deep - learning model to perform feature mapping and upsampling on the mask token code, and restores the feature mapping result to a reconstructed image with the same resolution as the input mask image.
[0014] Further, it also includes: constructing a mask reconstruction loss function to train the mask tokenizer. The mask reconstruction loss function includes a reconstruction error term and a codebook commitment term. The reconstruction error term calculates the binary cross-entropy loss and the region overlap loss between the reconstructed mask image and the original mask image. The region overlap loss term calculates the intersection over union (IoU) score between the predicted mask and the ground truth mask at different hierarchical scales. The codebook commitment term calculates the Euclidean distance between the quantized features of the mask tokens and the nearest neighbor codebook vectors. Based on the Adam optimizer, optimize the mask reconstruction loss function and update the network parameters of the mask tokenizer and the codebook vectors;
[0015] Map the mask token codes to token embedding vectors with a fixed dimension, expand the token embedding matrix of the large multimodal model to accommodate the token embedding vectors, construct training samples containing mask tokens based on low-resolution images. The input sequence of the training samples consists of image feature tokens and mask tokens. The large multimodal model uses a cross-entropy loss function and a hierarchical mask loss function to train the training samples, enabling the model to learn the semantic correspondence between mask tokens and other modal tokens. The hierarchical mask loss function performs supervised training on the mask token sequence at different hierarchical levels. Each hierarchical level of the token sequence is a sub-token sequence counted from the front. The sub-token sequence passes through a mask decoder to obtain the predicted mask at the corresponding level. The loss function at each hierarchical level calculates the binary cross-entropy loss and the region overlap loss term between the predicted mask and the ground truth mask at the corresponding level. The final hierarchical mask loss is equal to the sum of the loss functions at each hierarchical level.
[0016] Further, it also includes: constructing a mixed training data batch, where each training batch contains a preset proportion of segmented data samples and general data samples, fine-tuning the model after the second-stage training using high-resolution image data. The large multimodal model treats ordinary text and mask tokens equally, and the training loss function uses cross-entropy loss, thereby obtaining the final mask semantic segmentation model.
[0017] Further, it also includes: using an image encoder to extract features from the input image to be processed to obtain an image feature sequence, passing the image feature sequence through a multi-layer perceptron module to obtain a visual representation sequence aligned with the input layer feature space of the large language model, and inputting the aligned visual representation sequence into the large language model network of the mask semantic segmentation model;
[0018] The large language model network of the mask semantic segmentation model uses an autoregressive decoding method to generate a sequence of mask tokens. For each decoding time step, the generated mask tokens are interacted with the visual feature representation in a multi-modal manner to obtain interaction features. Based on the interaction features, the probability distribution of the next mask token is predicted, and a mask token is selected from the probability distribution through a sampling or greedy decoding strategy until a complete sequence of mask tokens is generated.
[0019] Furthermore, it further includes: inputting the generated sequence of mask tokens into a mask decoder, where the mask decoder is a Transformer model with a bidirectional attention mechanism. Along with the sequence of mask tokens, several initial hidden states are also input. After passing through the mask decoder, the sequence of mask tokens is mapped to an initial feature map. The spatial resolution of the feature map is restored through an upsampling module. Each upsampling module includes a nearest neighbor upsampling layer and a convolutional layer. The upsampled feature map is normalized to obtain a probability mask map, and the probability mask map is converted into a binary mask to obtain a segmented mask image.
[0020] Morphological processing and boundary optimization are performed on the segmented mask image. The morphological processing includes performing opening and closing operations on the segmented mask to remove noise and smooth the boundary. The boundary optimization locally adjusts the segmentation boundary based on image gradient information. The optimized segmented mask is fused with the original image at the pixel level, and the contour and boundary of the target area are marked on the image to be processed, generating a segmented result image containing the target area identifier.
[0021] In a second aspect, the present application provides a multi-modal large model image segmentation device based on hierarchical token representation, including:
[0022] A token representation module for encoding an input mask image into a one-dimensional sequence of mask tokens through a mask tokenizer. The mask tokenizer generates the current mask token using a causal attention mechanism based on the mask blocks corresponding to the mask image and the generated mask tokens. The front tokens of the sequence of mask tokens represent the position and shape prototype of the target area, and the rear tokens represent the local detail features of the target area. Each mask token in the sequence of mask tokens is generated conditionally based on its previous mask token. The sequence of mask tokens is input into a vector quantization layer for token quantization to obtain mask token codes, and the mask decoder reconstructs the mask image based on the mask token codes.
[0023] A model training module, configured to perform a first-stage training on the mask tokenizer through a mask reconstruction task, perform a second-stage training using a large multi-modal model, expand the mask token code into the vocabulary of the large multi-modal model, integrate the mask token into the large multi-modal model based on low-resolution image data, perform joint training according to a segmentation data set and a general data set, perform supervised training on the mask token sequence at different hierarchical levels using a hierarchical mask loss function, and perform a third-stage fine-tuning training using high-resolution image data to obtain a mask semantic segmentation model;
[0024] An image segmentation module, configured to receive a to-be-processed image and a corresponding target object description in a referential expression segmentation request, input the to-be-processed image and the target object description into the mask semantic segmentation model, the mask semantic segmentation model generates a corresponding mask token sequence based on the target object description, the mask decoder generates a corresponding segmentation mask image according to the mask token sequence, and the mask semantic segmentation model identifies a target area referred to by the target object description in the to-be-processed image according to the segmentation mask image.
[0025] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the multi-modal large model image segmentation method based on hierarchical token representation are implemented.
[0026] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-modal large model image segmentation method based on hierarchical token representation are implemented.
[0027] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the multi-modal large model image segmentation method based on hierarchical token representation are implemented.
[0028] As can be seen from the above technical solutions, the present application provides a multi-modal large model image segmentation method and device based on hierarchical token representation. By designing a mask tokenizer, a mask image is encoded into a token sequence, and progressive generation from a shape prototype to local details is realized through a causal attention mechanism. A three-stage training strategy is adopted. First, the tokenizer is trained through a mask reconstruction task, then the mask token is integrated into a large multi-modal model and joint training is performed, and finally fine-tuning is performed using high-resolution data. Multi-level supervision is performed based on a hierarchical mask loss function to achieve accurate segmentation of natural language description targets. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a schematic flow chart of the multi-modal large model image segmentation method based on hierarchical token representation in the embodiments of the present application;
[0031] Figure 2 It is a structural diagram of the multi-modal large model image segmentation device based on hierarchical token representation in the embodiments of the present application;
[0032] Figure 3 It is a schematic structural diagram of the electronic device in the embodiments of the present application.
[0033] Reference Numerals:
[0034] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Embodiments
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0036] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0037] In view of the problems existing in the prior art, the present application provides a multi-modal large model image segmentation method and device based on hierarchical token representation. By designing a mask tokenizer to encode the mask image into a token sequence, and implementing progressive generation from shape prototypes to local details through a causal attention mechanism. Adopting a three-stage training strategy, first training the tokenizer through a mask reconstruction task, then integrating the mask tokens into a large multi-modal model for joint training, and finally fine-tuning with high-resolution data. Conducting multi-level supervision based on a hierarchical mask loss function to achieve precise segmentation of natural language description targets. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0038] In order to effectively solve the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improve the performance of multi-modal image segmentation, an embodiment of a multi-modal large model image segmentation method based on hierarchical token representation is provided in the present application. Refer to Figure 1 , the multi-modal large model image segmentation method based on hierarchical token representation specifically includes the following content:
[0039] Step S101: Encode the input mask image into a one-dimensional mask token sequence through a mask tokenizer. The mask tokenizer generates the current mask token by using a causal attention mechanism based on the mask block corresponding to the mask image and the generated mask tokens. The front tokens of the mask token sequence represent the position and shape prototype of the target area, and the rear tokens represent the local detail features of the target area. Each mask token in the mask token sequence is generated conditionally based on its previous mask token. Input the mask token sequence into a vector quantization layer for token quantization to obtain mask token codes, and the mask decoder reconstructs the mask image according to the mask token codes;
[0040] Optionally, in view of the challenges of mask representation in the image segmentation task, this embodiment innovatively designs a mask encoding method based on hierarchical tokens. Traditional methods directly perform pixel-level modeling on the mask image, making it difficult to effectively capture the hierarchical semantic information of the target area. Therefore, this embodiment constructs a dedicated mask tokenizer to convert the two-dimensional mask image into a one-dimensional token sequence representation with clear semantic levels.
[0041] The mask tokenizer of this embodiment first performs uniform block processing on the input mask image. The block result forms a mask block sequence, and each mask block contains position coordinate and pixel value information.
[0042] This embodiment innovatively uses a causal attention mechanism in the mask tokenizer to generate mask tokens. The attention calculation uses the following formula:
[0043] Attention(Q, K, V) = softmax(QK^T / √d)V
[0044] Among them, Q, K, and V respectively represent the query, key-value pair matrix, and d is the dimension of the attention head. By using a causal mask, it is ensured that the current position can only attend to the information of previous positions, thereby realizing autoregressive generation. This design enables each masked token to make full use of the context information in the already generated tokens.
[0045] This embodiment implements a hierarchical encoding strategy for the masked token sequence. The front tokens of the sequence focus on capturing macroscopic features such as the position and shape of the target area. These tokens obtain the rough contour and spatial distribution information of the target area by downsampling the masked image and extracting the contour. The rear tokens are responsible for encoding local detailed features, including microscopic information such as the exact position of the boundary and texture changes. This hierarchical token organization is similar to the human visual perception process, first obtaining the overall and then focusing on the details.
[0046] This embodiment innovatively designs a conditional generation mechanism. The generation of each masked token is conditioned on its previous tokens. An information of the previous tokens is converted into a conditional vector through a conditional embedding network. This conditional vector is fused with the features at the current position to guide the generation of new tokens. This design ensures the coherence of the generation process and maintains semantic consistency between adjacent tokens.
[0047] This embodiment constructs an efficient vector quantization layer. This layer maintains a dictionary containing K codebook vectors, and the size of K is set according to the complexity of the application scenario. The quantization process calculates the Euclidean distance between the token features and the codebook vectors, and selects the nearest neighbor codebook vector as the quantization result. This discrete representation not only simplifies the feature complexity but also retains the necessary semantic information.
[0048] This embodiment implements an accurate mask reconstruction mechanism. The mask decoder uses a Transformer model with a bidirectional attention mechanism. Along with the masked token sequence, several initial hidden states are input. After passing through the mask decoder, the masked token sequence is mapped to an initial feature map. The spatial resolution of the feature map is restored through an upsampling module. Each upsampling module contains a nearest neighbor upsampling layer and a convolutional layer. The upsampled feature map is normalized to obtain a probability mask map, and the probability mask map is converted into a binary mask to obtain a segmented mask image.
[0049] Through the above technical innovations, this embodiment effectively solves the deficiencies of traditional mask representation methods in semantic expression and computational efficiency. This scheme can compress two-dimensional mask information into a compact one-dimensional sequence while maintaining a clear semantic hierarchical structure. In practical applications, this representation method significantly reduces the storage and computational overhead and improves the training efficiency of the model.
[0050] This embodiment has important practical significance in the field of image segmentation. By converting the mask information into a sequence of tokens, it lays a foundation for the subsequent integration with large language models. This design enables the model to better understand and generate segmentation results that conform to natural language descriptions, providing a new solution for cross-modal image understanding. Through testing and verification in different scenarios, this solution demonstrates good generalization ability and practical value.
[0051] Step S102: Conduct the first-stage training on the mask tokenizer through a mask reconstruction task, conduct the second-stage training using a large multi-modal model, expand the mask token code into the vocabulary of the large multi-modal model, integrate the mask tokens into the large multi-modal model based on low-resolution image data, conduct joint training according to the segmentation dataset and the general dataset, use a hierarchical mask loss function to conduct supervised training on the mask token sequence at different hierarchical levels, and use high-resolution image data for the third-stage fine-tuning training to obtain a mask semantic segmentation model;
[0052] Optionally, this embodiment designs an innovative three-stage training strategy, effectively solving the challenge of the integration of mask tokens and large multi-modal models. First, construct a mask reconstruction task to pre-train the mask tokenizer. The mask reconstruction loss function L_r is defined as:
[0053] L_r = MSE(M_rec, M_orig) + λ1||z_q - sg(e)||2 + λ2||sg(z_q) - e||2
[0054] where M_rec is the reconstructed mask, M_orig is the original mask, z_q is the quantization feature, e is the codebook vector, sg represents the gradient truncation operation, and λ1 and λ2 are balance parameters. This loss function simultaneously optimizes the reconstruction quality and the codebook usage efficiency.
[0055] This embodiment implements an adaptive learning rate adjustment mechanism in the first-stage training. The training process monitors the change trend of the reconstruction error. When the error decline rate slows down, the learning rate is adjusted through the cosine annealing strategy. At the same time, an early stopping mechanism is introduced. When the performance on the validation set no longer improves, the training is stopped in a timely manner to avoid overfitting. This training strategy ensures that the mask tokenizer can learn a stable and effective mask coding representation.
[0056] This embodiment innovatively designs an integration method for mask tokens and large multi-modal models. First, map the mask token code to a vector representation with the same dimension as the model word embedding. The mapping function F_map adopts a multi-layer perceptron structure:
[0057] F_map(x) = MLP(Embedding(x))
[0058] where x is the token code, Embedding is a learnable embedding layer, and MLP is a non-linear transformation network. This design ensures that the masked tokens can be represented in the same semantic space as the original tokens of the model.
[0059] This embodiment implements an efficient training strategy based on low-resolution images. A dedicated training data generation pipeline is constructed to downsample the original images to an appropriate resolution while maintaining the integrity of the masked information. The training samples include image feature sequences and corresponding masked token sequences, and the model learns the semantic correspondence between the two modalities through the masked language modeling task.
[0060] This embodiment innovatively designs a joint training mechanism. Each training batch contains a preset proportion of segmentation data and general data, and the joint loss function L_joint is defined as:
[0061] L_joint = αL_mask + γ * L_lm
[0062] where L_mask is the masked prediction loss, L_lm is the language modeling loss, and α and γ are weight coefficients. This multi-task learning framework ensures both the segmentation ability of the model and the general language understanding ability.
[0063] This embodiment constructs a hierarchical masked loss function. The difference between the predicted mask and the ground truth mask is calculated at different spatial scales, and the loss function L_hier contains multiple hierarchical terms:
[0064] L_hier = Σw_i * (IoU_i + BCE_i)
[0065] where i represents different scale levels, w_i is the hierarchical weight, IoU_i measures the region overlap, i.e., Dice Loss, and BCE_i evaluates the accuracy of pixel-level prediction, i.e., BCE Loss. This multi-scale supervision ensures that the model can generate segmentation results with different granularities using masked tokens of different lengths.
[0066] This embodiment implements a fine-tuning training strategy for high-resolution images. A progressive learning scheme is adopted, where training is first conducted at a lower resolution and then the image resolution is gradually increased. Each time the resolution is increased, the model parameters from the previous stage are inherited, and the learning rate and batch size are adjusted accordingly. This strategy ensures both training efficiency and the model's ability to process high-resolution inputs.
[0067] Through the above technological innovations, this embodiment effectively addresses the limitations of traditional image segmentation models in semantic understanding and cross-modal fusion. This training scheme enables the model to not only have accurate segmentation capabilities but also understand natural language descriptions, providing the possibility for flexible interactive image segmentation. In practical applications, this scheme significantly improves the model's understanding ability of complex scenes and segmentation accuracy.
[0068] The innovation points of this embodiment are mainly reflected in aspects such as training strategy design and loss function construction. Through a carefully designed three-stage training process, the effective fusion of masked tokens and large multi-modal models is achieved. This scheme provides a new solution for improving the semantic understanding ability of image segmentation models and has important application value in the field of computer vision. Experimental verification shows that this training strategy can effectively improve the model performance, especially showing obvious advantages in complex scenes that require precise semantic understanding.
[0069] Step S103: Receive the image to be processed and the corresponding target object description in the referential expression segmentation request, input the image to be processed and the target object description into the masked semantic segmentation model. The masked semantic segmentation model generates the corresponding masked token sequence based on the target object description. The masked decoder generates the corresponding segmentation mask image according to the masked token sequence. The masked semantic segmentation model identifies the target area referred to by the target object description in the image to be processed according to the segmentation mask image.
[0070] Optionally, this embodiment designs a complete processing flow for the language-guided segmentation problem in multi-modal image segmentation tasks. In practical application scenarios, users often need to specify the target object to be segmented through natural language descriptions, which requires the model to accurately understand the language description and locate the corresponding area in the image.
[0071] This embodiment first implements a preprocessing mechanism for multi-modal input. When receiving a segmentation request, the image to be processed is normalized, including size adjustment, color space conversion, and pixel value normalization. For the target object description text, a pre-trained language model is used for word segmentation and feature encoding to obtain the semantic representation of the text. This standardized preprocessing ensures the consistency of the input data and lays a foundation for subsequent feature extraction and multi-modal unified modeling.
[0072] This embodiment constructs an innovative masked token generation strategy. Based on the fused cross-modal features, the model uses an autoregressive decoding method to generate masked tokens one by one. The beam search strategy is used during the generation process to maintain multiple candidate sequences and select the optimal sequence according to the cumulative probability score. This design improves the quality and diversity of the generation results and can better adapt to different segmentation scenarios.
[0073] This embodiment optimizes the mask decoding process. The mask decoder adopts a multi-scale feature fusion architecture, which includes multiple feature reconstruction modules. Each module gradually restores the spatial resolution of the feature map through transposed convolution operations and retains the detailed information using residual connections. A spatial attention mechanism is also introduced during the decoding process to dynamically adjust the weight distribution of feature reconstruction according to the context information.
[0074] This embodiment implements an accurate boundary optimization mechanism. After generating the initial segmentation mask, conditional random field (CRF) is used for post-processing. CRF takes into account the spatial relationship and color similarity between pixels to optimize the smoothness and accuracy of the segmentation boundary. At the same time, the segmentation result is locally adjusted in combination with the edge information of the image to improve the accuracy of boundary localization.
[0075] This embodiment designs a visualization annotation scheme. The optimized segmentation mask is fused with the original image for display, and the target area is highlighted using a semi-transparent overlay method. The boundary of the target area is stroked to improve visual recognition. At the same time, annotation information corresponding to the language description is added to the image to facilitate users to understand and verify the segmentation result.
[0076] Through the above technological innovations, this embodiment effectively solves the deficiencies of traditional image segmentation methods in language understanding and cross-modal alignment. This solution can accurately understand the user's language description and precisely locate the corresponding target area in the image. In practical applications, this method significantly improves the accuracy and user-friendliness of interactive image segmentation.
[0077] This embodiment has important value in the fields of intelligent image editing and visual content understanding. Through the precise segmentation ability guided by language, it provides a reliable basis for subsequent tasks such as image editing and content analysis. This solution shows good application prospects in professional fields such as medical image analysis and remote sensing image processing, and makes important contributions to improving the automation level of related fields.
[0078] As can be seen from the above description, the multi-modal large model image segmentation method based on hierarchical token representation provided by the embodiments of this application can encode the mask image into a token sequence by designing a mask tokenizer, and achieve progressive generation from shape prototypes to local details through the causal attention mechanism. Adopting a three-stage training strategy, first train the tokenizer through the mask reconstruction task, then integrate the mask tokens into a large multi-modal model for joint training, and finally fine-tune using high-resolution data. Based on the hierarchical mask loss function for multi-level supervision, accurate segmentation of natural language description targets is achieved. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0079] In an embodiment of the multi-modal large model image segmentation method based on hierarchical token representation of the present application, the following specific contents may further be included:
[0080] Step S201: Perform block processing on the input mask image to obtain a sequence of mask image blocks, and input the sequence of mask image blocks and a number of fixed hidden vectors into the encoder network of the mask tokenizer for joint self-attention calculation. After the above encoding process, the hidden vectors obtain target mask tokens. The encoder network adopts conditional autoregressive modeling. Each attention layer has bidirectional attention among the mask image blocks and unidirectional attention among the hidden vectors, which ensures that the hidden vectors can represent the mask image from coarse-grained to fine-grained, so that the hidden vectors are encoded into the accurate one-dimensional mask token sequence.
[0081] Optionally, this embodiment innovatively realizes the hierarchical encoding process of the mask image. A processing scheme based on block division and tokenization is designed for the limitation of directly processing pixel-level masks in traditional methods. First, a uniform block division strategy is adopted for the input mask image, and each block includes pixel information and position coordinate information.
[0082] This embodiment innovatively designs a special mask encoder network structure. The network adopts a standard Transformer structure, and the input is mask image blocks and a number of hidden vectors. Each attention layer has bidirectional attention among the mask image blocks and unidirectional attention among the hidden vectors, which ensures that the hidden vectors can represent the mask image from coarse-grained to fine-grained, so that the hidden vectors are encoded into accurate mask tokens.
[0083] This embodiment adopts an attention mask mechanism. The calculation of the attention mask matrix M adopts the following formula: M = softmax(WQ * WK^T / √d)
[0084] where WQ and WK are feature projection matrices, and d is the feature dimension. This mechanism enables the model to selectively focus on important image block features and suppress the interference of irrelevant regions at the same time. The distribution of attention weights reflects the importance of different image blocks to the target representation.
[0085] This embodiment innovatively realizes the combination strategy of the token sequence. The combination of the front token and the rear token considers semantic coherence, and ensures the expression ability of the mask image from coarse to fine-grained through conditional autoregressive modeling and hierarchical mask loss function.
[0086] Through the above technological innovations, this embodiment effectively solves the limitations of traditional mask representation methods in semantic expression and computational efficiency. This solution can convert two-dimensional mask information into a one-dimensional sequence with clear semantic levels, significantly improving the efficiency of subsequent processing. In practical applications, this representation method provides a new solution for the compression and transmission of mask information.
[0087] The innovations of this embodiment are mainly reflected in aspects such as the chunking strategy, feature extraction, and token generation. Through a carefully designed processing flow, efficient encoding of mask information is achieved. This solution provides important support for improving the performance of image segmentation models and has broad application prospects in the field of computer vision. Experimental verification shows that this method significantly reduces the computational and storage overhead while maintaining the segmentation accuracy.
[0088] In an embodiment of the multi-modal large model image segmentation method based on hierarchical token representation of this application, the following content may also be specifically included:
[0089] Step S301: Perform chunking processing on the input mask image to obtain a sequence of mask image chunks, and input the sequence of mask image chunks and a number of fixed hidden vectors into the encoder network of the mask tokenizer for joint self-attention calculation. After the above encoding process, the hidden vectors obtain target mask tokens. The encoder network adopts conditional autoregressive modeling. Each attention layer is bidirectional attention between mask image chunks and unidirectional attention between hidden vectors, which ensures that the hidden vectors can represent the mask image from coarse-grained to fine-grained, so that the hidden vectors are encoded into the accurate one-dimensional mask token sequence;
[0090] Step S302: Input the sequence of conditionally generated mask tokens into the vector quantization layer. The vector quantization layer maintains a preset number of codebook vectors. By calculating the Euclidean distance between each mask token in the sequence of mask tokens and the codebook vectors, select the codebook vector with the smallest distance as the corresponding mask token code. The mask decoder uses a Transformer deep learning model to perform feature mapping and upsampling on the mask token code, and restores the feature mapping result to a reconstructed image with the same resolution as the input mask image.
[0091] Optionally, for the problem of generating the sequence of mask tokens in this embodiment, a conditional autoregressive modeling scheme is innovatively designed. Traditional methods often ignore the dependencies between tokens, resulting in a lack of coherence in the generated results. For this reason, this embodiment constructs a conditional autoregressive model based on the transformer architecture, and captures the long-range dependencies between tokens through a position-aware attention mechanism.
[0092] This embodiment implements an accurate position encoding mechanism. For each position i in the sequence, calculate its d-dimensional position encoding PE:
[0093] PE(i, 2k) = sin(i / 10000 ^ (2k / d))
[0094] PE(i, 2k + 1) = cos(i / 10000 ^ (2k / d))
[0095] where k is the dimension index. This encoding method enables the model to perceive the relative positional relationship of the tokens, which helps to maintain the spatial coherence of the generated sequence.
[0096] In this embodiment, a conditional autoregressive modeling for the masked token sequence is innovatively designed. Each masked token is obtained through a Transformer with a causal attention mechanism based on the masked image patch and its previous subsequence of masked tokens:
[0097] p(m_1,..., m_K | M) = ∏ p(m_k | M, m_1,…, m_(k - 1))
[0098] where p represents the probability distribution, m_i represents the i-th masked token, K is the total number of masked tokens, M represents the masked image patch, and ∏ represents the product (k ranges from 1 to K).
[0099] In this embodiment, an efficient vector quantization layer is constructed. This layer maintains a dictionary containing K codebook vectors, and the value of K is dynamically set according to the complexity of the application scenario. The quantization process uses nearest neighbor search and selects the most matching codebook vector through Euclidean distance measurement. To improve the search efficiency, a fast approximate nearest neighbor search based on locality-sensitive hashing is implemented.
[0100] Through the above technological innovations in this embodiment, the deficiencies of traditional mask generation methods in coherence and accuracy are effectively solved. This scheme can generate a semantically coherent and detail-rich mask representation, providing a reliable basis for subsequent image segmentation tasks. In practical applications, this method significantly improves the quality and efficiency of mask generation.
[0101] The innovation points of this embodiment are mainly reflected in conditional modeling, feature quantization, decoding and reconstruction, etc. Through a carefully designed processing flow, high-quality generation and reconstruction of mask information are achieved. This scheme provides important support for improving the performance of image segmentation models and has broad application prospects in the field of computer vision. Experimental verification shows that this method significantly improves the efficiency and stability of the generation process while maintaining the reconstruction quality.
[0102] In an embodiment of the multi-modal large model image segmentation method based on hierarchical token representation in this application, the following content may also be specifically included:
[0103] Step S401: Construct a mask reconstruction loss function to train the mask tokenizer. The mask reconstruction loss function includes a reconstruction error term and a codebook commitment term. The reconstruction error term calculates the binary cross-entropy loss and the region overlap loss between the reconstructed mask image and the original mask image. The region overlap loss term calculates the intersection over union (IoU) score between the predicted mask and the ground truth mask at different hierarchical scales. The codebook commitment term calculates the Euclidean distance between the quantized features of the mask tokens and the nearest neighbor codebook vectors. Optimize the mask reconstruction loss function based on the Adam optimizer, and update the network parameters of the mask tokenizer and the codebook vectors.
[0104] Step S402: Map the mask token codes to fixed-dimensional token embedding vectors, expand the token embedding matrix of the large multi-modal model to accommodate the token embedding vectors, and construct training samples containing mask tokens based on low-resolution images. The input sequence of the training samples consists of image feature tokens and mask tokens. The large multi-modal model uses the cross-entropy loss function and the hierarchical mask loss function to train the training samples, enabling the model to learn the semantic correspondence between mask tokens and other modal tokens. The hierarchical mask loss function performs supervised training on the mask token sequence at different hierarchical levels. Each hierarchical level of the token sequence is a sub-token sequence counted from the front. The sub-token sequence passes through the mask decoder to obtain the predicted mask at the corresponding level. The loss function at each hierarchical level calculates the binary cross-entropy loss and the region overlap loss term between the predicted mask and the ground truth mask at the corresponding level. The final hierarchical mask loss is equal to the sum of the loss functions at each hierarchical level.
[0105] Optionally, this embodiment innovatively designs a mask reconstruction loss function with dual constraints. This loss function L_total consists of two key components:
[0106] L_total = λ_r * L_recon + λ_c * L_commit
[0107] where L_recon is the reconstruction error term, L_commit is the codebook commitment term, and λ_r and λ_c are balance factors. This design ensures both the reconstruction quality and constrains the distribution of the quantized features.
[0108] This embodiment realizes adaptive reconstruction error calculation. The reconstruction error adopts the form of weighted mean square error: L_recon = -Σm_i(M_i log(M'_i)+(1 - M_i)log(1 - M'_i))
[0109] where M is the target mask, M' is the generated mask, and m is the position weight. The weight value is dynamically assigned according to the semantic importance of the pixels, enabling the model to pay more attention to the reconstruction quality of the key regions.
[0110] In this embodiment, the calculation method of the codebook commitment term is optimized. The commitment loss is calculated by the following formula: L_commit = λ1||z_q - sg(e)||2 + λ2||sg(z_q) - e||2
[0111] where z_q is the feature before quantization, e is the nearest neighbor codebook vector, sg represents the gradient clipping operation, and λ1 and λ2 are balance parameters. This design encourages the quantized features to be close to the codebook vectors and avoids the codebook collapse problem through gradient clipping.
[0112] In this embodiment, the Adam optimizer is used to update the model parameters, and the learning rate is dynamically adjusted according to the training progress:
[0113] lr = lr_0 * cos(π * t / T)
[0114] where lr_0 is the initial learning rate, t is the current step, and T is the total number of steps. This cosine annealing strategy makes the training process more stable.
[0115] In this embodiment, an efficient token mapping mechanism is implemented. The masked token code is converted into an embedding vector of a fixed dimension through a learnable mapping network:
[0116] E = MLP(OneHot(code))
[0117] where code is the token code and MLP is a multi-layer perceptron. The parameters of the mapping network are optimized synchronously through backpropagation to ensure that the generated embedding vectors have good semantic representation capabilities.
[0118] In this embodiment, a dynamic expansion scheme for the word embedding matrix is constructed. While keeping the original word embeddings unchanged, new embedding spaces are allocated for masked tokens. The expansion process uses orthogonal initialization to ensure that the newly added embedding vectors are properly distinguishable from the original embeddings. At the same time, a dynamic adjustment mechanism for the embedding space is implemented to adaptively allocate the embedding dimensions according to the training requirements.
[0119] In this embodiment, the construction strategy of training samples is optimized. The training samples constructed based on low-resolution images contain various modal information:
[0120] S = [T_img; T_mask; T_text]
[0121] where T_img is the image feature token, T_mask is the masked token, and T_text is the text token. The temporal correspondence relationship between different modalities is considered in the sample construction process to ensure that the model can learn the correct cross-modal mapping.
[0122] In this embodiment, the cross-entropy loss function is used to optimize the class distribution of masked tokens:
[0123] L_ce = -Σy_i * log(p_i)
[0124] where y_i is the true label and p_i is the predicted probability. The overfitting problem is alleviated through label smoothing technology, improving the generalization ability of the model.
[0125] Through the above technological innovations in this embodiment, the deficiencies of traditional methods in masked representation learning and multimodal fusion are effectively solved. This solution can learn high-quality masked representations and successfully integrate them into large multimodal models. In practical applications, this method significantly improves the model's ability to understand and generate masked information.
[0126] The innovations of this embodiment are mainly reflected in aspects such as loss function design, parameter optimization, and multimodal modeling. Through a carefully designed training strategy, the effective integration of masked information with other modalities is achieved. This solution provides a new solution for improving the semantic understanding ability of image segmentation models and has important application value in the field of computer vision. Experimental verification shows that this method can effectively improve the segmentation performance of the model, especially showing obvious advantages in tasks that require understanding complex semantic scenes.
[0127] In an embodiment of the multimodal large model image segmentation method based on hierarchical token representation in this application, the following content may also be specifically included:
[0128] Step S501: Construct a mixed training data batch. Each training batch contains a preset proportion of segmentation data samples and general data samples. Use high-resolution image data to fine-tune the model after the second-stage training. The large multimodal model treats ordinary text and masked tokens equally, and the training loss function uses cross-entropy loss to obtain the final masked semantic segmentation model.
[0129] Optionally, this embodiment innovatively designs a mixed batch training strategy. In each training batch, segmentation data and general data are mixed in the ratio of r:(1 - r), where r is a dynamically adjusted mixing ratio: r = r_0 * (1 + α * cos(2πt / T))
[0130] where r_0 is the basic mixing ratio, t is the current iteration number, T is the total iteration number, and α is the adjustment coefficient. This periodic adjustment strategy enables the model to balance between professional segmentation ability and general language understanding ability.
[0131] This embodiment implements a multi-task joint training mechanism. The construction of the joint training objective function L_joint is as follows: L_joint = αL_mask + γ * L_lm
[0132] Among them, L_mask is the mask prediction loss, L_lm is the language modeling loss, and α and γ are weight coefficients. This multi-task learning framework not only ensures the segmentation ability of the model but also maintains the general language understanding ability.
[0133] This embodiment innovatively designs a mask reconstruction loss. The pixel-level difference is calculated by the binary cross-entropy loss function:
[0134] L_recon = -Σm_i(M_i log(M'_i)+(1-M_i)log(1-M'_i))
[0135] Where M is the target mask, M' is the generated mask, and m is the position weight. The weight value is dynamically assigned according to the semantic importance of the pixel, enabling the model to pay more attention to the reconstruction quality of the key area.
[0136] This embodiment implements a hierarchical mask loss function. This loss function L_hier comprehensively considers multiple levels of supervision signals:
[0137] L_hier = Σw_i*(IoU_i + BCE_i)
[0138] Where i represents different scale levels, w_i is the hierarchical weight, IoU_i measures the region overlap degree, that is, Dice Loss, and BCE_i evaluates the accuracy of pixel-level prediction, that is, BCE Loss. This multi-scale supervision ensures that the model can generate segmentation results with different granularities using mask tokens of different lengths.
[0139] This embodiment constructs an adaptive IoU calculation mechanism. At each scale level, the region overlap degree is calculated by soft IoU:
[0140] IoU = (P∩G + ε) / (P∪G + ε)
[0141] Where P is the predicted mask, G is the ground truth mask, and ε is the smoothing factor. The use of soft IoU makes the loss function differentiable, facilitating backpropagation optimization.
[0142] This embodiment optimizes the evaluation method of boundary accuracy. The boundary distance metric uses a weighted Hausdorff distance:
[0143] D_b = max{sup_x∈P inf_y∈G d(x,y),sup_y∈G inf_x∈P d(x,y)}
[0144] Where d(x,y) is the Euclidean distance between pixel points. This metric can accurately reflect the deviation degree of the segmentation boundary.
[0145] This embodiment realizes an efficient fine-tuning training strategy. When fine-tuning on high-resolution data, a progressive learning scheme is adopted to gradually increase the resolution of the input image. Each resolution stage inherits the model parameters of the previous stage and adjusts the learning rate and batch size accordingly. This strategy not only guarantees the training efficiency but also maintains the model performance.
[0146] Through the above technological innovations, this embodiment effectively solves the deficiencies of traditional image segmentation models in multi-task learning and fine-grained segmentation. This solution can take into account both segmentation accuracy and model generality, providing a new solution for achieving high-quality semantic segmentation. In practical applications, this method significantly improves the segmentation performance and generalization ability of the model.
[0147] The innovations of this embodiment are mainly reflected in aspects such as training strategy design and loss function construction. Through a carefully designed training process and supervision mechanism, the overall performance of the model is improved. This solution provides important support for improving the accuracy and robustness of image segmentation models and has broad application prospects in the field of computer vision. Experimental verification shows that this method can effectively improve the model performance, especially showing obvious advantages in complex scenarios that require precise segmentation.
[0148] In an embodiment of the multi-modal large model image segmentation method based on hierarchical token representation of the present application, the following content may also be specifically included:
[0149] Step S601: Use an image encoder to extract image feature sequences from the input image to be processed, pass the image feature sequences through a multi-layer perceptron module to obtain a visual representation sequence aligned with the input layer feature space of the large language model, and input the aligned visual representation sequence into the large language model network of the masked semantic segmentation model;
[0150] Step S602: The large language model network of the masked semantic segmentation model uses an autoregressive decoding method to generate a masked token sequence. For each decoding time step, perform multi-modal interaction between the generated masked tokens and the visual feature representations to obtain interaction features, predict the probability distribution of the next masked token based on the interaction features, and select a masked token from the probability distribution through a sampling or greedy decoding strategy until a complete masked token sequence is generated.
[0151] The decoder of this embodiment uses an autoregressive transformer architecture, including multiple layers of self-attention and cross-modal attention modules. The calculation process of each layer includes feature update and normalization:
[0152] H' = LayerNorm(H + MultiHead(H))
[0153] Where H is the hidden state feature and MultiHead is the multi - head attention operation. This design enables the model to effectively integrate multimodal information.
[0154] This embodiment implements an adaptive masked token generation mechanism. At each decoding time step, the generated masked token sequence S_t is subjected to multimodal interaction with the visual feature V:
[0155] F_inter = Interaction(S_t, V)
[0156] Where Interaction is the interaction module, which includes attention calculation and feature fusion. The interaction feature carries rich context information and provides guidance for subsequent token generation.
[0157] This embodiment constructs a probability prediction network. Based on the interaction feature, the probability distribution of the next masked token is predicted:
[0158] P(t + 1) = softmax(MLP(F_inter) / τ)
[0159] Where MLP is the multi - layer perceptron and τ is the temperature parameter. The adjustment of the temperature parameter affects the diversity of sampling. A lower temperature tends to select high - probability tokens.
[0160] This embodiment optimizes the decoding strategy selection mechanism. Dynamically select sampling or greedy decoding according to the task requirements: token = argmax(P) # greedy decoding
[0161] token ~ multinomial(P) # random sampling
[0162] Greedy decoding is suitable for scenarios that require deterministic results, while random sampling helps to generate diverse segmentation results.
[0163] Through the above - mentioned technological innovations, this embodiment effectively solves the deficiencies of traditional image segmentation methods in semantic understanding and feature fusion. This solution can accurately understand text descriptions and locate and segment corresponding target regions in images. In practical applications, this method significantly improves the accuracy and user experience of interactive image segmentation.
[0164] The innovation points of this embodiment are mainly reflected in aspects such as feature alignment, decoding mechanism, and strategy selection. Through a carefully designed processing flow, precise segmentation guided by text is achieved. This solution provides a new solution for improving the semantic understanding ability of image segmentation models and has important application value in the field of computer vision. Experimental verification shows that this method can effectively improve the segmentation performance, especially showing obvious advantages in complex scenarios that require precise semantic understanding.
[0165] In an embodiment of the multi-modal large model image segmentation method based on hierarchical token representation of the present application, the following content may also be specifically included:
[0166] Step S701: Input the generated masked token sequence into a masked decoder, which is a Transformer model with a bidirectional attention mechanism. Along with the masked token sequence, several initial hidden states are also input. After passing through the masked decoder, the masked token sequence is mapped to an initial feature map. The spatial resolution of the feature map is restored through an upsampling module. Each upsampling module includes a nearest neighbor upsampling layer and a convolutional layer. The upsampled feature map is normalized to obtain a probability mask map, and the probability mask map is converted into a binary mask to obtain a segmentation mask image;
[0167] Step S702: Perform morphological processing and boundary optimization on the segmentation mask image. The morphological processing includes performing opening and closing operations on the segmentation mask to remove noise and smooth the boundary. The boundary optimization locally adjusts the segmentation boundary based on image gradient information. The optimized segmentation mask is pixel-level fused with the original image, and the contour and boundary of the target region are marked on the image to be processed, generating a segmentation result image containing the target region identifier.
[0168] Optionally, this embodiment implements an innovative masked reconstruction mechanism. The input of the masked decoder is a masked token sequence and several initial hidden state sequences. The generation of the initial feature map is completed by a Transformer with bidirectional attention:
[0169] F_init = Transformer(Concat(tokens, Embed))
[0170] Where tokens is the masked token sequence, Embed is the initial hidden state sequence, and Concat is the sequence concatenation operation. F_init is the feature corresponding to Embed after passing through the Transformer.
[0171] The upsampling module used in this embodiment is a nearest neighbor upsampling and convolution module. The upsampling operation is expressed as: F_up = Conv(Upsample(F_in))
[0172] Where F_in is the input feature, and Upsample and Conv are the nearest neighbor upsampling layer and the convolutional layer respectively. This design not only ensures the smooth upsampling of features but also maintains the transmission of detailed information.
[0173] This embodiment implements a generation strategy for probability masks. The feature map is normalized to convert it into a probability distribution:
[0174] P_mask = sigmoid(F_final)
[0175] Among them, F_final is the final feature map. The probability mask reflects the confidence of each pixel belonging to the target area, providing a reliable basis for subsequent binarization.
[0176] This embodiment constructs an adaptive binarization method. The binarization of the probability mask adopts a dynamic threshold strategy: M_binary = threshold(P_mask, t)
[0177] Among them, t is the adaptive threshold, which is dynamically determined according to the probability distribution by the OTSU algorithm. This method can better adapt to the segmentation requirements of different scenarios.
[0178] This embodiment optimizes the morphological processing flow. The kernel sizes of the opening operation and the closing operation are dynamically adjusted according to the image resolution:
[0179] k = max(3, round(min(H, W) / 100))
[0180] Among them, H and W are the height and width of the image. The morphological operations effectively remove the noise and holes in the segmentation result, improving the quality of the segmentation mask.
[0181] This embodiment implements an accurate boundary optimization mechanism. A boundary energy function is constructed based on the image gradient information:
[0182] Among them is the image gradient, smoothness is the smoothing term, and w is the weight coefficient. The segmentation boundary is optimized by minimizing the energy function to make it better fit the target contour.
[0183] This embodiment innovatively designs a mask fusion strategy. The optimized segmentation mask is alpha-blended with the original image:
[0184] I_result = α * I_orig + (1 - α) * M_color
[0185] Among them, I_orig is the original image, M_color is the colored mask, and α is the transparency parameter. This visualization method intuitively shows the segmentation result.
[0186] Through the above technological innovations, this embodiment effectively solves the deficiencies of traditional segmentation methods in mask reconstruction and boundary optimization. This scheme can generate high-quality segmentation results and further improve the segmentation quality through post-processing. In practical applications, this method significantly improves the visual quality and practicality of the segmentation results.
[0187] The innovations of this embodiment are mainly reflected in feature reconstruction, boundary optimization, and visual presentation. Through a carefully designed processing flow, a high-quality conversion from masked tokens to accurate segmentation results is achieved. This solution provides important support for improving the performance of image segmentation models and has broad application prospects in the field of computer vision. Experimental verification shows that this method can generate segmentation results with accurate boundaries and excellent visual effects, especially showing obvious advantages in application scenarios that require fine segmentation.
[0188] In order to effectively address the deficiencies of traditional technologies in complex scene understanding, model training, and mask generation, and significantly improve the performance of multi-modal image segmentation, this application provides an embodiment of a multi-modal large model image segmentation device based on hierarchical token representation for implementing all or part of the content of the multi-modal large model image segmentation method based on hierarchical token representation. Refer to Figure 2 The multi-modal large model image segmentation device based on hierarchical token representation specifically includes the following:
[0189] The token representation module 10 is used to encode the input masked image into a one-dimensional masked token sequence through a mask tokenizer. The mask tokenizer generates the current masked token using a causal attention mechanism based on the mask blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target area, and the rear tokens represent the local detail features of the target area. Each masked token in the masked token sequence is generated conditionally based on its previous masked token. The masked token sequence is input into a vector quantization layer for token quantization to obtain masked token codes, and a mask decoder reconstructs the masked image according to the masked token codes;
[0190] The model training module 20 is used to perform the first-stage training on the mask tokenizer through a mask reconstruction task, perform the second-stage training using a large multi-modal model, expand the masked token codes into the vocabulary of the large multi-modal model, integrate the masked tokens into the large multi-modal model based on low-resolution image data, perform joint training according to the segmentation dataset and the general dataset, use a hierarchical mask loss function to perform supervised training on the masked token sequence at different hierarchical levels, and perform the third-stage fine-tuning training using high-resolution image data to obtain a mask semantic segmentation model;
[0191] The image segmentation module 30 is configured to receive the image to be processed and the corresponding target object description in the referential expression segmentation request, input the image to be processed and the target object description into the masked semantic segmentation model. The masked semantic segmentation model generates the corresponding masked token sequence based on the target object description. The masked decoder generates the corresponding segmentation mask image according to the masked token sequence. The masked semantic segmentation model identifies the target area referred to by the target object description in the image to be processed according to the segmentation mask image.
[0192] As can be seen from the above description, the multi-modal large model image segmentation device based on hierarchical token representation provided by the embodiments of the present application can encode the mask image into a token sequence by designing a masked tokenizer, and achieve progressive generation from shape prototypes to local details through the causal attention mechanism. Adopting a three-stage training strategy, first training the tokenizer through a mask reconstruction task, then integrating the masked tokens into a large multi-modal model and performing joint training, and finally fine-tuning using high-resolution data. Hierarchical masked loss functions are used for multi-level supervision to achieve accurate segmentation of natural language description targets. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0193] At the hardware level, in order to effectively solve the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improve the performance of multi-modal image segmentation, the present application provides an embodiment of an electronic device for implementing all or part of the content in the multi-modal large model image segmentation method based on hierarchical token representation. The electronic device specifically includes the following:
[0194] A processor, a memory, a communication interface, and a bus; wherein, the processor, the memory, and the communication interface complete communication with each other through the bus; the communication interface is used to implement information transmission between the multi-modal large model image segmentation device based on hierarchical token representation and related devices such as the core business system, the user terminal, and the relevant database. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the multi-modal large model image segmentation method based on hierarchical token representation and the embodiments of the multi-modal large model image segmentation device based on hierarchical token representation, and the content is incorporated herein, and the repeated parts will not be elaborated.
[0195] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, smart watches, smart bracelets, etc.
[0196] In practical applications, part of the multimodal large model image segmentation method based on hierarchical token representation can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation in this regard. If all operations are completed in the client device, the client device may further include a processor.
[0197] The above-mentioned client device may have a communication module (i.e., a communication unit), and can communicate with a remote server to realize data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server on the intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.
[0198] Figure 3 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 3 shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It should be noted that this Figure 3 is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0199] In one embodiment, the function of the multimodal large model image segmentation method based on hierarchical token representation can be integrated into the central processor 9100. Among them, the central processor 9100 can be configured to perform the following controls:
[0200] Step S101: Encode the input masked image into a one-dimensional masked token sequence by a masked tokenizer. The masked tokenizer generates the current masked token by using a causal attention mechanism based on the masked blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target region, and the rear tokens represent the local detail features of the target region. Each masked token in the masked token sequence is generated conditionally based on its previous masked token. Input the masked token sequence into a vector quantization layer for token quantization to obtain masked token codes, and a masked decoder reconstructs the masked image according to the masked token codes;
[0201] Step S102: Conduct the first-stage training on the masked tokenizer through a masked reconstruction task, conduct the second-stage training by using a large multi-modal model, expand the masked token codes into the vocabulary of the large multi-modal model, integrate the masked tokens into the large multi-modal model based on low-resolution image data, conduct joint training according to the segmentation dataset and the general dataset, use a hierarchical masked loss function to conduct supervised training on the masked token sequence at different hierarchical levels, and conduct the third-stage fine-tuning training by using high-resolution image data to obtain a masked semantic segmentation model;
[0202] Step S103: Receive the image to be processed and the corresponding target object description in the referential expression segmentation request, input the image to be processed and the target object description into the masked semantic segmentation model. The masked semantic segmentation model generates the corresponding masked token sequence based on the target object description, the masked decoder generates the corresponding segmentation masked image according to the masked token sequence, and the masked semantic segmentation model identifies the target region referred to by the target object description in the image to be processed according to the segmentation masked image.
[0203] As can be seen from the above description, the electronic device provided in the embodiment of the present application encodes the masked image into a token sequence by designing a masked tokenizer, and realizes the progressive generation from the shape prototype to the local details through a causal attention mechanism. Adopt a three-stage training strategy. First, train the tokenizer through a masked reconstruction task, then integrate the masked tokens into a large multi-modal model and conduct joint training, and finally fine-tune by using high-resolution data. Conduct multi-level supervision based on a hierarchical masked loss function to achieve accurate segmentation of the natural language description target. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, masked generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0204] In another embodiment, the multi-modal large model image segmentation device based on hierarchical token representation can be separately configured from the central processing unit 9100. For example, the multi-modal large model image segmentation device based on hierarchical token representation can be configured as a chip connected to the central processing unit 9100, and the functions of the multi-modal large model image segmentation method based on hierarchical token representation can be realized through the control of the central processing unit.
[0205] As Figure 3 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 3 all the components shown in Figure 3 ; in addition, the electronic device 9600 may further include
[0206] As Figure 3 shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control unit, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.
[0207] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0208] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.
[0209] The memory 9140 can be a solid-state memory, for example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be a memory that stores information even when the power is off, can be selectively erased and has more data. An example of this memory is sometimes referred to as an EPROM, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage unit 9142, which is used to store application programs and function programs or the processes for operating the electronic device 9600 by the central processing unit 9100.
[0210] The memory 9140 can also include a data storage unit 9143, which is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 can include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0211] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.
[0212] Based on different communication technologies, multiple communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, so as to implement normal telecommunication functions. The audio processor 9130 can include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processing unit 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.
[0213] An embodiment of the present application also provides a computer-readable storage medium capable of implementing all steps of the multi-modal large model image segmentation method based on hierarchical token representation with the execution entity being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements all steps of the multi-modal large model image segmentation method based on hierarchical token representation with the execution entity being a server or a client in the above embodiments. For example, when the processor executes the computer program, the following steps are implemented:
[0214] Step S101: Encode the input masked image into a one-dimensional masked token sequence through a mask tokenizer. The mask tokenizer generates the current masked token by using a causal attention mechanism according to the mask blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target region, and the rear tokens represent the local detail features of the target region. Each masked token in the masked token sequence is generated conditionally based on its previous masked token. Input the masked token sequence into a vector quantization layer for token quantization to obtain masked token codes, and a mask decoder reconstructs the masked image according to the masked token codes;
[0215] Step S102: Conduct the first-stage training on the mask tokenizer through a mask reconstruction task, conduct the second-stage training by using a large multi-modal model, expand the masked token codes into the vocabulary of the large multi-modal model, integrate the masked tokens into the large multi-modal model based on low-resolution image data, conduct joint training according to the segmentation data set and the general data set, use a hierarchical mask loss function to conduct supervised training on the masked token sequence at different hierarchical levels, and conduct the third-stage fine-tuning training by using high-resolution image data to obtain a mask semantic segmentation model;
[0216] Step S103: Receive the image to be processed and the corresponding target object description in the referential expression segmentation request, input the image to be processed and the target object description into the mask semantic segmentation model. The mask semantic segmentation model generates the corresponding masked token sequence based on the target object description. The mask decoder generates the corresponding segmentation mask image according to the masked token sequence. The mask semantic segmentation model identifies the target region referred to by the target object description in the image to be processed according to the segmentation mask image.
[0217] As can be seen from the above description, the computer-readable storage medium provided by the embodiments of the present application encodes a masked image into a sequence of tokens through a designed mask tokenizer, and realizes progressive generation from a shape prototype to local details through a causal attention mechanism. A three-stage training strategy is adopted. First, the tokenizer is trained through a mask reconstruction task, then the masked tokens are integrated into a large multi-modal model and jointly trained, and finally fine-tuned using high-resolution data. Hierarchical mask loss functions are used for multi-level supervision to achieve accurate segmentation of natural language description targets. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0218] An embodiment of the present application also provides a computer program product that can implement all the steps of the multi-modal large model image segmentation method based on hierarchical token representation in which the execution subject in the above embodiment is a server or a client. When the computer program / instructions are executed by a processor, the steps of the multi-modal large model image segmentation method based on hierarchical token representation are implemented. For example, the computer program / instructions implement the following steps:
[0219] Step S101: Encode the input masked image into a one-dimensional masked token sequence through a mask tokenizer. The mask tokenizer generates the current masked token using a causal attention mechanism based on the masked blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target area, and the rear tokens represent the local detail features of the target area. Each masked token in the masked token sequence is generated conditionally based on its previous masked token. Input the masked token sequence into a vector quantization layer for token quantization to obtain masked token codes, and a mask decoder reconstructs the masked image according to the masked token codes;
[0220] Step S102: Perform the first-stage training on the mask tokenizer through a mask reconstruction task, perform the second-stage training using a large multi-modal model, expand the masked token codes into the vocabulary of the large multi-modal model, integrate the masked tokens into the large multi-modal model based on low-resolution image data, perform joint training according to the segmentation dataset and the general dataset, use hierarchical mask loss functions to perform supervised training on the masked token sequence at different hierarchical levels, and perform the third-stage fine-tuning training using high-resolution image data to obtain a mask semantic segmentation model;
[0221] Step S103: Receive the image to be processed and the corresponding target object description in the referring expression segmentation request, input the image to be processed and the target object description into the mask semantic segmentation model. The mask semantic segmentation model generates the corresponding mask token sequence based on the target object description. The mask decoder generates the corresponding segmentation mask image according to the mask token sequence. The mask semantic segmentation model identifies the target area referred to by the target object description in the image to be processed according to the segmentation mask image.
[0222] As can be seen from the above description, the computer program product provided by the embodiments of the present application encodes the mask image into a token sequence by designing a mask tokenizer, and realizes the progressive generation from the shape prototype to the local details through the causal attention mechanism. Adopt a three-stage training strategy. First, train the tokenizer through the mask reconstruction task, then integrate the mask tokens into a large multi-modal model and conduct joint training, and finally fine-tune with high-resolution data. Based on the hierarchical mask loss function for multi-level supervision, the accurate segmentation of the natural language description target is realized. This method effectively solves the deficiencies of traditional technologies in complex scene understanding, model training, mask generation, etc., and significantly improves the performance of multi-modal image segmentation.
[0223] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0224] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0225] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0226] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0227] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A multimodal large model image segmentation method based on hierarchical token representation, characterized in that, The method includes: Encoding an input masked image into a one-dimensional masked token sequence by a masked tokenizer, where the masked tokenizer generates a current masked token using a causal attention mechanism based on the masked blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target region, and the rear tokens represent the local detail features of the target region. Each masked token in the masked token sequence is conditionally generated based on its previous masked token. Inputting the masked token sequence into a vector quantization layer for token quantization to obtain masked token codes, and a masked decoder reconstructs the masked image according to the masked token codes; Performing a first-stage training on the masked tokenizer through a masked reconstruction task, performing a second-stage training using a large multi-modal model, expanding the masked token codes into the vocabulary of the large multi-modal model, integrating the masked tokens into the large multi-modal model based on low-resolution image data, performing joint training according to a segmentation dataset and a general dataset, using a hierarchical masked loss function to perform supervised training on the masked token sequence at different hierarchical levels, and performing a third-stage fine-tuning training using high-resolution image data to obtain a masked semantic segmentation model; Receiving a to-be-processed image and a corresponding target object description in a referential expression segmentation request, inputting the to-be-processed image and the target object description into the masked semantic segmentation model. The masked semantic segmentation model generates a corresponding masked token sequence based on the target object description. The masked decoder generates a corresponding segmentation masked image according to the masked token sequence, and the masked semantic segmentation model identifies the target region referred to by the target object description in the to-be-processed image according to the segmentation masked image.
2. The multi-modal large model image segmentation method based on hierarchical token representation according to claim 1, wherein The encoding of the input masked image into a one-dimensional masked token sequence by the masked tokenizer, where the masked tokenizer generates a current masked token using a causal attention mechanism based on the masked blocks corresponding to the masked image and the generated masked tokens. The front tokens of the masked token sequence represent the position and shape prototype of the target region, and the rear tokens represent the local detail features of the target region, includes: Performing a chunking process on the input masked image to obtain a sequence of masked image chunks, and jointly inputting the sequence of masked image chunks and several fixed hidden vectors into the encoder network of the masked tokenizer for joint self-attention calculation. After the above encoding process, the hidden vectors obtain target masked tokens. The encoder network adopts conditional autoregressive modeling. Each attention layer has bidirectional attention between masked image chunks and unidirectional attention between hidden vectors, which ensures that the hidden vectors can represent the masked image from coarse-grained to fine-grained, so that the hidden vectors are encoded into the accurate one-dimensional masked token sequence.
3. The multi-modal large model image segmentation method based on hierarchical token representation according to claim 1, characterized in that Each masked token in the masked token sequence is input into a vector quantization layer for token quantization to obtain masked token codes, and a masked decoder reconstructs the masked image according to the masked token codes, includes: Input the mask token sequence generated by conditioning into a vector quantization layer. The vector quantization layer maintains a preset number of codebook vectors. By calculating the Euclidean distance between each mask token in the mask token sequence and the codebook vectors, select the codebook vector with the smallest distance as the corresponding mask token code. The mask decoder uses a Transformer deep learning model to perform feature mapping and upsampling on the mask token code, and restore the feature mapping result to a reconstructed image with the same resolution as the input mask image.
4. The multimodal large model image segmentation method based on hierarchical token representation according to claim 1, characterized in that, The first stage of training for the mask tokenizer is performed through a mask reconstruction task, and the second stage of training is performed using a large multi-modal model. Expand the vocabulary of the large multi-modal model with the mask token code, and integrate the mask token into the large multi-modal model based on low-resolution image data, including: Construct a mask reconstruction loss function to train the mask tokenizer. The mask reconstruction loss function includes a reconstruction error term and a codebook commitment term. The reconstruction error term calculates the binary cross-entropy loss and the region overlap loss between the reconstructed mask image and the original mask image. The region overlap loss term calculates the intersection over union score between the predicted mask and the true mask at different hierarchical scales. The codebook commitment term calculates the Euclidean distance between the quantization feature of the mask token and the nearest neighbor codebook vector. Optimize the mask reconstruction loss function based on the Adam optimizer, and update the network parameters of the mask tokenizer and the codebook vectors. Map the mask token code to a word embedding vector with a fixed dimension, expand the word embedding matrix of the large multi-modal model to accommodate the word embedding vector, and construct a training sample containing mask tokens based on low-resolution images. The input sequence of the training sample consists of image feature tokens and mask tokens. The large multi-modal model uses a cross-entropy loss function and a hierarchical mask loss function to train the training sample, enabling the model to learn the semantic correspondence between mask tokens and other modal tokens. The hierarchical mask loss function performs supervised training on the mask token sequence at different hierarchical levels. Each hierarchical level of the token sequence is a sub-token sequence counted from the front. The sub-token sequence passes through the mask decoder to obtain the predicted mask at the corresponding level. The loss function at each hierarchical level calculates the binary cross-entropy loss and the region overlap loss term between the predicted mask and the true mask at the corresponding level. The final hierarchical mask loss is equal to the sum of the loss functions at each hierarchical level.
5. The multimodal large model image segmentation method based on hierarchical token representation according to claim 1, wherein, The joint training is performed according to the segmentation dataset and the general dataset, and the third stage of fine-tuning training is performed using high-resolution image data to obtain a mask semantic segmentation model, including: Construct a mixed training data batch. Each training batch contains a preset proportion of segmentation data samples and general data samples. Use high-resolution image data to fine-tune the model after the second stage of training. The large multi-modal model treats ordinary text and mask tokens equally, and the training loss function uses cross-entropy loss, thereby obtaining the final mask semantic segmentation model.
6. The multimodal large model image segmentation method based on hierarchical token representation according to claim 1, characterized in that The receiving refers to representing the image to be processed and the corresponding target object description in the expression segmentation request, and inputting the image to be processed and the target object description into the mask semantic segmentation model. The mask semantic segmentation model generates the corresponding mask token sequence based on the target object description, including: Using an image encoder to extract image feature sequences from the input image to be processed, passing the image feature sequences through a multi-layer perceptron module to obtain a visual representation sequence aligned with the input layer feature space of the large language model, and inputting the aligned visual representation sequence into the large language model network of the mask semantic segmentation model; The large language model network of the mask semantic segmentation model generates a mask token sequence using an autoregressive decoding method. For each decoding time step, the generated mask tokens are interacted with the visual feature representation to obtain interaction features, and the probability distribution of the next mask token is predicted based on the interaction features. The mask token is selected from the probability distribution through a sampling or greedy decoding strategy until a complete mask token sequence is generated.
7. The multimodal large model image segmentation method based on hierarchical token representation according to claim 1, characterized in that, The mask decoder generates a corresponding segmentation mask image according to the mask token sequence, and the mask semantic segmentation model identifies the target area referred to by the target object description in the image to be processed according to the segmentation mask image, including: Inputting the generated mask token sequence into the mask decoder. The mask decoder is a Transformer model with a bidirectional attention mechanism. Along with the mask token sequence, several initial hidden states are also input. After passing through the mask decoder, the mask token sequence is mapped to an initial feature map, and the spatial resolution of the feature map is restored through an upsampling module. Each upsampling module includes a nearest neighbor upsampling layer and a convolutional layer. The upsampled feature map is normalized to obtain a probability mask map, and the probability mask map is converted into a binary mask to obtain a segmentation mask image; Performing morphological processing and boundary optimization on the segmentation mask image. The morphological processing includes performing opening and closing operations on the segmentation mask to remove noise and smooth the boundary. The boundary optimization locally adjusts the segmentation boundary based on image gradient information, and the optimized segmentation mask is pixel-level fused with the original image to label the contour and boundary of the target area on the image to be processed, generating a segmentation result image containing the target area identification.
8. An image segmentation device for a multimodal large model based on hierarchical token representation, characterized in that, The device includes: A word representation module for encoding the input mask image into a one-dimensional mask token sequence through a mask tokenizer. The mask tokenizer generates the current mask token using a causal attention mechanism according to the mask blocks corresponding to the mask image and the generated mask tokens. The front tokens of the mask token sequence represent the position and shape prototype of the target area, and the rear tokens represent the local detail features of the target area. Each mask token in the mask token sequence is conditionally generated based on its previous mask token. The mask token sequence is input into a vector quantization layer for token quantization to obtain mask token codes, and the mask decoder reconstructs the mask image according to the mask token codes; A model training module for performing a first-stage training on the mask tokenizer through a mask reconstruction task, a second-stage training using a large multi-modal model, expanding the mask token code into the vocabulary of the large multi-modal model, integrating the mask token into the large multi-modal model based on low-resolution image data, performing joint training according to a segmentation data set and a general data set, using a hierarchical mask loss function to perform supervised training on the mask token sequence at different hierarchical levels, and performing a third-stage fine-tuning training using high-resolution image data to obtain a mask semantic segmentation model; An image segmentation module for receiving a to-be-processed image and a corresponding target object description in a referential expression segmentation request, inputting the to-be-processed image and the target object description into the mask semantic segmentation model, the mask semantic segmentation model generating a corresponding mask token sequence based on the target object description, the mask decoder generating a corresponding segmentation mask image according to the mask token sequence, and the mask semantic segmentation model identifying a target area referred to by the target object description in the to-be-processed image according to the segmentation mask image.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multi-modal large model image segmentation method based on hierarchical token representation according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-modal large model image segmentation method based on hierarchical token representation according to any one of claims 1 to 7.
Citation Information
Patent Citations
Zero sample image segmentation model training method and device based on multiple modes
CN117788981A
Visual localization and anaphora segmentation method, system and device based on mask anaphora modeling and storage medium
CN118734091A
Multimodality Image Segmentation of Volumetric Data Sets
US20140003686A1
Support of multi-mode extraction for multi-layer video codecs
US20150103888A1
Data processing method and apparatus
WO2024213099A1
Cited By
Feature matching remote sensing image small sample semantic segmentation model evolution method
CN120747517A
A feature matching remote sensing image small sample semantic segmentation model evolution method
CN120747517B
Conversation content generation method and electronic equipment
CN121524331A
A dialogue content generation method and an electronic device
CN121524331B
Segmentation model training method and device, segmentation method and device, equipment and medium
CN121982324A