Representation image segmentation method and system based on repetition-interleaving
Through the reaffirming-interleaving processing mechanism, the problems of insufficient text semantics and low computing efficiency in the prior art are solved, and the accuracy and robustness of image segmentation are improved, and efficient image segmentation is adapted to complex scenarios.
Patent Information
- Application Number
- CN202510320698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-18
AI Technical Summary
When the prior art processes complex and lengthy text descriptions, it is difficult to fully understand the multi-level semantics in the text, resulting in a decrease in image segmentation accuracy, and the computational efficiency of traditional architectures is low, making it difficult to adapt to real-time application requirements.
The reaffirmation-interleaving reference image segmentation method is adopted, through grouping, stacking, splicing, scanning fusion and aggregation of visual features and text features, combined with multi-layer convolutional neural network, the model is trained using the gradient descent method to achieve deep interaction between images and text.
It improves the accuracy and robustness of image segmentation, ensures the stability and accuracy of text semantics, enhances cross-modal understanding capabilities, and adapts to image segmentation in complex scenarios.
Smart Images

Figure CN120388038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to a reference image segmentation method and system based on reiteration-interleaving. Background Art
[0002] Image segmentation is a fundamental task in computer vision, aiming to classify and label different regions in an image. Traditional image segmentation tasks mainly focus on the visual information of the image itself, while in recent years, multi-modal learning combining language information has gradually become a research hotspot. In more challenging problems, the model not only needs to process image data but also accurately segment the image according to the text description. This problem is called reference image segmentation.
[0003] To solve this problem, certain progress has been made in the research on cross-modal alignment mechanisms. Chinese Patent Application Publication No. CN116912837A discloses a reference image segmentation method based on detail and boundary driving, which enhances feature alignment by using boundary, detail, and saliency detection methods; Chinese Patent Application Publication No. CN116704506A proposes a reference image segmentation method that adaptively adjusts the multi-modal correspondence relationship according to different global semantic features, enhancing the model's ability to understand cross-modal information. However, it usually assumes that the text description is a short phrase or a simple structure. This assumption overly simplifies the diversity and flexibility of the description language in actual situations, resulting in limited understanding ability for complex semantic expressions and thus restricting its practical applications. In actual scenarios, such methods still have significant defects:
[0004] (1) Existing models overly simplify the processing of text features and lack the ability to model complex syntactic structures and semantic dependencies, resulting in the implicit context relationships in the description not being fully utilized;
[0005] (2) There is a serious imbalance in the sequence lengths between visual and text features. The visual features of long sequences are likely to obscure the semantic information of short texts, weakening the fine-grainedness of cross-modal interaction;
[0006] (3) Traditional architectures (such as Transformer) have problems with low computational efficiency when processing long sequences and are difficult to meet the requirements of real-time applications. The above problems cause the segmentation accuracy of existing methods to significantly decline and the robustness to be insufficient when facing complex and lengthy text descriptions in real scenarios.
[0007] Therefore, there is an urgent need for a technical solution that can efficiently balance multi-modal sequence interaction, deeply analyze complex text semantics, and improve the quality of cross-modal alignment to break through the performance bottleneck of current reference image segmentation. Summary of the Invention
[0008] The objective of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a reference image segmentation method and system based on reiteration-interleaving, so as to solve or partially solve the problem that the implicit context relationship in the text description is not fully utilized, thereby affecting the segmentation effect.
[0009] The objective of the present invention can be achieved through the following technical solutions:
[0010] In one aspect of the present invention, there is provided a reference image segmentation method based on reiteration-interleaving. Based on the input image and text, the segmentation result is obtained by using the trained reference image segmentation model. Among them, the training process of the trained reference image segmentation model includes the following steps:
[0011] Obtain the target picture and the attribute text;
[0012] Through visual feature extraction, encode the target picture into visual features;
[0013] Through text feature extraction, encode the attribute text into text features;
[0014] Based on the visual features, perform grouping processing, based on the text features, perform stacking processing, complete the reiteration operation. Based on the grouped visual features and the stacked text features, perform splicing, scan fusion, rearrangement, and aggregation processing based on learnable proportional parameters to obtain new visual features and text features, complete the interleaving operation. By repeating the reiteration-interleaving operation multiple times, obtain the visual features after each reiteration-interleaving operation, and obtain the segmentation result according to the probability map;
[0015] Based on the segmentation result and the obtained true segmentation mask, train the reference image segmentation model.
[0016] As a preferred technical solution, the process of encoding the target picture into visual features includes the following steps:
[0017] Cut the target picture into multiple non-overlapping slices, map each slice to the feature space through linear projection or convolution to obtain a feature sequence;
[0018] Based on the learnable relative position bias, process the feature sequence;
[0019] For the feature sequence after bias processing, use the sliding window self-attention model, through multi-layer window self-attention, window sum, and progressive downsampling, to obtain the output visual features.
[0020] As a preferred technical solution, the process of encoding the attribute text into text features includes the following steps:
[0021] For the described text features, a tokenizer is used to generate sub-word level tokenization encodings;
[0022] Convert the tokenization encodings into embedding vectors, and add them to the absolute position encodings and segment encodings;
[0023] Based on the added features, use a text sub-attention model to obtain word-level features describing the text encoding as the text features.
[0024] As a preferred technical solution, the process of grouping processing based on the visual features includes the following steps:
[0025] Use the visual features as the input of a two-dimensional state space model to obtain updated visual features with invariant shape;
[0026] Based on the updated visual features, re-group by window and flatten into a sequence by group to obtain re-grouped visual features.
[0027] As a preferred technical solution, the process of stacking processing based on the text features includes the following steps:
[0028] Repeat the text feature encoding features several times and stack them to obtain new text features, where the number of repetitions matches the number of groups of visual features.
[0029] As a preferred technical solution, the process of splicing, scan fusion, reordering, and aggregation processing based on the learnable proportional parameter for the grouped visual features and the stacked text features includes the following steps:
[0030] Splice the grouped visual features and the stacked text features to obtain a hybrid multi-modal sequence;
[0031] Input the hybrid multi-modal sequence into a two-dimensional state space model to obtain an updated hybrid multi-modal sequence with invariant shape after cross-modal fusion, and obtain updated visual features and updated text features;
[0032] Reorder the updated visual features to obtain visually fused visual features as new visual features;
[0033] Aggregate the updated text features according to the learnable proportional parameter to obtain aggregated text features after modal fusion as new text features.
[0034] As a preferred technical solution, the process of obtaining the visual features after each reiteration-interleaving operation and obtaining the segmentation result according to the probability map by repeating the reiteration-interleaving operation multiple times includes the following steps:
[0035] By repeating the reiteration-interleaving operation multiple times, the visual features after each reiteration-interleaving operation are obtained, constituting a visual feature sequence;
[0036] Taking the visual feature sequence as the input of a multi-layer convolutional neural network, the visually characterized multi-modal fusion is obtained;
[0037] Based on the visually characterized multi-modal visual features, it is restored to the size matching the target picture by interpolation to obtain a segmentation probability map, and the segmentation result is obtained through threshold conversion.
[0038] As a preferred technical solution, the process of training the referential image segmentation model based on the segmentation result and the obtained true segmentation mask includes the following steps:
[0039] Based on the segmentation result and the obtained true segmentation mask, the Dice loss and the focal loss are calculated to obtain a comprehensive loss;
[0040] Based on the comprehensive loss, the referential image segmentation model is trained by gradient descent.
[0041] As a preferred technical solution, the comprehensive loss is:
[0042] L = Dice(y, m) + Focal(y, m)
[0043] Among them, y is the predicted segmentation result, m is the true segmentation result of the picture, Dice represents the Dice loss function, and Focal represents the focal loss function.
[0044] Another aspect of the present invention provides a reiteration-interleaving-based referential image segmentation system for implementing the foregoing reiteration-interleaving-based referential image segmentation method. The referential image segmentation system includes:
[0045] A visual feature extraction module for encoding the obtained target picture into visual features through visual feature extraction;
[0046] A text feature extraction module for encoding the obtained attribute text into text features through text feature extraction;
[0047] A reiteration-interleaving processing module for performing grouping processing based on the visual features, performing stacking processing based on the text features to complete the reiteration operation, and performing splicing, scan fusion, rearrangement, and aggregation processing based on learnable proportional parameters on the grouped visual features and the stacked text features to obtain new visual features and text features, completing the interleaving operation, and obtaining the visual features after each reiteration-interleaving operation by repeating the reiteration-interleaving operation multiple times, and obtaining the segmentation result according to the probability map;
[0048] A training module for training a referential image segmentation model based on the segmentation result and the obtained true segmentation mask;
[0049] A referential image segmentation module for obtaining a segmentation result by using the trained referential image segmentation model based on the input image and text.
[0050] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0051] (1) Fully understanding the multi-level semantics in the text: For long sentence texts with complex structures or ambiguous semantics, the present invention can better understand the multi-level semantics in the text through the reiteration and interleaving processing mechanism, make full use of the implicit context relationship in the text description, thus avoiding the drawback of over-simplifying the text in the traditional method. By integrating the image and text into a unified multi-modal sequence, the depth interaction between the image and text is further enhanced. This innovative reiteration and interleaving mechanism enables the image and text information to be more closely combined, thereby improving the accuracy and robustness of image segmentation.
[0052] (2) Good stability and accuracy of semantic text: Aiming at the problem of unbalanced proportion of image and text information, the present invention performs splicing, scanning fusion, reordering, and aggregation processing based on the learnable proportion parameter on the grouped visual features and stacked text features to obtain new visual features and text features. By balancing the proportion of the image and text in the multi-modal sequence, the phenomenon that the image information covers the text semantics is avoided, thereby ensuring the stability and accuracy of the text semantics.
[0053] (3) Strong cross-modal understanding ability: The present invention processes the visual features and text features through the reiteration and interleaving processing mechanism, enhances the cross-modal understanding ability for descriptions containing references, context dependencies, or ambiguities, shows higher robustness, and can accurately perform image segmentation in complex scenarios. Description of the Drawings
[0054] Figure 1 It is a flowchart of the reiteration-interleaving based referential image segmentation method in the embodiment;
[0055] Figure 2 It is a schematic diagram of the referential image segmentation model in the embodiment;
[0056] Figure 3 It is a schematic diagram of the reiteration-interleaving based referential image segmentation system in the embodiment;
[0057] Figure 4 It is a schematic diagram of the electronic device in the embodiment. Detailed Embodiment
[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] In view of the problems existing in the foregoing prior art, this embodiment provides a reference image segmentation method based on reiteration-interleaving. By constructing and training a reference image segmentation model as shown in Figure 2 , the trained model is used to implement reference image segmentation. Refer to Figure 1 , the method includes the following steps:
[0061] Step S1, obtain a target picture and attribute text.
[0062] Step S2, use a visual feature extractor to encode the target picture into visual features. Specifically, step S2 may include steps S201 - S203:
[0063] Step S201: Split the target picture x ∈ R H×W×3 into non-overlapping 2D slices, each slice with a size of P×P. Map each slice to a D-dimensional feature space through linear projection or 1×1 convolution to obtain serialized features.
[0064] Step S202: Combine the serialized features with a learnable relative position bias to obtain new features.
[0065] Step S203: Input the above features into a sliding window self-attention model (Swin Transformer). After multiple layers of window self-attention, shifted windows, and hierarchical downsampling (Patch Merging), output the visual features V ∈ R h×w×d , where h and w are the length and width dimensions of the output visual features, and d is the final number of channels.
[0066] Step S3, use a text feature extractor to encode the attribute text into text features. Specifically, step S3 may include steps S301 - S303:
[0067] Step S301: Send the description text into a tokenizer to generate a subword-level token sequence and add special tokens.
[0068] Step S302: Convert the token encoding into an embedding vector, and add it to the absolute position encoding (PositionEmbeddings) and the segment encoding (Segment Embeddings).
[0069] Step S303: Feed the above features into the text self-attention model (text transformer) to obtain the word-level features T ∈ R N×d , where N is the number of word-level features in the description text encoding.
[0070] Step S4: Feed the visual features and the text features into the modality interleaving network to obtain the segmentation probability map, and convert it into the segmentation result using the threshold method. Step S4 may include S401 - S410:
[0071] Step S401: First, feed the visual features V ∈ R h×w×d encoded from the target image into a two-dimensional state space model (SS2D) to obtain the updated visual features V with unchanged shape.
[0072] Step S402: Re-group the above updated visual features V by window and flatten them into a sequence by group to obtain the re-grouped visual features where P is the total number of window groups and L is the number of elements in each window.
[0073] Step S403: Repeat the description text encoding feature T ∈ R N×d several times, and stack them to obtain a new description text where P is the total number of window groups. Steps S401 - S403 implement the "reiteration" operation.
[0074] Step S404: Concatenate the processed visual features and the text features together into a mixed multi-modal sequence H ∈ R P×(L+N)×d .
[0075] Step S405: Feed the above mixed multi-modal sequence H into another two-dimensional state space model (SS2D) to obtain the updated mixed multi-modal sequence H with unchanged shape, and further obtain the updated visual features and the updated text features
[0076] Step S406: Rearrange the updated visual features to obtain the visually fused visual features V’ ∈ R h ×w×d , and use it as the visual features for the new round of iteration.
[0077] Step S407: Aggregate the updated text features according to the learnable proportion parameter a to obtain the aggregated text feature T’ ∈ R N×d , which can be expressed by the formula and use it as the text feature for the new round of iteration. Steps S404 - S407 implement the "interleaving" operation.
[0078] Step S408: Repeat Steps S401 - S407 for M times, and obtain the intermediate process output results {V’1, V’2, …, V’ M}} according to the final visual feature of each round.
[0079] Step S409: Input the above multi - level intermediate results into a multi - layer convolutional neural network (Convolution) and fuse them from top to bottom to obtain the final visual feature V after the reiteration - interleaving network multi - modal fusion out .
[0080] Step S410: For the final visual feature V out , use the interpolation method to restore it to the original picture length and width to obtain the segmentation probability map, and further convert it into the segmentation result y using the threshold method.
[0081] Step S5, calculate the loss between the predicted segmentation result and the true segmentation mask, and use the gradient descent method for training. Specifically, Step S5 can include Steps S501 - S502:
[0082] Step S501: Calculate the loss function L = Dice(y, m)+Focal(y, m), where Dice represents the Dice Loss function, Focal represents the Focal Loss function, y is the predicted segmentation result, and m is the true segmentation result of the picture.
[0083] Step S502: Use the gradient descent method to update L.
[0084] Step S6: Use the trained model to infer the test scene image and the test description text to obtain the test image segmentation result.
[0085] In summary, this method designs a referential image segmentation framework. Specifically, in the stage of processing complex text descriptions, first, for long sentence texts with complex structures or ambiguous semantics, through the "reiteration and interweaving" mechanism, it is possible to better understand the multi-level semantics in the text, thus avoiding the drawback of over-simplifying the text in traditional methods. Second, for the situation where the proportion of image and text information is unbalanced, by balancing the proportion of image and text in the multi-modal sequence, the phenomenon that image information covers text semantics is avoided, thus ensuring the stability and accuracy of text semantics. For descriptions containing references, context dependencies, or ambiguities, the cross-modal understanding ability is enhanced, showing higher robustness, and accurate image segmentation can be performed in complex scenarios. In the multi-modal interaction stage, this method further enhances the deep interaction between images and texts by integrating images and texts into a unified multi-modal sequence. This innovative reiteration and interweaving mechanism enables image and text information to be more closely combined, thereby improving the accuracy and robustness of image segmentation.
[0086] Example 2
[0087] Based on Example 1, refer to Figure 3 , this example provides a reiteration-interweaving based referential image segmentation system for implementing the reiteration-interweaving based referential image segmentation method as described in Example 1. The referential image segmentation system includes:
[0088] (1) Visual feature extraction module, which is used to encode the acquired target picture into visual features through visual feature extraction.
[0089] (2) Text feature extraction module, which is used to encode the acquired attribute text into text features through text feature extraction.
[0090] (3) Reiteration-interweaving processing module, which is used to perform grouping processing based on the visual features, perform stacking processing based on the text features to complete the reiteration operation, and perform splicing, scanning fusion, rearrangement, and aggregation processing based on the learnable ratio parameter on the grouped visual features and stacked text features to obtain new visual features and text features, complete the interweaving operation, and obtain the visual features after each reiteration-interweaving operation by repeating the reiteration-interweaving operation multiple times, and obtain the segmentation result according to the probability map.
[0091] (4) Training module, which is used to train the referential image segmentation model based on the segmentation result and the acquired true segmentation mask.
[0092] (5) Referential image segmentation module, which is used to obtain the segmentation result by using the trained referential image segmentation model based on the input image and text.
[0093] Example 3
[0094] Based on the foregoing embodiments, referring to Figure 4 , this embodiment provides an electronic device, including: one or more processors and a memory, where one or more programs are stored in the memory, and the one or more programs include instructions for executing the reiteration-interleaving based referential image segmentation method as described in Embodiment 1.
[0095] As Figure 4 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 described method. Of course, in addition to the software implementation, the present invention does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0096] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0097] Computer-readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0098] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A reiteration-interleaving based referential image segmentation method, characterized in that Based on the input image and text, the segmentation result is obtained by using the trained anaphora image segmentation model. The training process of the trained anaphora image segmentation model includes the following steps: Obtain the target picture and the attribute text; Through visual feature extraction, encode the target picture into visual features; Through text feature extraction, encode the attribute text into text features; Based on the visual features, perform grouping processing, based on the text features, perform stacking processing, complete the reiteration operation. Based on the grouped visual features and the stacked text features, perform splicing, scan fusion, rearrangement, and aggregation processing based on learnable scale parameters to obtain new visual features and text features, complete the interleaving operation. By repeating the reiteration-interleaving operation multiple times, obtain the visual features after each reiteration-interleaving operation, and obtain the segmentation result according to the probability map; Based on the segmentation result and the obtained true segmentation mask, train the anaphora image segmentation model.
2. The method for segmenting a reference image based on reiteration-interleaving according to claim 1, wherein The process of encoding the target picture into visual features includes the following steps: Slice the target picture into multiple non-overlapping slices, and map each slice to the feature space through linear projection or convolution to obtain a feature sequence; Based on the learnable relative position bias, process the feature sequence; For the feature sequence after bias processing, use the sliding window self-attention model, through multi-layer window self-attention, a window, and progressive downsampling, to obtain the output visual features.
3. A method for segmenting a referential image based on reiteration-interleaving according to claim 1, characterized in that, The process of encoding the attribute text into text features includes the following steps: For the text features, use a tokenizer to generate sub-word level token encodings; Convert the token encodings into embedding vectors, and add them to the absolute position encoding and segment encoding; Based on the added features, use the text sub-attention model to obtain the word-level features describing the text encoding as text features.
4. A reiteration-interleaving based referential image segmentation method according to claim 1, characterized in that The process of performing grouping processing based on the visual features includes the following steps: Use the visual features as the input of the two-dimensional state space model to obtain the updated visual features with unchanged shape; Based on the updated visual features, re-group by window and flatten into a sequence by group to obtain the re-grouped visual features.
5. A reiteration-interleaving-based referential image segmentation method according to claim 1, characterized in that The process of performing stacking processing based on the text features includes the following steps: Repeat the text feature encoding features several times and stack them to obtain new text features, where the number of repetitions matches the number of groups of visual features.
6. The method for segmenting a reference image based on reiteration-interleaving according to claim 1, characterized in that, The process of performing splicing, scan fusion, rearrangement, and aggregation processing based on learnable scale parameters based on the grouped visual features and the stacked text features includes the following steps: Splice the grouped visual features and the stacked text features to obtain a mixed multi-modal sequence; Input the mixed multi-modal sequence into the two-dimensional state space model to obtain the cross-modal fusion mixed multi-modal sequence with unchanged shape, and obtain the updated visual features and updated text features; Rearrange the updated visual features to obtain the modality-fused visual features as the new visual features; Aggregate the updated text features according to the learnable proportional parameters to obtain the aggregated text features after modality fusion, which are used as the new text features.
7. A reiteration-interleaving based referential image segmentation method according to claim 1, characterized in that Through multiple repetitions of the reiteration-interleaving operation, the visual features after each reiteration-interleaving operation are obtained. The process of obtaining the segmentation result according to the probability map includes the following steps: Through multiple repetitions of the reiteration-interleaving operation, the visual features after each reiteration-interleaving operation are obtained, forming a visual feature sequence; Use the visual feature sequence as the input of a multi-layer convolutional neural network to obtain the visual features after multi-level multi-modal fusion; Based on the visual features after multi-level multi-modal, interpolate it to restore it to the size matching the target image to obtain a segmentation probability map, and convert it through the threshold method to obtain the segmentation result.
8. A reiteration-interleaving-based referential image segmentation method according to claim 1, wherein The process of training the referring image segmentation model based on the segmentation result and the obtained true segmentation mask includes the following steps: Based on the segmentation result and the obtained true segmentation mask, calculate the Dice loss and the focal loss to obtain the comprehensive loss; Based on the comprehensive loss, train the referring image segmentation model through gradient descent.
9. A reiteration-interleaving based referential image segmentation method according to claim 8, characterized in that The comprehensive loss is: L = Dice(y, m) + Focal(y, m) where y is the predicted segmentation result, m is the true segmentation result of the image, Dice represents the Dice loss function, and Focal represents the focal loss function.
10. A reiteration-interleaving based anaphoric image segmentation system, characterized in that, A referring image segmentation system for implementing the reiteration-interleaving based referring image segmentation method as described in any one of claims 1-9 includes: A visual feature extraction module for encoding the obtained target image into visual features through visual feature extraction; A text feature extraction module for encoding the obtained attribute text into text features through text feature extraction; A reiteration-interleaving processing module for performing grouping processing based on the visual features, stacking processing based on the text features to complete the reiteration operation, and performing splicing, scan fusion, rearrangement, and aggregation processing based on the learnable proportional parameters on the grouped visual features and the stacked text features to obtain new visual features and text features, completing the interleaving operation. Through multiple repetitions of the reiteration-interleaving operation, the visual features after each reiteration-interleaving operation are obtained, and the segmentation result is obtained according to the probability map; A training module for training the referring image segmentation model based on the segmentation result and the obtained true segmentation mask; A referring image segmentation module for obtaining the segmentation result by using the trained referring image segmentation model based on the input image and text.
Citation Information
Patent Citations
Anaphora image segmentation method based on cross environment attention
CN116704506A
Anaphora target image segmentation method and system based on detail and boundary driving
CN116912837A
Two-stage scene text erasing method based on text segmentation
CN116012835A
Image segmentation method and system based on multi-modal dialogue language model
CN117036706A