Anaphora image segmentation method and device based on bimodal interaction, equipment and medium

Through the reference image segmentation method based on bimodal interaction, the image and text features are extracted and information alignment and interactive guidance are carried out, the problem of insufficient fusion of text semantic clues in the prior art is solved, and the segmentation effect of the reference target is improved.

CN120298685APending Publication Date: 2025-07-11GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298791.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, only visual prediction is supervised during neural network training, resulting in the semantic cues described in text cannot be fully integrated with visual features, affecting the segmentation effect of the referential target.

Method used

Using a reference image segmentation method based on dual-modal interaction, features are extracted through the image and text encoding module, information alignment and interaction guidance is performed using the image and text interaction guidance module, and interaction characteristics are decoded by the image and text decoding module to locate the reference targets contained in the image.

Benefits of technology

The segmentation effect of referential targets in the image is improved, and the positioning accuracy of referential targets is enhanced by fully fusion of text semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298685A_ABST
    Figure CN120298685A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image segmentation, and provides an anaphora image segmentation method based on bimodal interaction, and the method comprises the steps: obtaining a to-be-segmented image and a text description corresponding to the to-be-segmented image, inputting the to-be-segmented image and the text description into a trained anaphora image segmentation model, and carrying out the segmentation of an anaphora image. Image features and text features are extracted through the image and text coding module, the image and text interaction guiding module carries out alignment and interaction guiding processing on image information and text semantic information, and interaction features between images and texts are extracted, so that the text semantic information described by the texts is fully fused; and decoding the interaction characteristics through an image and text decoding module so as to position the reference target contained in the to-be-segmented image, thereby obtaining a reference target segmentation result of the to-be-segmented image. Due to the fact that text semantic information of text description can be fully fused according to image features through the image and text interaction guiding module, the reference target segmentation effect of the to-be-segmented image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and particularly to a referential image segmentation method, device, equipment and medium based on bimodal interaction. Background Art

[0002] The referential image segmentation technology is a technology that combines computer vision and natural language processing, aiming to accurately segment target objects according to natural language expressions. Currently, the referential image segmentation technology is applied in multiple fields. Exemplarily, in the field of intelligent security, security personnel use natural language instructions to let the system accurately locate and segment specific target persons or objects to assist in monitoring and analysis; in the intelligent shopping scenario, consumers use natural language to describe the characteristics of goods, and the system can quickly segment the corresponding goods from the goods image library to improve the shopping search efficiency.

[0003] In the prior art, the method for realizing referential image segmentation involves using a neural network to extract image features and analyzing text descriptions through natural language processing technology to obtain text features. Subsequently, these image features and text features are fused to achieve accurate image segmentation of the referential target.

[0004] However, when using the prior art, during the training of the neural network, usually only visual predictions are supervised, which results in the semantic clues of the text description not being fully fused with the visual features, making the visual features have an incomplete understanding of the referential target contained in the segmented image, that is, the text semantic information cannot be well considered, thus affecting the segmentation effect of the referential target. Summary of the Invention

[0005] Based on this, it is necessary to provide a referential image segmentation method based on bimodal interaction for the above technical problems.

[0006] An embodiment of the present invention provides a referential image segmentation method based on bimodal interaction, and the method includes:

[0007] Obtain the image to be segmented and the text description corresponding to the image to be segmented;

[0008] Input the image to be segmented and the text description into a trained referential image segmentation model to obtain the segmentation result of the referential target of the image to be segmented;

[0009] Among them, the referential image segmentation model includes: an image and text encoding module, an image and text interaction guiding module, and an image and text decoding module; the image and text encoding module is used to extract image features and text features, the image and text interaction guiding module is used to perform alignment and interaction guiding processing on image information and text semantic information, and extract interaction features between the image and the text, and the image and text decoding module is used to decode the interaction features to locate the referential target included in the image to be segmented.

[0010] In one embodiment, before inputting the image to be segmented and the text description into the trained referential image segmentation model to obtain the segmentation result of the referential target of the image to be segmented, it further includes:

[0011] Obtain a sample training set, where the sample training set includes: multiple groups of training samples, and each group of training samples includes a training picture, a text description, and the segmentation position parameters of the referential target included in the training picture;

[0012] Input the sample training set into the initial referential image segmentation model, and adjust the weight parameters of the model according to a preset loss function until the model converges, so as to obtain the trained referential image segmentation model.

[0013] In one embodiment, inputting the image to be segmented and the text description into the trained referential image segmentation model to obtain the segmentation result of the referential target of the image to be segmented includes:

[0014] Input the image to be segmented and the text description into the image and text encoding module for feature extraction processing, and extract the image features and the text features;

[0015] Input the image features and the text features into the image and text interaction guiding module for alignment and interaction guiding processing of image information and text semantic information, and extract the interaction features between the image and the text;

[0016] Input the interaction features into the image and text decoding module for decoding processing to determine the position of the referential target included in the image to be segmented, so as to obtain the segmentation result of the referential target of the image to be segmented.

[0017] In one embodiment, the image and text encoding module includes: an image encoding module and a text encoding module. Inputting the image to be segmented and the text description into the image and text encoding module for feature extraction processing, and extracting the image features and the text features includes:

[0018] Input the image to be segmented into the image encoding module for feature extraction processing to extract the image features, where the image features include: multiple image sub-features of different scales;

[0019] Input the text description into the text encoding module for feature extraction processing to extract the text features.

[0020] In one embodiment, the image and text interaction guidance module includes: an image-text interaction guidance sub-module and a text-image interaction guidance sub-module. The interaction features include: text-image interaction features and image-text interaction features. Input the image features and the text features into the image and text interaction guidance module for alignment and interaction guidance processing of image information and text semantic information to extract the interaction features between the image and the text, including:

[0021] Perform splicing processing on multiple image sub-features of different scales to obtain the spliced image features;

[0022] Input the spliced image features and the text features into the text-image interaction guidance sub-module for alignment and interaction guidance processing to obtain text-image interaction features;

[0023] Input the spliced image features and the text-image interaction features into the image-text interaction guidance sub-module for alignment and interaction guidance processing to obtain image-text interaction features.

[0024] In one embodiment, the step of inputting the spliced image features and the text features into the text-image interaction guidance sub-module for alignment and interaction guidance processing to obtain text-image interaction features includes:

[0025] For the text features, among the multiple first reference targets included in the image features, determine the first reference targets aligned with each second reference target included in the text features, calculate the first correlation degree between the first reference target and the second reference target, and guide the image information to the text features according to the first correlation degree to obtain the text-image interaction features;

[0026] The step of inputting the spliced image features and the text-image interaction features into the image-text interaction guidance sub-module for alignment and interaction guidance processing to obtain image-text interaction features includes:

[0027] For the image features, among the multiple second reference targets included in the text features, determine the second reference targets that are aligned with each of the first reference targets included in the image features, calculate the second correlation degree between the first reference targets and the second reference targets, and guide the text semantic information to the image features according to the second correlation degree to obtain an initial image-text interaction feature, where the initial image-text interaction feature is composed of splicing multiple initial image-text interaction sub-features of different scales;

[0028] Perform a normalized size process on the multiple initial image-text interaction sub-features of different scales to obtain multiple initial image-text interaction sub-features of the same size, and perform a convolution operation on the multiple initial image-text interaction sub-features of the same size to obtain the image-text interaction feature.

[0029] In one embodiment, the reference image segmentation model further includes: a text feature masking module and a context semantic text reconstruction module; wherein, the text feature masking module is used to perform a masking process on the text features to obtain masked text features, and the context semantic text reconstruction module is used to reconstruct the masked text features to update the text features.

[0030] In a second aspect, an embodiment of the present invention provides a reference image segmentation device based on bimodal interaction, including:

[0031] An acquisition module, configured to acquire an image to be segmented and a text description corresponding to the image to be segmented;

[0032] A segmentation module, configured to input the image to be segmented and the text description into a trained reference image segmentation model to obtain a reference target segmentation result of the image to be segmented;

[0033] Wherein, the reference image segmentation model includes: an image and text encoding module, an image and text interaction guiding module, and an image and text decoding module; the image and text encoding module is used to extract image features and text features, the image and text interaction guiding module is used to perform alignment and interaction guiding processing on image information and text semantic information, and extract interaction features between the image and the text, and the image and text decoding module is used to decode the interaction features to locate the reference targets included in the image to be segmented.

[0034] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the reference image segmentation method based on bimodal interaction described in the first aspect are implemented.

[0035] Fourth aspect, a computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method for anaphoric image segmentation based on bimodal interaction described in the first aspect are implemented.

[0036] The technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0037] For an anaphoric image segmentation method based on bimodal interaction provided by an embodiment of the present invention, in this way, by obtaining an image to be segmented and a text description corresponding to the image to be segmented, and inputting the image to be segmented and the text description into a trained anaphoric image segmentation model, the anaphoric image segmentation model includes: an image and text encoding module, an image and text interaction guiding module, and an image and text decoding module. In this way, image features and text features can be extracted through the image and text encoding module, and alignment and interaction guiding processing of image information and text semantic information can be performed through the image and text interaction guiding module to extract interaction features between the image and the text, so as to fully integrate the text semantic information of the text description. Further, the interaction features are decoded through the image and text decoding module to locate the anaphoric target included in the image to be segmented, so as to obtain the segmentation result of the anaphoric target of the image to be segmented. Since the image and text interaction guiding module can fully integrate the text semantic information of the text description for the image features, the segmentation effect of the anaphoric target of the image to be segmented is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a schematic flowchart of an anaphoric image segmentation method based on bimodal interaction provided by an embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of an image to be segmented provided by an embodiment of the present invention;

[0042] Figure 3 It is a schematic structural diagram of an anaphoric image segmentation model provided by an embodiment of the present invention;

[0043] Figure 4Another structural schematic diagram of the referential image segmentation model provided by the embodiment of the present invention;

[0044] Figure 5 A structural schematic diagram of a referential image segmentation device based on bimodal interaction provided by the embodiment of the present invention. Detailed implementation manners

[0045] In order to more clearly understand the above objects, features and advantages of the present invention, the solution of the present invention will be further described below. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0046] In the following description, many specific details are set forth to fully understand the present invention, but the present invention may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.

[0047] The referential image segmentation technology is a technology that combines computer vision and natural language processing technologies, aiming to accurately segment the target object according to the natural language expression. At present, the referential image segmentation technology is applied in multiple fields. Exemplarily, in the field of intelligent security, security personnel can use natural language instructions to let the system accurately locate and segment specific target persons or objects to assist in monitoring and analysis; in the intelligent shopping scenario, consumers use natural language to describe the features of goods, and the system can quickly segment the corresponding goods from the goods image library to improve the shopping search efficiency.

[0048] In the prior art, the method for realizing referential image segmentation involves using a neural network to extract image features and analyzing the text description through natural language processing technology to obtain text features. Subsequently, these image features and text features are fused to achieve accurate image segmentation of the referential target.

[0049] However, when using the prior art, during the training of the neural network, usually only visual prediction is supervised, which results in the semantic clues of the text description not being fully fused with the visual features, making the visual features have an incomplete understanding of the referential target contained in the segmented image, that is, the text semantic information cannot be well considered, thus affecting the segmentation effect of the referential target.

[0050] Therefore, the present invention provides a method for referential image segmentation based on bimodal interaction. By obtaining an image to be segmented and a text description corresponding to the image to be segmented, and inputting the image to be segmented and the text description into a trained referential image segmentation model, the referential image segmentation model includes: an image and text encoding module, an image and text interaction guiding module, and an image and text decoding module. In this way, image features and text features can be extracted through the image and text encoding module, and alignment and interaction guiding processing of image information and text semantic information can be carried out through the image and text interaction guiding module to extract interaction features between the image and the text, so as to fully integrate the text semantic information of the text description. Further, the interaction features are decoded through the image and text decoding module to locate the referential object included in the image to be segmented, thereby obtaining the referential object segmentation result of the image to be segmented. Since the image and text interaction guiding module can fully integrate the text semantic information of the text description for the image features, the segmentation effect of the referential object of the image to be segmented is improved.

[0051] In one embodiment, as Figure 1 shown, Figure 1 FIG. is a schematic flowchart of a method for referential image segmentation based on bimodal interaction provided by an embodiment of the present invention, which specifically includes the following steps:

[0052] S10: Obtain an image to be segmented and a text description corresponding to the image to be segmented.

[0053] Among them, the image to be segmented refers to the target image to be segmented for the referential object, and the text description refers to the text for describing the image to be segmented. Through the text description, each referential object to be segmented included in the image to be segmented can be more determined. Exemplarily, as Figure 2 shown, for the image to be segmented, the corresponding text description may be, for example: "The piece of cake near the fork tip", but not limited thereto. The present invention does not specifically limit, and those skilled in the art can set according to the actual situation.

[0054] Specifically, obtain the image to be segmented for the referential object segmentation and the text description corresponding to the image to be segmented.

[0055] Exemplarily, the image to be segmented can be obtained by camera shooting, and further described to obtain the corresponding text description, or called from a historical dataset, but not limited thereto. The present invention does not specifically limit, and those skilled in the art can set according to the actual situation.

[0056] S11: Input the image to be segmented and the text description into a trained referential image segmentation model to obtain the referential object segmentation result of the image to be segmented.

[0057] Figure 3The following is a schematic structural diagram of a referring image segmentation model provided by an embodiment of the present invention. Refer to Figure 3 As shown, the referring image segmentation model includes: an image and text encoding module 31, an image and text interaction guiding module 32, and an image and text decoding module 33.

[0058] Among them, the image and text encoding module 31 is used to extract image features and text features. The image features are used to represent the image information of the image to be segmented, and the text features are used to represent the context text semantic information of the text description for the image to be segmented.

[0059] The image and text interaction guiding module 32 is used to perform alignment and interaction guiding processing on the image information and the text semantic information, so as to extract the interaction features between the image and the text. In this way, the context text semantic information of the text description can be fully considered, which is convenient for improving the segmentation effect of the referring target in the image to be segmented.

[0060] The image and text decoding module 33 is used to decode the interaction features to locate the referring targets included in the image to be segmented. Specifically, the coordinate position parameters of the referring targets included in the image to be segmented can be obtained, so as to determine the corresponding segmentation positions of each referring target in the image to be segmented.

[0061] Specifically, the obtained image to be segmented and the text description corresponding to the image to be segmented are input into the trained referring image segmentation model, and the referring target segmentation result corresponding to the image to be segmented is output through the referring image segmentation model.

[0062] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, before executing S11, it further includes:

[0063] S01: Obtain a sample training set.

[0064] Among them, the sample training set includes: multiple groups of training samples. Each group of training samples includes a training picture, a text description, and the segmentation position parameters of the referring targets included in the training picture. Exemplarily, the sample training set is, for example, the training pictures, text descriptions, and mask annotation data (i.e., the segmentation position parameters of the referring targets) included in the public datasets: RefCOCO, RefCOCO+ and RefCOCOg. However, it is not limited thereto. The present invention does not specifically limit, and those skilled in the art can set according to the actual situation.

[0065] S02: Input the sample training set into the initial referring image segmentation model, and adjust the weight parameters of the model according to a preset loss function until the model converges, so as to obtain a trained referring image segmentation model.

[0066] Among them, the preset loss function can adopt existing loss functions such as the cross-entropy loss function or the mean square error loss function, etc., but not limited to this. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.

[0067] Specifically, a sample training set is obtained. The sample training set includes: multiple groups of training samples, and each group of training samples includes a training picture, a text description, and segmentation position parameters of the reference object included in the training picture. For the initial reference image segmentation model, the initial reference image segmentation model is trained using the sample training set. The sample training set is input into the initial reference image segmentation model. During the training process, the weight parameters of the model are adjusted according to the preset loss function until the model converges, so as to obtain the trained reference image segmentation model.

[0068] Furthermore, during the training process of the initial reference image segmentation model, the weight parameters of the model can also be adjusted according to the preset loss function, and the training is ended when the number of iterative training times is reached, so as to obtain the trained reference image segmentation model.

[0069] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, one implementation manner of S11 can be:

[0070] S111: Input the image to be segmented and the text description into the image and text encoding module for feature extraction processing, and extract image features and text features.

[0071] Specifically, after obtaining the image to be segmented and the text description corresponding to the image to be segmented, the obtained image to be segmented and the text description corresponding to the image to be segmented are input into the image and text encoding module. The image to be segmented and the text description are subjected to feature extraction processing through the image and text encoding module, and the image features corresponding to the image to be segmented and the text features corresponding to the text description are extracted.

[0072] S112: Input the image features and the text features into the image and text interaction guidance module for alignment and interaction guidance processing of the image information and the text semantic information, and extract the interaction features between the image and the text.

[0073] Among them, alignment refers to matching each reference object corresponding to the text features and each reference object corresponding to the image features.

[0074] Interaction guidance means that the image features can be integrated into the text semantic information, and the text features can be integrated into the image information, so as to improve the positioning of each reference object included in the image to be segmented, thereby improving the segmentation effect of the reference object.

[0075] Specifically, after obtaining the image features and text features, the obtained image features and text features are input into the image and text interaction guidance module. The image and text interaction guidance module processes the image features and text features to achieve the alignment and interaction guidance of the image information and text semantic information, thereby obtaining the interaction features between the image and the text.

[0076] S113: Input the interaction features into the image and text decoding module for decoding processing to determine the position of the referential target included in the image to be segmented, so as to obtain the segmentation result of the referential target of the image to be segmented.

[0077] Specifically, after obtaining the interaction features, the obtained interaction features are input into the image and text decoding module. The image and text decoding module processes the interaction features to determine the position of the referential target included in the image to be segmented. According to the position of the referential target, the image to be segmented is classified through convolution operations, and then the segmentation result of the referential target of the image to be segmented is obtained.

[0078] In this way, the referential image segmentation method based on bimodal interaction provided in this embodiment inputs the image to be segmented and the corresponding text description into the trained referential image segmentation model by obtaining the image to be segmented and the text description corresponding to the image to be segmented. The referential image segmentation model includes: an image and text encoding module, an image and text interaction guidance module, and an image and text decoding module. In this way, the image features and text features can be extracted through the image and text encoding module, and the alignment and interaction guidance processing of the image information and text semantic information can be performed through the image and text interaction guidance module to extract the interaction features between the image and the text, thereby fully integrating the text semantic information of the text description. Further, the interaction features are decoded through the image and text decoding module to locate the referential target included in the image to be segmented, so as to obtain the segmentation result of the referential target of the image to be segmented. Since the image and text interaction guidance module can fully integrate the text semantic information of the text description for the image features, the segmentation effect of the referential target of the image to be segmented is improved.

[0079] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, Figure 4 is a schematic structural diagram of another referential image segmentation model provided by an embodiment of the present invention. Refer to Figure 4 shown. The image and text encoding module includes: an image encoding module 311 and a text encoding module 312. Based on this, one implementation manner of S111 can be:

[0080] S1111: Input the image to be segmented into the image encoding module for feature extraction processing to extract image features.

[0081] Among them, the image features include: multiple image sub-features of different scales. Each image sub-feature at a scale can extract different image information. By obtaining multiple image sub-features of different scales, image information can be fully obtained in this way.

[0082] Specifically, after obtaining the image to be segmented and the corresponding text description of the image to be segmented, the obtained image to be segmented is input into the image encoding module, and the image encoding module performs feature extraction processing on the image to be segmented to extract the image features corresponding to the image to be segmented.

[0083] Optionally, the image encoding module 311 can adopt the Swin Transformer network model. Swintransformer is a Transformer model architecture applied in the field of computer vision and is widely used in tasks such as image classification, object detection, and semantic segmentation. For the input image to be segmented, the image encoding module can extract multiple visual features, namely image features: Among them, (H i , W i ) = (H / 2 i+1 , W / 2 i+1 ) represents the spatial resolution of the i-th visual feature, and C i represents its channel dimension. In this way, the image encoding module can perform feature extraction on the image to be segmented at multiple different scales, extract the image features corresponding to the image to be segmented, and thus obtain rich image information.

[0084] S1112: Input the text description into the text encoding module for feature extraction processing to extract text features.

[0085] Specifically, after obtaining the image to be segmented and the corresponding text description of the image to be segmented, the obtained text description is input into the text encoding module, and the text encoding module performs feature extraction processing on the text description to extract the text features corresponding to the text description.

[0086] Optionally, the text encoding module 312 can be a model with BERT as the backbone network. Specifically, by performing word segmentation and padding processing on the text description, text embeddings are obtained Among them, T is the maximum length after word segmentation, and D l represents the dimension of the text features. And a label such as [CLS] is added at the beginning of the text as the global representation of the sentence. Further, the text embeddings are input into BERT to obtain text features, thereby obtaining the context semantic information of the text description for the reference target.

[0087] In this way, in this embodiment, through the image encoding module and the text encoding module, the image features corresponding to the image to be segmented and the text features corresponding to the text description are respectively obtained, which facilitates subsequent interaction guidance for the image features and text features, thereby more fully considering the text semantic information and improving the segmentation effect of the reference target.

[0088] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, continue to refer to Figure 4 As shown, the image-text interaction guidance module 32 includes: an image-text interaction guidance sub-module 321 and a text-image interaction guidance sub-module 322. The interaction features include: text-image interaction features and image-text interaction features. Based on this, one implementation manner of S112 may be:

[0089] S1121: Perform splicing processing on multiple image sub-features of different scales to obtain the spliced image features.

[0090] Specifically, since the image features include multiple image sub-features of different scales, based on this, perform splicing processing on multiple image sub-features of different scales to obtain the spliced image features.

[0091] S1122: Input the spliced image features and text features into the text-image interaction guidance sub-module for alignment and interaction guidance processing to obtain text-image interaction features.

[0092] Among them, the text-image interaction features are text features that further fuse image information for the text features.

[0093] Specifically, after performing splicing processing on multiple image sub-features of different scales, input the spliced image features and text features into the text-image interaction guidance sub-module, and through the text-image interaction guidance sub-module, perform alignment and interaction guidance processing on the image features and text features to obtain text-image interaction features.

[0094] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, one implementation manner of S1122 may be:

[0095] For the text features, among the multiple first reference targets included in the image features, determine the first reference targets that are aligned with each of the second reference targets included in the text features, calculate the first correlation degree between the first reference targets and the second reference targets, and guide the image information to the text features according to the first correlation degree to obtain text-image interaction features.

[0096] Among them, the first correlation degree is used to express the similarity between each first reference target included in the image feature and each second reference target included in the text feature. For the text feature, through this first correlation degree, the image information can be guided to the text feature, so that the text feature can fuse the image information. In this way, both the image information and the text semantic information can be fully considered at the same time, and the segmentation effect of the image to be segmented can be improved.

[0097] Specifically, for each second reference target included in the text feature, match and align it among the multiple first reference targets included in the image feature, determine the first reference target aligned with each second reference target, further calculate the first correlation degree between the first reference target and the second reference target, and guide the image information to the text feature according to the first correlation degree to obtain the text-image interaction feature.

[0098] S1123: Input the spliced image feature and the text-image interaction feature into the image-text interaction guiding sub-module for alignment and interaction guiding processing to obtain the text-image interaction feature.

[0099] Among them, the text-image interaction feature is a text-image feature that further fuses the text semantic information for the image feature.

[0100] Specifically, after performing splicing processing on multiple image sub-features of different scales, input the spliced image feature and the text-image interaction feature into the image-text interaction guiding sub-module, and the image-text interaction guiding sub-module performs alignment and interaction guiding processing on the image feature and the text-image interaction feature to obtain the text-image interaction feature.

[0101] Optionally, on the basis of the above embodiments, in some embodiments of the present invention, one implementation manner of S1123 may be:

[0102] For the image feature, among the multiple second reference targets included in the text feature, determine the second reference target aligned with each first reference target included in the image feature, calculate the second correlation degree between the first reference target and the second reference target, and guide the text semantic information to the image feature according to the second correlation degree to obtain the initial text-image interaction feature.

[0103] Among them, the initial text-image interaction feature is composed of splicing multiple initial text-image interaction sub-features of different scales.

[0104] Specifically, for each first reference target included in the image feature, match and align it among the multiple second reference targets included in the text feature, determine the second reference target aligned with each first reference target, further calculate the second correlation degree between the first reference target and the second reference target, and guide the text semantic information to the image feature according to the second correlation degree to obtain the initial text-image interaction feature.

[0105] Normalize the sizes of multiple initial image-text interaction sub-features with different scales to obtain multiple initial image-text interaction sub-features of the same size, and perform a convolution operation on the multiple initial image-text interaction sub-features of the same size to obtain image-text interaction features.

[0106] Specifically, since the initial image-text interaction feature is composed of the splicing of multiple initial image-text interaction sub-features with different scales, normalize the sizes of the multiple initial image-text interaction sub-features of the initial image-text interaction feature to obtain multiple initial image-text interaction sub-features of the same size, and perform a convolution operation on the multiple initial image-text interaction sub-features of the same size through a convolution layer to obtain an image-text interaction feature.

[0107] In this way, this embodiment can fully fuse the different image information corresponding to the multi-size image sub-features and the text semantic information corresponding to the text features respectively, thereby improving the segmentation effect of the image to be segmented.

[0108] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, referring to Figure 3 as shown, the image segmentation model further includes: a text feature masking module 34 and a context semantic text reconstruction module 35. Among them, the text feature masking module is used to perform masking processing on the text feature to obtain a masked text feature, and the context semantic text reconstruction module is used to reconstruct the masked text feature to update the text feature.

[0109] Specifically, after extracting the text feature through the text encoding module, input the text feature into the text feature masking module, and the text feature masking module performs masking processing on the text feature to obtain a masked text feature. Input the masked text feature into the context semantic text reconstruction module, and use the text-image interaction sub-feature to perform text feature reconstruction on the masked text feature, and further return the reconstructed text feature to update the corresponding text feature.

[0110] In this way, this embodiment can use the text-image interaction sub-feature that fuses image information to re-update the text feature, which is convenient for more fully obtaining image information and text semantic information subsequently, and improving the segmentation effect of the image to be segmented.

[0111] It should be understood that although Figures 1 to 4 the steps in the flowchart of Figures 1 to 4At least a part of the steps may include multiple sub - steps or multiple stages. These sub - steps or stages do not necessarily need to be completed at the same time and can be executed at different times. The execution order of these sub - steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub - steps or stages of other steps.

[0112] In one embodiment, as Figure 5 shown, a referential image segmentation device based on bimodal interaction is provided, including: an acquisition module 10 and a segmentation module 11

[0113] Among them, the acquisition module 10 is used to acquire the image to be segmented and the text description corresponding to the image to be segmented.

[0114] The segmentation module 11 is used to input the image to be segmented and the text description into a trained referential image segmentation model to obtain the referential target segmentation result of the image to be segmented. Among them, the referential image segmentation model includes: an image - text encoding module, an image - text interaction guiding module, and an image - text decoding module; the image - text encoding module is used to extract image features and text features, the image - text interaction guiding module is used to perform alignment and interaction guiding processing on image information and text semantic information, and extract the interaction features between the image and the text, and the image - text decoding module is used to decode the interaction features to locate the referential target included in the image to be segmented.

[0115] In the above - mentioned embodiment, the acquisition module is used to acquire the image to be segmented and the text description corresponding to the image to be segmented. The segmentation module inputs the image to be segmented and the text description into a trained referential image segmentation model to obtain the referential target segmentation result of the image to be segmented. The referential image segmentation model includes: an image - text encoding module, an image - text interaction guiding module, and an image - text decoding module. In this way, the image - text encoding module can extract image features and text features, the image - text interaction guiding module can perform alignment and interaction guiding processing on image information and text semantic information, and extract the interaction features between the image and the text, so as to fully integrate the text semantic information of the text description. Further, the image - text decoding module decodes the interaction features to locate the referential target included in the image to be segmented, so as to obtain the referential target segmentation result of the image to be segmented. Since the image - text interaction guiding module can fully integrate the text semantic information of the text description for the image features, the segmentation effect of the referential target of the image to be segmented is improved.

[0116] For the specific limitations of the referential image segmentation device based on bimodal interaction, reference may be made to the limitations of the referential image segmentation method based on bimodal interaction in the foregoing text, which will not be elaborated herein. Each module in the above server can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0117] An embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the referential image segmentation method provided by the embodiment of the present invention can be implemented. For example, when the processor executes the computer program, it can implement Figures 1 to 4 the technical solutions of any of the method embodiments shown. The implementation principles and technical effects are similar, and will not be elaborated herein.

[0118] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is implemented by the processor, the referential image segmentation method provided by the embodiment of the present invention can be implemented. For example, when the computer program is executed by the processor, it can implement Figures 1 to 4 the technical solutions of any of the method embodiments shown. The implementation principles and technical effects are similar, and will not be elaborated herein.

[0119] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the method embodiments as described above. Among them, any reference to the memory, database, or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM).

[0120] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0121] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A reference image segmentation method based on bimodal interaction, characterized in that, Including: Obtain the image to be segmented and the corresponding text description of the image to be segmented; Input the image to be segmented and the text description into a trained anaphoric image segmentation model to obtain the anaphoric target segmentation result of the image to be segmented; Wherein, the anaphoric image segmentation model includes: an image and text encoding module, an image and text interaction guidance module, and an image and text decoding module; the image and text encoding module is used to extract image features and text features, and the image and text interaction guidance module is used to perform alignment and interaction guidance processing on image information and text semantic information, and extract interaction features between the image and the text, and the image and text decoding module is used to decode the interaction features to locate the anaphoric target included in the image to be segmented.

2. The method according to claim 1, wherein Before inputting the image to be segmented and the text description into a trained anaphoric image segmentation model to obtain the anaphoric target segmentation result of the image to be segmented, it further includes: Obtain a sample training set, wherein the sample training set includes: multiple groups of training samples, and each group of training samples includes a training picture, a text description, and the segmentation position parameters of the anaphoric target included in the training picture; Input the sample training set into an initial anaphoric image segmentation model, and adjust the weight parameters of the model according to a preset loss function until the model converges to obtain the trained anaphoric image segmentation model.

3. The method according to claim 1, wherein Inputting the image to be segmented and the text description into a trained anaphoric image segmentation model to obtain the anaphoric target segmentation result of the image to be segmented includes: Input the image to be segmented and the text description into the image and text encoding module for feature extraction processing to extract the image features and the text features; Input the image features and the text features into the image and text interaction guidance module for alignment and interaction guidance processing of image information and text semantic information to extract interaction features between the image and the text; Input the interaction features into the image and text decoding module for decoding processing to determine the position of the anaphoric target included in the image to be segmented, so as to obtain the anaphoric target segmentation result of the image to be segmented.

4. The method according to claim 3, characterized in that, The image and text encoding module includes: an image encoding module and a text encoding module. Inputting the image to be segmented and the text description into the image and text encoding module for feature extraction processing to extract the image features and the text features includes: Input the image to be segmented into the image encoding module for feature extraction processing to extract the image features, wherein the image features include: multiple image sub-features of different scales; Input the text description into the text encoding module for feature extraction processing to extract the text features.

5. The method according to claim 4, wherein The image-text interaction guidance module includes: an image-text interaction guidance sub-module and a text-image interaction guidance sub-module. The interaction features include: text-image interaction features and image-text interaction features. Inputting the image features and the text features into the image-text interaction guidance module for alignment and interaction guidance processing of image information and text semantic information, and extracting the interaction features between the image and the text includes: Performing splicing processing on multiple image sub-features of different scales to obtain the spliced image features; Inputting the spliced image features and the text features into the text-image interaction guidance sub-module for alignment and interaction guidance processing to obtain text-image interaction features; Inputting the spliced image features and the text-image interaction features into the image-text interaction guidance sub-module for alignment and interaction guidance processing to obtain image-text interaction features.

6. The method according to claim 5, wherein The step of inputting the spliced image features and the text features into the text-image interaction guidance sub-module for alignment and interaction guidance processing to obtain text-image interaction features includes: For the text features, among the multiple first referential targets included in the image features, determining the first referential targets aligned with each second referential target included in the text features, calculating the first correlation degree between the first referential targets and the second referential targets, and guiding the image information to the text features according to the first correlation degree to obtain the text-image interaction features; The step of inputting the spliced image features and the text-image interaction features into the image-text interaction guidance sub-module for alignment and interaction guidance processing to obtain image-text interaction features includes: For the image features, among the multiple second referential targets included in the text features, determining the second referential targets aligned with each first referential target included in the image features, calculating the second correlation degree between the first referential targets and the second referential targets, and guiding the text semantic information to the image features according to the second correlation degree to obtain initial image-text interaction features, where the initial image-text interaction features are composed of splicing multiple initial image-text interaction sub-features of different scales; Performing normalized size processing on the multiple initial image-text interaction sub-features of different scales to obtain multiple initial image-text interaction sub-features of the same size, and performing convolution operations on the multiple initial image-text interaction sub-features of the same size to obtain the image-text interaction features.

7. The method according to claim 1, wherein The referential image segmentation model further includes: a text feature masking module and a context semantic text reconstruction module; wherein, the text feature masking module is used to perform masking processing on the text features to obtain masked text features, and the context semantic text reconstruction module is used to reconstruct the masked text features to update the text features.

8. An anaphoric image segmentation device based on bimodal interaction, characterized in that It includes: An acquisition module, used to acquire the image to be segmented and the text description corresponding to the image to be segmented; A segmentation module, used to input the image to be segmented and the text description into the trained referential image segmentation model to obtain the referential target segmentation result of the image to be segmented. Among them, the referential image segmentation model includes: an image and text encoding module, an image and text interaction guiding module, and an image and text decoding module; the image and text encoding module is used to extract image features and text features, the image and text interaction guiding module is used to perform alignment and interaction guiding processing on image information and text semantic information, and extract interaction features between the image and the text, and the image and text decoding module is used to decode the interaction features to locate the referential target included in the image to be segmented.

9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the referential image segmentation method based on bimodal interaction according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the referential image segmentation method based on bimodal interaction according to any one of claims 1 to 7 are implemented.