Method and device for constructing a reference expression positioning and segmentation model
By using modal intrinsic relationship perception network and cross-modal fusion network in the reference positioning and segmentation model, and combining text and image features for intrinsic relationship learning and cross-modal fusion, the problem of poor learning and fusion performance in the prior art is solved, and more efficient information utilization and feature selection is achieved, which significantly improves the accuracy of positioning and segmentation.
Patent Information
- Application Number
- CN202111136455.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-09-27
AI Technical Summary
In the positioning and segmentation method, the feature learning and fusion of the two-modal data of image and text have poor performance, insufficient information utilization, insufficient distinction between prospect and background, and information redundancy, resulting in a degradation of positioning and segmentation performance.
By constructing a reference positioning and segmentation model, a modal intrinsic relationship perception network and a cross-modal fusion network are used, and intrinsic relationship learning and cross-modal fusion network are used to combine text and image features for intrinsic relationship learning and cross-modal fusion network, filtering and selection are used for image-text collaborative features, and feature selection and transformation are used for multi-scale fusion network to improve the accuracy of positioning and segmentation.
By making full use of image and text information, we can improve the effect of feature learning and fusion, enhance the distinction between foreground and background, reduce information redundancy, and significantly improve the performance and accuracy of positioning and segmentation of the indicators.
Smart Images

Figure CN114048284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image and text understanding, and in particular to a method and device for constructing a referential expression localization and segmentation model. Background Art
[0002] The positioning and segmentation method of referential expression refers to a method of finding the object referred to or related entities in the image according to the content of the text, and locating and segmenting them at the same time, given a descriptive text and an image. This technical content is often seen in the interaction between people and robots. People send relevant instructions through language description (language can be recognized as text), and the robot collects environmental information through the camera to generate an image. The image and text are processed using the pre-set positioning and segmentation method of referential expression, and the referent in the interactive scene is found and located and segmented. The referent is grasped and other related operations are performed, thereby realizing services for people. With the help of this technology, the interaction between robots and people and objects can be carried out more flexibly and intelligently, improving the intelligence of robots and the reliability and accuracy of services for people.
[0003] At present, in the localization and segmentation methods of referential expressions, the feature learning processes of the two modal data, image and text, are often carried out in isolation, and the obtained image features and text features are fused between the modal data using a feature fusion network.
[0004] The feature learning and fusion process in the above methods has the following defects: (1) Since feature learning is performed in isolation between modalities, the feature learning performance is poor. (2) The two modal features are only fused in the modal fusion process, resulting in insufficient utilization of the information of the two modalities, which in turn limits the actual performance of positioning and segmentation. (3) Current modal fusion is often based on the assumption that text information is uniformly distributed in the image feature position space. To a certain extent, it weakens the distinction between the foreground (the object referred to) and the background when the image and text are fused for positioning and segmentation of referential expressions, further limiting the performance of positioning and segmentation. (4) When fusing multi-scale features, information of all scales is often fused indiscriminately, resulting in information redundancy. In addition, the redundant information may have the opposite inhibitory effect on the actual performance of the method, resulting in a decrease in the performance of positioning and segmentation. Summary of the invention
[0005] In view of the problems existing in the prior art, the embodiments of the present invention provide a method and device for constructing a referential expression positioning and segmentation model that overcomes the above problems or at least partially solves the above problems.
[0006] In a first aspect, an embodiment of the present invention provides a method for constructing a reference expression positioning and segmentation model, including:
[0007] Step 1: construct a referential expression location and segmentation database; wherein the database samples include: images with location and segmentation annotations for referents, and texts describing referents;
[0008] Step 2: construct a preprocessed image backbone network and a preprocessed text backbone network; wherein the preprocessed image / text backbone network is used for feature extraction to obtain image / text preprocessed features; the image preprocessing features are feature pyramids composed of image features of different scales;
[0009] Step 3: For each scale image feature, a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visual-guided text perception subnetwork is constructed accordingly; wherein the text-guided visual perception subnetwork / visual-guided text perception subnetwork is used to combine text preprocessing features / corresponding scale image features, learn the corresponding scale image features / text preprocessing features, and obtain image features / text features at the corresponding scale;
[0010] Step 4: constructing each cross-modal fusion network corresponding to each modality intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales;
[0011] Step 5: construct a first multi-scale fusion network and a second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select and transform the image-text collaborative features at each scale to obtain the target features; the target features are used to realize the positioning / segmentation of the referent;
[0012] Step six: Using the database, the network composed of the preprocessed image backbone network, the preprocessed text backbone network, the cross-modal fusion networks, the intrinsic relationship perception network of each modality, the first multi-scale fusion network and the second multi-scale fusion network is optimized and trained to obtain the reference expression positioning and segmentation model.
[0013] According to the method for constructing a referential expression positioning and segmentation model provided by the present invention, the positioning and marking of the referent specifically includes: marking the position of the referent detection frame;
[0014] The segmenting and labeling of the referent specifically includes: labeling the pixel points covered by the referent.
[0015] According to the method for constructing a referential expression localization and segmentation model provided by the present invention, the method combines text preprocessing features to learn image features of corresponding scales to obtain image features at corresponding scales, including:
[0016] Based on the text self-attention mechanism, the text preprocessing features are learned to obtain the first intermediate features;
[0017] Mapping the first intermediate feature to the image space to obtain a second intermediate feature; wherein the number of channels of the second intermediate feature is consistent with the number of channels of the image feature of the corresponding scale;
[0018] The second intermediate features are copied and spliced to obtain a third intermediate feature having the same scale as the feature of the corresponding scale image;
[0019] The third intermediate feature is multiplied element by element with the image feature of the corresponding scale to obtain the text-image fusion feature at the corresponding scale;
[0020] Based on the channel attention mechanism and the spatial attention mechanism, the text-image fusion features at the corresponding scale are learned to obtain the image features at the corresponding scale.
[0021] According to the method for constructing a referential expression localization and segmentation model provided by the present invention, the text preprocessing features are learned by combining the corresponding scale image features to obtain the text features at the corresponding scale, specifically:
[0022] Perform global pooling operation on the image features of corresponding scales to obtain the corresponding global visual features;
[0023] Based on the visual self-attention mechanism, the global visual features are aggregated into one feature;
[0024] The aggregated features are multiplied element by element with the text preprocessing features to obtain the image-text fusion features at the corresponding scale;
[0025] Based on the text self-attention mechanism, the image-text fusion features at the corresponding scale are learned to obtain the text features at the corresponding scale.
[0026] According to the method for constructing a referential expression localization and segmentation model provided by the present invention, the feature similarity is considered, the image features and text features at the corresponding scale are fused to obtain the image-text collaborative features at the corresponding scale, specifically:
[0027] Calculate the similarity between image features and text features at the corresponding scale;
[0028] The similarity is used to fuse the image features and text features at the corresponding scale to obtain the image-text collaborative features at the corresponding scale.
[0029] According to the method for constructing a referential expression localization and segmentation model provided by the present invention, the image-text collaborative features at each scale are sampled, spliced, feature selected and transformed to obtain target features, specifically:
[0030] If the current network is the first multi-scale fusion network, the image-text collaborative features at each scale are downsampled based on the minimum scale, and are spliced in descending order of the scales before sampling; if the current network is the second multi-scale fusion network, the image-text collaborative features at each scale are upsampled based on the maximum scale, and are spliced in ascending order of the scales before sampling;
[0031] The concatenated features are sequentially subjected to feature selection and feature transformation to obtain the target features.
[0032] According to the method for constructing a referential expression localization and segmentation model provided by the present invention, the image-text collaborative feature at the corresponding scale is specifically calculated by the following formula:
[0033]
[0034]
[0035]
[0036] In the above formula, F vl,i represents the image-text collaborative feature at the i-th scale, represents the weighted image features at the i-th scale, represents the weighted text features at the i-th scale, S i represents the similarity between image features and text features at the i-th scale, V c,i represents the image features at the i-th scale, V l,i It means that the text features at the i-th scale are copied and concatenated to obtain text features with the same scale as the image features at the i-th scale.
[0037] According to the method for constructing the reference expression positioning and segmentation model provided by the present invention, the feature selection is performed on the spliced features, and is specifically calculated by the following formula:
[0038]
[0039]
[0040] F pool =MaxPooling(F sum )
[0041] F t =W 2 W 1 F pool
[0042]
[0043]
[0044] In the above formula, F represents the concatenated feature, Ω k (·) is the convolutional layer with the kth receptive field, It represents the features obtained by processing the concatenated features using the convolutional layer of the kth receptive field. represents the element-by-element addition operation, F sum represents the feature obtained by adding the features obtained by processing the concatenated features using the convolutional layers of each receptive field element by element, MaxPooling(·) represents the global pooling operation, and F pool Indicates F sum The feature obtained after the global pooling operation, W 2 represents the parameters of the second fully connected layer, W 1 represents the parameters of the first fully connected layer, F t Represents the global information after conversion, G k (·) represents the weight function of the kth receptive field branch, f k represents the selection weight coefficient of the kth receptive field branch after activation, F s It represents the features obtained by selecting the concatenated features, and K represents the total number of receptive fields.
[0045] According to the construction method of the reference expression positioning and segmentation model provided by the present invention, the preprocessing image backbone network includes: a VGG network and a ResNet network;
[0046] The preprocessing text backbone network includes: a recursive neural network and a long short-term memory network.
[0047] In a second aspect, an embodiment of the present invention further provides a device for constructing a reference expression positioning and segmentation model, including:
[0048] A referential expression positioning and segmentation database construction module is used to construct a referential expression positioning and segmentation database; wherein the database samples include: images with positioning and segmentation annotations for referents, and texts describing referents;
[0049] The preprocessing image backbone network and the preprocessing text backbone network construction module are used to construct the preprocessing image backbone network and the preprocessing text backbone network; wherein the preprocessing image / text backbone network is used for feature extraction to obtain image / text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales;
[0050] A modality intrinsic relationship perception network construction module is used to construct a modality intrinsic relationship perception network including a text-guided visual perception subnetwork and a visual-guided text perception subnetwork for image features of each scale; wherein the text-guided visual perception subnetwork / visual-guided text perception subnetwork is used to learn the image features of the corresponding scale / text preprocessing features in combination with the text preprocessing features / image features of the corresponding scale to obtain image features / text features at the corresponding scale;
[0051] A cross-modal fusion network construction module is used to construct each cross-modal fusion network corresponding to each modality intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales;
[0052] The first multi-scale fusion network and the second multi-scale fusion network construction module are used to construct the first multi-scale fusion network and the second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select features and transform features of image-text collaborative features at each scale to obtain target features; the target features are used to realize the positioning / segmentation of the referent;
[0053] The optimization training module is used to use the database to optimize the network composed of the preprocessed image backbone network, the preprocessed text backbone network, the cross-modal fusion networks, the intrinsic relationship perception networks of each modality, the first multi-scale fusion network and the second multi-scale fusion network to obtain the reference expression positioning and segmentation model.
[0054] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for constructing a positioning and segmentation model as described in the first aspect are implemented.
[0055] In a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for constructing a positioning and segmentation model as described in the first aspect.
[0056] The method and device for constructing a referential expression positioning and segmentation model provided by the present invention are aimed at image pyramid and text preprocessing features. The modal intrinsic relationship perception network uses text / image features as auxiliary information to assist in image / text feature learning, obtains image / text features related to the referent, makes more full use of image / text information, and improves the effect of image / text feature learning; the cross-modal fusion network maps the features of the two modalities of image and text to the same public space to calculate the similarity, uses the similarity to filter the image and text features, and then obtains the image-text collaborative features, which establish the collaboration of the two modalities of image and text in semantic and position space, further distinguishes the referent from the background, and improves the fusion effect of the two modalities; the first multi-scale fusion network and the second multi-scale fusion network fuse the image-text collaborative features under multi-scales with a feature selection strategy, and improves the accuracy of feature selection by selecting feature information related to the referent and suppressing other background information. The above points are combined to achieve the effect of improving the positioning and segmentation accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0058] Figure 1 is a flow chart of a method for constructing a positioning and segmentation model provided by the present invention;
[0059] Figure 2 It is a schematic diagram of the structure of the reference expression positioning and segmentation model provided by the present invention;
[0060] Figure 3 It is a schematic diagram of the structure of the modal intrinsic relationship perception network provided by the present invention;
[0061] Figure 4 It is a schematic diagram of the cross-modal fusion network structure provided by the present invention;
[0062] Figure 5 is a schematic diagram of the first / second multi-scale fusion network structure provided by the present invention;
[0063] Figure 6 It is a structural diagram of a construction device for expressing positioning and segmentation model provided by the present invention;
[0064] Figure 7 It is a structural schematic diagram of an electronic device for implementing a method for constructing a reference expression positioning and segmentation model provided by the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] Combine the following Figure 1-Figure 7 The invention describes a method and device for constructing a reference expression positioning and segmentation model.
[0067] First, as Figure 1 As shown, the present invention provides a method for constructing a reference expression positioning and segmentation model, comprising:
[0068] Step 1: construct a referential expression location and segmentation database; wherein the database samples include: images with location and segmentation annotations for referents, and texts describing referents;
[0069] The present invention can collect samples from image and text understanding data sets (including images and corresponding text description sentences), or can collect samples by itself, that is, using a camera to collect image data or crawling image data from the Internet, and artificially generate description text for a certain object in the image. A large number of samples are collected to generate a database.
[0070] Step 2: construct a preprocessed image backbone network and a preprocessed text backbone network; wherein the preprocessed image / text backbone network is used for feature extraction to obtain image / text preprocessed features; the image preprocessing features are feature pyramids composed of image features of different scales;
[0071] In this technical field, it is clearly known that the preprocessing image backbone network is a deep learning network. When preprocessing the image, the intermediate layer will output image features with different scales, such as 1 / 8, 1 / 16, 1 / 32 of the original image, etc. The output layer will combine these image features in order from small to large and from top to bottom to obtain a feature pyramid.
[0072] It is also clearly known that the preprocessing text backbone network is a natural language feature network. When preprocessing the text, the text is first segmented, and then the segmented words are arranged according to their front and back positions in the text to obtain the word vectors of the text. Finally, the word vectors of the text are time-series processed to obtain the text preprocessing features, that is, the text preprocessing features are actually time series sequences.
[0073] Step 3: For each scale image feature, a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visual-guided text perception subnetwork is constructed accordingly; wherein the text-guided visual perception subnetwork / visual-guided text perception subnetwork is used to combine text preprocessing features / corresponding scale image features, learn the corresponding scale image features / text preprocessing features, and obtain image features / text features at the corresponding scale;
[0074] In the existing technology, the feature learning of images and texts is often carried out independently, and there is no information interaction between the two modalities, resulting in poor feature learning ability. In this regard, the present invention designs a modal intrinsic relationship perception network to learn image and text features. The network includes a text-guided visual perception subnetwork and a visual-guided text perception subnetwork. The former uses the text preprocessing features as auxiliary information to guide the learning of image features at various scales, and learns image features at various scales (visual features related to text information); the latter uses the image preprocessing features as auxiliary information to guide the learning of text preprocessing features, and learns text features at various scales (text features related to visual information); based on this, information interaction between image and text modal information is achieved when learning their respective modal features, which promotes feature learning and makes up for the problem of insufficient utilization of information to a certain extent.
[0075] Step 4: constructing each cross-modal fusion network corresponding to each modality intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales;
[0076] In the prior art, the fusion of image and text modes is achieved only through one fusion network, and the mining and utilization of information are insufficient. In addition, the fusion network often assumes that the text is evenly distributed in the image position space, that is, the influence of the text on the image is independent of the image spatial position, which limits the fusion performance to a certain extent. Based on this, the present invention configures each cross-modal fusion network for image features and text features at each scale to fuse image features and text features at each scale, so as to solve the problem of insufficient mining and utilization of information when the fusion of image and text modes is achieved only through one fusion network. In addition, each cross-modal fusion network maps the image features and text features to the same common space, performs similarity calculation, obtains the text-image similarity matrix, establishes the connection between the two modes in terms of semantics and spatial position, and uses the similarity matrix to filter both the image features and the text features to obtain weighted image features and text features, and then fuses them to obtain the collaborative features of the two modes. The size of the similarity value here reflects the similarity between the text information and the image information. The text and image features are weighted by the similarity value, so that the utilization of text and image information in the position space changes with the similarity value, solving the problem that the utilization of text information is independent of the image position space.
[0077] Step 5: construct a first multi-scale fusion network and a second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select and transform the image-text collaborative features at each scale to obtain the target features; the target features are used to realize the positioning / segmentation of the referent;
[0078] It should be noted that the target feature is used to realize the positioning / segmentation of the referent, specifically, the target feature is used to locate and segment the referent through the positioning network and the segmentation network;
[0079] In addition, the present invention substitutes the features after feature selection into several convolutional layers for feature transformation to increase the depth, improve the transformation capability and model performance, but it should be noted that the number of layers should not be too many.
[0080] Step six: Using the database, the network composed of the preprocessed image backbone network, the preprocessed text backbone network, the cross-modal fusion networks, the intrinsic relationship perception network of each modality, the first multi-scale fusion network and the second multi-scale fusion network is optimized and trained to obtain the reference expression positioning and segmentation model.
[0081] Figure 2The schematic diagram of the structure of the referential expression positioning and segmentation model is illustrated. Those skilled in the art should be able to understand that the image and text are respectively input into the preprocessing image backbone network and the preprocessing text backbone network as the input of the entire network. The modal intrinsic relationship perception network processes the image pyramid and text preprocessing features output by the preprocessing image backbone network and the preprocessing text backbone network to obtain the processed image pyramid and text features; the cross-modal collaborative network performs feature fusion on the image pyramid and text features to obtain image-text collaborative features at various scales; the first multi-scale fusion network / the second multi-scale fusion network constructs an image-text collaborative feature information transmission path; it samples the image-text collaborative features of various scales to the same scale through up / down sampling, and then transmits them through the transmission path, and then splices them to obtain multi-scale fusion features, which are used as the input of the positioning branch (in the first multi-scale fusion network) / segmentation branch (in the second multi-scale fusion network); in the positioning branch, the target features pass through several convolutional layers and output predicted positioning data; in the segmentation branch, the target features pass through several layers of convolutional layers and hollow space pyramid pooling layers to obtain predicted segmentation data. These predicted positioning data and segmentation data serve as the output of the entire network. Together with the positioning data and segmentation data labeled in the dataset, they constitute the calculation process of the loss function to realize the training of the entire network.
[0082] When there is an actual need, the image and the text describing the referent are input into the referential expression localization and segmentation model, and the model will automatically output the localization and / or segmentation results, where the localization and segmentation are selected according to the current needs.
[0083] The present invention targets image pyramid and text preprocessing features. The modal intrinsic relationship perception network uses text / image features as auxiliary information to assist in image / text feature learning, obtains image / text features related to the referent, makes more effective use of image / text information, and improves the effect of image / text feature learning. The cross-modal fusion network maps the features of the two modalities of image and text to the same common space to calculate the similarity, uses the similarity to filter the image and text features, and then obtains image-text collaborative features, which establish the collaboration of the two modalities of image and text in semantic and position space, further distinguish the referent from the background, and improves the fusion effect of the two modalities. The first multi-scale fusion network and the second multi-scale fusion network fuse the image-text collaborative features at multiple scales with a feature selection strategy, and improves the accuracy of feature selection by selecting feature information related to the referent and suppressing other background information. The above points are combined to achieve the effect of improving the positioning and segmentation accuracy of the model.
[0084] Based on the above embodiments, as an optional embodiment, the positioning and marking of the referent is specifically: marking the position of the referent detection frame;
[0085] The segmenting and labeling of the referent specifically includes: labeling the pixel points covered by the referent.
[0086] It should be understood that segmentation and annotation of the referent can also be performed by contour annotation of the referent, and the portion within the default contour is the referent to be segmented; the position of the referent detection box is the coordinate data of the box surrounding the referent;
[0087] In the present invention, the position (coordinates) of the detector frame of the pointer is used as the positioning mark of the pointer, and the pixels / contours covered by the pointer are used as the segmentation marks of the pointer. It can be clearly seen that segmentation is more precise and has stricter requirements than positioning.
[0088] In this embodiment, the selected positioning annotation and segmentation annotation are only a preferred method, which can make the positioning and segmentation of the referent more accurate in terms of experimental effect.
[0089] Based on the above embodiments, as an optional embodiment, Figure 3 The schematic diagram of the structure of the modal intrinsic relationship perception network is illustrated, wherein the text-guided visual perception subnetwork is used to realize the function of learning the image features of the corresponding scale in combination with the text preprocessing features to obtain the image features at the corresponding scale. The learning of the image features of the corresponding scale in combination with the text preprocessing features to obtain the image features at the corresponding scale includes:
[0090] Based on the text self-attention mechanism, the text preprocessing features are learned to obtain the first intermediate features;
[0091] In this embodiment, the image feature of the first scale is The image features of the second scale are The other layers are similar, where w, h, and d represent width, height, and channel, respectively; the jth element in the preprocessed text feature is represented by q j (j∈{1,2,…,L}), L is the sentence length, the process of using the text self-attention mechanism to process the text preprocessing features is as follows: the text features are processed by the fully connected layer, activated by the hyperbolic tangent function, the obtained value is multiplied by the original text features, activated by the flexible maximum transfer function, and the weight is obtained. The original text features are weighted with their weights, and the first intermediate features are added; the text self-attention mechanism, the calculation process is as follows:
[0092] u j =tanh(W q q j )
[0093]
[0094]
[0095] Among them, q j is the text feature vector, W q are the fully connected layer parameters, This is the first intermediate feature.
[0096] Mapping the first intermediate feature to the image space to obtain a second intermediate feature; wherein the number of channels of the second intermediate feature is consistent with the number of channels of the image feature of the corresponding scale;
[0097] Among them, the first intermediate feature is mapped to the image space through the fully connected layer.
[0098] The second intermediate features are copied and spliced to obtain a third intermediate feature having the same scale as the feature of the corresponding scale image;
[0099] wherein the second intermediate feature is copied and spliced along the channel;
[0100] The third intermediate feature is multiplied element by element with the image feature of the corresponding scale to obtain the text-image fusion feature at the corresponding scale;
[0101] Based on the channel attention mechanism and the spatial attention mechanism, the text-image fusion features at the corresponding scale are learned to obtain the image features at the corresponding scale.
[0102] In this embodiment, the main idea of the text-guided visual perception subnetwork is to use the input text preprocessing features as auxiliary information, further extract the input image features, and obtain the corresponding image features, thereby realizing the information interaction between the image and text modal information when learning their respective modal features, promoting feature learning, and enhancing the utilization of information.
[0103] Based on the above embodiments, as an optional embodiment, Figure 3 The visually guided text perception subnetwork in the exemplary modal intrinsic relationship perception network structure diagram is used to realize the function of learning the text preprocessing features in combination with the corresponding scale image features to obtain the text features at the corresponding scale. The text preprocessing features are learned in combination with the corresponding scale image features to obtain the text features at the corresponding scale, specifically:
[0104] Perform global pooling operation on the image features of corresponding scales to obtain the corresponding global visual features;
[0105] Among them, the global visual features are obtained through the global pooling operation (max pooling) (d i is the number of image feature channels);
[0106] Based on the visual self-attention mechanism, the global visual features are aggregated into one feature;
[0107] Among them, the visual self-attention mechanism is similar to the text self-attention mechanism, which aggregates the global image into a feature f v ;
[0108] The aggregated features are multiplied element by element with the text preprocessing features to obtain the image-text fusion features at the corresponding scale;
[0109] Based on the text self-attention mechanism, the image-text fusion features at the corresponding scale are learned to obtain the text features at the corresponding scale.
[0110] In this embodiment, the main idea of the visually guided text perception subnetwork is to use the input image features as auxiliary information, further extract features from the text preprocessing features, and obtain corresponding text features, thereby realizing information interaction between the image and text modal information when learning their respective modal features. Compared with traditional methods, the information of images and texts is more fully utilized, the feature learning effect is better, and the performance of the model can be improved.
[0111] Based on the above embodiments, as an optional embodiment, Figure 4 The schematic diagram of the cross-modal fusion network structure is illustrated. The structure realizes the function of considering feature similarity, fusing image features and text features at corresponding scales, and obtaining image-text collaborative features at corresponding scales. The considering feature similarity, fusing image features and text features at corresponding scales, and obtaining image-text collaborative features at corresponding scales are specifically:
[0112] Calculate the similarity between image features and text features at the corresponding scale;
[0113] In this embodiment, the cross-modal fusion network maps the learned image and text features to the same common space through the convolution layer and the fully connected layer. and The subscript a here represents the level in the pyramid; the mapping space formed after the image and text features are mapped to the same space is the same size as the image feature map space;
[0114] The image features With text features c Multiply the vectors and activate the flexible maximum transfer function to get the text-image similarity value, which is:
[0115] m c,x =f c,x l c
[0116]
[0117] Among them, f c,x Yes V c The characteristics of the grid point in the cth row and xth column above, C represents V c The number of rows included, X represents V c The number of columns to include, s c,x is the similarity between the image feature and the text feature at the grid point in the cth row and the xth column of the mapping space. Correspondingly, it is easy to perform the same operation on each grid point of the image feature map to determine the similarity matrix of the image feature and the text feature of the entire mapping space
[0118] The similarity is used to fuse the image features and text features at the corresponding scale to obtain the image-text collaborative features at the corresponding scale.
[0119] In this embodiment, the method of calculating the similarity of the two modalities by mapping the features of the two modalities to the same common space improves the accuracy of similarity calculation, and weights the text and image features with the similarity values, so that the utilization of the text and image information in the position space changes according to the similarity value, thereby solving the problem that the utilization of text information is independent of the image position space.
[0120] Based on the above embodiments, as an optional embodiment, the image-text collaborative features at each scale are sampled, spliced, feature selected and transformed to obtain target features, specifically:
[0121] If the current network is the first multi-scale fusion network, the image-text collaborative features at each scale are downsampled based on the minimum scale, and are spliced in descending order of the scales before sampling; if the current network is the second multi-scale fusion network, the image-text collaborative features at each scale are upsampled based on the maximum scale, and are spliced in ascending order of the scales before sampling;
[0122] This embodiment provides a preferred splicing sequence. It is understandable that the splicing operation can be performed using any splicing sequence, and the corresponding splicing effect is also predictable.
[0123] The concatenated features are sequentially subjected to feature selection and feature transformation to obtain the target features.
[0124] In the prior art, multi-scale image features of feature pyramids are often fused without selection, resulting in the fusion of unfavorable information in image features of each scale, which affects the performance of multi-scale fusion to a certain extent. Based on this, the present invention targets the multi-scale features of feature pyramids, up / down samples them to the same scale, then splices them, and then uses a feature selection strategy to select the spliced features, selects features related to the object being referred to, suppresses irrelevant background feature information, and solves the problem of effective fusion of multi-scale features.
[0125] Based on the above embodiments, as an optional embodiment, the image-text collaborative feature at the corresponding scale is specifically calculated by the following formula:
[0126]
[0127]
[0128]
[0129] In the above formula, F vl,i represents the image-text collaborative feature at the i-th scale, represents the weighted image features at the i-th scale, represents the weighted text features at the i-th scale, S i represents the similarity between image features and text features at the i-th scale, V c,i represents the image features at the i-th scale, V l,i It means that the text features at the i-th scale are copied and concatenated to obtain text features with the same scale as the image features at the i-th scale.
[0130] In the present invention, the similarity matrix of image features and text features is used to filter the input image features and text features to obtain weighted image features and text features, which are then multiplied and fused element by element to obtain two-modal collaborative features; by establishing the collaboration of image and text modalities in semantic and positional space, the image-text collaborative features are obtained. Compared with traditional methods, the image-text collaborative features can better distinguish the object referred to and the background, which is beneficial to the fusion of the two modalities and the improvement of model performance.
[0131] Based on the above embodiments, as an optional embodiment, Figure 5 The schematic diagram of the first / second multi-scale fusion network structure is illustrated, where the first / second multi-scale fusion network is used to analyze the image-text collaborative features at each scale. After downsampling / upsampling to the minimum / maximum scale, each feature is then concatenated into a feature vector, i.e., F = [F 1 ,…,F n], where F is the obtained feature and [·] is the concatenation operation.
[0132] The feature selection of the spliced features is specifically calculated by the following formula:
[0133]
[0134]
[0135] F pool =MaxPooling(F sum )
[0136] F t =W 2 W 1 F pool
[0137]
[0138]
[0139] In the above formula, F represents the concatenated feature, Ω k (·) is the convolutional layer with the kth receptive field, It represents the features obtained by processing the concatenated features using the convolutional layer of the kth receptive field. represents the element-by-element addition operation, F sum represents the feature obtained by adding the features obtained by processing the concatenated features using the convolutional layers of each receptive field element by element, MaxPooling(·) represents the global pooling operation, and F pool Indicates F sum The feature obtained after the global pooling operation, W 2 represents the parameters of the second fully connected layer, W 1 represents the parameters of the first fully connected layer, F t Represents the global information after conversion, G k (·) represents the weight function of the kth receptive field branch, f k represents the selection weight coefficient of the kth receptive field branch after activation, F s It represents the features obtained by selecting the concatenated features, and K represents the total number of receptive fields.
[0140] In this embodiment, the global feature information is transformed by two fully connected layers and activated by the flexible maximum transfer function. The feature is then used to select the features processed by the original different receptive fields, select the features related to the object being referred to, and filter out the less relevant background information, thereby improving the positioning and segmentation accuracy of the model.
[0141] Based on the above embodiments, as an optional embodiment, the pre-processed image backbone network includes: a VGG network and a ResNet network;
[0142] The preprocessing text backbone network includes: a recursive neural network and a long short-term memory network.
[0143] In this embodiment, the VGG network and the ResNet network are only two preferred methods for preprocessing the image backbone network. Similarly, the recursive neural network and the long short-term memory network are only two preferred methods for preprocessing the text backbone network. The purpose of the optimization is to ensure the preprocessing effect of the image / text.
[0144] It should be understood that the level of the preprocessing image backbone network is determined according to actual needs. For example, in actual operation, the preprocessing image backbone network can use a 101-layer deep residual backbone network (ResNet-101) to perform image preprocessing and construct an image feature pyramid.
[0145] Similarly, the actual choice of the preprocessing text backbone network is also determined according to actual needs. For example, the glove word vector corresponds to the input word vector of each word in the text, and then a two-layer bidirectional gated recurrent unit (Bi-GRU) network is used to perform basic text data processing.
[0146] In the second aspect, the construction device of the reference expression positioning and segmentation model provided by the present invention is described, and the construction device of the reference expression positioning and segmentation model described below and the construction method of the reference expression positioning and segmentation model described above can be referenced to each other. Figure 6 A schematic diagram of a structure of a device for constructing a positioning and segmentation model is shown as an example. Figure 6 As shown, the device includes: a reference expression positioning and segmentation database construction module 21, a preprocessing image backbone network and a preprocessing text backbone network construction module 22, a modality intrinsic relationship perception network construction module 23, a cross-modality fusion network construction module 24, a first multi-scale fusion network and a second multi-scale fusion network construction module 24 and an optimization training module 26;
[0147] Among them, the referential expression positioning and segmentation database construction module 21 is used to construct a referential expression positioning and segmentation database; wherein the database samples include: images with the referents positioned and segmented and annotated, and texts describing the referents; the preprocessed image backbone network and preprocessed text backbone network construction module 22 are used to construct the preprocessed image backbone network and the preprocessed text backbone network; wherein the preprocessed image / text backbone network is used for feature extraction to obtain image / text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales; the modal intrinsic relationship perception network construction module 23 is used to construct a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visually guided text perception subnetwork for image features of each scale; wherein the text-guided visual perception subnetwork / visually guided text perception subnetwork is used to combine text preprocessing features / corresponding scale image features to learn the corresponding scale image features / text preprocessing features to obtain images at the corresponding scale. Image features / text features; a cross-modal fusion network construction module 24, used to construct each cross-modal fusion network corresponding to each modal intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales; a first multi-scale fusion network and a second multi-scale fusion network construction module 25, used to construct a first multi-scale fusion network and a second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select features and transform features of image-text collaborative features at each scale to obtain target features; the target features are used to realize the positioning / segmentation of the referent; an optimization training module 26, used to use the database to optimize the network composed of the pre-processed image backbone network, the pre-processed text backbone network, each cross-modal fusion network, each modal intrinsic relationship perception network, the first multi-scale fusion network and the second multi-scale fusion network to obtain a referential expression positioning and segmentation model.
[0148] The device for constructing the reference expression positioning and segmentation model provided in the embodiment of the present invention specifically executes the above-mentioned construction method embodiments of the reference expression positioning and segmentation model. For details, please refer to the contents of the above-mentioned construction method embodiments of the reference expression positioning and segmentation model, which will not be repeated here.
[0149] The device for constructing a referential expression positioning and segmentation model provided by an embodiment of the present invention, for image pyramid and text preprocessing features, the modal intrinsic relationship perception network uses text / image features as auxiliary information to assist in image / text feature learning, obtains image / text features related to the referent, makes more effective use of image / text information, and improves the effect of image / text feature learning; the cross-modal fusion network maps the features of the two modalities of image and text to the same common space to calculate the similarity, uses the similarity to filter the image and text features, and then obtains the image-text collaborative features, which establish the collaboration of the two modalities of image and text in the semantic and position space, further distinguishes the referent from the background, and improves the fusion effect of the two modalities; the first multi-scale fusion network and the second multi-scale fusion network fuse the image-text collaborative features at multiple scales with a feature selection strategy, and improves the accuracy of feature selection by selecting feature information related to the referent and suppressing other background information. The above points are combined to achieve the effect of improving the positioning and segmentation accuracy of the model.
[0150] Thirdly, Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute the construction method of the referential expression positioning and segmentation model, which method includes: step 1: constructing a referential expression positioning and segmentation database; wherein the database samples include: images with positioning and segmentation annotations for the referents, and texts describing the referents; step 2: constructing a preprocessed image backbone network and a preprocessed text backbone network; wherein the preprocessed image / text backbone network is used for feature extraction to obtain image / text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales; step 3: for each scale image feature, a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visually guided text perception subnetwork is correspondingly constructed; wherein the text-guided visual perception subnetwork / visually guided text perception subnetwork is used to combine the text preprocessing features / corresponding scale image features to extract the corresponding scale image features / text preprocessing features. The method comprises the following steps: learning the rational features to obtain the image features / text features at the corresponding scale; step four: constructing each cross-modal fusion network corresponding to each modal intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider the feature similarity, fuse the image features and text features at the corresponding scale, and obtain the image-text collaborative features at the corresponding scale; step five: constructing a first multi-scale fusion network and a second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select and transform the image-text collaborative features at each scale to obtain the target features; the target features are used to realize the positioning / segmentation of the referent; step six: using the database, optimizing and training the network composed of the pre-processed image backbone network, the pre-processed text backbone network, each cross-modal fusion network, each modal intrinsic relationship perception network, the first multi-scale fusion network and the second multi-scale fusion network to obtain the referential expression positioning and segmentation model.
[0151] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0152] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for constructing a referential expression localization and segmentation model provided above, the method comprising: step one: constructing a referential expression localization and segmentation database; wherein the database samples include: images in which referents are located and segmented and annotated, and texts describing referents; step two: constructing a preprocessed image backbone network and a preprocessed text backbone network; wherein the preprocessed image / text backbone network is used for feature extraction to obtain image / text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales; step three: for each scale image feature, constructing a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visually guided text perception subnetwork; wherein the text-guided visual perception subnetwork / visually guided text perception subnetwork is used to combine text preprocessing features / corresponding scale images Features, learn the image features / text preprocessing features of the corresponding scale to obtain the image features / text features at the corresponding scale; Step 4: Construct each cross-modal fusion network corresponding to the intrinsic relationship perception network of each modality; wherein the cross-modal fusion network is used to consider feature similarity, fuse the image features and text features at the corresponding scale, and obtain the image-text collaborative features at the corresponding scale; Step 5: Construct a first multi-scale fusion network and a second multi-scale fusion network; wherein the first / second multi-scale fusion network is used to sample, splice, select features and transform features of the image-text collaborative features at each scale to obtain the target features; the target features are used to realize the positioning / segmentation of the referent; Step 6: Using the database, the network composed of the preprocessed image backbone network, the preprocessed text backbone network, each cross-modal fusion network, each modal intrinsic relationship perception network, the first multi-scale fusion network and the second multi-scale fusion network is optimized and trained to obtain a referential expression positioning and segmentation model.
[0153] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0154] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a reference expression localization and segmentation model, It is characterized in that include: Step 1: construct a referential expression location and segmentation database; wherein the samples of the database include: images with location and segmentation annotations of referents, and texts describing referents; Step 2: construct a preprocessed image backbone network and a preprocessed text backbone network; wherein the preprocessed image backbone network is used for feature extraction to obtain image preprocessing features, and the preprocessed text backbone network is used for feature extraction to obtain text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales; Step 3: For each scale image feature, a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visual-guided text perception subnetwork is constructed accordingly; wherein the text-guided visual perception subnetwork is used to learn the image features of the corresponding scale in combination with the text preprocessing features to obtain the image features at the corresponding scale, and the visual-guided text perception subnetwork is used to learn the text preprocessing features in combination with the image features of the corresponding scale to obtain the text features at the corresponding scale; Step 4: constructing each cross-modal fusion network corresponding to each modality intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales; Step 5: construct a first multi-scale fusion network and a second multi-scale fusion network; wherein the first multi-scale fusion network and the second multi-scale fusion network are used to sample, splice, select and transform the image-text collaborative features at each scale to obtain target features; the target features are used to realize the positioning and segmentation of the referent; Step six: Using the database, train the network composed of the preprocessed image backbone network, the preprocessed text backbone network, the cross-modal fusion networks, the intrinsic relationship perception network of each modality, the first multi-scale fusion network and the second multi-scale fusion network to obtain the reference expression positioning and segmentation model.
2. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The positioning and marking of the referent specifically includes: marking the position of the referent detection frame; The segmenting and labeling of the referent specifically includes: labeling the pixel points covered by the referent.
3. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The method combines the text preprocessing features to learn the image features of the corresponding scale to obtain the image features at the corresponding scale, including: Based on the text self-attention mechanism, the text preprocessing features are learned to obtain the first intermediate features; Mapping the first intermediate feature to the image space to obtain a second intermediate feature; wherein the number of channels of the second intermediate feature is consistent with the number of channels of the image feature of the corresponding scale; The second intermediate features are copied and spliced to obtain a third intermediate feature having the same scale as the feature of the corresponding scale image; The third intermediate feature is multiplied element by element with the image feature of the corresponding scale to obtain the text-image fusion feature at the corresponding scale; Based on the channel attention mechanism and the spatial attention mechanism, the text-image fusion features at the corresponding scale are learned to obtain the image features at the corresponding scale.
4. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The text preprocessing features are learned by combining the image features of the corresponding scale to obtain the text features at the corresponding scale, specifically: Perform global pooling operation on the image features of corresponding scales to obtain the corresponding global visual features; Based on the visual self-attention mechanism, the global visual features are aggregated into one feature; The aggregated features are multiplied element by element with the text preprocessing features to obtain the image-text fusion features at the corresponding scale; Based on the text self-attention mechanism, the image-text fusion features at the corresponding scale are learned to obtain the text features at the corresponding scale.
5. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The feature similarity is considered to fuse the image features and text features at the corresponding scale to obtain the image-text collaborative features at the corresponding scale, specifically: Calculate the similarity between image features and text features at the corresponding scale; The similarity is used to fuse the image features and text features at the corresponding scale to obtain the image-text collaborative features at the corresponding scale.
6. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The image-text collaborative features at each scale are sampled, spliced, feature selected and transformed to obtain target features, specifically: If the current network is the first multi-scale fusion network, the image-text collaborative features at each scale are downsampled based on the minimum scale, and are spliced in descending order of the scales before sampling; if the current network is the second multi-scale fusion network, the image-text collaborative features at each scale are upsampled based on the maximum scale, and are spliced in ascending order of the scales before sampling; The concatenated features are sequentially subjected to feature selection and feature transformation to obtain the target features.
7. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The image-text collaborative feature at the corresponding scale is specifically calculated by the following formula: In the above formula, F vl,i represents the image-text collaborative feature at the i-th scale, represents the weighted image features at the i-th scale, represents the weighted text features at the i-th scale, S i represents the similarity between image features and text features at the i-th scale, V c,i represents the image features at the i-th scale, V l,i It means that the text features at the i-th scale are copied and concatenated to obtain text features with the same scale as the image features at the i-th scale.
8. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The feature selection of the spliced features is specifically calculated by the following formula: F pool =MaxPooling(F sum ) F t =W 2 W 1 F pool In the above formula, F represents the concatenated feature, Ω k (·) is the convolutional layer with the kth receptive field, It represents the features obtained by processing the concatenated features using the convolutional layer of the kth receptive field. represents the element-by-element addition operation, F sum represents the feature obtained by adding the features obtained by processing the concatenated features using the convolutional layers of each receptive field element by element, MaxPooling(·) represents the global pooling operation, and F pool Indicates F sum The feature obtained after the global pooling operation, W 2 represents the parameters of the second fully connected layer, W 1 represents the parameters of the first fully connected layer, F t Represents the global information after conversion, G k (·) represents the weight function of the kth receptive field branch, f k represents the selection weight coefficient of the kth receptive field branch after activation, F s It represents the features obtained by selecting the concatenated features, and K represents the total number of receptive fields.
9. The method for constructing a reference expression positioning and segmentation model according to claim 1, It is characterized in that The preprocessing image backbone network includes: a VGG network and a ResNet network; The preprocessing text backbone network includes: a recursive neural network and a long short-term memory network.
10. A device for constructing a reference expression positioning and segmentation model, It is characterized in that include: A referential expression location and segmentation database construction module is used to construct a referential expression location and segmentation database; wherein the samples of the database include: images with location and segmentation annotations for referents, and texts describing referents; The preprocessing image backbone network and the preprocessing text backbone network construction module are used to construct the preprocessing image backbone network and the preprocessing text backbone network; wherein the preprocessing image backbone network is used for feature extraction to obtain image preprocessing features, and the preprocessing text backbone network is used for feature extraction to obtain text preprocessing features; the image preprocessing features are feature pyramids composed of image features of different scales; The modal intrinsic relationship perception network construction module is used to construct a modal intrinsic relationship perception network including a text-guided visual perception subnetwork and a visual-guided text perception subnetwork for image features of each scale; wherein the text-guided visual perception subnetwork is used to learn the image features of the corresponding scale in combination with the text preprocessing features to obtain the image features at the corresponding scale, and the visual-guided text perception subnetwork is used to learn the text preprocessing features in combination with the image features of the corresponding scale to obtain the text features at the corresponding scale; A cross-modal fusion network construction module is used to construct each cross-modal fusion network corresponding to each modality intrinsic relationship perception network; wherein the cross-modal fusion network is used to consider feature similarity, fuse image features and text features at corresponding scales, and obtain image-text collaborative features at corresponding scales; A first multi-scale fusion network and a second multi-scale fusion network construction module are used to construct the first multi-scale fusion network and the second multi-scale fusion network; wherein the first multi-scale fusion network and the second multi-scale fusion network are used to sample, splice, select features and transform features of image-text collaborative features at each scale to obtain target features; the target features are used to realize the positioning and segmentation of the referent; The optimization training module is used to use the database to train a network composed of a preprocessed image backbone network, a preprocessed text backbone network, each cross-modal fusion network, each modality intrinsic relationship perception network, a first multi-scale fusion network and a second multi-scale fusion network to obtain a reference expression positioning and segmentation model.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the method for constructing a positioning and segmentation model as described in any one of claims 1 to 9 are implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method for constructing a positioning and segmentation model as described in any one of claims 1 to 9 are implemented.