Image semantic segmentation method and device, electronic equipment and readable medium

By embedding trust tokens in the CLIP model and performing feature fusion, the shortcomings of the CLIP model in dealing with unknown categories and accurately distinguishing known categories are solved, and the accuracy and stability of image semantic segmentation are improved.

CN119942109APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411978931.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In actual application of the CLIP model, there are still some problems with predicted categories, which are difficult to deal with unknown categories and lack the ability to accurately distinguish known categories, resulting in omissions or incorrect labeling of segmentation results, affecting the accuracy and stability of the actual business.

Method used

By embedding a trust token in the text token, it is used to identify the known category label and unknown category label in the text to be predicted, and the matching image token is fused with the text token embedded in the trust token through the trust learner to obtain the fusion feature. Then, semantically segmenting the fusion feature through the semantic segmentation network to obtain the result mask, and finally allocating the category label to each pixel in the predicted image.

Benefits of technology

Improve the adaptability and identification capabilities of the model, avoid omissions or labeling errors caused by unknown categories, thereby improving the accuracy and stability of the actual business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942109A_ABST
    Figure CN119942109A_ABST
Patent Text Reader

Abstract

The invention provides an image semantic segmentation method and device, electronic equipment and a readable medium. The method comprises the following steps: inputting a to-be-predicted image and a to-be-predicted text into a multi-modal model containing a text encoder and an image encoder for feature extraction, and respectively obtaining an image token and a text token; embedding a trust token in the text token, wherein the trust token is used for identifying a known category label and an unknown category label in the to-be-predicted text; performing feature fusion on the matched image token and the text token embedded with the trust token through a trust learner to obtain a fusion feature; performing semantic segmentation on the fusion features through a semantic segmentation network to obtain a result mask; and allocating a category label to each pixel in the to-be-predicted image through the result mask to obtain an image segmentation result of the to-be-predicted image. According to the method, the adaptability and the recognition capability of the model can be improved, and omission or marking errors caused by unknown categories are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image semantic segmentation method, device, electronic device and readable medium. Background Art

[0002] In today's era of surging digitalization, images are important carriers of information, and their semantic segmentation technology has become a key driving force for the development of many industries. From intelligent driving systems accurately identifying roads, vehicles, pedestrians and traffic signs, to medical images accurately analyzing lesions and organ tissues, to intelligent monitoring systems capturing abnormal behaviors and target objects in real time, image semantic segmentation is everywhere.

[0003] In related technologies, multimodal learning has become a research hotspot, especially in combining text and image information. The CLIP (Contrastive Language-Image Pre-training) model jointly trains text and image encoders, and learns the semantic association between text and image through large-scale text-image data comparison learning, thereby continuing the image semantic segmentation task.

[0004] However, in practical applications, the CLIP model still has some problems in predicting categories. It is difficult to handle unknown categories and lacks the ability to accurately distinguish known categories, which leads to omissions or incorrect labeling of segmentation results, affecting the accuracy and stability of actual business. Summary of the invention

[0005] Based on the above technical problems, the present application provides an image semantic segmentation method, device, electronic device and readable medium to improve the adaptability and recognition ability of the model and avoid omissions or labeling errors caused by unknown categories.

[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0007] According to one aspect of an embodiment of the present application, a method for image semantic segmentation is provided, comprising:

[0008] The image to be predicted and the text to be predicted are input into a multimodal model including a text encoder and an image encoder for feature extraction, and image tokens and text tokens are obtained respectively;

[0009] Embedding a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted;

[0010] Performing feature fusion on the matched image token and the text token embedded in the trust token through a trust learner to obtain a fused feature;

[0011] Performing semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask;

[0012] A category label is assigned to each pixel in the image to be predicted by using the result mask to obtain an image segmentation result of the image to be predicted.

[0013] According to one aspect of an embodiment of the present application, there is provided an image semantic segmentation apparatus, comprising:

[0014] A feature extraction module is configured to input the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder for feature extraction, and obtain image tokens and text tokens respectively;

[0015] A token embedding module, configured to embed a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted;

[0016] A feature fusion module is configured to perform feature fusion on the matched image token and the text token embedded in the trust token through a trust learner to obtain a fused feature;

[0017] A semantic segmentation module, configured to perform semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask;

[0018] The category assignment module is configured to assign a category label to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be predicted.

[0019] In some embodiments of the present application, based on the above technical solution, the trust learner includes a linear mapping layer, an attention layer and a normalization layer; the feature fusion module is configured as follows:

[0020] Performing linear mapping on the text token after embedding the trust token through the linear mapping layer of the trust learner and the image token to obtain a similarity matrix, wherein the similarity matrix at least includes a first vector, a second vector and a third vector, the first vector is generated based on the text token mapping, and the second vector and the third vector are generated based on the image embedding representation mapping;

[0021] The similarity matrix is ​​sequentially input into the attention layer and the normalization layer for processing to obtain fusion features.

[0022] In some embodiments of the present application, based on the above technical solution, the image semantic segmentation device further includes a model training module configured to:

[0023] Acquire images with category annotations and their corresponding text description data as training data, wherein the category annotations include reference masks and reference categories for semantic segmentation, and the training data include a training set, a validation set, and a test set;

[0024] Extracting image and text data pairs in batches from the training set, training through a multimodal model to be trained, a trust learner to be trained, and a semantic segmentation network to be trained, to obtain output original masks and trust masks;

[0025] Performing loss calculation based on the original mask and the difference between the trusted mask and the reference mask to obtain a first loss;

[0026] Performing loss calculation according to the difference between the category recognition result output by the trust learner to be trained and the benchmark category to obtain a second loss;

[0027] According to the first loss and the second loss, parameters of the multimodal model to be trained, the trust learner to be trained, and the semantic segmentation network to be trained are adjusted to obtain a trained multimodal model, a trust learner, and a semantic segmentation network.

[0028] In some embodiments of the present application, based on the above technical solution, the image encoder encodes the image to be predicted into a matrix of a fixed size of 2048*2048 dimensions.

[0029] In some embodiments of the present application, based on the above technical solution, the text encoder converts the text to be predicted into a 768-dimensional vector.

[0030] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the image semantic segmentation method in the above technical solution by executing the executable instructions.

[0031] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the image semantic segmentation method in the above technical solution is implemented.

[0032] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the image semantic segmentation method provided in the above-mentioned various optional implementations.

[0033] In an embodiment of the present application, the system inputs the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder for feature extraction, and obtains image tokens and text tokens respectively, and then embeds a trust token in the text token, and the trust token is used to identify the known category labels and unknown category labels in the text to be predicted, and then the matching image token and the text token embedded with the trust token are subjected to feature fusion through a trust learner to obtain a fusion feature, and then the fusion feature is subjected to semantic segmentation through a semantic segmentation network to obtain a result mask, and finally the category label is assigned to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be tested. In this way, a trust token is added to the text token, and the category label in the text token is marked with the matching image through the trust token, so as to identify whether a new category label appears in the known category, thereby improving the adaptability and recognition ability of the model, avoiding omissions or labeling errors caused by unknown categories, and thus improving the accuracy and stability of actual business.

[0034] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 The image semantic segmentation method of the embodiment of the present application is applied to the system structure of the image semantic segmentation system.

[0037] Figure 2 Flowchart of an image semantic segmentation method according to an embodiment of the present application.

[0038] Figure 3 It is a structural schematic diagram of the image semantic segmentation system in an embodiment of the present application.

[0039] Figure 4 Schematic diagram of the structure of the trust learner in the embodiment of the present application.

[0040] Figure 5 The block diagram schematically shows the composition of the image semantic segmentation device in an embodiment of the present application.

[0041] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION

[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0043] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0044] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function, and works together with other related parts to achieve the predetermined function, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0045] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0046] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0047] It should be understood that the solution of the present application can be applied in the field of image recognition, and specifically in the scene of image semantic segmentation. Semantic segmentation is to assign category labels to pixels in an image to achieve classification and regional division of different objects in the image. It focuses on the semantic meaning of different regions in the image, without distinguishing different instances of the same category. The solution of the present application can be used in various scenarios such as autonomous driving, medical image analysis, robot vision, and smart city management. Specifically, in the autonomous driving scenario, the autonomous driving vehicle needs to perceive the surrounding environment in real time and identify objects such as roads, pedestrians, and traffic signs. Semantic segmentation technology can help vehicles accurately understand and interpret images captured by the camera, improving driving safety and reliability. By introducing trust tokens, the model can show higher robustness in the face of complex road conditions and unknown objects, reducing the risk of misjudgment. In the medical image analysis scenario, semantic segmentation is used to automatically identify and annotate lesion areas to assist doctors in diagnosis. For example, in CT, MRI and other images, accurately segmenting structures such as tumors and blood vessels is crucial for early detection and treatment of diseases. Trust tokens can enhance the model's ability to distinguish known and unknown lesion types and improve the accuracy of diagnosis. In robot vision scenarios, robots need to have strong visual perception capabilities to perform tasks in complex environments. Semantic segmentation can help robots identify and understand surrounding objects and scenes, so as to make reasonable decisions and actions. Trust tokens can improve the adaptability of robots when facing unknown objects or new environments, and enhance their ability to navigate and operate autonomously. In smart city management scenarios, semantic segmentation technology can be used for tasks such as intelligent traffic monitoring and crowd behavior analysis during city monitoring and management. By segmenting and classifying objects in video streams in real time, functions such as traffic flow monitoring and abnormal behavior detection can be achieved. Trust tokens can improve the stability and accuracy of models in complex backgrounds and changing environments.

[0048] In today's era of surging digitalization, images are important carriers of information, and their semantic segmentation technology has become a key driving force for the development of many industries. From intelligent driving systems accurately identifying roads, vehicles, pedestrians and traffic signs, to medical images accurately analyzing lesions, organ tissues, and intelligent monitoring systems capturing abnormal behaviors and target objects in real time, image semantic segmentation is everywhere. In related technologies, multimodal learning has become a research hotspot, especially in combining text and image information. The CLIP (Contrastive Language-Image Pre-training) model jointly trains text and image encoders, and learns the semantic association between text and image through large-scale text-image data comparison learning, thereby continuing the image semantic segmentation task. However, in practical applications, the CLIP model still has some problems in predicting categories. It is difficult to handle unknown categories and lacks the ability to accurately distinguish known categories, resulting in omissions or incorrect annotations in segmentation results, affecting the accuracy and stability of actual business.

[0049] Based on this, the technical solution of the embodiment of the present application proposes an image semantic segmentation solution. Figure 1 According to the image semantic segmentation method of the embodiment of the present application, the system structure of the image semantic segmentation system can mainly include two parts, namely the user terminal 110 and the server 130. Among them, the devices of each part in the system structure can include smart phones, tablet computers, laptops, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The device of the blockchain can also be a server that provides various services. It can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network) and big data and artificial intelligence platforms. The user terminal 110 and the server 130 communicate through the network 120. The network 120 can be a communication medium of various connection types of communication links, for example, it can be a wired communication link or a wireless communication link.

[0050] According to the implementation requirements, the system architecture in the embodiment of the present application can have any number of terminal devices and servers. For example, the devices of the server 130 and other parts can be a server group composed of multiple server devices. In addition, the technical solution provided in the embodiment of the present application can be applied to a computing platform, or can be implemented by various parts in the system, and the present application does not make any special restrictions on this.

[0051] like Figure 1As shown, an image semantic segmentation system is deployed on the server 130. A client for accessing the image semantic segmentation system on the server 130 is installed on the terminal device 110. The terminal device 110 inputs the image to be predicted and the text to be predicted into the image semantic segmentation system on the server 130 for feature extraction, and obtains an image token and a text token respectively. The server 130 embeds a trust token in the text token, and the trust token is used to identify the known category label and the unknown category label in the text to be predicted. Subsequently, the server 130 performs feature fusion on the matched image token and the text token embedded with the trust token to obtain a fused feature, and then inputs the fused feature into the semantic segmentation network for segmentation to obtain a result mask, and finally assigns a category label to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be tested. The image segmentation result will be sent to the terminal device 110 by the server 130, and the terminal device 110 will perform further business processing according to the labeling results of each part in the image segmentation result, such as obstacle avoidance for vehicles or target recognition for robots.

[0052] The implementation details of the technical solution of the embodiment of the present application are described in detail below: Figure 2 A flowchart of an image semantic segmentation method according to an embodiment of the present application is shown. The image semantic segmentation method can be executed by a device having a computing and processing function, such as a server or a terminal device of an image semantic segmentation system. Figure 2 As shown, the image semantic segmentation includes at least steps S210 to S250, which are described in detail as follows:

[0053] Step S210: Input the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder for feature extraction to obtain image tokens and text tokens respectively.

[0054] For the image to be predicted and the text to be predicted, the image semantic segmentation system inputs them into a multimodal model including a text encoder and an image encoder. The image encoder is responsible for encoding the image to be predicted into a fixed-size 2048*2048-dimensional matrix. This specific dimension setting is obtained after a large number of experimental optimizations. It can fully retain the key feature information of the image and facilitate the subsequent fusion calculation with text features. The text encoder converts the text to be predicted into a 768-dimensional vector, which adapts to the expression requirements of text features and effectively extracts the semantic connotation of the text. Through the collaborative work of these two encoders, image tokens and text tokens are obtained respectively. In one embodiment, the image encoder can use a pre-trained convolutional neural network (CNN) or Transformer architecture, such as a pre-trained ResNet50 or ViT-L / 14 model, and the text encoder can use a pre-trained BERT-base or RoBERTa-large model. In some embodiments, the image semantic segmentation system pre-processes the image to be predicted and the text to be predicted. For example, the image is pre-processed, including operations such as cropping, scaling, and normalization, to ensure that all images are of the same size for subsequent processing. For text, you can perform preprocessing operations such as word segmentation and stop word removal to ensure text quality.

[0055] Step S220: embedding a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted.

[0056] Specifically, text tokens are usually represented by vectors, and trust tokens can be added to the back of the vector of text tokens. The role of trust tokens is to help the model more accurately determine the semantic category of pixels in the inference stage, especially when facing unknown categories. Trust tokens are a learnable parameter vector, and the initial value can be set by random initialization or preset. During the training process, trust tokens are dynamically adjusted according to the feedback of the loss function, and their role in the model is gradually optimized. Trust tokens are not only used to enhance the model's ability to distinguish between known and unknown categories, but also serve as additional supervisory signals to help the model better understand the relationship between text and images, thereby improving the accuracy of semantic segmentation. Trust tokens have a unique identification function, which is used to identify known category labels and unknown category labels in the text to be predicted. For example, when the text to be predicted is described as "there is a colorful insect that has never been seen in the flowers", the trust token can identify "flowers" as a known category label, and the part corresponding to "colorful insects that have never been seen" is marked as an unknown category label, thereby guiding the model to pay special attention to unknown categories in subsequent processing, while strengthening semantic understanding with the help of known category information.

[0057] Step S230, using a trust learner, feature fusion is performed on the matched image token and the text token embedded with the trust token to obtain a fused feature.

[0058] The image semantic segmentation system reintroduces a specially designed trust learner for feature fusion. The trust learner maps image tokens and text tokens to a unified feature space. Specifically, the image semantic segmentation system first matches the input image tokens and text tokens, and performs feature mapping on the matched image tokens and text tokens, thereby ensuring that the image tokens and the text tokens embedded with the trust tokens can effectively interact and fuse in the same feature space. Subsequently, the trust learner captures the relationship between text and image, and focuses on different features through the attention mechanism. Finally, the feature data output by the attention mechanism is normalized to obtain the final feature fusion result.

[0059] In some embodiments of the present application, based on the above technical solution, the trust learner includes a linear mapping layer, an attention layer and a normalization layer; in the process of fusing the matched image token and the text token embedded in the trust token to obtain the fused feature, the image semantic segmentation system will perform linear mapping through the linear mapping layer of the trust learner and the text token embedded in the trust token to obtain a similarity matrix, wherein the similarity matrix at least includes a first vector, a second vector and a third vector, the first vector is generated based on the text token mapping, the second vector and the third vector are generated based on the image embedding representation mapping, and then the similarity matrix is ​​sequentially input into the attention layer and the normalization layer for processing to obtain the fused feature. Specifically, in this embodiment, the trust learner includes a linear mapping layer, an attention layer and a normalization layer. The linear mapping layer performs linear mapping on the image token and the text token embedded in the trust token to generate a similarity matrix. The similarity matrix at least includes a first vector (Q), a second vector (K) and a third vector (V). The first vector is generated based on the text token mapping, and the second vector and the third vector are generated based on the image embedding representation mapping. Specifically, the text token is linearly mapped to generate a Q (Query) matrix. The specific operation is to multiply the text token vector by a trainable weight matrix W_Q to obtain the Q matrix: Q = W_Q * text token vector. Linearly map the image token vector to generate K (key) and V (Value) matrices. Multiply the image token vector by the trainable weight matrices W_K and W_V respectively to obtain the K matrix and V matrix: K = W_K * image token vector, V = W_V * image token vector. Calculate the similarity score between Q and K. Usually use the dot product operation, such as calculating the dot product of the transpose of Q and K to obtain the similarity matrix S: S = Q * K ^T.

[0060] The attention weight A is calculated based on the similarity matrix S. S is usually scaled (for example, divided by where d k is the dimension of the K matrix), and then normalized by the softmax function to obtain the attention weight A: The attention output O is calculated based on the attention weight A and the V matrix: O = A * V. In the multi-head attention mechanism, there will be multiple sets of different linear mapping matrices (i.e. different W_Q, W_K, W_V), which calculate multiple sets of attention outputs respectively, and then these outputs are spliced ​​and linearly mapped again to obtain the final multi-head attention output. The multi-head attention output is normalized, for example, using Layer Normalization. Through normalization, the model convergence can be accelerated and the stability of the model can be improved.

[0061] Step S240, performing semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask.

[0062] The image semantic segmentation system inputs the fused features into the semantic segmentation network for segmentation to obtain the result mask. The semantic segmentation network can use architectures such as U-Net, DeepLab, and Mask R-CNN to assign a semantic label to each pixel in the image to indicate the semantic category to which it belongs. The result mask is used to guide the model to more accurately judge the semantic category of the pixel during the inference stage, especially when facing unknown categories. The trust mask is used to enhance the model's ability to distinguish between known and unknown categories and reduce the possibility of misclassification.

[0063] Step S250 , assigning a category label to each pixel in the image to be predicted by using the result mask to obtain an image segmentation result of the image to be predicted.

[0064] The result mask has the same dimension as the image to be predicted. The image semantic segmentation system determines the category label of each pixel in the image to be predicted based on the category label of the pixel at the corresponding position in the result mask, thereby obtaining the image segmentation result of the image to be predicted. The image segmentation result will contain information such as the identified object category, the location of the object in the image, and the contour range.

[0065] In an embodiment of the present application, the system inputs the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder for feature extraction, and obtains image tokens and text tokens respectively, and then embeds a trust token in the text token, and the trust token is used to identify the known category labels and unknown category labels in the text to be predicted, and then the matching image token and the text token embedded with the trust token are subjected to feature fusion through a trust learner to obtain a fusion feature, and then the fusion feature is subjected to semantic segmentation through a semantic segmentation network to obtain a result mask, and finally the category label is assigned to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be tested. In this way, a trust token is added to the text token, and the category label in the text token is marked with the matching image through the trust token, so as to identify whether a new category label appears in the known category, thereby improving the adaptability and recognition ability of the model, avoiding omissions or labeling errors caused by unknown categories, and thus improving the accuracy and stability of actual business.

[0066] In some embodiments of the present application, based on the above technical solution, the image semantic segmentation system will also obtain images with category annotations and their corresponding text description data as training data, the category annotations include reference masks and reference categories for semantic segmentation, and the training data include training sets, validation sets and test sets, and then extract image and text data pairs in batches based on the training set, train through the multimodal model to be trained, the trust learner to be trained and the semantic segmentation network to be trained, and obtain the output original mask and trust mask, and then calculate the loss based on the difference between the original mask and the trust mask and the reference mask to obtain the first loss, and calculate the loss based on the difference between the category recognition result output by the trust learner to be trained and the reference category to obtain the second loss, and finally adjust the parameters of the multimodal model to be trained, the trust learner to be trained and the semantic segmentation network to be trained according to the first loss and the second loss to obtain the trained multimodal model, trust learner and semantic segmentation network. In this embodiment, the multimodal model, trust learner and semantic segmentation network in the image semantic segmentation system can be trained as components of the same model. The training process can be iterative training. The image semantic segmentation system will initialize the weights, biases and other parameters of the multimodal model, trust learner and semantic segmentation network according to the set parameters, laying the foundation for the start of training. In each round of training, image and text data pairs are extracted from the training set in batches, and feature extraction, trust token embedding, trust learning and feature fusion are performed through the multimodal model to be trained and the trust learner to be trained in strict accordance with the above scheme, and the obtained fusion features are input into the semantic segmentation network to be trained. The semantic segmentation network will output the original mask and the trust mask. According to the difference between the original mask and the trust mask and the reference mask, and the difference between the category recognition result output by the trust learner and the reference category, the first loss and the second loss are calculated, and the loss is back-propagated using the optimization algorithm to update the model parameters.

[0067] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application with specific examples. Figure 3 , Figure 3 FIG. 1 is a schematic diagram of the structure of the image semantic segmentation system in the embodiment of the present application. Figure 3As shown in the figure, the image semantic segmentation system mainly includes modules such as image encoder, text encoder, trust learner, trusted token, and segmentation head. Among them, the image encoder is used to encode the image to be predicted into a matrix of fixed size 2048*2048 dimensions. The text encoder is used to encode the category to be predicted into a 768-dimensional vector. The trust learner includes a linear mapping layer, a multi-head attention layer, and a normalization layer, and is used to convert the 768-dimensional text vector into a Q (Query) matrix and input it into the segmentation head. The trusted token is an additional trusted token introduced as a parameter to optimize the trust learner, which is used to distinguish known categories and unknown categories in the segmentation task. The segmentation head is used to assign a semantic label to each pixel in the image to indicate the semantic category to which it belongs. As shown in Figure 3 As shown in the figure, the system inputs the image and text prompt words into the CLIP network framework (including text encoder and image encoder) to obtain text token and category token vectors respectively; a trust token for learning is added after the text token vector. The trust token is connected with the text token by matrix addition, and the connected token is marked with the matching image to identify whether new category labels appear in the known categories that have not been recognized before. The matched image token and category token are input into the proposed trust learner through a trust learner. Please refer to Figure 4 , Figure 4 Schematic diagram of the structure of the trust learner in the embodiment of the present application. Figure 4 As shown in the figure, the trust learner includes a linear mapping layer, a multi-head attention layer, and a normalization layer; the linear mapping layer is applied to generate the Q (Query), K (key), and V (Value) similarity matrices. Q is the matrix of the linear mapping of the text vector, and K and V are the matrices of the linear mapping after the image embedding. Among them, the generation method of the Q (Query), K (key), and V (Value) similarity matrices is as follows:

[0068]

[0069] The similarity matrix QKV is input into the multi-head attention to enhance the accuracy of the semantic segmentation task. Q represents the current information to be paid attention to. Each attention head generates its own query vector based on the input sequence. K represents the characteristics of each element in the input sequence. Each input element has a K vector. The similarity between Q and K determines the importance of the element to the current query. V represents the actual information contained. Each K vector corresponds to a V vector.

[0070] The output of the trust learner is input into the semantic segmentation network to obtain two masks, the trust mask and the original mask. The training phase includes training trust labels and loss functions, and the difference between the original mask and the trust mask is used to enhance the learning ability and generalization of the model.

[0071] Loss function L = L mask +γL A , L mask Represents the loss between the mask and the true value of the inference output, L A is the loss of the trustworthy learner and γ is a hyperparameter.

[0072] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.

[0073] The following introduces the device implementation of the present application, which can be used to execute the image semantic segmentation method in the above embodiments of the present application. Figure 5 The block diagram of the image semantic segmentation device in the embodiment of the present application is schematically shown. Figure 5 As shown, the image semantic segmentation device 500 may mainly include:

[0074] A feature extraction module 510 is configured to input the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder to extract features, and obtain image tokens and text tokens respectively;

[0075] A token embedding module 520, configured to embed a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted;

[0076] A feature fusion module 530 is configured to perform feature fusion on the matched image token and the text token embedded in the trust token through a trust learner to obtain a fused feature;

[0077] A semantic segmentation module 540 is configured to perform semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask;

[0078] The category assignment module 550 is configured to assign a category label to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be predicted.

[0079] In some embodiments of the present application, based on the above technical solution, the trust learner includes a linear mapping layer, an attention layer and a normalization layer; the feature fusion module 530 is configured as follows:

[0080] Performing linear mapping on the text token after embedding the trust token through the linear mapping layer of the trust learner and the image token to obtain a similarity matrix, wherein the similarity matrix at least includes a first vector, a second vector and a third vector, the first vector is generated based on the text token mapping, and the second vector and the third vector are generated based on the image embedding representation mapping;

[0081] The similarity matrix is ​​sequentially input into the attention layer and the normalization layer for processing to obtain fusion features.

[0082] In some embodiments of the present application, based on the above technical solution, the image semantic segmentation device further includes a model training module configured to:

[0083] Acquire images with category annotations and their corresponding text description data as training data, wherein the category annotations include reference masks and reference categories for semantic segmentation, and the training data include a training set, a validation set, and a test set;

[0084] Extracting image and text data pairs in batches from the training set, training through a multimodal model to be trained, a trust learner to be trained, and a semantic segmentation network to be trained, to obtain output original masks and trust masks;

[0085] Performing loss calculation based on the original mask and the difference between the trusted mask and the reference mask to obtain a first loss;

[0086] Performing loss calculation according to the difference between the category recognition result output by the trust learner to be trained and the benchmark category to obtain a second loss;

[0087] According to the first loss and the second loss, parameters of the multimodal model to be trained, the trust learner to be trained, and the semantic segmentation network to be trained are adjusted to obtain a trained multimodal model, a trust learner, and a semantic segmentation network.

[0088] In some embodiments of the present application, based on the above technical solution, the image encoder encodes the image to be predicted into a matrix of a fixed size of 2048*2048 dimensions.

[0089] In some embodiments of the present application, based on the above technical solution, the text encoder converts the text to be predicted into a 768-dimensional vector.

[0090] It should be noted that the apparatus provided in the above embodiment and the method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module performs the operation has been described in detail in the method embodiment and will not be repeated here.

[0091] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.

[0092] It should be noted that Figure 6 The computer system 600 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0093] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 to the random access memory (RAM) 603. Various programs and data required for system operation are also stored in the RAM 603. The CPU 601, the ROM 602 and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0094] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom is installed into the storage section 608 as needed.

[0095] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 609, and / or installed from a removable medium 611. When the computer program is executed by a central processing unit (CPU) 601, various functions defined in the system of the present application are executed.

[0096] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CompactDisc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0097] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0098] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0099] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation method of the present application.

[0100] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.

[0101] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for image semantic segmentation, characterized in that: include: The image to be predicted and the text to be predicted are input into a multimodal model including a text encoder and an image encoder for feature extraction, and image tokens and text tokens are obtained respectively; Embedding a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted; Performing feature fusion on the matched image token and the text token embedded in the trust token through a trust learner to obtain a fused feature; Performing semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask; A category label is assigned to each pixel in the image to be predicted by using the result mask to obtain an image segmentation result of the image to be predicted.

2. The method according to claim 1, characterized in that: The trust learner includes a linear mapping layer, an attention layer and a normalization layer; the matched image token and the text token embedded in the trust token are subjected to feature fusion to obtain a fusion feature, including: Performing linear mapping on the text token after embedding the trust token through the linear mapping layer of the trust learner and the image token to obtain a similarity matrix, wherein the similarity matrix at least includes a first vector, a second vector and a third vector, the first vector is generated based on the text token mapping, and the second vector and the third vector are generated based on the image embedding representation mapping; The similarity matrix is ​​sequentially input into the attention layer and the normalization layer for processing to obtain fusion features.

3. The method according to claim 1, characterized in that The method further comprises: Acquire images with category annotations and their corresponding text description data as training data, wherein the category annotations include reference masks and reference categories for semantic segmentation, and the training data include a training set, a validation set, and a test set; Extracting image and text data pairs in batches from the training set, training through a multimodal model to be trained, a trust learner to be trained, and a semantic segmentation network to be trained, to obtain output original masks and trust masks; Performing loss calculation based on the original mask and the difference between the trusted mask and the reference mask to obtain a first loss; Performing loss calculation according to the difference between the category recognition result output by the trust learner to be trained and the benchmark category to obtain a second loss; According to the first loss and the second loss, parameters of the multimodal model to be trained, the trust learner to be trained, and the semantic segmentation network to be trained are adjusted to obtain a trained multimodal model, a trust learner, and a semantic segmentation network.

4. The method according to claim 1, characterized in that: The image encoder encodes the image to be predicted into a matrix of a fixed size of 2048*2048 dimensions.

5. The method according to claim 1, characterized in that The text encoder converts the text to be predicted into a 768-dimensional vector.

6. An image semantic segmentation device, characterized in that: include: A feature extraction module is configured to input the image to be predicted and the text to be predicted into a multimodal model including a text encoder and an image encoder for feature extraction, and obtain image tokens and text tokens respectively; A token embedding module, configured to embed a trust token in the text token, wherein the trust token is used to identify known category labels and unknown category labels in the text to be predicted; A feature fusion module is configured to perform feature fusion on the matched image token and the text token embedded in the trust token through a trust learner to obtain a fused feature; A semantic segmentation module, configured to perform semantic segmentation on the fused features through a semantic segmentation network to obtain a result mask; The category assignment module is configured to assign a category label to each pixel in the image to be predicted through the result mask to obtain the image segmentation result of the image to be predicted.

7. The device according to claim 6, characterized in that The trust learner includes a linear mapping layer, an attention layer and a normalization layer; the feature fusion module is configured as follows: Performing linear mapping on the text token after embedding the trust token through the linear mapping layer of the trust learner and the image token to obtain a similarity matrix, wherein the similarity matrix at least includes a first vector, a second vector and a third vector, the first vector is generated based on the text token mapping, and the second vector and the third vector are generated based on the image embedding representation mapping; The similarity matrix is ​​sequentially input into the attention layer and the normalization layer for processing to obtain fusion features.

8. The device according to claim 6, characterized in that The image semantic segmentation device also includes a model training module, which is configured to: Acquire images with category annotations and their corresponding text description data as training data, wherein the category annotations include reference masks and reference categories for semantic segmentation, and the training data include a training set, a validation set, and a test set; Extracting image and text data pairs in batches from the training set, training through a multimodal model to be trained, a trust learner to be trained, and a semantic segmentation network to be trained, to obtain output original masks and trust masks; Performing loss calculation based on the original mask and the difference between the trusted mask and the reference mask to obtain a first loss; Performing loss calculation according to the difference between the category recognition result output by the trust learner to be trained and the benchmark category to obtain a second loss; According to the first loss and the second loss, parameters of the multimodal model to be trained, the trust learner to be trained, and the semantic segmentation network to be trained are adjusted to obtain a trained multimodal model, a trust learner, and a semantic segmentation network.

9. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to perform the image semantic segmentation method described in any one of claims 1 to 5 by executing the executable instructions.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image semantic segmentation method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Image segmentation method and system based on adaptive alignment technology

    CN120912881A

  • Image segmentation and mask optimization model training method and device, and electronic equipment

    CN121353661A

  • Lightweight public sign intelligent identification and evaluation system

    CN121938012A