Model training method, image segmentation method and device

By using the encoding network to encode and feature fusion of prompt information in image segmentation, the complexity of the decoding network is reduced, and the model's understanding of semantics is improved, the problems of large computing resources and poor image segmentation in the prior art are solved, and more efficient and accurate image segmentation is achieved.

CN120014410APending Publication Date: 2025-05-16VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094755.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing interactive segmentation method needs to rely on decoding networks with high computational complexity, which leads to a large amount of computing resources and affects the semantic understanding of the model, resulting in poor image segmentation effect.

Method used

By obtaining sample images, prompt information and labeling information, the sample images and prompt information are input to the encoded network, fusion features are obtained, and input them to the decoding network for segmentation. Based on the actual segmentation results and labeling information, the encoded network and the decoded network are trained to reduce the complexity of the decoded network and improve the model's understanding of semantics.

Benefits of technology

Save computing resources, improve the accuracy of image segmentation, reduce the complexity of the decoding network, and enable prompt information to fully participate in the model's processing process, thereby improving the model's semantic understanding ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014410A_ABST
    Figure CN120014410A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, and an image segmentation method and device. Belongs to the technical field of artificial intelligence. The method comprises the steps that a sample image, prompt information and annotation information are acquired, the prompt information is used for indicating a target area in the sample image, and the annotation information is used for indicating a target segmentation result of the sample image; inputting the sample image and the prompt information into a coding network of the first model to obtain a fusion feature of the sample image and the prompt information output by the coding network; inputting the fusion feature into a decoding network of the first model to obtain an actual segmentation result of the sample image output by the decoding network; and training the coding network and the decoding network based on the actual segmentation result and the annotation information to obtain a second model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of communication technology, and specifically to a model training method, an image segmentation method and a device thereof. Background Art

[0002] Interactive Segmentation (IS) is a method of image segmentation based on prompt information provided by the user. The user guides the algorithm to segment specific areas through prior knowledge.

[0003] In the prior art, interactive segmentation methods usually need to encode the prompt information input by the user into an input sequence, which is then integrated into the decoding network. This method relies on decoding networks with high computational complexity such as Vision Transformer (ViT), which requires a large amount of computing resources. In addition, the prompt information is not obtained by the encoding network, which affects the overall semantic understanding of the model and leads to poor image segmentation results. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, an image segmentation method and a device thereof, which can save computing resources and improve the accuracy of image segmentation.

[0005] In a first aspect, an embodiment of the present application provides a model training method, the method comprising: obtaining a sample image, prompt information and annotation information, the prompt information being used to indicate a target area in the sample image, and the annotation information being used to indicate a target segmentation result of the sample image; inputting the sample image and the prompt information into an encoding network of a first model to obtain a fusion feature of the sample image and the prompt information output by the encoding network; inputting the fusion feature into a decoding network of the first model to obtain an actual segmentation result of the sample image output by the decoding network; and training the encoding network and the decoding network based on the actual segmentation result and the annotation information to obtain a second model.

[0006] In a second aspect, an embodiment of the present application provides an image segmentation method, the method comprising: obtaining an image to be segmented and prompt information, wherein the prompt information is used to indicate a target area in the image to be segmented; inputting the image to be segmented and the prompt information into a second model trained by the model training method described in the first aspect above to obtain an actual segmentation result.

[0007] In a third aspect, an embodiment of the present application provides a model training device, which includes: an acquisition unit, used to acquire a sample image, prompt information and annotation information, wherein the prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image; an encoding unit, used to input the sample image and the prompt information into an encoding network of a first model, and obtain a fusion feature of the sample image and the prompt information output by the encoding network; a decoding unit, used to input the fusion feature into a decoding network of the first model, and obtain an actual segmentation result of the sample image output by the decoding network; a training unit, used to train the encoding network and the decoding network based on the actual segmentation result and the annotation information, and obtain a second model.

[0008] In a fourth aspect, an embodiment of the present application provides a model training device, which includes: an acquisition unit, used to acquire an image to be segmented and prompt information, wherein the prompt information is used to indicate a target area in the image to be segmented; an image segmentation unit, used to input the image to be segmented and the prompt information into a second model trained using the model training method in the first aspect above, to obtain an actual segmentation result.

[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0010] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.

[0011] In a seventh aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method described in the first aspect.

[0012] In an eighth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0013] In an embodiment of the present application, a sample image, prompt information and annotation information are first obtained, the prompt information is used to indicate the target area in the sample image, and the annotation information is used to indicate the target segmentation result of the sample image; then the sample image and the prompt information are input into the encoding network of the first model to obtain the fusion feature of the sample image and the prompt information output by the encoding network; then the fusion feature is input into the decoding network of the first model to obtain the actual segmentation result of the sample image output by the decoding network; finally, based on the actual segmentation result and the annotation information, the encoding network and the decoding network are trained to obtain the second model. During the model training process, the encoding network encodes, extracts features and fuses features for the prompt information. On the one hand, it can reduce the complexity of the decoding network, thereby saving computing resources. On the other hand, it can enable the prompt information to fully participate in the processing process of the network structure of each part of the model, improve the overall understanding ability of the model for semantics, and thus improve the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flow chart of the model training method provided in an embodiment of the present application;

[0015] Figure 2A It is one of the model structure schematic diagrams of the model training method provided in the embodiment of the present application;

[0016] Figure 2B This is the second model structure schematic diagram of the model training method provided in the embodiment of the present application;

[0017] Figure 3 is a flow chart of the image segmentation method provided in an embodiment of the present application;

[0018] Figure 4 It is a structural schematic diagram of a model training device provided in an embodiment of the present application;

[0019] Figure 5 is a structural schematic diagram of an image segmentation device provided in an embodiment of the present application;

[0020] Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0021] Figure 7 It is a schematic diagram of the hardware structure of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0024] The model training method and device provided in the embodiments of the present application are described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0025] Please refer to Figure 1 , which shows one of the flow charts of the model training method provided in the embodiment of the present application. The model training method provided in the embodiment of the present application can be applied to electronic devices. In practice, the above-mentioned electronic devices can be electronic devices such as smart phones, tablet computers, laptops, and wearable devices.

[0026] The process of the model training method provided in the embodiment of the present application includes the following steps:

[0027] Step 101: Obtain a sample image, prompt information, and annotation information. The prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image.

[0028] In this embodiment, image segmentation is a technique and process for dividing an image into several specific regions with unique properties and extracting objects of interest. The image segmentation process is also a labeling process, that is, assigning the same label to pixels belonging to the same region.

[0029] In this embodiment, a sample set may be stored in advance. The sample set may include a large number of sample images. Each sample image may have corresponding annotation information. The annotation information is used to indicate the target segmentation result of the sample image, which is an accurate segmentation result, and may include a label for each pixel in the sample image. Pixels in areas where different objects are located may have different labels, and pixels in areas where the same object is located may have the same label. For example, if a sample image includes the sky and grass, the pixels in the sky area of ​​the sample image may be marked as 0, and the pixels in the grass area may be marked as 1. The annotation information may be represented in the form of a mask image. In practice, the annotation information may be set by a technician or generated by a professional annotation tool.

[0030] In this embodiment, the acquisition method and the number of sample images are not limited. Specifically, sample images can be randomly extracted from the sample set, or sample images with better clarity can be extracted from the sample set, or sample images can be extracted from the sample set in a specified order. The sample images obtained can be one or more sample images in the sample set. For example, there can be B, where B is a pre-set batch size and B is a positive integer. The sample image input to the encoding network can be recorded as [B, C i ,H i ,W i ]. Among them, C i H is the number of image channels input to the network, usually 3 channels; i is the height of the sample image, W i is the width of the sample image.

[0031] In this embodiment, the prompt information can also be called a prompt word (Prompt), which can be used to indicate the target area in the sample image. The target area is the area containing the object to be segmented, that is, the region of interest. The prompt information can be used as prior knowledge to guide the model to focus on a specific area or target. The prompt information can be recorded as [B, C p ]. Among them, C p To indicate the number of channels, take a box as an example, it is generally 4 channels. As an example, the prompt information can be used to indicate the position of a rectangular box in an image, which can be [x1, y1, x2, y2]. Among them, [x1, y1] is the coordinate of the upper left vertex of the rectangular box. [x2, y2] is the coordinate of the lower right vertex of the rectangular box.

[0032] In this embodiment, during the model training process, the prompt information can be obtained by fine-tuning the annotation information. For example, the perturbed bounding box of the rectangular box indicated by the annotation information can be used as the prompt information. In addition, it can also be manually input by a technician, for example, manually marking points, boxes, masks, etc. of the target area in the sample image.

[0033] By obtaining prompt information, interactive image segmentation can be achieved. Through interactive segmentation, users can provide prior input to participate in the segmentation process, improving the accuracy and customizability of image segmentation. Compared with traditional segmentation methods, interactive segmentation makes full use of the user's subjective prior knowledge. Through the prompt information or marks entered by the user, the algorithm can better understand the image and generate segmentation results that are more in line with the user's expectations.

[0034] Step 102: input the sample image and the prompt information into the encoding network of the first model to obtain the fusion features of the sample image and the prompt information output by the encoding network.

[0035] In this embodiment, the first model may include an encoding network (encoder). The encoding network may be used to perform operations such as encoding, feature extraction, and feature fusion on the sample image and the prompt information. The encoding network may extract image features and semantic features of the sample image, and may also extract prompt features of the prompt information. The encoding network may also fuse the extracted image features, semantic features, and prompt features during or after the feature extraction process to obtain a fused feature.

[0036] In practice, the encoding network can have multiple encoders and fusion modules. The above-mentioned image features, semantic features and prompt features can be extracted by different encoders and fused by the fusion module. Each encoder can adopt a lightweight network structure, for example, a lightweight CNN (Convolutional Neural Network), to save computing resources.

[0037] In addition, each encoder may include multiple encoding modules to output multi-scale image features, semantic features, and prompt features. Similarly, after feature fusion, multi-scale fusion features can be obtained. This can improve the richness of features and thus improve the accuracy of image segmentation.

[0038] By integrating the prompt features of the prompt information during the encoding process, the model can obtain the prompt information input by the user during encoding, so that the prompt information can fully participate in the processing process of the network structure of each part of the model, thereby improving the model's overall understanding of semantics and thus improving the accuracy of image segmentation.

[0039] Step 103: input the fused features into the decoding network of the first model to obtain the actual segmentation result of the sample image output by the decoding network.

[0040] In this embodiment, the first model may further include a decoding network (decoder). The decoding network may be constructed based on a lightweight CNN structure. The decoding network may decode the fused features by deconvolution processing and the like, and finally output an actual segmentation result. The actual segmentation result is the image segmentation result actually output by the decoding network in the first model, including the actual labeling of each pixel in the sample image. The actual segmentation result may be represented in the form of a mask image. During the model training phase, the actual segmentation result is usually different from the target segmentation result. The purpose of model training is to make the actual segmentation result close to or equal to the target segmentation result.

[0041] In practice, the decoding network may also include multiple decoding modules to decode the multi-scale fusion features. Each decoding module may include a set number of deconvolution layers to decode the features and gradually fuse the low-resolution features with the high-resolution features. The decoding network may eventually output a mask image of a set size as the actual segmentation result. For example, a mask image of 1 / 4 the size of the sample image may be output.

[0042] By using a lightweight network structure to build a decoding network, the decoder of the structure with high computational complexity such as visual converter commonly used in interactive segmentation is abandoned, which reduces the computational complexity and improves the running efficiency of the second model on mobile devices such as mobile phones, tablets, and wearable devices. At the same time, it can output more refined segmentation mask edges.

[0043] Step 104: Based on the actual segmentation results and the annotation information, the encoding network and the decoding network are trained to obtain a second model.

[0044] In this embodiment, the loss value of the first model can be calculated based on the actual segmentation result and annotation information output by the decoding network. The above loss value can be calculated by a loss function. The loss function is a non-negative real-valued function that can be used to characterize the difference between the detection result and the true result. In general, the smaller the loss value, the better the robustness of the model. The loss function can be set according to actual needs. For example, it can be set to a cross entropy loss function. After calculating the loss value of the first model, the back propagation algorithm can be used to obtain the gradient of the loss value relative to the model parameters, and then the gradient descent algorithm can be used to update the parameters of the encoding network and the decoding network based on the above gradient to achieve one-time training of the first model.

[0045] By iteratively executing the above steps 101 to 104, the parameters of the encoding network and the decoding network in the first model can be gradually optimized. When the loss value of the first model tends to be stable, it can be determined that the model training is completed. At this time, the second model can be obtained.

[0046] The method provided by the above-mentioned embodiment of the present application first obtains a sample image, prompt information and annotation information, wherein the prompt information is used to indicate the target area in the sample image, and the annotation information is used to indicate the target segmentation result of the sample image; then the sample image and the prompt information are input into the encoding network to obtain the fusion feature of the sample image and the prompt information output by the encoding network; then the fusion feature is input into the decoding network to obtain the actual segmentation result of the sample image output by the decoding network; finally, based on the actual segmentation result and the annotation information, the encoding network and the decoding network are trained to obtain the second model. During the model training process, the encoding network encodes, extracts features and fuses features for the prompt information, which, on the one hand, can reduce the complexity of the decoding network, thereby saving computing resources. On the other hand, it can enable the prompt information to fully participate in the processing process of the network structure of each part of the model, improve the overall understanding ability of the model for semantics, and thus improve the accuracy of image segmentation.

[0047] In some optional embodiments, the encoding network may include a first encoder, a second encoder, and a third encoder. The first encoder may be used to perform fine-grained encoding of a sample image to extract image features. The second encoder may be used to perform fine-grained semantic encoding of a sample image to extract semantic features. The third encoder may be used to encode prompt information to extract prompt features. In practice, the first encoder, the second encoder, and the third encoder may all be constructed based on CNN. As an example, a lightweight network architecture such as RepVGG (Re-parameterization Visual Geometry Group, re-parameterized visual geometry group network) may be used. By using a lightweight first encoder, a second encoder, and a third encoder to extract image features, semantic features, and prompt features, respectively, the computational complexity may be reduced and the operating efficiency of the model on mobile devices such as mobile phones, tablet computers, and wearable devices may be improved.

[0048] In some optional embodiments, the model structure can be found in Figure 2A . Based on this structure, the fusion features can be obtained according to the following steps: input the sample image into the first encoder to obtain the image features output by the first encoder; input the prompt information into the third encoder to obtain the prompt features output by the third encoder; input the sample image and the prompt features into the second encoder to obtain the semantic features output by the second encoder; fuse the image features and the semantic features to obtain the fusion features of the sample image and the prompt information. In the above process, by fusing the prompt features of the prompt information in the process of extracting the semantic features of the sample image, the prompt information can be fully involved in the processing process of the network structure of each part of the model, thereby improving the overall understanding ability of the model on semantics, thereby improving the accuracy of image segmentation.

[0049] In some optional embodiments, see Figure 2A , the first encoder may further include N first encoding modules. The image features may further include image sub-features output by each of the N first encoding modules, where N is a positive integer. The value of N can be preset as needed, for example, it can be set to 4. The second encoder may further include a first initial encoding module and N second encoding modules corresponding one-to-one to the N first encoding modules. The semantic features may further include semantic sub-features output by each of the N second encoding modules. Each encoding module corresponds to a feature extraction stage. Each encoding module may include at least one convolutional layer for feature extraction. For example, each encoding module may include two convolutional layers. It should be noted that the number of convolutional layers included in different encoding modules may be the same or different, and is not specifically limited here. In addition, in addition to the encoding modules listed above, the second encoder may also include a first initial encoding module to perform preliminary feature encoding before the N second encoding modules.

[0050] In the sample image [B,C i ,H i ,W i ] is input to the first encoder, the first first encoding module can output an image sub-feature with a resolution of 1 / 8 of the original image, and the resolution of the image sub-feature output by each subsequent first encoding module can be 1 / 2 of the resolution of the image sub-feature output by the previous first encoding module. For example, if N=4, then 4 image sub-features can be output in the end, and the sizes of the 4 image sub-features can be [B, C f1 ,H i / 8,W i / 8]、[B,C f2 ,H i / 16,W i / 16]、[B,C f3 ,H i / 32,W i / 32] and [B,C f4 ,H i / 64,W i / 64] where C f1 , C f2 , C f3 , C f4 The number of channels of image sub-features encoded at different stages that can be predefined.

[0051] When the prompt information [B,C p] is input to the third encoder, the third encoder can first initialize a single-channel matrix with the same length and width as the sample image and all element values ​​are 0. Each element in the matrix corresponds to a pixel in the sample image. Then, according to the target area indicated by the prompt information, the elements corresponding to the target area and all the pixels in the target area are updated to 1, and the elements corresponding to the pixels in the remaining areas are set to 0, so as to obtain the matrix [1,H i ,W i Then, batch expansion is performed to obtain the expanded matrix [B,1,H i ,W i After that, two convolutional layers can be used to transform the matrix [B,1,H i ,W i ] is encoded, and the output size is [B,C p ,H p ,W p ] prompt features. Among them, C p is the number of channels of the encoded prompt feature, H p is the height of the encoded prompt feature, W p is the width of the encoded hint feature.

[0052] In the sample image [B,C i ,H i ,W i ] is input to the second encoder, the second encoder can firstly encode [B,C i ,H i ,W i ] for semantic encoding, the output size is [B,C s0 ,H i / 64,W i 64] of the semantic sub-feature. Then, the semantic sub-feature is combined with the output of the third encoder to form an output shape of [B,C p ,H p ,W p ] are fused to obtain the size of [B,C s0 ,H i / 64,W i 64] after fusion. Among them, C p =C s0 , H p =H i / 64,W p =W i / 64. The two can be directly added together to obtain the fused feature. After that, the fused feature is input into the first second encoding module of the N second encoding modules. The N second encoding modules perform feature processing in sequence to obtain N semantic sub-features output by the N second encoding modules. For example, if N = 4, the sizes of the N semantic sub-features are [B, C s1 ,H i / 64,W i 64], [B,C s2 ,H i / 64,W i 64], [B,C s3 ,H i / 64,W i 64] and [B,C s4 ,H i / 64,W i 64]. Among them, C s0 , C s1 , C s2 , C s3 , C s4 is the number of channels of semantic sub-features encoded at different stages that can be predefined. In practice, the second encoder can be pre-trained in advance so that it has stronger semantic feature extraction ability during initialization.

[0053] In the feature fusion process, for each of the N first encoding modules, the image sub-feature output by the first encoding module may be fused with the semantic sub-feature output by the second encoding module corresponding to the first encoding module in the N second encoding modules to obtain a fused sub-feature. Then, the obtained N fused sub-features are summarized to obtain a fused feature.

[0054] In practice, the encoding network may further include a fusion module for feature fusion. The fusion module may include N fusion submodules. The above-mentioned N fusion submodules correspond one-to-one to the N first encoding modules, and also correspond one-to-one to the N second encoding modules, to receive the corresponding image subfeatures and semantic subfeatures and fuse the two. Specifically, each fusion submodule may first resize the semantic subfeatures input to the fusion submodule, and then fuse them with the image subfeatures input to the fusion submodule. Continuing with the above example, if N=4, the first fusion submodule can convert the size of the output of the first second encoding module to [B,C s1 ,H i / 64,W i 64] to the size [B,C f1 ,H i / 8,W i / 8], the size of the semantic sub-feature after size conversion is combined with the size of the output of the first encoding module [B,C f1 ,H i / 8,W i / 8] image sub-features are stacked to obtain a size of [B,2C f1 ,H i / 8,W i / 8] features. Then, two convolutional layers are used to transform the above-mentioned image with size [B,2C f1 ,H i / 8,W i / 8] features are convolved to obtain a size of [B,C f1 ,H i / 8,W i / 8], which has the same size as the image sub-feature output by the first encoding module. Similarly, the size of the fusion sub-feature output by the second fusion sub-module is [B,C f2 ,H i / 16,W i / 16], the size of the fusion sub-feature output by the third fusion sub-module is [B,C f3 ,H i / 32,W i / 32], the size of the fusion sub-feature output by the fourth fusion sub-module is [B,C f4 ,H i / 64,W i / 64]. The above four fusion sub-features are aggregated to obtain the fusion feature.

[0055] By integrating the hint features of the hint information in the process of extracting semantic features from sample images, the hint information can be fully involved in the processing of the network structures of various parts of the model, thus improving the overall understanding of the model on semantics and thus improving the accuracy of image segmentation.

[0056] In some optional embodiments, the model structure can be found in Figure 2B . Based on this structure, the fusion features can be obtained according to the following steps: input the sample image to the first encoder and the second encoder respectively, and obtain the image features output by the first encoder and the semantic features output by the second encoder; input the prompt information to the third encoder, and obtain the prompt features output by the third encoder; fuse the image features, semantic features and prompt features to obtain the fusion features of the sample image and the prompt information. In the above process, by simultaneously fusing the image features, semantic features and prompt features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0057] In some optional embodiments, see Figure 2B , the first encoder may further include N first encoding modules. The image features may further include image sub-features output by each of the N first encoding modules, where N is a positive integer. The second encoder may further include N second encoding modules corresponding one-to-one to the N first encoding modules. The semantic features may further include semantic sub-features output by each of the N second encoding modules. The third encoder may further include N third encoding modules corresponding one-to-one to the N first encoding modules to perform fine-grained encoding on the prompt information. The prompt features may further include prompt sub-features output by each of the N third encoding modules.

[0058] Among them, each encoding module corresponds to a feature extraction stage. Each encoding module may include at least one convolutional layer for feature extraction. For example, each encoding module may include two convolutional layers. It should be noted that the number of convolutional layers included in different encoding modules may be the same or different, and is not specifically limited here. In addition, in addition to the encoding modules listed above, the second encoder may also include a first initial encoding module to perform preliminary feature encoding before the N second encoding modules. Similarly, the third encoder may also include a second initial encoding module to perform preliminary feature encoding before the N third encoding modules.

[0059] In the sample image [B,C i ,H i ,W i ] is input to the first encoder, the first first encoding module in the first encoder can output an image sub-feature with a resolution of 1 / 8 of the original image, and the resolution of the image sub-feature output by each subsequent first encoding module can be 1 / 2 of the resolution of the image sub-feature output by the previous first encoding module. For example, if N=4, then 4 image sub-features can be output in the end, and the sizes of the 4 image sub-features are [B, C f1 ,H i / 8,W i / 8]、[B,C f2 ,H i / 16,W i / 16]、[B,C f3 ,H i / 32,W i / 32] and [B,C f4 ,H i / 64,W i / 64] where C f1 , C f2 , C f3 , C f4 The number of channels of image sub-features encoded at different stages that can be predefined.

[0060] In the sample image [B,C i ,H i ,W i ] is input to the second encoder, the second encoder can firstly encode [B,C i ,H i ,W i ] for semantic encoding, the output size is [B,C s0 ,H i / 64,W i 64]. Then, the semantic sub-feature is input into the first second encoding module of the N second encoding modules, and the N second encoding modules perform feature processing in sequence, thereby obtaining N semantic sub-features output by the N second encoding modules. For example, if N = 4, the sizes of the four semantic sub-features are [B, C s1 ,H i / 64,W i 64], [B,C s2 ,H i / 64,W i 64], [B,C s3 ,H i / 64,W i 64] and [B,C s4 ,H i / 64,W i 64]. Among them, C s0 , C s1 , C s2 , C s3 , C s4 is the number of channels of semantic sub-features encoded at different stages that can be predefined. In practice, the second encoder can be pre-trained in advance so that it has stronger semantic feature extraction ability during initialization.

[0061] When the prompt information [B,C p ] is input to the third encoder, the third encoder can first initialize a single-channel matrix with the same length and width as the sample image and all element values ​​are 0 through the second initial encoding module. Each element in the matrix corresponds to a pixel in the sample image. Then, according to the target area indicated by the prompt information, the elements corresponding to the target area and all the pixels in the target area are updated to 1, and the elements corresponding to the pixels in the remaining areas are set to 0, so as to obtain the matrix [1,H i ,W i Then, batch expansion is performed to obtain the expanded matrix [B,1,H i ,W i After that, two convolutional layers can be used to transform the matrix [B,1,Hi ,W i ] is encoded, and the output size is [B,C p ,H p ,W p ] prompt sub-feature. Then, the prompt sub-feature is input to the first third encoding module of the subsequent N third encoding modules, and the N third encoding modules perform feature processing in sequence to obtain N prompt sub-features output by the N third encoding modules. For example, if N=4, the sizes of the N prompt sub-features are [B, C f1 ,H i / 8,W i 8], [B,C f2 ,H i / 16,W i 16], [B,C f3 ,H i / 32,W i 32] and [B,C f4 ,H i / 64,W i 64].

[0062] When performing feature fusion, for each of the N first encoding modules, the image sub-feature output by the first encoding module, the semantic sub-feature output by the second encoding module corresponding to the first encoding module in the N second encoding modules, and the prompt sub-feature output by the third encoding module corresponding to the first encoding module in the N third encoding modules may be fused to obtain a fused sub-feature. Then, the obtained N fused sub-features are summarized to obtain a fused feature.

[0063] In practice, the encoding network may further include a fusion module for feature fusion. The fusion module may include N fusion submodules. The N fusion submodules correspond one-to-one to the N first encoding modules, one-to-one to the N second encoding modules, and one-to-one to the N third encoding modules, to receive corresponding image subfeatures, semantic subfeatures, and prompt subfeatures. Each fusion submodule may resize the semantic subfeatures input to the fusion submodule, and then fuse them with the image subfeatures and prompt subfeatures input to the fusion submodule.

[0064] Continuing with the above example, if N = 4, the first fusion submodule can pass through two layers of deconvolution layers to convert the size of the output of the first second encoding module to [B, C s1 ,H i / 64,W i 64] to the size [B,C f1 ,H i / 8,W i / 8], the size of the semantic sub-feature after size conversion and the output of the first encoding module is [B,C f1 ,H i / 8,W i / 8], and the size of the output of the first third encoding module is [B,C f1 ,H i / 8,W i 8] are stacked to obtain a size of [B,3C f1 ,H i / 8,W i / 8] features. Then, two convolutional layers are used to transform the above-mentioned image with size [B,2C f1 ,H i / 8,W i / 8] features are convolved to obtain a size of [B,C f1 ,H i / 8,W i / 8] fusion sub-feature. Similarly, the size of the fusion sub-feature output by the second fusion sub-module is [B,C f2 ,H i / 16,W i / 16], the size of the fusion sub-feature output by the third fusion sub-module is [B,C f3 ,H i / 32,W i / 32], the size of the fusion sub-feature output by the fourth fusion sub-module is [B,C f4 ,H i / 64,W i / 64]. The above four fusion sub-features are aggregated to obtain the fusion feature.

[0065] By simultaneously fusing fine-grained image sub-features, semantic sub-features, and cue sub-features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0066] In some optional embodiments, the decoding network may include N decoding modules, and the N decoding modules may correspond one-to-one to the N fused sub-features. For each fused sub-feature in the N fused sub-features, the fused sub-feature may be input into the decoding module corresponding to the fused sub-feature in the N decoding modules to obtain the actual segmentation result of the sample image output by the decoding network.

[0067] Each decoding module of the decoding network may correspond to a decoding stage. Each decoding module may include two deconvolution layers to decode the fused sub-features input to the decoding module, and gradually fuse the low-resolution fused sub-features with the high-resolution fused sub-features, and finally output a mask image with a size of 1 / 4 of the sample image.

[0068] As an example, if N = 4, the fourth decoding module can first pass through two deconvolution layers to convert the output of the fourth fusion submodule into a size of [B, C f4 ,H i / 64,W i / 64] fusion sub-features are deconvolved, and the output size is [B,C f3 ,H i / 32,W i / 32] features and inputs them to the third decoding module. The third decoding module receives the inputs of the fourth decoding module and the third fusion submodule at the same time, stacks the two in the channel dimension, and obtains a size of [B,2C f3 ,H i / 32,W i / 32] features. Then, two deconvolution layers are used to transform the image of size [B,2C f3 ,H i / 32,W i / 32] features are deconvolved, and the output size is [B,C f2 ,H i / 16,W i / 16] features and inputs them to the second decoding module. Similarly, the second decoding module simultaneously receives the inputs of the third decoding module and the second fusion submodule, and stacks them in the channel dimension to obtain a size of [B,2C f2 ,H i / 16,W i / 16] features. Then, two deconvolution layers are used to transform the image of size [B,C f2 ,H i / 16,W i / 16] features are deconvolved, and the output size is [B,C f1 ,H i / 8,W i / 8] features and inputs them to the first decoding module. Similarly, the first decoding module simultaneously receives the inputs of the second decoding module and the first fusion submodule, and stacks them in the channel dimension to obtain a size of [B,2C f1 ,H i / 8,W i / 8] features. Then, two deconvolution layers are used to transform the image of size [B,2Cf1 ,H i / 8,W i / 8] features are deconvolved, and the output size is [B,1,H i / 4,W i / 4] as the actual segmentation result.

[0069] By using a lightweight network structure to build a decoding network and designing a lightweight CNN module to participate in the decoding process, the decoder with a high computational complexity of the visual converter commonly used in interactive segmentation is abandoned, which reduces the computational complexity and improves the running efficiency of the model on mobile devices such as mobile phones, tablets, and wearable devices. At the same time, it can output more refined segmentation mask edges.

[0070] Please refer to Figure 3 , which shows one of the flow charts of the image segmentation method provided in the embodiment of the present application. The image segmentation method provided in the embodiment of the present application can be applied to electronic devices. In practice, the above electronic devices can be electronic devices such as smart phones, tablet computers, laptops, and wearable devices.

[0071] The process of the image segmentation method provided in the embodiment of the present application includes the following steps:

[0072] Step 301: Acquire an image to be segmented and prompt information, where the prompt information is used to indicate a target area in the image to be segmented.

[0073] In this embodiment, the prompt information can be input by the user. In practice, the user can input the prompt information in a variety of ways, including but not limited to manually marking points, boxes, masks, etc.

[0074] Step 302: input the image to be segmented and the prompt information into the second model to obtain the actual segmentation result.

[0075] The second model in this embodiment can be trained using the model training method in any of the above embodiments, which will not be repeated here.

[0076] Due to the adoption Figure 1 The second model trained in the corresponding embodiment has lower computational complexity and better image segmentation effect. Therefore, the image segmentation method based on the above-mentioned second model can be applied to electronic devices of mobile terminals, and the actual segmentation results obtained can be more accurate.

[0077] It should be noted that the image segmentation method of this embodiment can be used to test the second model trained in the above embodiment, and then the second model can be continuously optimized according to the test results. The image segmentation method of this embodiment can also be a practical application method of the second model trained in the above embodiment. Using the second model trained in the above embodiment, image segmentation can be performed to obtain high-quality actual segmentation results.

[0078] It should be noted that the model training method provided in the embodiment of the present application can be executed by a model training device. In the embodiment of the present application, the model training device executing the model training method is taken as an example to illustrate the model training device provided in the embodiment of the present application.

[0079] like Figure 4 As shown, the model training device 400 of this embodiment includes: an acquisition unit 401, used to acquire a sample image, prompt information and annotation information, wherein the prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image; an encoding unit 402, used to input the sample image and the prompt information into an encoding network of a first model, and obtain a fusion feature of the sample image and the prompt information output by the encoding network; a decoding unit 403, used to input the fusion feature into a decoding network of the first model, and obtain an actual segmentation result of the sample image output by the decoding network; a training unit 404, used to train the encoding network and the decoding network based on the actual segmentation result and the annotation information, and obtain a second model.

[0080] In some optional implementations of this embodiment, the encoding network includes a first encoder, a second encoder and a third encoder; the encoding unit 402 is further used to: input the sample image to the first encoder to obtain the image features output by the first encoder; input the prompt information to the third encoder to obtain the prompt features output by the third encoder; input the sample image and the prompt features to the second encoder to obtain the semantic features output by the second encoder; fuse the image features and the semantic features to obtain the fusion features of the sample image and the prompt information. By using the lightweight first encoder, the second encoder and the third encoder to extract the image features, the semantic features and the prompt features respectively, the computational complexity can be reduced and the running efficiency of the second model on mobile devices such as mobile phones, tablet computers and wearable devices can be improved. In addition, by fusing the prompt features of the prompt information in the process of extracting the semantic features of the sample image, the prompt information can be fully involved in the processing process of the network structure of each part of the second model, thereby improving the overall understanding ability of the second model for semantics, thereby improving the accuracy of image segmentation.

[0081] In some optional implementations of this embodiment, the first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules, and the semantic features include semantic sub-features output by each of the N second encoding modules; the encoding unit 402 is further used to: for each of the N first encoding modules, fuse the image sub-features output by the first encoding module with the semantic sub-features output by the second encoding module corresponding to the first encoding module in the N second encoding modules to obtain a fused sub-feature; and summarize the obtained N fused sub-features to obtain a fused feature. By fusing the prompt features of the prompt information in the process of extracting the semantic features of the sample image, the prompt information can be fully involved in the processing process of the network structure of each part of the model, thereby improving the overall understanding ability of the model for semantics, thereby improving the accuracy of image segmentation.

[0082] In some optional implementations of this embodiment, the encoding network includes a first encoder, a second encoder, and a third encoder; the encoding unit 402 is further used to: input the sample image to the first encoder and the second encoder respectively, obtain the image features output by the first encoder and the semantic features output by the second encoder; input the prompt information to the third encoder, obtain the prompt features output by the third encoder; fuse the image features, the semantic features, and the prompt features to obtain the fusion features of the sample image and the prompt information. By simultaneously fusing image features, semantic features, and prompt features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0083] In some optional implementations of this embodiment, the first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one-to-one, the semantic features include semantic sub-features output by each of the N second encoding modules; the third encoder includes N third encoding modules corresponding to the N first encoding modules one-to-one, and the prompt features include prompt sub-features output by each of the N third encoding modules; the encoding unit 402 is further used to: for each first encoding module in the N first encoding modules, fuse the image sub-features output by the first encoding module, the semantic sub-features output by the second encoding module corresponding to the first encoding module in the N second encoding modules, and the prompt sub-features output by the third encoding module corresponding to the first encoding module in the N third encoding modules to obtain a fused sub-feature; and summarize the obtained N fused sub-features to obtain a fused feature. By simultaneously fusing fine-grained image sub-features, semantic sub-features, and cue sub-features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0084] In some optional implementations of this embodiment, the decoding network includes N decoding modules, and the N decoding modules correspond one to one with the N fusion sub-features; the decoding unit 403 is also used to: for each fusion sub-feature of the N fusion sub-features, input the fusion sub-feature into the decoding module corresponding to the fusion sub-feature in the N decoding modules, and obtain the actual segmentation result of the sample image output by the decoding network. By using a lightweight network structure to construct a decoding network, a lightweight CNN module is designed to participate in the decoding process, and the Vision Transformer decoder with high computational complexity commonly used in interactive segmentation is abandoned, the computational complexity is reduced, and the running efficiency of the model on mobile devices such as mobile phones, tablets, and wearable devices is improved. At the same time, it can output finer segmentation mask edges.

[0085] The device provided by the above-mentioned embodiment of the present application first obtains a sample image, prompt information and annotation information, wherein the prompt information is used to indicate the target area in the sample image, and the annotation information is used to indicate the target segmentation result of the sample image; then the sample image and the prompt information are input into the encoding network to obtain the fusion feature of the sample image and the prompt information output by the encoding network; then the fusion feature is input into the decoding network to obtain the actual segmentation result of the sample image output by the decoding network; finally, based on the actual segmentation result and the annotation information, the encoding network and the decoding network are trained to obtain the second model. During the model training process, the encoding network encodes, extracts features and fuses features for the prompt information, which, on the one hand, can reduce the complexity of the decoding network, thereby saving computing resources. On the other hand, it can enable the prompt information to fully participate in the processing process of the network structure of each part of the model, improve the overall understanding ability of the model for semantics, and thus improve the accuracy of image segmentation.

[0086] like Figure 5 As shown, the image segmentation device 500 of this embodiment includes: an acquisition unit 501, used to acquire the image to be segmented and prompt information, wherein the prompt information is used to indicate the target area in the image to be segmented; an image segmentation unit 502, used to input the image to be segmented and the prompt information into the second model to obtain the actual segmentation result.

[0087] It is understandable that the units recorded in the device 500 correspond to the steps in the above-mentioned image segmentation method. Therefore, the operations, features and beneficial effects described above for the image segmentation method are also applicable to the device 500 and the units contained therein, and will not be repeated here.

[0088] The model training device in the embodiment of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than a terminal. Exemplary, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, a super mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., and the embodiment of the present application is not specifically limited.

[0089] The model training device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0090] The model training device provided in the embodiment of the present application can achieve Figure 1 or Figure 3 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0091] Alternatively, if Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 601 and a memory 602, wherein the memory 602 stores programs or instructions that can be executed on the processor 601, and when the program or instructions are executed by the processor 601, the various steps of the above-mentioned model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.

[0092] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0093] Figure 7 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of the present application.

[0094] The electronic device 700 includes but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710.

[0095] Those skilled in the art will appreciate that the electronic device 700 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 710 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Figure 7 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.

[0096] Among them, the processor 710 is used to obtain a sample image, prompt information and annotation information, wherein the prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image; the sample image and the prompt information are input into the encoding network of the first model to obtain a fusion feature of the sample image and the prompt information output by the encoding network; the fusion feature is input into the decoding network of the first model to obtain an actual segmentation result of the sample image output by the decoding network; based on the actual segmentation result and the annotation information, the encoding network and the decoding network are trained to obtain a second model.

[0097] During the model training process, the coding network encodes, extracts and fuses the prompt information. On the one hand, it can reduce the complexity of the decoding network, thereby saving computing resources. On the other hand, it can enable the prompt information to fully participate in the processing process of the network structure of each part of the model, improve the model's overall understanding of semantics, and thus improve the accuracy of image segmentation.

[0098] Optionally, the encoding network includes a first encoder, a second encoder, and a third encoder; the processor 710 is further used to input the sample image into the first encoder to obtain the image features output by the first encoder; input the prompt information into the third encoder to obtain the prompt features output by the third encoder; input the sample image and the prompt features into the second encoder to obtain the semantic features output by the second encoder; fuse the image features and the semantic features to obtain the fusion features of the sample image and the prompt information. By using the lightweight first encoder, the second encoder, and the third encoder to extract image features, semantic features, and prompt features respectively, the computational complexity can be reduced and the running efficiency of the model on mobile devices such as mobile phones, tablet computers, and wearable devices can be improved.

[0099] Optionally, the first encoder includes N first encoding modules, the image features include image sub-features output by each first encoding module in the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one by one, and the semantic features include semantic sub-features output by each second encoding module in the N second encoding modules; the processor 710 is also used to fuse the image sub-features output by each first encoding module in the N first encoding modules with the semantic sub-features output by the second encoding module corresponding to the first encoding module in the N second encoding modules to obtain fused sub-features; and the obtained N fused sub-features are summarized to obtain fused features. By fusing the prompt features of the prompt information in the process of extracting the semantic features of the sample image, the prompt information can be fully involved in the processing process of the network structure of each part of the model, thereby improving the overall understanding ability of the model for semantics, thereby improving the accuracy of image segmentation. In addition, by simultaneously fusing image features, semantic features and prompt features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0100] Optionally, the encoding network includes a first encoder, a second encoder, and a third encoder; the processor 710 is further used to input the sample image to the first encoder and the second encoder respectively, obtain the image features output by the first encoder and the semantic features output by the second encoder; input the prompt information to the third encoder, obtain the prompt features output by the third encoder; fuse the image features, the semantic features, and the prompt features to obtain the fusion features of the sample image and the prompt information. By fusing image features, semantic features, and prompt features at the same time, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0101] Optionally, the first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding one-to-one to the N first encoding modules, and the semantic features include semantic sub-features output by each of the N second encoding modules; the third encoder includes N third encoding modules corresponding one-to-one to the N first encoding modules, and the prompt features include prompt sub-features output by each of the N third encoding modules; the processor 710 is also used to, for each first encoding module in the N first encoding modules, fuse the image sub-features output by the first encoding module, the semantic sub-features output by the second encoding module corresponding to the first encoding module in the N second encoding modules, and the prompt sub-features output by the third encoding module corresponding to the first encoding module in the N third encoding modules to obtain a fused sub-feature; and summarize the obtained N fused sub-features to obtain a fused feature. By simultaneously fusing fine-grained image sub-features, semantic sub-features, and cue sub-features, the information interaction between different encoding networks can be further improved, which helps to capture finer segmentation edges and output more accurate actual segmentation results.

[0102] Optionally, the decoding network includes N decoding modules, and the N decoding modules correspond one to one with the N fused sub-features; the processor 710 is also used to input each fused sub-feature of the N fused sub-features into the decoding module corresponding to the fused sub-feature in the N decoding modules, so as to obtain the actual segmentation result of the sample image output by the decoding network. By using a lightweight network structure to construct a decoding network, a lightweight CNN module is designed to participate in the decoding process, and the VisionTransformer decoder with high computational complexity commonly used in interactive segmentation is abandoned, the computational complexity is reduced, and the running efficiency of the model on mobile devices such as mobile phones, tablets, and wearable devices is improved. At the same time, it can output finer segmentation mask edges.

[0103] In addition, the processor 710 can also be used to obtain the image to be segmented and prompt information, where the prompt information is used to indicate the target area in the image to be segmented; input the image to be segmented and the prompt information into the second model to obtain an actual segmentation result.

[0104] It should be understood that in the embodiment of the present application, the input unit 704 may include a graphics processor (Graphics Processing Unit, GPU) 7041 and a microphone 7042, and the graphics processor 7041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes a touch panel 7071 and at least one of other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0105] The memory 709 can be used to store software programs and various data. The memory 709 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 709 may include a volatile memory or a non-volatile memory, or the memory 709 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 709 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0106] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 710.

[0107] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned model training method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0108] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0109] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned model training method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0110] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0111] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0112] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0113] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0114] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A model training method, characterized in that: The method comprises: Acquire a sample image, prompt information, and annotation information, wherein the prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image; Inputting the sample image and the prompt information into the encoding network of the first model to obtain a fusion feature of the sample image and the prompt information output by the encoding network; Inputting the fused features into a decoding network of the first model to obtain an actual segmentation result of the sample image output by the decoding network; Based on the actual segmentation result and the annotation information, the encoding network and the decoding network are trained to obtain a second model.

2. The method according to claim 1, characterized in that: The encoding network includes a first encoder, a second encoder, and a third encoder; the sample image and the prompt information are input into the encoding network of the first model, and a fusion feature of the sample image and the prompt information output by the encoding network is obtained, including: Inputting the sample image into the first encoder to obtain image features output by the first encoder; Inputting the prompt information into the third encoder to obtain the prompt feature output by the third encoder; Inputting the sample image and the prompt feature into the second encoder to obtain the semantic feature output by the second encoder; The image feature and the semantic feature are fused to obtain a fusion feature of the sample image and the prompt information.

3. The method according to claim 2, characterized in that The first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one by one, and the semantic features include semantic sub-features output by each of the N second encoding modules; the image features and the semantic features are fused to obtain fused features of the sample image and the prompt information, including: For each first encoding module among the N first encoding modules, fusing the image sub-feature output by the first encoding module with the semantic sub-feature output by a second encoding module among the N second encoding modules corresponding to the first encoding module to obtain a fused sub-feature; The obtained N fusion sub-features are aggregated to obtain a fusion feature.

4. The method according to claim 1, characterized in that: The encoding network includes a first encoder, a second encoder, and a third encoder; the sample image and the prompt information are input into the encoding network of the first model, and a fusion feature of the sample image and the prompt information output by the encoding network is obtained, including: Inputting the sample image into the first encoder and the second encoder respectively, obtaining image features output by the first encoder and semantic features output by the second encoder; Inputting the prompt information into the third encoder to obtain the prompt feature output by the third encoder; The image feature, the semantic feature and the prompt feature are fused to obtain a fusion feature of the sample image and the prompt information.

5. The method according to claim 4, characterized in that The first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one by one, the semantic features include semantic sub-features output by each of the N second encoding modules; the third encoder includes N third encoding modules corresponding to the N first encoding modules one by one, and the prompt features include prompt sub-features output by each of the N third encoding modules; The fusing the image feature and the semantic feature to obtain a fusion feature of the sample image and the prompt information includes: For each first encoding module in the N first encoding modules, an image sub-feature output by the first encoding module, a semantic sub-feature output by a second encoding module in the N second encoding modules corresponding to the first encoding module, and a prompt sub-feature output by a third encoding module in the N third encoding modules corresponding to the first encoding module are fused to obtain a fused sub-feature; The obtained N fusion sub-features are aggregated to obtain a fusion feature.

6. An image segmentation method, characterized in that: The method comprises: Acquire an image to be segmented and prompt information, wherein the prompt information is used to indicate a target area in the image to be segmented; The image to be segmented and the prompt information are input into a second model trained by the model training method described in any one of claims 1 to 5 to obtain an actual segmentation result.

7. A model training device, characterized in that: The device comprises: An acquisition unit, used to acquire a sample image, prompt information and annotation information, wherein the prompt information is used to indicate a target area in the sample image, and the annotation information is used to indicate a target segmentation result of the sample image; An encoding unit, used for inputting the sample image and the prompt information into an encoding network of a first model, and obtaining a fusion feature of the sample image and the prompt information output by the encoding network; A decoding unit, used for inputting the fused features into a decoding network of the first model to obtain an actual segmentation result of the sample image output by the decoding network; A training unit is used to train the encoding network and the decoding network based on the actual segmentation result and the annotation information to obtain a second model.

8. The device according to claim 7, characterized in that The encoding network includes a first encoder, a second encoder and a third encoder; the encoding unit is further used for: Inputting the sample image into the first encoder to obtain image features output by the first encoder; Inputting the prompt information into the third encoder to obtain the prompt feature output by the third encoder; Inputting the sample image and the prompt feature into the second encoder to obtain the semantic feature output by the second encoder; The image feature and the semantic feature are fused to obtain a fusion feature of the sample image and the prompt information.

9. The device according to claim 8, characterized in that The first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one by one, and the semantic features include semantic sub-features output by each of the N second encoding modules; the encoding unit is further used to: For each first encoding module among the N first encoding modules, fusing the image sub-feature output by the first encoding module with the semantic sub-feature output by a second encoding module among the N second encoding modules corresponding to the first encoding module to obtain a fused sub-feature; The obtained N fusion sub-features are aggregated to obtain a fusion feature.

10. The device according to claim 7, characterized in that The encoding network includes a first encoder, a second encoder and a third encoder; the encoding unit is further used for: Inputting the sample image into the first encoder and the second encoder respectively, obtaining image features output by the first encoder and semantic features output by the second encoder; Inputting the prompt information into the third encoder to obtain the prompt feature output by the third encoder; The image feature, the semantic feature and the prompt feature are fused to obtain a fusion feature of the sample image and the prompt information.

11. The device according to claim 10, characterized in that The first encoder includes N first encoding modules, the image features include image sub-features output by each of the N first encoding modules, and N is a positive integer; the second encoder includes N second encoding modules corresponding to the N first encoding modules one by one, the semantic features include semantic sub-features output by each of the N second encoding modules; the third encoder includes N third encoding modules corresponding to the N first encoding modules one by one, and the prompt features include prompt sub-features output by each of the N third encoding modules; the encoding unit is further used to: For each first encoding module in the N first encoding modules, an image sub-feature output by the first encoding module, a semantic sub-feature output by a second encoding module in the N second encoding modules corresponding to the first encoding module, and a prompt sub-feature output by a third encoding module in the N third encoding modules corresponding to the first encoding module are fused to obtain a fused sub-feature; The obtained N fusion sub-features are aggregated to obtain a fusion feature.

12. An image segmentation device, characterized in that: The device comprises: An acquisition unit, used for acquiring an image to be segmented and prompt information, wherein the prompt information is used for indicating a target area in the image to be segmented; An image segmentation unit is used to input the image to be segmented and the prompt information into a second model trained by the model training method described in one of claims 1 to 5 to obtain an actual segmentation result.