Image segmentation method, apparatus, device, and medium

By acquiring and fusing the coding features of multimodal information, the accuracy problem of medical robot image segmentation is solved, the accuracy of image segmentation is improved and the review cost of planning schemes is reduced.

CN119228822BActive Publication Date: 2025-10-17BEIJING NATONG MEDICAL ROBOT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411243484.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-10-17
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

In existing technologies, medical robots have low accuracy in image segmentation, resulting in inaccurate planning plans, which require manual correction or re-labeling, increasing review costs.

Method used

By obtaining the image position code of the original image, the prompt position code of the user-triggered information, and the description text, the encoder is used to extract features and perform feature fusion, and the decoder is used to generate the target segmentation result, and image segmentation is performed by combining multimodal information.

Benefits of technology

Improved image segmentation accuracy leads to improved planning accuracy and reduced review costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228822B_ABST
    Figure CN119228822B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image segmentation method, device, equipment and medium. Including: acquiring image position encoding of an original image, prompt position encoding of user trigger information on the original image and description text of the original image; using a pre-acquired encoder to process the image position encoding, the prompt position encoding and the description text to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding and text encoding features corresponding to the description text; performing feature fusion processing on the image encoding features, the prompt encoding features and the text encoding features to generate fused encoding features; using a pre-acquired decoder to process the fused encoding features to generate a target segmentation result of the original image. In this way, the multi-modal information of the image is utilized for image segmentation, the accuracy of image segmentation is improved, and the accuracy of the planning scheme is improved and the auditing cost of the planning scheme is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image segmentation, and particularly relates to an image segmentation method, device, equipment and medium. BACKGROUND

[0002] In the field of medical robot control, the robot usually needs to segment the collected image to formulate a planning scheme, and then the planning scheme is audited by manual auditing to ensure the accuracy of the planning scheme and the safety of the work based on the planning scheme.

[0003] At present, the robot usually segments the image only according to the image features of the collected image, and the segmentation accuracy of this segmentation method is low, which leads to inaccurate planning scheme, and then the planning scheme needs to be manually corrected or re-labeled. Therefore, it is necessary to provide a new image segmentation method to improve the segmentation accuracy, and then improve the accuracy of the planning scheme and reduce the auditing cost of the planning scheme. SUMMARY

[0004] In order to solve the above technical problems, the present disclosure provides an image segmentation method, device, equipment and medium.

[0005] In a first aspect, the present disclosure provides an image segmentation method, comprising:

[0006] obtaining image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image;

[0007] using a pre-acquired encoder to process the image position encoding, the prompt position encoding and the description text, to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text;

[0008] performing feature fusion processing on the image encoding features, the prompt encoding features and the text encoding features to generate fusion encoding features;

[0009] using a pre-acquired decoder to process the fusion encoding features to generate a target segmentation result of the original image.

[0010] In a second aspect, the present disclosure provides an image segmentation device, comprising:

[0011] an information acquisition module configured to obtain image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image;

[0012] an encoding module, configured to process the image position encoding, the prompt position encoding, and the description text by using a pre-acquired encoder to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text;

[0013] a feature fusion module, configured to perform feature fusion processing on the image encoding features, the prompt encoding features, and the text encoding features to generate fused encoding features;

[0014] a decoding module, configured to process the fused encoding features by using a pre-acquired decoder to generate a target segmentation result of the original image.

[0015] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:

[0016] one or more processors;

[0017] a storage device configured to store one or more programs,

[0018] when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect.

[0019] In a fourth aspect, the embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method provided in the first aspect.

[0020] Compared with the prior art, the technical solutions provided by the embodiments of the present disclosure have the following advantages:

[0021] The image segmentation method, device, equipment, and medium provided by the embodiments of the present disclosure comprise the following steps: obtaining image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image; processing the image position encoding, the prompt position encoding, and the description text by using a pre-acquired encoder to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text; performing feature fusion processing on the image encoding features, the prompt encoding features, and the text encoding features to generate fused encoding features; and processing the fused encoding features by using a pre-acquired decoder to generate a target segmentation result of the original image. Thus, the multi-modal information composed of the original image, the annotation information of the original image, and the description information of the original image is first encoded and processed, and then the encoding results are fused and decoded to implement image segmentation. Therefore, the multi-modal information of the image is utilized to implement image segmentation, the accuracy of image segmentation is improved, and the accuracy of a planning scheme is improved and the auditing cost of the planning scheme is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings required by the embodiments or the prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0024] Figure 1 A flowchart of an image segmentation method provided by an embodiment of the present disclosure;

[0025] Figure 2 A structural diagram of a feature fusioner provided by an embodiment of the present disclosure;

[0026] Figure 3 A structural diagram of a decoder provided by an embodiment of the present disclosure;

[0027] Figure 4 A logic diagram of an image segmentation method provided by an embodiment of the present disclosure;

[0028] Figure 5 A structural diagram of an image segmentation device provided by an embodiment of the present disclosure;

[0029] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings required by the embodiments or the prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0031] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some of the embodiments of the present disclosure, not all the embodiments.

[0032] In order to improve the accuracy of image segmentation, the following will be combined Figures 1 to 4 The image segmentation method provided by the embodiment of the present disclosure is described. In the embodiment of the present disclosure, the image segmentation method can be executed by an electronic device. The electronic device can be a control device of a medical robot.

[0033] Figure 1A flowchart of an image segmentation method provided by an embodiment of the present disclosure is shown.

[0034] As shown in the figure, the image segmentation method can include the following steps. Figure 1

[0035] S110, obtaining image position encoding of the original image, prompt position encoding of user trigger information on the original image, and description text of the original image.

[0036] In this embodiment, the electronic device first obtains an image that needs to be segmented as an original image, then performs image block embedding processing on the original image, so that the original image is divided into a plurality of image blocks, and then the electronic device respectively encodes the positions of all image blocks of the original image to obtain the position encodings of all image blocks of the original image.

[0037] In order to improve the segmentation accuracy of the original image, the multi-modal information of the original image can also be obtained. The multi-modal information includes determining the user trigger information on the original image based on the user's trigger operation on the interactive interface, and includes the description text of the original image. Further, the electronic device encodes the positions of the user trigger information to obtain the prompt position encoding of the user trigger information.

[0038] The image position encoding refers to the specific position corresponding to each pixel in the original image, which is used to enhance the understanding ability of the image segmentation algorithm for the pixel position in the original image. Optionally, the original image includes but is not limited to a tooth image, a bone image, a lung image, etc.

[0039] The user trigger information refers to the content marked by the user on the original image. Optionally, the user trigger information includes one or both of the key points and the detection frame. Correspondingly, the prompt position encoding refers to the specific position corresponding to the pixels contained in the user trigger information, which is used to enhance the understanding ability of the image segmentation algorithm for the pixel position in the user trigger information.

[0040] The description text refers to the language description information of the original image, which can be a description of the image features of the original image, so as to facilitate the image segmentation algorithm to understand the original image based on the description text.

[0041] S120, using the pre-obtained encoder to process the image position encoding, the prompt position encoding, and the description text to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text.

[0042] ​In this embodiment, the electronic device calls the pre-acquired encoder, processes the image position encoding based on the original image encoding network in the encoder to obtain image encoding features, processes the prompt position encoding based on the prompt encoding network in the encoder to obtain prompt encoding features, and processes the description text based on the text encoding network in the encoder to obtain text encoding features.

[0043] Optionally, the original image encoding network in the encoder includes but is not limited to a masked autoencoder (MAE).

[0044] The prompt encoding network in the encoder is configured to accurately segment the prompt position encoding of the user trigger information. Specifically, the electronic device can first acquire the classification vectors of the key points and the classification vectors of the detection boxes contained in the user trigger information, select two key points (for example, the upper left corner point and the lower right corner point) from the detection boxes, and then use the prompt encoding network to encode the prompt position encoding of the key points and the prompt position encoding of the two key points in the detection boxes using the same encoding method, and distinguish the prompt encoding features corresponding to the prompt position encoding of the key points and the prompt encoding features corresponding to the prompt position encoding of the detection boxes based on the classification vectors.

[0045] The text encoding network in the encoder is configured to extract features in the description text to guide accurate segmentation of the original image. Optionally, the text encoding network includes but is not limited to a Chinese language model (such as Taiyi-CLIP-Roberta-large-326M-Chinese).

[0046] In this way, the encoder can assist the encoder to improve the understanding ability of the original image based on the multi-modal information of the original image, which is beneficial to improve the segmentation accuracy of the original image.

[0047] S130, performing feature fusion processing on the image encoding features, the prompt encoding features and the text encoding features to generate fusion encoding features.

[0048] In this embodiment, in order to realize accurate segmentation of the image, the electronic device can also call a specific feature fusion algorithm to perform feature fusion processing on the multi-modal features including the image encoding features, the prompt encoding features and the text encoding features, so as to efficiently fuse the multi-modal features to obtain the fusion encoding features.

[0049] The fusion encoding features refer to a feature vector that fuses the image features of the original image, the features of the user trigger information and the features of the description text representation, i.e., the fusion encoding features are a feature vector that fuses multi-modal features.

[0050] S140, processing the fusion coding features by using the pre-acquired decoder to generate the target segmentation result of the original image.

[0051] In the embodiment, the electronic device invokes the pre-acquired decoder to decode the fusion coding features to segment the target object from the original image as the target segmentation result.

[0052] The target segmentation result refers to the target object that is focused on in the original image. For example, the original image is a tooth image, and the target segmentation result is a root canal of the tooth image. For another example, the original image is a lung image, and the target segmentation result is a cancer cell in the lung image.

[0053] An image segmentation method according to an embodiment of the present disclosure includes: acquiring image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image; processing the image position encoding, the prompt position encoding, and the description text by using a pre-acquired encoder to acquire image coding features corresponding to the image position encoding, prompt coding features corresponding to the prompt position encoding, and text coding features corresponding to the description text; performing feature fusion processing on the image coding features, the prompt coding features, and the text coding features to generate fusion coding features; and processing the fusion coding features by using a pre-acquired decoder to generate a target segmentation result of the original image. Thus, the multi-modal information composed of the original image, the annotation information of the original image, and the description information of the original image is first processed by encoding, and then the encoding results are fused and decoded to implement image segmentation. Therefore, the multi-modal information of the image is used for image segmentation, the accuracy of image segmentation is improved, and the accuracy of the planning scheme is improved and the auditing cost of the planning scheme is reduced.

[0054] To expand the form of the multi-modal information input to the encoder, the method further includes:

[0055] Acquiring an initial segmentation result corresponding to the original image, wherein the initial segmentation result is obtained by processing the original image by using a preset segmentation algorithm;

[0056] Processing the image position encoding, the initial segmentation result, the prompt position encoding, and the description text by using the pre-acquired encoder to acquire image coding features corresponding to the image position encoding, segmentation coding features corresponding to the initial segmentation result, prompt coding features corresponding to the prompt position encoding, and text coding features corresponding to the description text;

[0057] Correspondingly, the method for implementing S130 includes:

[0058] Adding the image coding features and the segmentation coding features to obtain image addition features;

[0059] The image addition feature, the prompt encoding feature, and the text encoding feature are fused to generate a fused encoding feature.

[0060] Specifically, the electronic device obtains an initial segmentation result obtained by segmenting an original image based on a preset segmentation algorithm, and then retrieves a pre-obtained encoder. The initial segmentation result is processed based on a segmentation image encoding network in the encoder to obtain a segmentation encoding feature. Meanwhile, the image position encoding is processed based on an original image encoding network in the encoder to obtain an image encoding feature. The prompt position encoding is processed based on a prompt encoding network in the encoder to obtain a prompt encoding feature. The description text is processed based on a text encoding network in the encoder to obtain a text encoding feature. Thus, more abundant encoding features are obtained.

[0061] Further, since the image encoding feature and the segmentation encoding feature can both represent the image features of the original image, the electronic device adds the image encoding feature and the segmentation encoding feature of the same pixel to obtain an image addition feature. Then, the electronic device calls a specific feature fusion algorithm to perform feature fusion processing on the multi-modal features including the image addition feature, the prompt encoding feature, and the text encoding feature, so that more multi-modal features are efficiently fused to obtain a fused encoding feature.

[0062] The preset segmentation algorithm is used to segment the original image based on the image features (such as pixel values and texture values) of the original image. Optionally, the preset segmentation algorithm includes but is not limited to a threshold-based segmentation algorithm, an edge-based segmentation algorithm, a region-based segmentation algorithm, and an energy-based segmentation algorithm.

[0063] In this way, since more multi-modal features are obtained, the segmentation accuracy of the original image can be improved based on more multi-modal features.

[0064] In another embodiment of the present disclosure, S110, S130, and S140 are explained in detail.

[0065] In some embodiments of the present disclosure, the specific method of obtaining the image position encoding of the original image and the prompt position encoding of the user trigger information in S110 includes but is not limited to the following method: the original image is position-encoded based on the coordinate data of the first pixel point in the original image, the dimension of the encoding vector corresponding to the first pixel point, and the dimension index in the encoding vector corresponding to the first pixel point to obtain the image position encoding; and the user trigger information is position-encoded based on the coordinate data of the second pixel point in the user trigger information, the dimension of the encoding vector corresponding to the second pixel point, and the dimension index in the encoding vector corresponding to the second pixel point to obtain the prompt position encoding.

[0066] Specifically, the electronic device takes each pixel point in the original image as a first pixel point, takes the number of channels as the dimension of the encoding vector corresponding to the first pixel point, and adopts the sinusoidal position encoding method to perform position encoding on the original image based on the coordinate data of the first pixel point, the dimension of the encoding vector corresponding to the first pixel point, and the dimension index in the encoding vector corresponding to the first pixel point, to obtain image position encoding. At the same time, the electronic device takes the pixel point contained in the user trigger information as a second pixel point, takes the number of channels as the position of the encoding vector corresponding to the second pixel point, and adopts the sinusoidal position encoding method to perform position encoding on the user trigger information based on the coordinate data of the second pixel point, the dimension of the encoding vector corresponding to the second pixel point, and the dimension index in the encoding vector corresponding to the second pixel point, to obtain prompt position encoding.

[0067] Optionally, the sinusoidal position encoding method can be implemented in one of the following ways:

[0068] PE(x, y, z, 2i) = sin(x / 10000^(6i / D))

[0069] PE(x, y, z, 2i+1) = cos(x / 10000^(6i / D))

[0070] PE(x, y, z, 2j+D / 3) = sin(y / 10000^(6j / D))

[0071] PE(x, y, z, 2j+1+D / 3) = cos(y / 10000^(6j / D))

[0072] PE(x, y, z, 2k+2D / 3) = sin(z / 10000^(6k / D))

[0073] PE(x, y, z, 2k+1+2D / 3) = cos(z / 10000^(6k / D))

[0074] wherein x, y, z are the coordinate data of the first pixel point or the second pixel point, D is the number of channels, i.e. the dimension of the encoding vector corresponding to the first pixel point or the second pixel point, 2i, 2i+1, 2j+D / 3, 2j+1+D / 3, 2k+2D / 3, 2k+1+2D / 3 are the dimension indexes in the encoding vector corresponding to the first pixel point or the second pixel point.

[0075] Therefore, by adopting the sinusoidal position encoding method, the image position encoding of the original image and the position encoding of the user trigger information on the original image are performed respectively, so that the encoder enhances the understanding ability of the pixel position in the original image based on the obtained image position encoding and prompt position encoding, which is further conducive to improving the segmentation accuracy of the image.

[0076] In some embodiments of the present disclosure, the specific implementation method of S130 includes but is not limited to the following methods:

[0077] S1301, stack processing the prompt encoding feature and the text encoding feature to obtain a prompt encoding feature;

[0078] S1302, using a pre-acquired feature fusioner to perform feature fusion processing on the prompt encoding feature and the image encoding feature to generate a fusion encoding feature.

[0079] Among them, the prompt encoding feature includes key point encoding feature and detection box encoding feature; Correspondingly, the specific implementation method of S1301 includes but is not limited to the following methods: stack processing the key point encoding feature and the detection box encoding feature along a first stacking direction to obtain a first stacked feature; stack processing the first stacked feature and the text encoding feature along a second stacking direction to obtain a second stacked feature, and taking the second stacked feature as the prompt encoding feature.

[0080] Among them, the first stacking direction refers to the dimension (such as token dimension) of the classification vector corresponding to the prompt encoding feature and the image encoding feature. The second stacking dimension refers to the dimension (such as token dimension) of the classification vector corresponding to the first stacked feature and the text encoding feature.

[0081] Among them, the specific implementation method of S1302 includes but is not limited to the following methods: based on the self-attention network in the feature fusioner, processing the prompt encoding feature to obtain a self-attention feature; based on the prompt-to-image attention network in the feature fusioner, processing the self-attention feature and the image encoding feature to obtain a first cross-attention feature; based on the first fully connected network in the feature fusioner, processing the first cross-attention feature to obtain a target prompt feature; based on the image-to-prompt attention network in the feature fusioner, processing the target prompt feature and the image encoding feature to obtain a second cross-attention feature, and taking the second cross-attention feature as a target image feature; wherein the target prompt feature and the target image feature constitute the fusion encoding feature.

[0082] For ease of understanding, see Figure 2The structure diagram of the feature fusioner shown. Specifically, first, the electronic device inputs the prompt encoding feature into the self-attention network for self-attention processing to obtain a self-attention feature, then inputs the self-attention feature and the image encoding feature into the prompt-to-image attention network for prompt-to-image cross-attention processing to obtain a first cross-attention feature, next inputs the first cross-attention feature into the first fully connected network to obtain an updated prompt feature as a target prompt feature, at the same time, inputs the updated prompt feature and the image encoding feature into the image-to-prompt attention network for image-to-prompt cross-attention processing to obtain an updated image feature as a target image feature, and the fusion encoding feature is composed of the target prompt feature and the target image feature.

[0083] In this way, the feature fusioner adopts the cross-attention mechanism to process the prompt encoding feature and the image encoding feature, not only retains the detail information of the original image itself, but also fully considers the user triggered information containing the key points and the detection frame, which is conducive to improving the image segmentation accuracy.

[0084] In some embodiments of the present disclosure, the specific implementation method of S140 includes but is not limited to the following method: based on the upsampling network in the decoder, processing the target image feature contained in the fusion encoding feature to obtain an upsampling image feature; based on the prompt-to-image attention network in the decoder, processing the target prompt feature contained in the fusion encoding feature to obtain a second cross-attention feature; based on the second fully connected network in the decoder, processing the second cross-attention feature to obtain the confidence of the target prompt feature; based on the third fully connected network in the decoder, processing the second cross-attention feature to obtain a fully connected feature; point multiplying the upsampling image feature and the fully connected feature to obtain a candidate segmentation result of the original image; if the confidence of the target prompt feature is greater than a preset confidence threshold, the candidate segmentation result is taken as the target segmentation result of the original image.

[0085] For ease of understanding, see Figure 3 The structure diagram of the decoder shown. Specifically, first, the electronic device inputs the updated image feature (i.e. the target image feature) into the upsampling network for upsampling processing to obtain an upsampling image feature; at the same time, the electronic device inputs the updated prompt feature (i.e. the target prompt feature) into the prompt-to-image attention network for cross-attention processing to obtain a second cross-attention feature, next inputs the second cross-attention feature into the second fully connected network and the third fully connected network for processing to obtain the confidence of the target prompt feature and a fully connected feature, next point multiplies the upsampling image feature and the fully connected feature to obtain a candidate segmentation result of the original image, and finally judges whether the confidence of the target prompt feature is greater than a preset confidence threshold, if yes, the candidate segmentation result is taken as the target segmentation result of the original image.

[0086] The confidence of the target prompt feature can be specifically an intersection over union, which is used to evaluate whether the candidate segmentation result is reliable.

[0087] The second full connection network and the third full connection network are full connection networks with different parameters and different functions.

[0088] In this way, the decoder processes the target prompt feature by using the cross attention mechanism, retains the detail information of the original image itself, fully considers the user trigger information containing the key points and the detection frame, and combines the up-sampled image feature corresponding to the target image feature to perform image segmentation, which is conducive to improving the image segmentation accuracy.

[0089] In another embodiment of the present disclosure, the overall logic of the image segmentation method is specifically explained.

[0090] Figure 4 A logic diagram of an image segmentation method provided by an embodiment of the present disclosure is shown.

[0091] S410, obtaining image position encoding of an original image, an initial segmentation result corresponding to the original image, prompt position encoding of user trigger information on the original image, and description text of the original image.

[0092] Optionally, the user trigger information includes key points and a detection frame.

[0093] S420, processing the image position encoding, the initial segmentation result, the prompt position encoding, and the description text by using a pre-acquired encoder to obtain image encoding features corresponding to the image position encoding, segmentation encoding features corresponding to the initial segmentation result, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text.

[0094] Specifically, the electronic device calls the pre-acquired encoder, processes the image position encoding based on an original image encoding network in the encoder to obtain image encoding features, processes the initial segmentation result based on a segmentation image encoding network in the encoder to obtain segmentation encoding features, processes the prompt position encoding based on a prompt encoding network in the encoder to obtain prompt encoding features, and processes the description text based on a text encoding network in the encoder to obtain text encoding features.

[0095] Optionally, the prompt position encoding includes key point encoding features and detection frame encoding features.

[0096] S430, adding the image encoding features and the segmentation encoding features to obtain image addition features.

[0097] Specifically, the electronic device adds the image encoding feature of the same pixel and the segmentation encoding feature to obtain an image addition feature.

[0098] S440, stack the key point encoding feature and the detection box encoding feature along a first stacking direction to obtain a first stacked feature.

[0099] In this way, by stacking the key point encoding feature and the detection box encoding feature, the key point information and the detection box information are fused into the first stacked feature.

[0100] S450, stack the first stacked feature and the text encoding feature along a second stacking direction to obtain a second stacked feature, and take the second stacked feature as a prompt encoding feature.

[0101] In this way, by stacking the first stacked feature and the text encoding feature, the key point information, the detection box information and the text information are fused into the second stacked feature as the prompt encoding feature.

[0102] S460, perform feature fusion processing on the image addition feature and the prompt encoding feature to generate a fusion encoding feature.

[0103] In this embodiment, the specific implementation method of S460 includes but is not limited to the following method: using a pre-acquired feature fusioner to perform feature fusion processing on the prompt encoding feature and the image encoding feature to generate the fusion encoding feature.

[0104] Specifically, the electronic device processes the prompt encoding feature based on a self-attention network in the feature fusioner to obtain a self-attention feature; processes the self-attention feature and the image encoding feature based on a prompt-to-image attention network in the feature fusioner to obtain a first cross-attention feature; processes the first cross-attention feature based on a first fully connected network in the feature fusioner to obtain a target prompt feature; processes the target prompt feature and the image encoding feature based on an image-to-prompt attention network in the feature fusioner to obtain an updated image feature, and takes the updated image feature as a target image feature; wherein the target prompt feature and the target image feature constitute the fusion encoding feature.

[0105] S470, using a pre-acquired decoder, processing the fusion encoding feature to generate a target segmentation result of the original image.

[0106] The specific implementation method of S470 includes but is not limited to the following method: based on the upsampling network in the decoder, processing the target image features contained in the fusion encoding features to obtain upsampling image features; based on the attention network from the prompt to the image in the decoder, processing the target prompt features contained in the fusion encoding features to obtain second cross-attention features; based on the second fully connected network in the decoder, processing the second cross-attention features to obtain the confidence of the target prompt features; based on the third fully connected network in the decoder, processing the second cross-attention features to obtain fully connected features; point multiplying the upsampling image features and the fully connected features to obtain the candidate segmentation result of the original image; if the confidence of the target prompt features is greater than a preset confidence threshold, the candidate segmentation result is taken as the target segmentation result of the original image.

[0107] In some embodiments, when the encoder and the decoder are model trained, the electronic device acquires pre-prepared training data, and the training data includes a reference image and annotation information of a segmented region in the reference image. For example, a pre-prepared oral and maxillofacial cone beam (oral CBCT) image is acquired, and the segmentation annotation of each tooth and each root canal is acquired. Specifically, first, the electronic device randomly selects a segmentation object in the reference image, such as tooth No. 11, tooth No. 23, and the like, then a user randomly selects a point on the segmentation object in the reference image as a key point, and randomly draws a detection frame on the segmentation object in the reference image, thereby acquiring user trigger information, and at the same time, a text is used to describe the reference image to obtain a description text of the reference image, such as the description text including the root canal in tooth No. 11 and tooth No. 23, and the like. Then, the electronic device encodes the position of part of the image blocks of the reference image and the user trigger information of the reference image to determine the image position encoding of the part of the image blocks of the reference image and the prompt position encoding of the user trigger information on the reference image. Then, the electronic device inputs the image position encoding of the part of the image blocks of the reference image, the prompt position encoding of the user trigger information on the reference image, and the description text of the reference image into the encoder, fuses the image features output by the encoder, and inputs the fused image features into the decoder together with the remaining image blocks of the reference image which are not encoded to obtain a predicted segmentation result. Then, the model is iteratively trained based on the predicted segmentation result and the annotation information. In the training process, a loss function is calculated, which includes two parts, including a dice loss and a focal loss, and the intersection over union part adopts a mean square error. The two loss functions are added in a 1:1 ratio. Finally, the parameters of the encoder and the decoder are optimized based on the loss function to obtain the trained encoder and decoder.

[0108] Optionally, the loss function includes but is not limited to a mean squared error (MSE) loss function, and can also be other loss functions.

[0109] The embodiment of the present disclosure further provides an image segmentation device for implementing the image segmentation method described above, which is configured in a control device of a robot. The following will be described in combination with Figure 5 The image segmentation device can be an electronic device. The electronic device can be a control device of a medical robot.

[0110] Figure 5 The structure of the image segmentation device provided by the embodiment of the present disclosure is shown.

[0111] As Figure 5 shown, the image segmentation device 500 can include:

[0112] An information acquisition module 510 is configured to acquire image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image.

[0113] An encoding module 520 is configured to process the image position encoding, the prompt position encoding, and the description text by using a pre-acquired encoder, to acquire image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text.

[0114] A feature fusion module 530 is configured to perform feature fusion processing on the image encoding features, the prompt encoding features, and the text encoding features, to generate fusion encoding features.

[0115] A decoding module 540 is configured to process the fusion encoding features by using a pre-acquired decoder, to generate a target segmentation result of the original image.

[0116] The image segmentation device provided by the embodiment of the present disclosure includes: acquiring image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image; processing the image position encoding, the prompt position encoding, and the description text by using a pre-acquired encoder, to acquire image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text; performing feature fusion processing on the image encoding features, the prompt encoding features, and the text encoding features, to generate fusion encoding features; and processing the fusion encoding features by using a pre-acquired decoder, to generate a target segmentation result of the original image. Thus, the multi-modal information composed of the original image, the annotation information of the original image, and the description information of the original image is first processed by encoding, and then the encoding results are fused and processed by decoding, to implement image segmentation. Therefore, the multi-modal information of the image is utilized for image segmentation, the accuracy of image segmentation is improved, and the accuracy of the planning scheme is improved and the auditing cost of the planning scheme is reduced.

[0117] In some embodiments of the present disclosure, the information acquisition module 510 comprises:

[0118] a first position encoding unit, configured to perform position encoding on the original image based on coordinate data of the first pixel point in the original image, a dimension of an encoding vector corresponding to the first pixel point, and a dimension index in the encoding vector corresponding to the first pixel point, to obtain the image position encoding;

[0119] a second position encoding unit, configured to perform position encoding on the user trigger information based on coordinate data of the second pixel point in the user trigger information, a dimension of an encoding vector corresponding to the second pixel point, and a dimension index in the encoding vector corresponding to the second pixel point, to obtain the prompt position encoding.

[0120] In some embodiments of the present disclosure, the feature fusion module 530 comprises:

[0121] a stacking processing unit, configured to perform stacking processing on the prompt encoding feature and the text encoding feature, to obtain a prompt encoding feature;

[0122] a feature fusion unit, configured to perform feature fusion processing on the prompt encoding feature and the image encoding feature by using a pre-acquired feature fusioner, to generate the fusion encoding feature.

[0123] In some embodiments of the present disclosure, the prompt encoding feature comprises a key point encoding feature and a detection box encoding feature.

[0124] Correspondingly, the stacking processing unit is specifically configured to:

[0125] perform stacking processing on the key point encoding feature and the detection box encoding feature along a first stacking direction, to obtain a first stacking feature;

[0126] perform stacking processing on the first stacking feature and the text encoding feature along a second stacking direction, to obtain a second stacking feature, and take the second stacking feature as the prompt encoding feature.

[0127] In some embodiments of the present disclosure, the feature fusion unit is specifically configured to:

[0128] perform processing on the prompt encoding feature based on a self-attention network in the feature fusioner, to obtain a self-attention feature;

[0129] perform processing on the self-attention feature and the image encoding feature based on a prompt-to-image attention network in the feature fusioner, to obtain a first cross-attention feature;

[0130] The first cross-attention feature is processed based on a first full connection network in the feature fusioner to obtain a target prompt feature;

[0131] The target prompt feature and the image encoding feature are processed based on an image-to-prompt attention network in the feature fusioner to obtain an updated image feature, and the updated image feature is taken as a target image feature.

[0132] The target prompt feature and the target image feature constitute the fusion encoding feature.

[0133] In some embodiments of the present disclosure, the decoding module 540 comprises:

[0134] The target image feature included in the fusion encoding feature is processed based on an up-sampling network in the decoder to obtain an up-sampled image feature;

[0135] The target prompt feature included in the fusion encoding feature is processed based on a prompt-to-image attention network in the decoder to obtain a second cross-attention feature;

[0136] The first processing unit is configured to process the second cross-attention feature based on a second full connection network in the decoder to obtain a confidence of the target prompt feature;

[0137] The second processing unit is configured to process the second cross-attention feature based on a third full connection network in the decoder to obtain a full connection feature;

[0138] The up-sampled image feature and the full connection feature are point-multiplied to obtain a candidate segmentation result of the original image;

[0139] The segmentation result determination unit is configured to, if the confidence of the target prompt feature is greater than a preset confidence threshold, take the candidate segmentation result as a target segmentation result of the original image.

[0140] In some embodiments of the present disclosure, the apparatus further comprises:

[0141] The initial segmentation result acquisition module is configured to acquire an initial segmentation result corresponding to the original image, wherein the initial segmentation result is obtained by processing the original image through a preset segmentation algorithm;

[0142] The feature coding module is configured to process the image position code, the initial segmentation result, the prompt position code and the description text by using the pre-acquired encoder, to acquire image coding features corresponding to the image position code, segmentation coding features corresponding to the initial segmentation result, prompt coding features corresponding to the prompt position code, and text coding features corresponding to the description text.

[0143] Correspondingly, the feature fusion module 530 comprises:

[0144] An adding unit is configured to add the image coding features and the segmentation coding features to obtain image addition features.

[0145] A feature fusion unit is configured to perform feature fusion processing on the image addition features, the prompt coding features and the text coding features to generate the fusion coding features.

[0146] It should be noted that, Figure 5 The image segmentation device 500 shown can perform each step in the method embodiment shown, and achieve each process and effect in the method embodiment shown, which will not be described here in detail. Figures 1 to 4 The image segmentation device 500 shown can perform each step in the method embodiment shown, and achieve each process and effect in the method embodiment shown, which will not be described here in detail. Figures 1 to 4 The image segmentation device 500 shown can perform each step in the method embodiment shown, and achieve each process and effect in the method embodiment shown, which will not be described here in detail.

[0147] Figure 6 A structural schematic diagram of an electronic device provided by the embodiment of the present disclosure is shown.

[0148] As Figure 6 shown, the electronic device can include a processor 601 and a memory 602 storing computer program instructions.

[0149] Specifically, the processor 601 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits of the embodiments of the present application.

[0150] Memory 602 may include a large-capacity memory for information or instructions. By way of example, and not limitation, memory 602 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway device. In certain embodiments, memory 602 is non-volatile solid-state memory. In certain embodiments, memory 602 includes read-only memory (ROM). Where appropriate, the ROM may be mask-programmed ROM, programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0151] The processor 601 reads and executes the computer program instructions stored in the memory 602 to perform the steps of the image segmentation method provided in the embodiment of the present disclosure.

[0152] In one example, the electronic device may further include a transceiver 603 and a bus 604. Figure 6 As shown, the processor 601 , the memory 602 and the transceiver 603 are connected via a bus 604 and communicate with each other.

[0153] Bus 604 includes a hardware, software, or both that couples components of computer system 600 to each other. By way of example and not limitation, bus 604 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side BUS (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 604 can include one or more buses. Although this application describes and illustrates a particular bus, this application contemplates any suitable bus or interconnect.

[0154] The following is an embodiment of the computer-readable storage medium provided by the embodiments of the present disclosure, which belongs to the same inventive concept as the image segmentation method of each of the above embodiments. Details not described in the embodiment of the computer-readable storage medium can be referred to the above embodiments of the image segmentation method.

[0155] The embodiment provides a storage medium containing computer executable instructions, which are used to execute an image segmentation method when executed by a computer processor, and the method is applied to an electronic device corresponding to a mechanical arm base. The method comprises the following steps:

[0156] Obtaining image position encoding of an original image, prompt position encoding of user trigger information on the original image, and description text of the original image;

[0157] Using a pre-acquired encoder to process the image position encoding, the prompt position encoding, and the description text, to obtain image encoding features corresponding to the image position encoding, prompt encoding features corresponding to the prompt position encoding, and text encoding features corresponding to the description text.

[0158] perform feature fusion processing on the image encoding features, the prompt encoding features and the text encoding features to generate fused encoding features;

[0159] processing the fused encoding features using a decoder obtained in advance to generate a target segmentation result of the original image.

[0160] Of course, the storage medium containing computer executable instructions provided by the embodiments of the present disclosure is not limited to the method operations as above, and can also perform related operations in the image segmentation method provided by any embodiment of the present disclosure.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the present disclosure can be realized by means of software and necessary universal hardware, and of course can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a FLASH memory, a hard disk or an optical disk, etc., including a number of instructions to make a computer cloud platform (which can be a personal computer, a server, or a network cloud platform, etc.) execute the image segmentation method provided by each embodiment of the present disclosure.

[0162] Note that the above is only the preferred embodiment of the present disclosure and the applied technical principles. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present disclosure. Therefore, although the present disclosure has been described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. An image segmentation method, characterized in that: include: Obtaining an image position code of the original image, a prompt position code of the user trigger information on the original image, and a description text of the original image, wherein the image position code is the position corresponding to each pixel in the original image, and the prompt position code is the position corresponding to the pixel included in the user trigger information; Using a pre-acquired encoder, the image position code, the prompt position code, and the description text are processed to obtain image coding features corresponding to the image position code, prompt coding features corresponding to the prompt position code, and text coding features corresponding to the description text; Performing feature fusion processing on the image coding feature, the prompt coding feature, and the text coding feature to generate a fused coding feature; Using a pre-acquired decoder, the fused coding features are processed to generate a target segmentation result of the original image; The step of obtaining the image position code of the original image and the prompt position code of the user trigger information on the original image includes: Performing position encoding on the original image based on coordinate data of a first pixel in the original image, a dimension of a coding vector corresponding to the first pixel, and a dimension index in the coding vector corresponding to the first pixel to obtain the image position code; Based on the coordinate data of the second pixel point in the user trigger information, the dimension of the coding vector corresponding to the second pixel point, and the dimension index in the coding vector corresponding to the second pixel point, the user trigger information is position-encoded to obtain the prompt position code; The performing feature fusion processing on the image coding feature, the prompt coding feature, and the text coding feature to generate a fused coding feature includes: Stacking the prompt coding feature and the text coding feature to obtain a prompt coding feature; Using a pre-acquired feature fuser, the prompt coding feature and the image coding feature are subjected to feature fusion processing to generate the fused coding feature; The step of using a pre-acquired feature fuser to perform feature fusion processing on the prompt coding feature and the image coding feature to generate the fused coding feature includes: Based on the self-attention network in the feature fuser, the prompt encoding feature is processed to obtain a self-attention feature; processing the self-attention feature and the image encoding feature based on a cue-to-image attention network in the feature fuser to obtain a first cross-attention feature; Processing the first cross-attention feature based on a first fully connected network in the feature fuser to obtain a target prompt feature; processing the target cue feature and the image encoding feature based on the image-to-cue attention network in the feature fuser to obtain an updated image feature, and using the updated image feature as the target image feature; The target prompt feature and the target image feature constitute the fused coding feature.

2. The method according to claim 1, characterized in that The prompt coding features include key point coding features and detection frame coding features; Accordingly, the prompt coding feature and the text coding feature are stacked to obtain the prompt coding feature, including: Stacking the key point coding features and the detection frame coding features along a first stacking direction to obtain a first stacking feature; The first stacking feature and the text coding feature are stacked along a second stacking direction to obtain a second stacking feature, and the second stacking feature is used as the prompt coding feature.

3. The method according to claim 1, characterized in that The method of processing the fused coding features using the pre-acquired decoder to generate a target segmentation result of the original image includes: Based on the upsampling network in the decoder, the target image features included in the fused coding features are processed to obtain upsampled image features; processing the target cue feature included in the fused encoding feature based on the cue-to-image attention network in the decoder to obtain a second cross-attention feature; processing the second cross-attention feature based on a second fully connected network in the decoder to obtain a confidence level of the target cue feature; Processing the second cross-attention features based on a third fully connected network in the decoder to obtain a fully connected feature; Performing a dot product on the upsampled image feature and the fully connected feature to obtain a candidate segmentation result of the original image; If the confidence of the target prompt feature is greater than a preset confidence threshold, the candidate segmentation result is used as the target segmentation result of the original image.

4. The method according to claim 1, wherein Also includes: Obtaining an initial segmentation result corresponding to the original image, wherein the initial segmentation result is obtained by processing the original image through a preset segmentation algorithm; Using the pre-acquired encoder, the image position code, the initial segmentation result, the prompt position code, and the description text are processed to obtain image coding features corresponding to the image position code, segmentation coding features corresponding to the initial segmentation result, prompt coding features corresponding to the prompt position code, and text coding features corresponding to the description text; Accordingly, the performing feature fusion processing on the image coding feature, the prompt coding feature and the text coding feature to generate a fused coding feature includes: Adding the image coding feature and the segmentation coding feature to obtain an image addition feature; The image addition feature, the prompt coding feature and the text coding feature are subjected to feature fusion processing to generate the fused coding feature.

5. An image segmentation device, characterized in that: include: An information acquisition module, configured to acquire an image position code of an original image, a prompt position code of user-triggered information on the original image, and a description text of the original image, wherein the image position code is the position corresponding to each pixel in the original image, and the prompt position code is the position corresponding to the pixel included in the user-triggered information; an encoding module, configured to process the image position code, the prompt position code, and the description text using a pre-acquired encoder to obtain image coding features corresponding to the image position code, prompt coding features corresponding to the prompt position code, and text coding features corresponding to the description text; A feature fusion module, configured to perform feature fusion processing on the image coding feature, the prompt coding feature, and the text coding feature to generate a fused coding feature; A decoding module, configured to process the fused coding features using a pre-acquired decoder to generate a target segmentation result of the original image; The information acquisition module includes: a first position encoding unit, configured to perform position encoding on the original image based on coordinate data of a first pixel point in the original image, a dimension of a coding vector corresponding to the first pixel point, and a dimension index in the coding vector corresponding to the first pixel point, to obtain the image position code; a second position encoding unit, configured to perform position encoding on the user trigger information based on the coordinate data of the second pixel point in the user trigger information, the dimension of the encoding vector corresponding to the second pixel point, and the dimension index in the encoding vector corresponding to the second pixel point, to obtain the prompt position code; The feature fusion module includes: a stacking processing unit, configured to stack the prompt coding feature and the text coding feature to obtain a prompt coding feature; a feature fusion unit, configured to perform feature fusion processing on the prompt coding feature and the image coding feature using a pre-acquired feature fusion device to generate the fused coding feature; a feature fusion unit, configured to process the prompt encoding feature based on the self-attention network in the feature fuser to obtain a self-attention feature; processing the self-attention feature and the image encoding feature based on a cue-to-image attention network in the feature fuser to obtain a first cross-attention feature; Processing the first cross-attention feature based on a first fully connected network in the feature fuser to obtain a target prompt feature; processing the target cue feature and the image encoding feature based on the image-to-cue attention network in the feature fuser to obtain an updated image feature, and using the updated image feature as the target image feature; The target prompt feature and the target image feature constitute the fused coding feature.

6. An electronic device, characterized in that: include: processor; a memory for storing executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Instance segmentation method and system, model training method, medium and electronic equipment

    CN116452600A