Image processing method and device, storage medium and electronic device

By introducing a type recognition mechanism and corresponding module set in the model, the problem that the model can only handle the segmentation task of a single type of medical image in the prior art is solved, and the multi-type segmentation task processing capability for 2D, 3D and sequence images is realized.

CN120013953APending Publication Date: 2025-05-16SHANGHAI SHANGTANG SHANCUI MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411159254.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, models can only handle segmentation tasks of 2D or 3D medical images, and cannot process two types of images at the same time.

Method used

By determining the type of the target image, the processing result of the image is predicted using a corresponding set of modules (first set of modules or second set of modules). The first set of modules is used for 2D images, and the second set of modules is used for sequence images (including 3D data and video data).

Benefits of technology

It realizes that the model can handle both 2D image segmentation tasks and 3D and sequential images segmentation tasks, solving the limitations of the model's single type image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013953A_ABST
    Figure CN120013953A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, a storage medium and an electronic device, and the method comprises the steps: determining the type of a target image; under the condition that the type of the target image is a first type, predicting a processing result of the target image through a first module set of a target model; and under the condition that the type of the target image is a second type, predicting a processing result of the target image through a second module set of a target model. By adopting the technical scheme, the problem that the model can only process the segmentation task of one type of image in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to an image processing method and device, a storage medium, and an electronic device. Background Art

[0002] Segmentation is a key step in medical image analysis, which helps in accurate disease diagnosis, treatment and monitoring. Manual segmentation, which is the gold standard, usually requires detailed annotation by highly skilled expert doctors, which is very time-consuming and labor-intensive. An intelligent segmentation system that can significantly reduce annotation costs and enable large-scale data analysis is urgently needed in clinical practice.

[0003] Thanks to the progress of deep learning, many fully automatic segmentation methods based on convolutional neural networks (CNN), recurrent neural networks (RNN) and transformers have been developed. At present, the success of the Segment Anything Model (SAM) based on prompt learning in the field of computer vision has inspired many studies on its application to general medical image segmentation tasks. For example, the first proposed 2D medical image segmentation model MedSAM based on bounding box prompts, the 2D medical image segmentation model SAM-Med2D trained on larger medical images based on both bounding box and point prompts, and the 3D point prompt-based segmentation model SAM-Med3D proposed specifically for 3D medical images. Although these methods can segment a wide range of organs or lesions across multiple medical imaging modalities (such as CT, MR, X-ray, ultrasound, pathology, etc.).

[0004] However, in the related art, the same model can either only handle 2D medical image segmentation tasks or only handle 3D medical image segmentation tasks.

[0005] Regarding related technologies, the model can only handle the problem of segmentation tasks for one type of image, and no effective solution has been proposed yet.

[0006] Therefore, it is necessary to improve the related technology to overcome the above-mentioned defects in the related technology. Summary of the invention

[0007] Embodiments of the present application provide an image processing method and device, a storage medium, and an electronic device.

[0008] According to one aspect of an embodiment of the present application, there is provided an image processing method, comprising: determining the type of a target image; when the type of the target image is a first type, predicting the processing result of the target image by a first module set of a target model; when the type of the target image is a second type, predicting the processing result of the target image by a second module set of a target model.

[0009] In an exemplary embodiment, predicting a processing result of the target image through a first module set of a target model includes: determining a first image encoding of the target image through an image encoding module, and determining a first prompt encoding corresponding to first prompt information through a first prompt encoding module, wherein the first prompt information is prompt information related to the received target image; receiving the first image encoding and the first prompt encoding through a first mask decoding module, and determining a first processing result of the target image; receiving the first processing result through an iteration module, and predicting a processing result of the target image based on the first processing result, wherein the first module set includes: the image encoding module, the first prompt encoding module, the first mask decoding module and the iteration module.

[0010] In an exemplary embodiment, the first processing result is received through an iteration module, and the processing result of the target image is predicted based on the first processing result, including: executing a first prediction step: receiving the Nth processing result and the N+1th second prompt information through a second prompt encoding module, and determining a second prompt encoding; receiving the second prompt encoding, the first image encoding and the first prompt encoding through a second mask decoding module, and determining the N+1th processing result of the target image, wherein the N+1th second prompt information is the corrected prompt information of the received Nth processing result, N is a positive integer, and N is 1, 2, 3... in sequence, and the iteration module includes: the second prompt encoding module and the second mask decoding module; looping the first prediction step until the accuracy of the N+1th processing result meets the preset conditions, and determining the N+1th processing result as the processing result of the target image.

[0011] In an exemplary embodiment, predicting the processing result of the target image through the second module set of the target model includes: in the case of predicting the processing result of a first frame image of the target image, determining the image coding of the first frame image through the image coding module, and determining the prompt coding of the first frame image through the first prompt coding module, wherein the prompt coding of the first frame image is the prompt coding corresponding to the received prompt information related to the first frame image; receiving the image coding of the first frame image and the prompt coding of the first frame image through the first mask decoding module, and determining the first processing result of the first frame image;

[0012] In the case of predicting other frame images of the target image, the second loop step is executed in a loop until the first processing result of each frame image of the target image is determined; the processing result of the target image is determined according to the first processing result of each frame image; wherein the second loop step includes: determining the image coding of the Mth frame image and the image coding of the M+1th frame image of the target image through the image coding module, wherein M is a positive integer; receiving the prompt coding of the Mth frame image, the image coding of the Mth frame image and the image coding of the M+1th frame image through the sequence feature transformation module, and determining the prompt coding of the M+1th frame image; receiving the image coding of the M+1th frame image and the prompt coding of the M+1th frame image through the first mask decoding module, and determining the first processing result of the M+1th frame image, wherein the second module set includes at least: the image coding module, the first prompt coding module, the sequence feature transformation module and the first mask decoding module.

[0013] In an exemplary embodiment, determining the processing result of the target image according to the first processing result of each frame image includes: determining the accuracy of the first processing result of each frame image in the target image; when it is determined that the accuracy of the first processing result of the Xth frame image in the target image does not meet the preset conditions, determining the third prompt information of the Xth frame image, wherein the third prompt information is the received corrected prompt information, and X is a positive integer; receiving the third prompt information and the first processing result of the Xth frame image through a second prompt encoding module, and determining the third prompt encoding of the Xth frame image; receiving the third prompt encoding of the Xth frame image, the prompt encoding of the Xth frame image and the image encoding of the Xth frame image through a second mask decoding module, and determining the second processing result of the Xth frame image; determining the processing result of the target image according to the second processing result of the Xth frame image and the first processing results of other frame images, wherein the second module set also includes: the second prompt encoding module and the second mask decoding module.

[0014] In an exemplary embodiment, determining the processing result of the target image based on the second processing result of the Xth frame image and the first processing results of other frame images includes: determining the second processing result of one or more first frame images of the target image, wherein the first frame image is a frame image after the Xth frame image; determining the processing result of the target image based on the second processing result of the Xth frame image, the second processing results of the one or more first frame images, and the first processing results of one or more second frame images, wherein the second frame image is a frame image before the Xth frame image.

[0015] In an exemplary embodiment, determining the second processing result of one or more first frame images of the target image includes at least one of the following: receiving the third prompt code of the Xth frame image, the image code of the Xth frame image and the image code of the X+1th frame image through the sequence feature transformation module, and determining the prompt code of the X+1th frame image; receiving the image code of the X+1th frame image and the prompt code of the X+1th frame image through the first mask decoding module, and determining the processing result of the X+1th frame image.

[0016] In an exemplary embodiment, before determining the image type of the target image, the method further includes: training an image encoding module, a first prompt encoding module and a first mask decoding module of an initial model based on training samples and initial prompt information of the training samples to obtain trained image encoding modules, trained first prompt encoding modules and trained first mask decoding modules; training a second prompt encoding module and a second mask decoding module of the initial model based on the training samples, the initial prompt information of the training samples, the Yth processing result of the training samples and the revised prompt information of the Yth processing result to obtain trained second prompt encoding modules and second mask decoding modules; when the training samples are sequence images, training a sequence feature transformation module based on the revised prompt information of the Zth frame image of the training samples, the image encoding of the Zth frame image and the image encoding of the Z+1th frame image to obtain a trained sequence feature transformation module; and determining the target model according to the trained image encoding module, the first prompt encoding module, the first mask decoding module, the second prompt encoding module, the second mask decoding module and the sequence feature transformation module.

[0017] According to another aspect of an embodiment of the present application, an image processing device is also provided, including: a determination module for determining the type of a target image; a first prediction module for predicting the processing result of the target image through a first module set of a target model when the type of the target image is the first type; and a second prediction module for predicting the processing result of the target image through a second module set of the target model when the type of the target image is the second type.

[0018] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned image processing method when it is run.

[0019] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the image processing method through the computer program.

[0020] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned image processing method is implemented.

[0021] Through the present application, when the type of the target image is the first type, the processing result of the target image is predicted by the first module set of the target model; when the type of the target image is the second type, the processing result of the target image is predicted by the second module set of the target model. That is, the model of the present application can handle both the segmentation task of the first type of image and the segmentation task of the second type of image; the above technical solution solves the problem that the model in the related art can only handle the segmentation task of one type of image. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The exemplary embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0023] Figure 1 is a hardware structure block diagram of a computer device of an image processing method according to an embodiment of the present application;

[0024] Figure 2 is a flowchart of an image processing method according to an embodiment of the present application;

[0025] Figure 3 is a schematic diagram of iterative interactive reasoning of a single image according to an embodiment of the present application;

[0026] Figure 4 is a schematic diagram of image segmentation according to an embodiment of the present application (I);

[0027] Figure 5 is a schematic diagram of image segmentation according to an embodiment of the present application (II);

[0028] Figure 6 Schematic diagram of image segmentation according to an embodiment of the present application (III);

[0029] Figure 7 It is a structural block diagram of an image processing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0031] It should be noted that the terms and "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] The method embodiments provided in the embodiments of the present application can be executed in a computer device or a similar computing device. Taking running on a computer device as an example, Figure 1 is a hardware structure block diagram of a computer device of the image processing method of an embodiment of the present application. Figure 1 As shown, the computer device may include one or more ( Figure 1Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a microprocessor (Microprocessor Unit, referred to as MPU) or a programmable logic device (Programmable logic device, referred to as PLD)) and a memory 104 for storing data. In an exemplary embodiment, the computer device may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer device. Figure 1 More or fewer components as shown, or with Figure 1 Equivalent functions or comparisons shown Figure 1 A different configuration with more features is shown.

[0033] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the image processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, which is equivalent to implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0034] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0035] In this embodiment, an image processing method is provided, which is applied to the above-mentioned computer device. Figure 2 is a flowchart of an image processing method according to an embodiment of the present application, the process comprising the following steps:

[0036] Step S202, determining the type of the target image;

[0037] Step S204, when the type of the target image is the first type, predicting the processing result of the target image by using the first module set of the target model;

[0038] It should be noted that the first type mentioned above can be understood as a 2D image type;

[0039] Step S206: When the type of the target image is the second type, predict the processing result of the target image by using the second module set of the target model.

[0040] The second type mentioned above can be understood as a sequence image type, and the sequence images include but are not limited to: front and back layers of a cross section of 3D data, and front and back frames of a time dimension of video data.

[0041] Through the above steps, when the type of the target image is the first type, the processing result of the target image is predicted by the first module set of the target model; when the type of the target image is the second type, the processing result of the target image is predicted by the second module set of the target model. That is, the model of the present application can handle both the segmentation task of the first type of image and the segmentation task of the second type of image; the above technical solution solves the problem that the model in the related art can only handle the segmentation task of one type of image.

[0042] Optionally, in order to better understand the above step S204, the above step S204 can be implemented in the following manner:

[0043] The first image coding of the target image is determined by an image coding module, and the first prompt coding corresponding to the first prompt information is determined by a first prompt coding module, wherein the first prompt information is prompt information related to the received target image; the first image coding and the first prompt coding are received by a first mask decoding module, and a first processing result of the target image is determined; the first processing result is received by an iteration module, and a processing result of the target image is predicted based on the first processing result, wherein the first module set includes: the image coding module, the first prompt coding module, the first mask decoding module and the iteration module.

[0044] In an embodiment of the present application, the first module set includes: an image encoding module, a first prompt encoding module, a first mask decoding module and an iteration module, wherein the image encoding module is used to determine the first image encoding of the target image; the first prompt encoding module is used to determine the first prompt encoding corresponding to the first prompt information of the target image; the first mask decoding module is used to determine the first processing result of the target image according to the first image encoding and the first prompt encoding; the iteration module is used to adjust the first processing result to determine the final processing result of the target image.

[0045] The first prompt information may be prompt information input by the target object (ie, the user), and the Nth processing result is used to indicate a processing result generated during the iterative processing of the 2D image.

[0046] Through the embodiments of the present application, the accuracy of the processing result of the target image can be improved.

[0047] Furthermore, the iterative module predicts the processing result of the target image according to the first processing result, including:

[0048] Execute the first prediction step: receive the Nth processing result and the N+1th second prompt information through the second prompt encoding module, and determine the second prompt code; receive the second prompt code, the first image code and the first prompt code through the second mask decoding module, and determine the N+1th processing result of the target image, wherein the N+1th second prompt information is the modified prompt information of the received Nth processing result, N is a positive integer, N is 1, 2, 3... in sequence, and the iteration module includes: the second prompt encoding module and the second mask decoding module; loop the first prediction step until the accuracy of the N+1th processing result meets the preset conditions, and determine the N+1th processing result as the processing result of the target image.

[0049] That is, the iteration module includes: a second prompt encoding module and a second mask decoding module, wherein the second prompt encoding module is used to receive the modified prompt information and determine the second prompt code corresponding to the modified prompt information; the second mask decoding module is used to determine the N+1th processing result of the target image according to the second prompt code, the first image code, the first prompt code and the Nth processing result; until the accuracy of the N+1th processing result meets the preset conditions, and the N+1th processing result is determined as the processing result of the target image.

[0050] When the accuracy of the N+1th processing result does not meet the preset conditions, the corrected prompt information of the N+1th processing result is obtained, so that the second prompt code determines the fourth prompt information according to the corrected prompt information of the N+1th processing result; the second mask decoding module determines the N+2th processing result of the target image according to the fourth prompt code, the first image code, the first prompt code and the N+1th processing result.

[0051] It should be noted that the accuracy of the above-mentioned N+1th processing result meets the preset conditions, which can be understood as the accuracy of the above-mentioned N+1th processing result is higher than the preset accuracy; the above-mentioned correction prompt information can be the prompt information input by the target object.

[0052] The N+1th second prompt code output by the second prompt coding module includes: a prompt code corresponding to the modified prompt information of the Nth processing result and a prompt code corresponding to the Nth processing result.

[0053] Optionally, in order to better understand the above step S206, the above step S206 can be implemented in the following manner:

[0054] In the case of predicting the processing result of the first frame image of the target image, the image coding of the first frame image is determined by the image coding module, and the hint coding of the first frame image is determined by the first hint coding module, wherein the hint coding of the first frame image is the hint coding corresponding to the hint information related to the first frame image received; the image coding of the first frame image and the hint coding of the first frame image are received by the first mask decoding module, and a first processing result of the first frame image is determined;

[0055] In the case of predicting other frame images of the target image, the second loop step is executed in a loop until the first processing result of each frame image of the target image is determined; the processing result of the target image is determined according to the first processing result of each frame image; wherein the second loop step includes: determining the image coding of the Mth frame image and the image coding of the M+1th frame image of the target image through the image coding module, wherein M is a positive integer; receiving the prompt coding of the Mth frame image, the image coding of the Mth frame image and the image coding of the M+1th frame image through the sequence feature transformation module, and determining the prompt coding of the M+1th frame image; receiving the image coding of the M+1th frame image and the prompt coding of the M+1th frame image through the first mask decoding module, and determining the first processing result of the M+1th frame image, wherein the second module set includes at least: the image coding module, the first prompt coding module, the sequence feature transformation module and the first mask decoding module.

[0056] For sequence images (such as the front and back layers of the cross-section of 3D data, and the front and back frames of the time dimension of video data) that have certain similarities, their prompts in the interactive segmentation scenario also have certain correlations. By mining this relationship mapping, the prompt interaction cost of sequence images can be greatly reduced.

[0057] In the embodiment of the present application, when predicting the processing result of the first frame image of the second type of image (i.e., sequence image), the implementation process adopted is consistent with the implementation process of predicting the processing result of the first type of image (i.e., 2D image); when predicting the processing results of other frame images of the sequence image, it is necessary to introduce a sequence feature transformation module, and predict the prompt coding of other frame images through the sequence feature transformation module, that is, in the embodiment of the present application, it is only necessary to obtain the prompt information of the first frame image provided externally, and then directly predict the prompt coding of the second frame image according to the image coding of the first frame image, the image coding of the second frame image and the prompt coding of the prompt information of the first frame image according to the sequence feature transformation module, so that it is no longer necessary to obtain the prompt coding corresponding to the prompt information of the second frame image from the outside, and so on, the prompt coding corresponding to the prompt information of each frame image of the target image can be determined. The above-mentioned first processing result is used to indicate the processing result generated during the iterative processing of the sequence image.

[0058] In the embodiment of the present application, the second module set includes at least: an image encoding module, a first prompt encoding module, a sequence feature transformation module and a first mask decoding module, wherein the image encoding module is used to determine the image encoding of each frame image of the target image; the first prompt encoding module is used to determine the prompt encoding of the first frame image according to the prompt information input by the user; the sequence feature transformation module is used to determine the prompt encoding of the M+1 frame image according to the prompt encoding of the M frame image, the image encoding of the M+1 frame image and the image encoding of the M+1 frame image; the first mask decoding module is used to determine the processing result of the M+1 frame image according to the image encoding of the M+1 frame image and the prompt encoding of the M+1 frame image.

[0059] It should be noted that the feature encoding of the Mth and M+1th frames in the known sequence of images is I m and I m+1 , the Mth frame image is given a hint code P m , then the prompt code P of the M+1th frame image can be directly predicted through the sequence feature transformation module m+1 , so there is no need for the prompt interaction of the M+1 frame image, which can be expressed as: P m+1 =F(I m+1 ,I m ,P m ).

[0060] The sequence feature transformation module can be implemented through affine transformations such as convolution, attention mechanism or optical flow. This application determines the prompt code of the M+1th image based on the cross attention mechanism. m As the Key, prompt code P m As Value, and for the image encoding I of the M+1th frame image m+1 As Query, its hint code P m+1 This can be achieved through cross-image interaction, namely:

[0061]

[0062] Among them, Q, K, and V are the mappings of Query, Key, and Value to feature space point tokens, respectively. k is the number of feature channels of Key. In order to further enhance the ability of information mining between sequences, the cross attention mechanism is formulated as a global-local dual-path method. The global path calculates the global correlation between each spatial feature point in Query and all spatial feature points in Key, while the local path calculates the correlation between each spatial feature point in Query and the feature points in the local space of its corresponding position point in Key.

[0063] Specifically, assuming that the spatial dimension of the global feature is H×W and the channel dimension is C, the dimensions of Q, K, and V in the global cross attention mechanism are all HW×C, and the dimension of the global attention Corr(Q,K) is HW×HW, which directly calculates the global prediction with the same dimension of HW×C. The local cross attention mechanism needs to calculate each feature point in the query separately, in which case the dimension of Q is 1×C, and the dimensions of K and V are λ 2 ×C, that is, the feature point token mapping in the λ×λ local area. The dimension of each local attention Corr(Q,K) is 1×λ 2 , each time a local prediction of dimension 1×C is calculated, wait for the prediction of all feature points in the query space dimension HW to be completed, and then splice them together to form a global prediction of HW×C.

[0064] Using local paths between sequence image features with small target changes can make feature matching and propagation faster and more effective. The hint codes predicted by the global and local paths can be directly added together 1:1, or they can be weighted by learning parameters. The complementary characteristics of the global-local dual paths achieve adaptive feature matching between sequence images, thereby realizing hint feature propagation between sequences and hint-driven interactive segmentation expansion.

[0065] Further, determining the processing result of the target image according to the first processing result of each frame image includes: determining the accuracy of the first processing result of each frame image in the target image; when it is determined that the accuracy of the first processing result of the Xth frame image in the target image does not meet the preset conditions, determining the third prompt information of the Xth frame image, wherein the third prompt information is the received corrected prompt information, and X is a positive integer; receiving the third prompt information and the first processing result of the Xth frame image through a second prompt encoding module, and determining the third prompt encoding of the Xth frame image; receiving the third prompt encoding of the Xth frame image, the prompt encoding of the Xth frame image and the image encoding of the Xth frame image through a second mask decoding module, and determining the second processing result of the Xth frame image; determining the processing result of the target image according to the second processing result of the Xth frame image and the first processing results of other frame images, wherein the second module set also includes: the second prompt encoding module and the second mask decoding module.

[0066] It should be noted that, when a processing result is obtained, the processing result of the image may need to be adjusted. The adjustment method is as follows: determine the Xth frame image in the target image that does not conform to the processing result; obtain prompt correction information for the Xth frame image, and then determine the second processing result of the Xth frame image based on the prompt code of the Xth frame image and the image code of the Xth frame image.

[0067] Optionally, after adjusting the processing result of the Xth frame image, it is also necessary to adjust the processing results of images after the Xth frame image, specifically:

[0068] Determine a second processing result of one or more first frame images of the target image, wherein the first frame image is a frame image after the X-th frame image; determine a processing result of the target image according to the second processing result of the X-th frame image, the second processing results of the one or more first frame images, and the first processing results of one or more second frame images, wherein the second frame image is a frame image before the X-th frame image.

[0069] Optionally, determining the second processing result of one or more first frame images of the target image includes at least one of the following: receiving the third prompt code of the Xth frame image, the image code of the Xth frame image and the image code of the X+1th frame image through the sequence feature transformation module, and determining the prompt code of the X+1th frame image; receiving the image code of the X+1th frame image and the prompt code of the X+1th frame image through the first mask decoding module, and determining the processing result of the X+1th frame image.

[0070] When adjusting the processing result of the image after the Xth frame image, it is necessary to determine the prompt code of the X+1th frame image according to the third prompt code corresponding to the modified prompt information of the Xth frame image, the image code of the Xth frame image, and the image code of the X+1th frame image through the sequence feature transformation module, and the first mask decoding module determines the processing result of the X+1th frame image according to the image code of the X+1th frame image and the prompt code of the X+1th frame image. And so on, until the processing result of the last frame image in the target image is determined, so as to complete the adjustment of the processing result of the image after the Xth frame image.

[0071] In an exemplary embodiment, before determining the image type of the target image, the method further includes: training an image encoding module, a first prompt encoding module and a first mask decoding module of an initial model based on training samples and initial prompt information of the training samples to obtain trained image encoding modules, trained first prompt encoding modules and trained first mask decoding modules; training a second prompt encoding module and a second mask decoding module of the initial model based on the training samples, the initial prompt information of the training samples, the Yth processing result of the training samples and the revised prompt information of the Yth processing result to obtain trained second prompt encoding modules and second mask decoding modules; when the training samples are sequence images, training a sequence feature transformation module based on the revised prompt information of the processing result of the Zth frame image of the training samples, the image encoding of the Zth frame image and the image encoding of the Z+1th frame image to obtain a trained sequence feature transformation module; and determining the target model according to the trained image encoding module, the first prompt encoding module, the first mask decoding module, the second prompt encoding module, the second mask decoding module and the sequence feature transformation module.

[0072] In the embodiment of the present application, a large model training method is provided, which is as follows:

[0073] In the initial stage, the training model learns the prediction ability of the overall target based on the initial prompt information. Specifically, given an image x and an initial prompt p0 for the target, the function f is learned based on the parameter θ θ (x, p0) to predict the overall segmentation y0′ of the target. The model is trained by minimizing the difference between the true segmentation annotation y and the predicted segmentation y0′ to obtain the optimal solution for the parameter θ, which is expressed as:

[0074]

[0075] For supervised segmentation losses, commonly used ones are dice loss, cross-entropy loss, and focal loss.

[0076] In the iteration phase, the training model learns the iterative correction capability for complex error regions based on more prompt information and the prediction processing results of the previous phase. Specifically, for the i-th iteration, given the image x, the initial prompt p0 for the target, and the predicted segmentation y of the previous phase i ' -1 , and the correction tips for the last prediction results p i , based on the parameter θ′, the function f is learned θ′ To predict the corrected target segmentation y i ′. Similarly, by minimizing the difference between the true segmentation annotation y and the i-th iteration predicted segmentation y i ′ to train the model to obtain the optimal solution for the parameters. Assuming that after all the iterative stages of learning are completed, a set of optimal solutions θ′ for the model parameters is obtained, then the entire iterative training process can be expressed as:

[0077] Among them, k is the total number of iterations.

[0078] In order to obtain a more hierarchical feature expression, an implementable solution is that θ and θ′ are two different sets of parameters, that is, two different deep models are trained to complete interactive tasks of different complexity. They can have two different sets of parameters in the hint encoding and prediction decoding modules, but share parameters in the image encoding module. Because the image encoding module is the one that takes the most time in the reasoning process, the θ′ model directly inherits the image features encoded by the θ model, which guarantees the timeliness of the joint reasoning of the two models to a certain extent.

[0079] In order to better understand the above embodiment, the iterative interactive reasoning implementation of the system for a single image is also given, such as Figure 3 As shown, the specific implementation is as follows:

[0080] Inputting the input image x to the image encoding module so that the image encoding module outputs the image encoding of the input image x;

[0081] The initial prompt p0 is sent to the prompt encoding module (equivalent to the first prompt encoding module in the above embodiment) so that the prompt encoding module outputs the initial prompt encoding of the input image x;

[0082] Enter the image code and hint code into Figure 3 The mask encoding module above (equivalent to the first mask encoding module in the above embodiment) is configured so that the mask encoding module outputs the predicted segmentation y0′ of the input image x according to the image encoding and the initial hint encoding;

[0083] Then the predicted segmentation y0′ and the iterative hint p i Input to Figure 3The prompt encoding module below (equivalent to the second prompt encoding module in the above embodiment) is used to output the iterative prompt code of the iterative prompt;

[0084] Figure 3 The mask decoding module below (equivalent to the second mask encoding module in the above embodiment) outputs the predicted segmentation y according to the iterative hint coding, image coding, initial hint coding and predicted segmentation y0′. i ′.

[0085] The specific embodiment of the iterative interactive segmentation of the general medical image interactive segmentation system proposed in this application is as follows: Figure 4 , Figure 5 , Figure 6 As shown, Figure 4 , Figure 5 , Figure 6 Each row in represents a set of iterative interactive segmentation cases. The first column is the input image, the second column is the true segmentation annotation (yellow), the third column is the initial interactive prompt (bounding box), the initial prediction result (yellow) and the initial prediction accuracy (such as 0.7156 in the first row), and the fourth to sixth columns are iterative interactive prompts (positive graffiti: green line, negative graffiti: red), continuously optimized iterative prediction results (yellow) and continuously improved iterative prediction accuracy.

[0086] That is, based on the above embodiments, the present application implements a prompt-driven general medical image interactive segmentation system. The system separates intra-image modeling from inter-image modeling, not only realizing prompt interactive segmentation at the 2D image level, but also realizing the extension of the given prompt of the current frame image to the adjacent continuous frame images in the 3D image or video sequence based on the sequence feature transformation module.

[0087] The models in the system are trained through three courses designed from different perspectives:

[0088] (1) Simple interactions from initial prompts (such as bounding boxes) to complex interactive corrections from iterative prompts (such as points, lines / doodles, masks);

[0089] (2) Continuous optimization of segmentation targets from simple (such as organs and bones) to difficult (such as lesions and cell structures) and then to fine (such as blood vessels and brain tissue);

[0090] (3) The gradual expansion of modeling within a single image to modeling between images in different sequence directions.

[0091] In order to balance the interactive segmentation efficiency and accuracy of objects of different sizes, the embodiment of the present application also distills the corresponding small-size input (such as 256×256, 128×128) model for the large-size input (such as 1024×1024) model. Different models can be automatically called according to the initial prompt information (such as the bounding box size) to complete the real-time reasoning of objects of different sizes.

[0092] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0093] In this embodiment, an image processing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0094] Figure 7 is a structural block diagram of an image processing device according to an embodiment of the present application, the device comprising:

[0095] A determination module 72, used to determine the type of the target image;

[0096] A first prediction module 74, configured to predict a processing result of the target image by using a first module set of a target model when the type of the target image is the first type;

[0097] The second prediction module 76 is used to predict the processing result of the target image by using the second module set of the target model when the type of the target image is the second type.

[0098] Through the present application, when the type of the target image is the first type, the processing result of the target image is predicted by the first module set of the target model; when the type of the target image is the second type, the processing result of the target image is predicted by the second module set of the target model. That is, the model of the present application can handle both the segmentation task of the first type of image and the segmentation task of the second type of image; the above technical solution solves the problem that the model in the related art can only handle the segmentation task of one type of image.

[0099] In an exemplary embodiment, a first prediction module 74 is used to determine a first image code of the target image through an image coding module, and to determine a first prompt code corresponding to first prompt information through a first prompt coding module, wherein the first prompt information is prompt information related to the received target image; to receive the first image code and the first prompt code through a first mask decoding module, and to determine a first processing result of the target image; to receive the first processing result through an iteration module, and to predict a processing result of the target image based on the first processing result, wherein the first module set includes: the image coding module, the first prompt coding module, the first mask decoding module and the iteration module.

[0100] In an exemplary embodiment, the first prediction module 74 is used to perform the first prediction step: receive the Nth processing result and the N+1th second prompt information through the second prompt encoding module, and determine the second prompt code; receive the second prompt code, the first image code and the first prompt code through the second mask decoding module, and determine the N+1th processing result of the target image, wherein the N+1th second prompt information is the corrected prompt information of the received Nth processing result, N is a positive integer, and N is 1, 2, 3... in sequence, and the iteration module includes: the second prompt encoding module and the second mask decoding module; loop the first prediction step until the accuracy of the N+1th processing result meets the preset conditions, and determine the N+1th processing result as the processing result of the target image.

[0101] In an exemplary embodiment, the second prediction module 76 is used to determine the image coding of the first frame image through the image coding module and determine the prompt coding of the first frame image through the first prompt coding module when predicting the processing result of the first frame image of the target image, wherein the prompt coding of the first frame image is the prompt coding corresponding to the prompt information related to the received first frame image; receive the image coding of the first frame image and the prompt coding of the first frame image through the first mask decoding module, and determine the first processing result of the first frame image; when predicting other frame images of the target image, loop the second loop step until the first processing result of each frame image of the target image is determined; determine the first processing result of each frame image according to the first processing result of each frame image. The processing result of the target image; wherein the second loop step includes: determining the image coding of the Mth frame image and the image coding of the M+1th frame image of the target image through the image coding module, wherein M is a positive integer; receiving the prompt coding of the Mth frame image, the image coding of the Mth frame image and the image coding of the M+1th frame image through the sequence feature transformation module, and determining the prompt coding of the M+1th frame image; receiving the image coding of the M+1th frame image and the prompt coding of the M+1th frame image through the first mask decoding module, and determining the first processing result of the M+1th frame image, wherein the second module set includes at least: the image coding module, the first prompt coding module, the sequence feature transformation module and the first mask decoding module.

[0102] In an exemplary embodiment, the determination module 72 is used to determine the accuracy of the first processing result of each frame image in the target image; when it is determined that the accuracy of the first processing result of the Xth frame image in the target image does not meet the preset conditions, the second prediction module 76 is used to determine the third prompt information of the Xth frame image, wherein the third prompt information is the received modified prompt information, and X is a positive integer; the third prompt information and the first processing result of the Xth frame image are received through the second prompt encoding module, and the third prompt encoding of the Xth frame image is determined; the third prompt encoding of the Xth frame image, the prompt encoding of the Xth frame image and the image encoding of the Xth frame image are received through the second mask decoding module, and the second processing result of the Xth frame image is determined; the processing result of the target image is determined according to the second processing result of the Xth frame image and the first processing results of other frame images, wherein the second module set also includes: the second prompt encoding module and the second mask decoding module.

[0103] In an exemplary embodiment, the second prediction module 76 is used to determine the second processing result of one or more first frame images of the target image, wherein the first frame image is a frame image after the Xth frame image; and determine the processing result of the target image according to the second processing result of the Xth frame image, the second processing results of the one or more first frame images, and the first processing results of one or more second frame images, wherein the second frame image is a frame image before the Xth frame image.

[0104] In an exemplary embodiment, the second prediction module 76 is used to perform at least one of the following: receiving the third hint code of the Xth frame image, the image code of the Xth frame image and the image code of the X+1th frame image through the sequence feature transformation module, and determining the hint code of the X+1th frame image; receiving the image code of the X+1th frame image and the hint code of the X+1th frame image through the first mask decoding module, and determining the processing result of the X+1th frame image.

[0105] In an exemplary embodiment, the above-mentioned device also includes: a training module, which is used to train the image encoding module, the first prompt encoding module and the first mask decoding module of the initial model based on the training sample and the initial prompt information of the training sample to obtain the trained image encoding module, the trained first prompt encoding module and the first mask decoding module; based on the training sample, the initial prompt information of the training sample, the Yth processing result of the training sample, and the revised prompt information of the Yth processing result, the second prompt encoding module and the second mask decoding module of the initial model are trained to obtain the trained second prompt encoding module and the second mask decoding module; when the training sample is a sequence image, based on the revised prompt information of the processing result of the Zth frame image of the training sample, the image encoding of the Zth frame image and the image encoding of the Z+1th frame image, the sequence feature transformation module is trained to obtain the trained sequence feature transformation module; the target model is determined according to the trained image encoding module, the first prompt encoding module, the first mask decoding module, the second prompt encoding module, the second mask decoding module and the sequence feature transformation module.

[0106] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when run.

[0107] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0108] S1, determine the type of target image;

[0109] S2, when the type of the target image is the first type, predicting the processing result of the target image by using a first module set of the target model;

[0110] S3: When the type of the target image is the second type, predict the processing result of the target image by using a second module set of the target model.

[0111] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0112] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0113] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0114] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0115] S1, determine the type of target image;

[0116] S2, when the type of the target image is the first type, predicting the processing result of the target image by using a first module set of the target model;

[0117] S3: When the type of the target image is the second type, predict the processing result of the target image by using a second module set of the target model.

[0118] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0119] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.

[0120] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0121] An embodiment of the present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in any one of the above method embodiments.

[0122] Optionally, in this embodiment, the computer program product may be configured to store a computer program for performing the following steps:

[0123] S1, determine the type of target image;

[0124] S2, when the type of the target image is the first type, predicting the processing result of the target image by using a first module set of the target model;

[0125] S3: When the type of the target image is the second type, predict the processing result of the target image by using a second module set of the target model.

[0126] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0127] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0128] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0129] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image processing method, characterized in that: include: Determine the type of target image; When the type of the target image is the first type, predicting a processing result of the target image by using a first module set of a target model; In a case where the type of the target image is the second type, a processing result of the target image is predicted by a second module set of a target model.

2. The image processing method according to claim 1, characterized in that: Predicting the processing result of the target image by using the first module set of the target model includes: Determine a first image code of the target image through an image coding module, and determine a first prompt code corresponding to first prompt information through a first prompt coding module, wherein the first prompt information is prompt information related to the received target image; receiving the first image code and the first hint code through a first mask decoding module, and determining a first processing result of the target image; The first processing result is received through an iteration module, and a processing result of the target image is predicted based on the first processing result, wherein the first module set includes: the image encoding module, the first hint encoding module, the first mask decoding module and the iteration module.

3. The image processing method according to claim 2, characterized in that: Receiving the first processing result through an iteration module, and predicting the processing result of the target image according to the first processing result, including: Execute the first prediction step: receive the Nth processing result and the N+1th second prompt information through the second prompt encoding module, and determine the second prompt encoding; receive the second prompt encoding, the first image encoding and the first prompt encoding through the second mask decoding module, and determine the N+1th processing result of the target image, wherein the N+1th second prompt information is the modified prompt information of the received Nth processing result, N is a positive integer, and N is 1, 2, 3... in sequence, and the iteration module includes: the second prompt encoding module and the second mask decoding module; The first prediction step is executed in a loop until the accuracy of the N+1th processing result meets a preset condition, and the N+1th processing result is determined as the processing result of the target image.

4. The image processing method according to claim 1, characterized in that: Predicting the processing result of the target image by the second module set of the target model includes: In the case of predicting the processing result of the first frame image of the target image, the image coding of the first frame image is determined by the image coding module, and the hint coding of the first frame image is determined by the first hint coding module, wherein the hint coding of the first frame image is the hint coding corresponding to the hint information related to the first frame image received; the image coding of the first frame image and the hint coding of the first frame image are received by the first mask decoding module, and a first processing result of the first frame image is determined; In the case of predicting other frame images of the target image, the second loop step is executed repeatedly until a first processing result of each frame image of the target image is determined; and a processing result of the target image is determined according to the first processing result of each frame image; The second loop step includes: determining the image coding of the Mth frame image and the image coding of the M+1th frame image of the target image by the image coding module, wherein M is a positive integer; Receiving the prompt code of the Mth frame image, the image code of the Mth frame image and the image code of the M+1th frame image through the sequence feature conversion module, and determining the prompt code of the M+1th frame image; The image encoding of the M+1th frame image and the hint encoding of the M+1th frame image are received through the first mask decoding module, and a first processing result of the M+1th frame image is determined, wherein the second module set includes at least: the image encoding module, the first hint encoding module, the sequence feature transformation module and the first mask decoding module.

5. The image processing method according to claim 4, characterized in that: Determining the processing result of the target image according to the first processing result of each frame of image includes: Determining the accuracy of the first processing result of each frame of the target image; When it is determined that the accuracy of the first processing result of the X-th frame image in the target image does not meet the preset condition, determine the third prompt information of the X-th frame image, wherein the third prompt information is the received correction prompt information, and X is a positive integer; Receiving the third prompt information and the first processing result of the X-th frame image through the second prompt encoding module, and determining the third prompt code of the X-th frame image; receiving, through a second mask decoding module, a third hint code of the X-th frame image, a hint code of the X-th frame image, and an image code of the X-th frame image, and determining a second processing result of the X-th frame image; The processing result of the target image is determined according to the second processing result of the Xth frame image and the first processing results of other frame images, wherein the second module set further includes: the second hint encoding module and the second mask decoding module.

6. The image processing method according to claim 5, characterized in that: Determining the processing result of the target image according to the second processing result of the Xth frame image and the first processing results of other frame images includes: Determine a second processing result of one or more first frame images of the target image, wherein the first frame image is a frame image after the Xth frame image; The processing result of the target image is determined according to the second processing result of the Xth frame image, the second processing results of the one or more first frame images, and the first processing results of the one or more second frame images, wherein the second frame image is a frame image before the Xth frame image.

7. The image processing method according to claim 6, characterized in that: Determining the second processing result of one or more first frames of the target image comprises at least one of the following: receiving, through the sequence feature transformation module, the third prompt code of the Xth frame image, the image code of the Xth frame image, and the image code of the X+1th frame image, and determining the prompt code of the X+1th frame image; The image code of the X+1th frame image and the hint code of the X+1th frame image are received through the first mask decoding module, and a processing result of the X+1th frame image is determined.

8. The image processing method according to claim 1, characterized in that: Before determining the image type of the target image, the method further includes: Based on the training samples and the initial prompt information of the training samples, the image encoding module, the first prompt encoding module and the first mask decoding module of the initial model are trained to obtain the trained image encoding module, the trained first prompt encoding module and the first mask decoding module; based on the training samples, the initial prompt information of the training samples, the Yth processing result of the training samples, and the corrected prompt information of the Yth processing result, the second prompt encoding module and the second mask decoding module of the initial model are trained to obtain the trained second prompt encoding module and the second mask decoding module; In the case where the training sample is a sequence image, the sequence feature transformation module is trained based on the correction prompt information of the processing result of the Zth frame image of the training sample, the image code of the Zth frame image and the image code of the Z+1th frame image to obtain a trained sequence feature transformation module; The target model is determined according to the trained image encoding module, the first prompt encoding module, the first mask decoding module, the second prompt encoding module, the second mask decoding module and the sequence feature transformation module.

9. The image processing method according to claim 1, characterized in that: The first type is: 2D image type, the second type is: sequence image type.

10. An image processing device, characterized in that: include: A determination module, used for determining the type of the target image; A first prediction module, used for predicting a processing result of the target image by using a first module set of a target model when the type of the target image is a first type; The second prediction module is used to predict the processing result of the target image through the second module set of the target model when the type of the target image is the second type.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 9 when executed.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 9 through the computer program.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.