A medical image segmentation method and related device based on spatiotemporal consistency constraints

By constructing superpixel masks and cross-frame matching tables, combining U-Net and spatiotemporal Transformer architectures, designing spatiotemporal consistency loss functions and performing conditional random field post-processing, the problems of temporal inconsistency, motion artifacts and edge blur in dynamic medical images are solved, and efficient medical image segmentation is achieved.

CN120580250BActive Publication Date: 2025-10-03ZHUHAI HENGQIN ALL-STAR MEDICAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511072662.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-03
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing medical image segmentation methods suffer from temporal inconsistency, motion artifact interference, and edge blurring effects in dynamic medical images, making it difficult to balance real-time performance and spatiotemporal consistency.

Method used

A medical image segmentation method based on spatiotemporal consistency constraints is adopted. By constructing superpixel masks and cross-frame superpixel matching tables, combining U-Net and spatiotemporal Transformer architectures, designing a spatiotemporal consistency loss function, and performing conditional random field post-processing, the inter-frame coherence and accuracy are improved.

Benefits of technology

It significantly improves the segmentation robustness of low-contrast tissues, reduces the computational complexity, maintains the sharpness of blood vessel/organ boundaries, eliminates inter-frame mutations, and achieves higher segmentation accuracy and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580250B_ABST
    Figure CN120580250B_ABST
Patent Text Reader

Abstract

The present application discloses a medical image segmentation method and related device based on spatiotemporal consistency constraints. The method includes: inputting medical image data, constructing superpixel masks and cross-frame matching tables to generate preprocessed data; constructing an image segmentation model containing a U-Net spatial branch and a spatiotemporal Transformer temporal branch; inputting the preprocessed data into the model, training and obtaining weights based on a spatiotemporal consistency loss function; using the weights to segment the data to obtain preliminary results; and performing spatiotemporal conditional random field post-processing on the preliminary results to obtain the final segmentation. The spatiotemporal loss function integrates a spatial smoothing term (optimizing adjacent superpixels within a frame) and a temporal continuity term (constraining cross-frame matching superpixels); the post-processing stage uses the segmentation probability map as a unary potential and the medical image Gaussian kernel and the cross-frame table as a binary potential for optimization. The present invention significantly improves the spatiotemporal consistency and edge continuity of dynamic medical image segmentation and can reduce inter-frame jitter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a medical image segmentation method, apparatus, device and storage medium based on spatiotemporal consistency constraints. Background Art

[0002] In the field of dynamic medical image segmentation (such as cardiac ultrasound and endoscopic videos), existing methods often face three major challenges: 1) temporal inconsistency—single-frame segmentation models ignore inter-frame correlations, resulting in non-physiological jitter in the segmentation results of consecutive frames, which affects the quantitative analysis of moving organs; 2) motion artifact interference—organ deformation, respiratory movement, and other factors cause pixel displacement, which easily causes traditional segmentation boundaries to break or adhere; and 3) edge blurring—in low-contrast tissue, spatial and temporal information are not co-optimized, resulting in insufficient edge continuity of key anatomical structures. Current mainstream solutions either rely on complex optical flow calculations, which increase time consumption, or use post-processing smoothing at the expense of detail accuracy, making it difficult to balance real-time performance with spatiotemporal consistency. A robust segmentation mechanism that integrates joint spatiotemporal constraints is urgently needed to ensure single-frame accuracy while enhancing inter-frame coherence. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention relates to a medical image segmentation method based on spatiotemporal consistency constraints and related devices, which include but are not limited to image segmentation equipment, electronic devices, computer-readable storage media and computer program products based on spatiotemporal consistency constraints.

[0004] In a first aspect, a medical image segmentation method based on spatiotemporal consistency constraints is provided, comprising:

[0005] a. Input medical image data, construct a superpixel mask and a cross-frame superpixel matching table for the medical image data and integrate it with the medical image data to obtain preprocessed data;

[0006] b. Constructing a spatiotemporal consistency constrained image segmentation model, wherein the image segmentation model includes a spatial branch and a temporal consistency branch;

[0007] c. inputting the preprocessed data into the image segmentation model, constructing a spatiotemporal consistency loss function according to the processing results, and training the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights;

[0008] d. Based on the training weights, use the image segmentation model to perform image segmentation on the preprocessed data to obtain a preliminary segmentation result;

[0009] e. Performing spatiotemporal consistency conditional random field post-processing on the preliminary segmentation results to obtain the final segmentation results.

[0010] In combination with any embodiment of the present application, constructing a superpixel mask and a cross-frame superpixel matching table for the medical image data and integrating them with the medical image data includes:

[0011] Performing simple linear iterative clustering on the first frame of the medical image data to obtain superpixels of the first frame;

[0012] Starting from the second frame, for each frame, the inter-frame optical flow is estimated using a bilateral Gaussian process, and the superpixel center of the previous frame is forward mapped to the next frame to obtain the label;

[0013] Use the online local K-means method to fine-tune the labels and process the missing and newly added superpixels to obtain the superpixel mask for each frame and cross-frame superpixel matching table ;

[0014] The superpixel mask, the cross-frame superpixel matching table and the medical image data are integrated to obtain preprocessed data.

[0015] In combination with any embodiment of the present application, the training of the image segmentation model includes:

[0016] Input the medical image data into the spatial branch to obtain coarse segmentation logits for each frame; the spatial branch uses a U-Net architecture;

[0017] The medical image data is input into the temporal consistency branch to obtain temporal enhancement features of consecutive k frames of images; the temporal consistency branch uses a spatiotemporal transformer architecture.

[0018] In conjunction with any embodiment of the present application, the spatiotemporal consistency loss function formula is:

[0019] ,

[0020] in is the spatial smoothing term, and the formula is:

[0021] ,

[0022] in is the set of all adjacent superpixel pairs (p,q) in the frame, determined by the superpixel mask, For edge rights, 、 are the corresponding coarse segmentation logits of adjacent superpixel pairs;

[0023] is a time-continuous term, and the formula is:

[0024] ,

[0025] in is the cross-frame superpixel matching table, (p,q) is the corresponding superpixel pair between frames, where p is located in frame t and q is located in frame t+1. 、 For coarse segmentation logits.

[0026] In combination with any embodiment of the present application, the image segmentation model is used to perform image segmentation on the preprocessed data based on the training weights to obtain a preliminary segmentation result, including:

[0027] Importing the training weights and initializing the image segmentation model;

[0028] Input the preprocessed data, start image segmentation, and obtain coarse segmentation logits;

[0029] The coarse segmentation logits are integrated with the preprocessed data to obtain the preliminary segmentation result.

[0030] In conjunction with any embodiment of the present application, the spatiotemporal consistency conditional random field post-processing includes:

[0031] The softmax probability map in the preliminary segmentation result is used as a unary potential;

[0032] Using the color or position Gaussian kernel of the medical image data as a spatial term, using the cross-frame superpixel matching table as a temporal term, and using the spatial term and the temporal term as a binary potential;

[0033] The unary potential and the binary potential are input into a conditional random field, and the preliminary segmentation result is post-processed to obtain a final segmentation result.

[0034] In a second aspect, a medical image segmentation device based on spatiotemporal consistency constraints is provided, the device comprising:

[0035] a preprocessing unit, configured to input medical image data, construct a superpixel mask and a cross-frame superpixel matching table for the medical image data, and integrate the superpixel mask and the cross-frame superpixel matching table with the medical image data to obtain preprocessed data;

[0036] A model loading unit, configured to construct a spatiotemporal consistency constrained image segmentation model, wherein the image segmentation model includes a spatial branch and a temporal consistency branch;

[0037] a model training unit, configured to input the preprocessed data into the image segmentation model, construct a spatiotemporal consistency loss function according to the processing result, and train the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights;

[0038] An image segmentation unit, configured to perform image segmentation on the preprocessed data using the image segmentation model based on the training weights to obtain a preliminary segmentation result;

[0039] The post-processing unit is used to perform spatiotemporal consistency conditional random field post-processing on the preliminary segmentation result to obtain a final segmentation result.

[0040] In a third aspect, an electronic device is provided, comprising: a processor, a communication module, a sensor, a user interface, and a storage unit, wherein the storage unit is configured to store computer program code, wherein the program code comprises computer instructions. When the processor executes these instructions, the electronic device performs the method described in the second aspect and any embodiment thereof.

[0041] In a fourth aspect, another electronic device is provided, comprising: a processor, a wireless communication module, a touch screen, a speaker, and a storage unit, wherein the storage unit is configured to store computer program code, wherein the program code comprises computer instructions. When the processor executes these instructions, the electronic device performs the method described in the second aspect and any embodiment thereof.

[0042] In a fifth aspect, a computer-readable storage medium is provided, wherein a computer program is stored, wherein the program includes program instructions. When these instructions are executed by a processor, the processor will perform the method described in the second aspect and any embodiment thereof.

[0043] In a sixth aspect, a computer program product is provided, wherein the computer program product comprises a computer program or instructions. When the computer program or instructions are run on a computer, the computer will execute the method described in the second aspect and any embodiment thereof.

[0044] It should be understood that the above general description and the following detailed description are only used as examples and explanations and do not limit the present application in any way.

[0045] Compared with the existing technology, the present invention provides a medical image segmentation method based on the collaboration of spatiotemporal dual branches. Its basic process is: first, pre-processed data is constructed through superpixel masks and cross-frame matching tables to establish pixel-level spatiotemporal associations; then, a joint architecture of the spatial branch (U-Net) and the temporal consistency branch (spatiotemporal Transformer) is designed. The former outputs single-frame coarse segmentation logits, and the latter extracts continuous frame temporal features; then, the spatial smoothing term (optimizing superpixel consistency within the frame) and the temporal continuity term (constraining the similarity of matching superpixel pairs) are integrated to construct a spatiotemporal loss function training model; finally, conditional random field post-processing is used to optimize the output with the preliminary segmentation probability map as the unary potential and the image Gaussian kernel and the cross-frame matching table as the binary potential. The innovation of this invention lies in: joint spatiotemporal modeling mechanism: the first dual-branch architecture driven by superpixel matching table, the spatial branch ensures the accuracy of single-frame anatomical structure, and the temporal branch eliminates motion artifacts through Transformer long-range dependency modeling, solving the jitter problem of isolated frame segmentation in traditional methods; designing dynamic spatiotemporal consistency loss to significantly improve the segmentation robustness of low-contrast tissue; lightweight post-processing framework: reusing the cross-frame matching table into the CRF time term binary potential to achieve superpixel-level spatiotemporal optimization, reducing the computational complexity by more than 70% compared with traditional pixel-level CRF, eliminating inter-frame mutations while maintaining the sharpness of blood vessel / organ boundaries. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0047] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0048] Figure 1 A schematic flow chart of a medical image segmentation method based on spatiotemporal consistency constraints proposed in an embodiment of the present application;

[0049] Figure 2 A schematic diagram of a medical image segmentation device based on spatiotemporal consistency constraints proposed in an embodiment of the present application;

[0050] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] In order to allow professionals in this technical field to more fully understand the technical solution of the present application, the technical solution of the present application will be explained in detail and clearly with the help of the accompanying drawings. It should be noted that the described embodiments are only some examples of the present application and do not represent all. Based on these embodiments, those skilled in the art can directly deduce all other possible implementation plans without engaging in creative thinking, and these are also included in the scope of protection of the present application.

[0052] In the specification, claims, and related drawings of this application, the terms "first," "second," and the like are used solely to distinguish between different elements and do not imply any particular order. Furthermore, the use of "including," "having," and their variations denotes non-exclusive inclusion. This means that if a process, method, system, product, or apparatus includes a series of steps or components, the process, method, system, product, or apparatus is not limited to the enumerated steps or components and may include other steps or components not listed, or other steps or units inherent to the process, method, system, product, or apparatus.

[0053] The “embodiment” mentioned in this document refers to any instance in which a particular feature, structure or characteristic is combined, and these instances may belong to at least one embodiment of the present application. The “embodiment” mentioned in this document does not necessarily refer to the same specific case, nor does it mean that they are independent or exclusive alternatives. It should be understood by those skilled in the art that the embodiments described herein can be used in conjunction with other embodiments. It should be understood that in this application, “at least one” includes one or more instances, “a plurality” means two or more instances, and “at least two” means two or more instances.

[0054] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The embodiment of the present application is described below in conjunction with the drawings in the embodiment of the present application.

[0055] See also Figure 1 , Figure 1 A flowchart of a medical image segmentation method based on spatiotemporal consistency constraints provided in an embodiment of the present application.

[0056] 101. Spatiotemporal perception data preprocessing: input medical imaging data, construct a superpixel mask and a cross-frame superpixel matching table for the medical imaging data, and integrate them with the medical imaging data to obtain preprocessed data.

[0057] In this embodiment, the medical imaging data is a series of 2D slices of the patient's target organ.

[0058] In this embodiment, the medical imaging data includes but is not limited to CT, MRI, etc.

[0059] In this embodiment, the step of constructing a superpixel mask and a cross-frame superpixel matching table for the medical image data and integrating the superpixel mask and the cross-frame superpixel matching table with the medical image data includes:

[0060] Performing simple linear iterative clustering on the first frame of the medical image data to obtain superpixels of the first frame;

[0061] Starting from the second frame, for each frame, the inter-frame optical flow is estimated using a bilateral Gaussian process, and the superpixel center of the previous frame is forward mapped to the next frame to obtain the label;

[0062] Use the online local K-means method to fine-tune the labels and process the missing and newly added superpixels to obtain the superpixel mask for each frame and cross-frame superpixel matching table ;

[0063] The superpixel mask, the cross-frame superpixel matching table and the medical image data are integrated to obtain preprocessed data.

[0064] In another possible implementation, other superpixel clustering methods may be used using the superpixels of the first frame, including but not limited to the Felzenszwalb superpixel algorithm.

[0065] In another possible implementation, the inter-frame mapping starting from the second frame may adopt other methods, including but not limited to the RAFT optical flow model.

[0066] In another possible implementation, a multi-scale superpixel fusion strategy is introduced: a fine scale is used for near-field organs and a sparse scale is used for far-field tissues.

[0067] In another possible implementation, other preprocessing may be performed on the medical image data, including at least resizing, cropping, etc.

[0068] 102. Import a dual-branch segmentation model: construct a spatiotemporal consistency constrained image segmentation model, wherein the image segmentation model includes a spatial branch and a temporal consistency branch.

[0069] In this embodiment, the spatial branch uses a U-Net architecture, and the temporal consistency branch uses a spatiotemporal transformer architecture.

[0070] In another possible implementation, the spatial branch may use other architectures, including but not limited to deeplabv3.

[0071] In another possible implementation, the temporal consistency branch may use other attention architectures, including but not limited to 3D-CNN, etc.

[0072] 103. Dynamic spatiotemporal consistency constraint learning: input the preprocessed data into the image segmentation model, construct a spatiotemporal consistency loss function according to the processing results, and train the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights.

[0073] In this embodiment, the spatiotemporal consistency loss function formula is:

[0074] ,

[0075] in is the spatial smoothing term, and the formula is:

[0076] ,

[0077] in is the set of all adjacent superpixel pairs (p,q) in the frame, determined by the superpixel mask, For edge rights, 、 are the corresponding coarse segmentation logits of adjacent superpixel pairs;

[0078] is a time-continuous term, and the formula is:

[0079] ,

[0080] in is the cross-frame superpixel matching table, (p,q) is the corresponding superpixel pair between frames, where p is located in frame t and q is located in frame t+1. 、 For coarse segmentation logits.

[0081] In another possible implementation, the spatiotemporal consistency loss function may also be combined with other loss functions.

[0082] 104. Image segmentation: Based on the training weights, use the image segmentation model to perform image segmentation on the preprocessed data to obtain a preliminary segmentation result.

[0083] In this embodiment, the spatial branch operates as follows: inputting the medical image data into the spatial branch to obtain coarse segmentation logits for each frame;

[0084] In this embodiment, the operation mode of the temporal consistency branch is as follows: inputting the medical image data into the temporal consistency branch to obtain temporal enhancement features of consecutive k frames of images;

[0085] In another possible implementation, for the segmentation task of dynamic organs (such as the heart), a phase alignment module is introduced in the preliminary result generation stage: the ECG signal is used to synchronize the multi-frame segmentation results and eliminate motion artifacts.

[0086] 105. Post-processing of spatiotemporal consistency conditional random fields: performing spatiotemporal consistency conditional random field post-processing on the preliminary segmentation result to obtain a final segmentation result.

[0087] In this embodiment, the input data of the spatiotemporal consistency conditional random field are unary potential and binary potential. The unary potential is the softmax probability map in the preliminary segmentation result, and the binary potential is the color or position Gaussian kernel (spatial term) and the cross-frame superpixel matching table (temporal term) of the medical imaging data.

[0088] In another possible implementation, the conditional random field is replaced by the graph cut algorithm: a global energy function containing spatial adjacent terms and temporal coherence terms is constructed, and the solution is obtained through maximum flow minimum cut.

[0089] In another possible implementation, for real-time application scenarios, a sliding window local CRF is used: optimization is performed only on the current frame and the adjacent ±2 frames, and the delay is controlled within 100ms.

[0090] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0091] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0092] See also Figure 2 , Figure 2 A schematic diagram of a medical image segmentation device based on spatiotemporal consistency constraints provided in an embodiment of the present application is provided. The image segmentation device 1 includes: a preprocessing unit 11, a model loading unit 12, a model training unit 13, an image segmentation unit 14, and a post-processing unit 15. Specifically:

[0093] A preprocessing unit 11 is configured to input medical image data, construct a superpixel mask and a cross-frame superpixel matching table for the medical image data, and integrate the superpixel mask and the cross-frame superpixel matching table with the medical image data to obtain preprocessed data;

[0094] A model loading unit 12 is used to construct a spatiotemporal consistency constrained image segmentation model, wherein the image segmentation model includes a spatial branch and a temporal consistency branch;

[0095] A model training unit 13 is configured to input the preprocessed data into the image segmentation model, construct a spatiotemporal consistency loss function according to the processing results, and train the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights;

[0096] An image segmentation unit 14 is configured to perform image segmentation on the preprocessed data using the image segmentation model based on the training weights to obtain a preliminary segmentation result;

[0097] The post-processing unit 15 is configured to perform spatiotemporal consistency conditional random field post-processing on the preliminary segmentation result to obtain a final segmentation result.

[0098] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0099] See also Figure 3 , Figure 3 The following is a schematic diagram of the hardware architecture of an electronic device described in an embodiment of the present application. The electronic device 2 is primarily composed of a processor 21 and a memory 22. In addition, the device may also include an input device 23 and an output device 24. The processor 21, memory 22, input device 23, and output device 24 are interconnected via connecting components, which may be various interfaces, data cables, or communication buses, and are not specifically specified in the present embodiment.

[0100] Processor 21 may be one or more graphics processing units (GPUs). If processor 21 is a GPU, the GPU may be single-core or multi-core. Optionally, processor 21 may comprise a processor group consisting of multiple GPUs, interconnected via one or more buses. Furthermore, the processor may be other types of processors, which are not specifically limited in this embodiment of the present application.

[0101] Memory 22 is designed to store computer program instructions and various program codes required to execute the present invention. Optionally, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which are used to store relevant instructions and data.

[0102] The input device 23 is used to input data and / or signals, and the output device 24 is used to output data and / or signals. The input device 23 and the output device 24 can be independent devices or an integrated device.

[0103] It should be appreciated that in the embodiment of the present application, the memory 22 can store not only relevant instructions but also relevant data. The embodiment of the present application does not specify the specific data content stored in the memory.

[0104] You should understand that Figure 3 Only a simplified design of an electronic device is shown. In actual use, the electronic device may also include other necessary components, such as different numbers of input / output devices, processors, memories, etc. All electronic devices that can implement the embodiments of this application are within the scope of protection of this application.

[0105] Those skilled in the art will recognize that, according to the components and algorithm steps of each example described in the embodiments disclosed herein, these functions can be implemented by electronic hardware or by combining computer software and electronic hardware. Whether these functions are performed by hardware or software will be determined based on the specific application requirements and design limitations of the technical solution. Technicians can adopt different implementation methods according to the requirements of each specific application, but such implementation methods should not be considered to exceed the scope of protection of this application.

[0106] Professionals should understand that, for the sake of ease of description and simplification, the specific operating procedures of the above-mentioned systems, devices, and components can refer to the corresponding steps in the previous method embodiments and will not be repeated here. At the same time, professionals should also understand that each embodiment in this application has its own focus. For the sake of ease of description and simplification, the same or similar content may not be repeated in different embodiments. Therefore, if a part is not mentioned or not explained in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0107] In the several embodiments provided in this application, it should be recognized that the disclosed systems, devices and methods can also be implemented in other ways. For example, the device embodiments described are only exemplary, in which the division of the units is only a division of logical functions, and there may be different division methods in actual implementation. For example, multiple units or components may be merged or integrated into another system, or certain features may be omitted, or certain steps may not be performed. In addition, the connections between each other shown or discussed, whether direct or indirect, whether coupling or communication connection, may be implemented in electrical, mechanical or other forms through interfaces, devices or units.

[0108] Units described as independent components may or may not actually be physically separate; parts presented as units may or may not be physical entities; that is, they may be centralized in one location or distributed across multiple network nodes. Depending on actual needs, some or all of these units may be selected to achieve the objectives of this embodiment.

[0109] Furthermore, in the various embodiments of the present application, the various functional units may be integrated into a single processing unit, physically exist independently, or two or more units may be combined into a single unit. In the aforementioned embodiments, the relevant functions may be implemented in whole or in part through software, hardware, firmware, or any combination thereof. If software implementation is chosen, it may be implemented in whole or in part in the form of a computer program product. This computer program product comprises one or more computer instructions. When these instructions are loaded and executed on a computer, they will generate, in whole or in part, the processes or functions described in the embodiments of this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. These computer instructions may be stored in a computer-readable storage medium or transmitted via such a medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic cable, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium may be any computer-accessible, usable medium, or a data storage facility such as a server or data center that integrates one or more usable media. These available media may include magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), semiconductor media (e.g., SSDs), etc. Those skilled in the art will appreciate that all or part of the process steps for implementing the above-described method embodiments can be accomplished through hardware associated with computer program instructions. These programs can be stored on computer-readable storage media. When executed, these programs will contain the processes for each of the above-described method embodiments. These storage media include, but are not limited to, various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A medical image segmentation method based on spatiotemporal consistency constraints, characterized in that: include: a. Input medical image data, perform simple linear iterative clustering on the first frame of the medical image data to obtain superpixels of the first frame; For the second to last frames of the medical image data, use a bilateral Gaussian process to estimate the inter-frame optical flow, and forward map the superpixel center of the previous frame to the next frame to obtain a label; Iteratively performing an online local K-means method to fine-tune the labels in frame number order and process missing and newly added superpixels to obtain a superpixel mask for each frame and a cross-frame superpixel matching table; integrating the superpixel mask, the cross-frame superpixel matching table, and the medical image data to obtain preprocessed data; b. Constructing a spatiotemporal consistency constraint image segmentation model, the image segmentation model includes a spatial branch and a temporal consistency branch; c. Inputting the preprocessed data into the image segmentation model, constructing a spatiotemporal consistency loss function based on the processing results, and training the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights; the spatiotemporal consistency loss function formula is: , in is the cross entropy loss, is the spatial smoothing term, is a time-continuous term, and is the weight coefficient; The spatial smoothness term The formula is: , in is the set of all adjacent superpixel pairs (p,q) in the frame, determined by the superpixel mask, For edge rights, 、 are the corresponding coarse segmentation logits of adjacent superpixel pairs; The time continuous term The formula is: , Where (p,q) is a pair of corresponding superpixels between frames, where p is located in frame t and q is located in frame t+1. is the cross-frame superpixel matching table, 、 For coarse segmentation logits; d. Based on the training weights, using the image segmentation model to perform image segmentation on the preprocessed data to obtain a preliminary segmentation result; e. Performing spatiotemporal consistency conditional random field post-processing on the preliminary segmentation result to obtain the final segmentation result.

2. The method according to claim 1, characterized in that The spatial branch uses a U-Net architecture, and the temporal consistency branch uses a spatiotemporal transformer architecture.

3. The method according to claim 1, wherein The training of the image segmentation model includes: inputting the medical image data into the spatial branch to obtain coarse segmentation logits for each frame; inputting the medical image data into the temporal consistency branch to obtain temporal enhancement features of consecutive k frames of images.

4. The method according to claim 1, wherein The spatial smoothness term It is used to measure the difference between adjacent superpixels in the current frame in the coarse segmentation logits. It is used to measure the difference in coarse segmentation logits between each frame and the next frame of the superpixels associated via the cross-frame superpixel matching table.

5. The method according to claim 1, characterized in that The step of performing image segmentation on the preprocessed data based on the training weights using the image segmentation model to obtain a preliminary segmentation result includes: Importing the training weights and initializing the image segmentation model; Input the preprocessed data, start image segmentation, and obtain coarse segmentation logits; The coarse segmentation logits are integrated with the preprocessed data to obtain the preliminary segmentation result.

6. The method according to claim 1, wherein The spatiotemporal consistency conditional random field post-processing includes: Using the softmax probability map in the preliminary segmentation result as a unary potential; Using the color or position Gaussian kernel of the medical image data as a spatial term, using the cross-frame superpixel matching table as a temporal term, and using the spatial term and the temporal term as a binary potential; The unary potential and the binary potential are input into a conditional random field, and the preliminary segmentation result is post-processed to obtain a final segmentation result.

7. An image segmentation device based on spatiotemporal consistency constraints, characterized in that: The device comprises: A preprocessing unit is configured to input medical image data, perform simple linear iterative clustering on the first frame of the medical image data to obtain superpixels of the first frame; estimate inter-frame optical flow using a bilateral Gaussian process for the second to last frames of the medical image data, and forward map the superpixel centers of the previous frame to the next frame to obtain labels; iteratively perform an online local K-means method to fine-tune the labels in frame number order, and process missing and newly added superpixels to obtain a superpixel mask and a cross-frame superpixel matching table for each frame; and integrate the superpixel mask with the cross-frame superpixel matching table and the medical image data to obtain preprocessed data; A model loading unit, configured to construct a spatiotemporal consistency constrained image segmentation model, wherein the image segmentation model includes a spatial branch and a temporal consistency branch; A model training unit is used to input the preprocessed data into the image segmentation model, construct a spatiotemporal consistency loss function according to the processing results, and train the image segmentation model based on the spatiotemporal consistency loss function to obtain training weights; the spatiotemporal consistency loss function formula is: , in is the cross entropy loss, is the spatial smoothing term, is a time-continuous term, and is the weight coefficient; The spatial smoothness term The formula is: , in is the set of all adjacent superpixel pairs (p,q) in the frame, determined by the superpixel mask, For edge rights, 、 are the corresponding coarse segmentation logits of adjacent superpixel pairs; The time continuous term The formula is: , Where (p,q) is a pair of corresponding superpixels between frames, where p is located in frame t and q is located in frame t+1. is the cross-frame superpixel matching table, 、 For coarse segmentation logits; An image segmentation unit, configured to perform image segmentation on the preprocessed data using the image segmentation model based on the training weights to obtain a preliminary segmentation result; A post-processing unit is used to perform spatiotemporal consistency conditional random field post-processing on the preliminary segmentation result to obtain a final segmentation result.

8. An electronic device, characterized in that: include: A processor and a storage unit, the storage unit is used to store computer program code, the code includes computer instructions, when the processor executes these instructions, the electronic device performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product comprises a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is caused to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Medical image segmentation method, electronic equipment and storage medium

    CN117808831A

  • Method for generating spatial-temporally consistent depth map sequences based on convolution neural networks

    US20190332942A1