An image data compression method, device and storage medium
By decoupling medical image data across layers, modalities, and time dimensions, extracting common and differential components, and generating compressed data streams, the problem of low compression efficiency in existing technologies is solved, achieving efficient multi-dimensional compression.
Patent Information
- Application Number
- CN202511848232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-12-09
AI Technical Summary
Existing medical image compression technologies fail to fully exploit the correlations of multi-dimensional data, resulting in low compression efficiency. Furthermore, deep learning-based methods suffer from long inference times.
A decoupling model is used to decouple the medical image dataset in the interlayer, intermodal and temporal dimensions, extract common and differential components respectively, and generate a compressed data stream through an encoding model.
It achieves efficient and multi-dimensional compression of medical image data without sacrificing diagnostic value, significantly improving the compression ratio and reducing storage and transmission pressure.
Smart Images

Figure CN121357335B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of medical image data compression, and in particular to an image data compression method, apparatus, and storage medium. Background Technology
[0002] With the rapid development of medical imaging technology, imaging techniques such as computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET) play an irreplaceable role in clinical diagnosis, treatment planning, and follow-up evaluation. However, the resulting data volume is growing explosively, placing enormous storage and transmission pressure on hospital medical image archiving and communication systems.
[0003] Currently, medical image compression technology is mainly divided into two categories: traditional compression methods and deep learning-based intelligent compression methods. Traditional compression methods can only process single images independently, resulting in insufficient redundancy utilization. They fail to fully exploit the multiple correlations within a single case, including the significant data correlations within a sequence (such as consecutive layers in the same scan), between modalities (such as images from different devices like CT, MRI, and PET), and over time (such as images from multiple follow-up visits), leading to low compression efficiency. Deep learning-based intelligent compression methods utilize convolutional neural networks to achieve efficient image encoding and decoding, but most work remains limited to compressing single images and fails to utilize multidimensional correlations at the system level. In particular, diffusion models can generate visually realistic images at extremely low bit rates, offering hope for medical image compression. However, this approach suffers from long inference times and typically focuses on the generation quality itself, failing to systematically and jointly mine the three dimensions of redundancy (intra-layer, inter-modal, and temporal), and also failing to deeply integrate generative priors with decoupled representations.
[0004] In view of this, it is necessary to provide an image data compression method, device and storage medium to perform multi-dimensional and efficient compression of massive medical image data without losing diagnostic value. Summary of the Invention
[0005] One aspect of this invention provides an image data compression method, the method comprising: acquiring a medical image dataset of a target object, the medical image dataset including multiple types of image data originating from different data sources; inputting the medical image dataset into a decoupling model, wherein the decoupling model decouples the medical image dataset in at least two different dimensions to obtain common components and differential components; inputting the common components into a first encoding model to generate a common component encoded stream; inputting the differential components into a second encoding model to generate a differential component encoded stream; and integrating the common component encoded stream and the differential component encoded stream to form a compressed data stream.
[0006] One aspect of the present invention provides an image data compression apparatus, the apparatus comprising at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least a portion of the computer instructions to implement an image data compression method.
[0007] One aspect of this invention provides a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes an image data compression method. Attached Figure Description
[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0009] Figure 1 These are schematic diagrams illustrating application scenarios of image data compression based on some embodiments of this specification;
[0010] Figure 2 This is an exemplary flowchart of an image data method according to some embodiments of this specification;
[0011] Figure 3 This is one of the exemplary schematic diagrams of an image data compression method according to some embodiments of this specification;
[0012] Figure 4 This is a second exemplary schematic diagram of an image data compression method according to some embodiments of this specification;
[0013] Figure 5 This is an exemplary schematic diagram illustrating joint training according to some embodiments of this specification. Detailed Implementation
[0014] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0015] The terms “system,” “device,” “unit,” and / or “module” as used herein are one method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0016] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0017] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0018] Figure 1 This is a schematic diagram illustrating the application scenario of the image data compression device according to some embodiments of this specification.
[0019] like Figure 1 As shown, the application scenario 100 involved in the embodiments of this specification may include an imaging device 110, a network 120, a processor 130, and a storage device 140.
[0020] In some embodiments, the image data compression apparatus may compress image data by implementing the methods and / or processes disclosed in this specification.
[0021] In some embodiments, the imaging device 110, network 120, processor 130, and storage device 140 may be interconnected and / or communicate via wireless, wired, or a combination thereof. The connections between the components of the image optimization system can be variable. As an example only, the imaging device 110 may be connected to the processor 130 via the network 120 or directly. Similarly, the storage device 140 may be connected to the processor 130 via the network 120 or directly.
[0022] Imaging device 110 can be used to generate medical image data or medical image datasets of a target object. Imaging device 110 can scan a target area of the target object and generate medical images. The target object can include biological objects (e.g., human bodies, animals, etc.), non-biological objects (e.g., phantoms), etc. In some embodiments, the target area of the target object can include a specific part of the object, an organ and / or tissue, and other organs and / or tissues within a certain range around it. For example, the target area of the target object can include the head, chest, legs, etc., or any combination thereof, without limitation.
[0023] In some embodiments, the processor 130 and the storage device 140 may be part of the imaging device 110.
[0024] In some embodiments, the imaging device 110 may include medical imaging devices, such as X-ray imaging devices, digital radiography (DR), computed radiography (CR), digital fluorography (DF), biochemical immunoassay analyzers, computed tomography (CT) devices, magnetic resonance imaging (MR) devices, positron emission tomography (PET) devices, emission computed tomography (ECT) devices, digital subtraction angiography (DSA) devices, electrocardiographs, C-arm devices, ultrasound imaging devices, fluorescence fluoroscopy imaging devices, etc., or any combination thereof. In some embodiments, the imaging device 110 may include industrial imaging devices, such as optical imaging devices, digital radiograph (DR) imaging devices, and industrial computed tomography (ICT) devices. In some embodiments, the imaging device 110 may transmit medical image datasets to the processor 130 and / or storage device 140 via the network 120 for further processing.
[0025] Network 120 may include any suitable network that can facilitate information and / or data exchange within the image data compression apparatus. In some embodiments, one or more components of the image data compression apparatus (e.g., imaging device 110, processor 130, and storage device 140) may be connected to and / or communicate with other components of the image data compression apparatus via network 120. For example, processor 130 may acquire a medical image dataset from imaging device 110 via network 120. As another example, processor 130 may acquire a medical image dataset from storage device 140 via network 120.
[0026] In some embodiments, network 120 can be any one or more of wired or wireless networks. For example, network 120 may include cable networks, fiber optic networks, telecommunications networks, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, near field communication (NFC), device internal buses, device internal wiring, cable connections, etc., or any combination thereof.
[0027] Processor 130 can be used to process data and / or information obtained from other devices / components or components. Processor 130 can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in the embodiments of this specification. By way of example only, processor 130 can be, but is not limited to, a central processing unit (CPU), a microprocessor (MCU), an application-specific integrated circuit (ASIC), a programmable logic device (FPGA), or any combination thereof. For example, processor 130 can acquire a medical image dataset of a target object from imaging device 110. As another example, processor 130 can compress the medical image dataset of the target object to form a compressed data stream. Yet another example, processor 130 can decompress the compressed data stream to generate restored medical image data.
[0028] In some embodiments, processor 130 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 130 may be local or remote. In some embodiments, processor 130 may be connected via network 120 or directly to imaging device 110 and storage device 140 to access information and / or data stored thereon. In some embodiments, processor 130 may be integrated into imaging device 110. In some embodiments, processor 130 may be implemented on a cloud platform. By way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-cloud, etc., or any combination thereof.
[0029] Storage device 140 may store data, instructions, and / or any other information. In some embodiments, storage device 140 may store data obtained from imaging device 110 or processor 130. In some embodiments, storage device 140 may store data and / or instructions used by processor 130 to perform or use the exemplary methods described herein.
[0030] In some embodiments, storage device 140 may be used to store computer instructions or computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the medical image optimization method in the embodiments of this specification. Processor 130 executes various functional applications and data processing by running the computer instructions or computer programs stored in storage device 140, thereby implementing the medical image optimization method described in the embodiments of this specification. Storage device 140 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, storage device 140 may further include memory remotely located relative to processor 130, and these remote memories can be connected to a terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and any combination thereof. In some embodiments, storage device 140 may be implemented on a cloud platform. In some embodiments, storage device 140 may be part of processor 130.
[0031] It should be noted that the application scenario 100 of the image data compression device is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can make various modifications or variations based on the description in this specification. For example, the application scenario 100 of the image data compression device may also include a database, information source, etc. Furthermore, the application scenario 100 of the image data compression device may be implemented on other devices to achieve similar or different functions. However, these variations and modifications will not depart from the scope of this specification.
[0032] Figure 2 This is an exemplary flowchart of an image data compression method according to some embodiments of this specification. In some embodiments, process 200 may be executed by a processor. Figure 2 As shown, process 200 includes the following steps.
[0033] Step 210: Obtain the medical image dataset of the target object.
[0034] A medical image dataset is a dataset consisting of a series of medical images. For example, medical images from different scans, or medical images from different imaging devices.
[0035] In some embodiments, a medical imaging dataset includes multiple types of image data originating from different data sources.
[0036] For example, medical imaging datasets can include raw data, imaging data, target scan data, and multi-timepoint data. Raw data refers to unreconstructed, original acquired data, such as K-space and Sinogram raw acquisition data. Imaging data refers to reconstructed images used for clinical diagnosis, such as reconstructed DICOM images. Target scan data refers to detailed scan data of a specific region. Multi-timepoint data refers to time-series data of patient follow-up.
[0037] For example, biological data may include medical image data acquired by X-ray imaging equipment, digital radiography (DR), computed radiography (CR), digital fluorography (DF), biochemical and immunoassay analyzers, computed tomography (CT) equipment, magnetic resonance imaging (MR) equipment, positron emission tomography (PET) imaging equipment, emission computed tomography (ECT) equipment, digital subtraction angiography (DSA) equipment, electrocardiographs, C-arm equipment, ultrasound imaging equipment, fluorescence fluoroscopy imaging equipment, etc., or any combination thereof.
[0038] In some implementations, medical image datasets can be obtained through user input (e.g., a doctor or nurse). For example, a user can upload medical image data from a terminal device.
[0039] In some embodiments, the medical image dataset may be medical image data stored in a Picture Archiving and Communication System (PACS). The processor can obtain the medical image dataset by communicating with the PACS.
[0040] In some embodiments, the medical image dataset can be read from a storage device. This storage device can be the storage device 140 integrated into the image data compression apparatus, or it can be an external storage device not part of the image data compression apparatus, such as a hard disk or optical disc.
[0041] In some embodiments, the medical image dataset can be read through an interface, which includes, but is not limited to, a program interface, a data interface, and a transmission interface. In some embodiments, when the image data compression device is operating, it can automatically extract the medical image dataset from the interface. In some embodiments, the medical image dataset can also be obtained using any method well known to those skilled in the art, and this specification does not limit this.
[0042] Step 220: Input the medical image dataset into the decoupling model. The decoupling model decouples the medical image dataset in at least two different dimensions to obtain common components and differential components.
[0043] Decoupling refers to the process of breaking down complex data.
[0044] In some embodiments, the decoupling process can separate highly redundant portions from low-redundant portions of a medical image dataset. The highly redundant portions are common components, and the low-redundant portions are differential components.
[0045] Common components refer to shared, common, and stable data information extracted from a medical imaging dataset. Examples include the skull and ventricle shapes common to all layers in a brain CT sequence, or the common tumor location in a patient's CT and MRI images. This part of the data has extremely high redundancy and is the primary target for compression.
[0046] In some embodiments, the shared component may represent stable background information (such as organ shape or basic anatomical structure) that recurs across multiple layers, modalities, or time points. This data can be highly compressed.
[0047] The differential component refers to the unique, specific, and variable data information of each data point in a medical imaging dataset. Examples include the subtle vascular structures unique to a particular CT scan layer, the soft tissue contrast that distinguishes MRI images from CT scans, or the change in lesion size in follow-up images. This part of the data is relatively small and requires processing with higher fidelity.
[0048] Dimension refers to a specific analytical level during decoupling. Through the analytical level, shared information (commonality) and unique information (difference) between different types of data in a medical imaging dataset can be discovered.
[0049] For each of at least two dimensions, the decoupling model can extract the common and dissimilar components corresponding to that dimension.
[0050] A decoupling model is a computer rule or algorithm model used to decouple medical image datasets.
[0051] In some embodiments, the decoupling model may employ a machine learning-based multi-branch feature extraction network. This multi-branch feature extraction network comprises a shared feature extraction backbone network and branch processing networks for each dimension to be decoupled. Each branch network is responsible for learning and separating common and dissimilar features along a specific dimension.
[0052] In some embodiments, the machine learning-based multi-branch feature extraction network can be a pre-trained machine learning model. The processor can directly call the pre-trained multi-branch feature extraction network to decouple the medical image dataset in at least two different dimensions to obtain common components and differential components.
[0053] In some embodiments, the decoupling model can also be based on an autoencoder structure, which separates common components from differential components by introducing a decoupling loss function; or it can be based on a conditional generative adversarial network, which uses its discriminator to guide the decoupling of common components from differential components.
[0054] In some embodiments, the decoupling model includes at least two decoupling modules corresponding to at least two dimensions, a first fusion module, and a second fusion module.
[0055] In some embodiments, each decoupling module corresponds to one of at least two dimensions and is configured to decouple the corresponding image data in the medical image dataset along that dimension to generate common sub-components and differential sub-components corresponding to that dimension.
[0056] Each decoupling module focuses on a specific dimension, with its inputs, outputs, and internal mechanisms varying depending on the dimension. Each decoupling module receives a subset of medical image data corresponding to its dimension and maps the input data (i.e., the medical image data subset corresponding to the dimension) to a feature space through an encoder network. In this space, a cleverly designed loss function forces the network to learn to separate the feature channels into two parts: common sub-components and differential sub-components.
[0057] Common subcomponents refer to the shared, common, and stable data information extracted from a subset of medical image data in a certain dimension.
[0058] Differential sub-components refer to the unique, specific, and variable data information extracted from a subset of medical image data in a certain dimension.
[0059] In some embodiments, at least two dimensions include at least two of the following: interlayer dimension, modal dimension, and temporal dimension.
[0060] Interlayer dimension focuses on the correlation in three-dimensional spatial coordinates. Decoupling models analyze the relationships between multiple spatially continuous or adjacent images (such as CT or MRI sequence slices) naturally generated by an imaging device in a single scan in the interlayer dimension.
[0061] Modal dimension focuses on the correlation between the physical properties reflected by different imaging principles (modals). Decoupling models, in the modal dimension, analyze the relationship between images obtained from different imaging techniques (such as CT, MRI, PET) for the same anatomical site.
[0062] In some embodiments, the imaging modality is divided at the data generation stage, including at least two of the following: raw data, imaging data, and target scan data.
[0063] Raw data is the unreconstructed, original acquired data, such as CT sine charts and MRI K-space data. It contains the object's most original projection or frequency information.
[0064] Imaging data refers to reconstructed images used for clinical diagnosis, obtained from raw data, such as reconstructed DICOM images. Examples include DICOM format CT images and MRI images.
[0065] Target scan data refers to detailed scan data for a specific area.
[0066] In some embodiments, imaging modalities are classified from the perspective of imaging principles, including at least two of the following: computed tomography (CT) data, X-ray imaging (DR or CR) data, magnetic resonance imaging (MOR) data, positron emission tomography (PET) data, and ultrasound imaging data.
[0067] The temporal dimension focuses on the correlations within a time series. Decoupling models, in the temporal dimension, analyze the relationships between imaging data of the same patient collected at different time points (such as before disease treatment and post-treatment follow-up).
[0068] In some embodiments, decoupling at the inter-layer dimension includes analyzing the correlation between image data from different layers in the 3D image data corresponding to the same scan in a medical image dataset. Image data from different layers in the 3D image data corresponding to the same scan are highly correlated in terms of anatomical structure. Image data from adjacent layers depict continuous cross-sections of the same anatomical location at different depths, and therefore contain a large amount of repetitive anatomical structural information (such as organ outlines and major blood vessels).
[0069] For example, we can analyze the shared, common, and stable data information across different layers of 3D image data from the same scan, and analyze the unique, specific, and variable data information for each data point within the same 3D image data from different layers. As an example, in a head CT sequence, major structures such as the skull, ventricles, and cerebral cortex appear in most slices; this is the "common" information. A specific small blood vessel section on a particular layer represents the "difference" information for that layer.
[0070] In some embodiments, decoupling at the inter-layer dimension can be achieved using a 3D convolutional neural network. The 3D convolutional neural network extracts features in the spatial dimension using 3D convolutional kernels and compresses sequence information through pooling operations or attention mechanisms, separating the stable macroscopic anatomical structures (common sub-components) from the subtle details unique to each layer (discretionary sub-components). Further details can be found in the following description.
[0071] In some embodiments, decoupling at the modal dimension includes analyzing the correlation between image data from different imaging modalities. Although images from different modalities differ greatly in grayscale, contrast, and resolution, they depict the same anatomical entity. Therefore, behind the differences in pixel values, there is shared semantic information (such as the location, size, and shape of lesions).
[0072] For example, we can analyze the shared, common, and stable data information across all data points in imaging data from different imaging modalities, and analyze the unique, specific, and variable data information for each data point within the imaging data from different imaging modalities. As an example, a patient's lung nodule appears as a high-density solid shadow on CT and as a bright area with high metabolism on PET. Their "common" information is the nodule's three-dimensional location and approximate shape; their "difference" information is the density of the CT value and the standardized uptake value of the PET.
[0073] In some embodiments, decoupling at the modality dimension can be achieved through a shared-private encoder architecture. The shared encoder (with weights shared across modalities) is trained to extract modality-invariant semantic features (common sub-components), while the private encoder (independent for each modality) is responsible for capturing modality-specific features (discrete sub-components) determined by physical imaging principles. By constraining the shared encoder through adversarial training, it can be ensured that its output does not contain modality-specific information. Further details are provided in the following description.
[0074] In some embodiments, decoupling in the temporal dimension includes analyzing the correlation between image data from different acquisition time points. During follow-up, the patient's background anatomical structures (such as organ shape and normal tissue) are usually relatively stable or change slowly, while changes in the true clinical points of interest (such as tumors and inflammatory areas) are more significant and rapid.
[0075] For example, we can analyze the shared, common, and stable data information across all images acquired at different time points, and analyze the unique, specific, and changing data information for each individual data point within the images acquired at different time points. As an example, in annual MRI follow-ups of Alzheimer's disease patients, the overall ventricular morphology and sulcular patterns of the brain are relatively stable "common" information, while the annual shrinkage of the hippocampus is "difference" information that requires attention.
[0076] In some embodiments, decoupling in the temporal dimension can be achieved using a deformation registration network and a recurrent neural network (RNN). For example, image data from different acquisition time points (also known as time-series images) are first registered to a common space to eliminate non-rigid deformation. Subsequently, the RNN analyzes the registered time-series images, whose final hidden state aggregates time-invariant information (common sub-components), while the residual between the output at each time point and the final state represents the amount of change (discrepancy sub-components). Further explanation is provided in the following description.
[0077] Some embodiments in this specification systematically mine the correlations between layers, modalities, and time by clearly defining the three key dimensions of "interlayer, modality, and time," accurately covering the most important and common sources of redundancy in medical imaging data. This achieves a leap from "image-level compression" to "patient-level compression," thereby obtaining an order-of-magnitude improvement in compression ratio. Simultaneously, it ensures that the invention can specifically mine the data correlations with the greatest compression potential, making the high compression ratio effect universal and reliable, and widely applicable to various complex clinical scenarios.
[0078] In some embodiments, the at least two decoupling modules include at least two of the following: an inter-layer decoupling module, a modal decoupling module, and a temporal decoupling module.
[0079] The inter-layer decoupling module is a module in the decoupling model used for decoupling processing in the inter-layer dimension.
[0080] In some embodiments, the interlayer decoupling module is configured to perform interlayer dimensional decoupling processing on the first image data in the medical image dataset to obtain a first common sub-component and a first differential sub-component, wherein the first image data is three-dimensional image data.
[0081] The first image data is the three-dimensional image data corresponding to the same scan. For example, a single three-dimensional scan (such as CT or MRI) produces dozens or even hundreds of consecutive two-dimensional slice images.
[0082] The first common subcomponent refers to the common subcomponent obtained through decoupling in the inter-layer dimension. The first differential subcomponent refers to the differential subcomponent obtained through decoupling in the inter-layer dimension.
[0083] The inter-layer decoupling module can achieve decoupling processing in the inter-layer dimension in a variety of ways.
[0084] In some embodiments, the interlayer decoupling module can model the spatial correlation between layers by performing 3D convolution on the first image data or by using a self-attention mechanism along the slice direction. The algorithm automatically identifies shared anatomical patterns (common sub-components) that occur stably in all slices and separates them from the unique, subtle anatomical details or noise (discrete sub-components) of each layer.
[0085] In some embodiments, the inter-layer decoupling module may employ an architecture of a 3D convolutional neural network and an attention branch network. The decoupling process performed by the inter-layer decoupling module may include S11-S12:
[0086] S11. Obtain the first common subcomponent using a 3D convolutional neural network: Through continuous 3D convolution and downsampling, a low-dimensional, global feature representation is extracted. This process naturally suppresses local, layer-specific details, thereby extracting the anatomical structure shared by all layers, i.e., the first common subcomponent.
[0087] S12. Introducing a parallel attention branch network in a 3D convolutional neural network to obtain the first dissimilar component: A spatial attention map is generated for each layer of image data using a parallel attention branch network. The first dissimilar component is calculated through residual connections, or the unique first dissimilar component for each layer is reconstructed using an attention mechanism. For example, the original layer features - (first common component) The result of the spatial attention map was determined as the first differential component.
[0088] The modal decoupling module is a module in the decoupling model used for decoupling in the modal dimension.
[0089] In some embodiments, the modal decoupling module is configured to perform modal dimension decoupling processing on the second image data corresponding to different imaging modalities in the medical image dataset to obtain a second common subcomponent and a second differential subcomponent.
[0090] Secondary imaging data consists of imaging data from different imaging modalities. For example, registered image pairs or groups of images from different devices or modalities (such as CT and MRI images of the same patient).
[0091] The second common subcomponent refers to the common subcomponent obtained through decoupling in the modal dimension. The second differential subcomponent refers to the differential subcomponent obtained through decoupling in the modal dimension.
[0092] The modal decoupling module can achieve decoupling processing in the modal dimension in a variety of ways.
[0093] In some embodiments, the modality decoupling module may employ cross-modal representation learning techniques, such as using a shared-weight encoder to map images from different modalities to the same feature space. Within this space, the model is forced to learn modality-independent semantic commonalities (i.e., a second common subcomponent) while preserving the unique features of each modality due to its physical imaging principles (i.e., a second differential subcomponent), such as the sensitivity of CT to bone or the display of metabolic activity by PET.
[0094] In some embodiments, the modal decoupling module may employ a pre-trained diffusion model to perform iterative denoising under the condition of a known second common subcomponent, thereby generating a second differential subcomponent that is complementary to the second common subcomponent through multiple iterations.
[0095] In some embodiments, the modal decoupling module may employ a shared-private encoder architecture, which includes a shared encoder and a private encoder. The weights of the shared encoder are shared across different modalities, and its goal is to learn modality-invariant feature maps. The private encoder is independently owned by each modality and is used to capture modality-specific information. The decoupling process performed by the modal decoupling module may include S21-S22:
[0096] S21. Extracting the Second Common Sub-component: Mapping images from different modalities to the same feature space. To force this space to contain only common information, a modality discriminator can be introduced for adversarial training. This discriminator attempts to distinguish which modality a feature comes from, while the shared encoder tries to "deceive" the discriminator, making it unable to distinguish. In this way, the encoder will remove modality-specific information from the features as much as possible, retaining only the second common sub-component (such as the semantic features of lesions).
[0097] S22. Extracting the Second Differential Sub-component: Extracting the second differential sub-component, determined by the physical imaging principle of each modality's data (such as the Heinz unit characteristics of CT and the relaxation time characteristics of MRI), using a proprietary encoder. The proprietary encoder is not subject to the constraints described in S11; its task is to extract the unique features of the modality to form the second differential sub-component.
[0098] The temporal decoupling module is a module in the decoupling model used for decoupling in the temporal dimension.
[0099] In some embodiments, the temporal decoupling module is configured to perform temporal decoupling processing on the third image data corresponding to different acquisition time points in the medical image data to obtain a third common sub-component and a third differential sub-component.
[0100] The third type of imaging data consists of images acquired at different time points. For example, imaging sequences from multiple follow-up examinations of a patient.
[0101] The third common subcomponent refers to the common subcomponent obtained through decoupling in the temporal dimension. The third differential subcomponent refers to the differential subcomponent obtained through decoupling in the temporal dimension.
[0102] The timing decoupling module can achieve decoupling processing in the timing dimension in a variety of ways.
[0103] In some embodiments, the temporal decoupling module may employ an architecture combining a deformation registration network and a recurrent neural network (RNN) / Transformer. The decoupling process performed by the temporal decoupling module may include S31-S33:
[0104] S31. Registration: Register the image with the time point to the reference space (common coordinate space).
[0105] S32. Common component extraction: Input the third image data into an RNN or Transformer, and use the RNN or Transformer to analyze the entire registered time series. Its final state or global representation is the third common sub-component.
[0106] S33. Difference Component Extraction: A deformation field estimation network (e.g., a CNN with a U-Net structure) is introduced. This network takes the images from the previous and current time points as input and calculates a dense optical flow or deformation field. This field quantifies the motion or deformation of each pixel from the baseline to the current time point. The third difference component can then be obtained by calculating the residual between the current image and the baseline image.
[0107] In some embodiments of this specification, the inter-slice decoupling module performs decoupling processing at the inter-slice dimension, which can efficiently eliminate intra-slice redundancy between consecutive slices in a single scan. By extracting shared anatomical structures, only one set of basic structural information and a small number of inter-slice detail differences need to be stored, thereby achieving extremely high compression efficiency for three-dimensional volumetric data generated by CT, MRI, etc., solving the problem of low efficiency when processing sequential data using traditional methods. The modal decoupling module performs decoupling processing at the modal dimension, breaking down data barriers between different imaging modalities and exploring the complementarity and correlation between modalities. For example, for CT and MRI images of the same patient, by extracting shared semantic features, the repeated storage of the same anatomical information is avoided, realizing collaborative compression of cross-modal data, which is particularly suitable for efficient data management in multimodal fusion diagnostic scenarios. The temporal decoupling module performs decoupling processing at the temporal dimension, which greatly eliminates redundancy in time-series data. By locating time-invariant baseline structures and storing only the changes over time, high compression ratios can be achieved for long-term follow-up (such as tumor treatment effect evaluation and chronic disease progression monitoring) imaging data, significantly reducing long-term storage costs.
[0108] The first fusion module is used to fuse the common sub-components obtained by decoupling the decoupling modules of different dimensions.
[0109] In some embodiments, the first fusion module is configured to fuse common sub-components corresponding to at least two dimensions to generate a common component. In some embodiments, the first fusion module may fuse at least two of the first common sub-component, the second common sub-component, and the third common sub-component to generate a common component.
[0110] The first fusion module can use various methods to fuse common sub-components corresponding to at least two dimensions to generate common components.
[0111] In some embodiments, the first fusion module may use feature-level fusion to fuse at least two of the first common sub-component, the second common sub-component, and the third common sub-component to generate a common component. For example, the first common sub-component, the second common sub-component, and the third common sub-component may be mapped to a feature space of a unified dimension through a fully connected layer or a convolutional layer, and then weighted averaging or attention-based fusion may be performed.
[0112] The second fusion module is used to fuse the differential sub-components obtained by decoupling the decoupling modules of different dimensions.
[0113] In some embodiments, the second fusion module is configured to fuse the difference sub-components corresponding to at least two dimensions to generate a difference component. In some embodiments, the second fusion module may fuse at least two of the first difference sub-component, the second difference sub-component, and the third difference sub-component to generate a difference component.
[0114] The second fusion module can use various methods to fuse the differential sub-components corresponding to at least two dimensions to generate differential components.
[0115] In some embodiments, the second fusion module may use a splicing method to fuse at least two of the first differential sub-component, the second differential sub-component, and the third differential sub-component to generate a differential component.
[0116] In some embodiments of this specification, the decoupling module is designed to include at least two decoupling modules corresponding to at least two dimensions, a first fusion module, and a second fusion module. The decoupling module focuses on feature separation for a specific dimension, while the fusion module focuses on information aggregation. This makes the decoupling model easier to train, debug, and optimize. Simultaneously, the processing dimensions can be easily added or reduced. For example, if a certain inspection only has sequence data and modal data, but no time-series data, the system can enable only the sequence and modal decoupling modules without affecting the overall architecture. Furthermore, a decoupling module dedicated to a specific dimension can more accurately learn the data characteristics of that dimension, thereby producing purer and more thoroughly decoupled sub-components. The fusion module can then perform effective information fusion on this basis, avoiding early interference between features from different sources.
[0117] In some embodiments, the processor can decouple the medical image dataset in three different dimensions using a decoupling model to obtain common and differential components.
[0118] In some embodiments, the processor performs inter-layer decoupling processing on the first image data using the inter-layer decoupling module in the decoupling model to obtain a first common sub-component and a first difference sub-component; performs modal decoupling processing on the second image data using the modal decoupling module in the decoupling model to obtain a second common sub-component and a second difference sub-component; and performs temporal decoupling processing on the third image data using the temporal decoupling module in the decoupling model to obtain a third common sub-component and a third difference sub-component. See the preceding descriptions for further details.
[0119] In some embodiments, the processor may also acquire clinical context information related to the target object and input the clinical context information into the first fusion module and the second fusion module respectively, wherein: the first fusion module is configured to perform weighted fusion of common sub-components corresponding to at least two dimensions based on the clinical context information and using an attention mechanism; the second fusion module is configured to perform weighted fusion of differential sub-components corresponding to at least two dimensions based on the clinical context information and using an attention mechanism.
[0120] Clinical context information refers to data related to the target patient's historical treatment. In some embodiments, clinical context information includes the target patient's basic information, examination and diagnostic information, and equipment parameter information. Basic patient information includes age, gender, and history of underlying diseases. Examination and diagnostic information includes the type of examination and body location (e.g., "plain chest CT scan," "enhanced brain MRI"), suspected or confirmed diseases (e.g., "follow-up of pulmonary nodules," "multiple sclerosis," "acute stroke"), and clinical concerns (e.g., "focusing on hippocampal volume," "assessing the solid component of nodules in the left lower lobe"). Equipment parameter information includes scan sequences and contrast agent usage.
[0121] Clinical context information can be obtained in various ways. In some embodiments, the processor can automatically obtain clinical context information from a medical information system. In some embodiments, the processor can also read clinical context information from a Radiology Information System (RIS) or a Hospital Information System (HIS). In some embodiments, the processor can parse clinical context information from a DICOM file header.
[0122] In some embodiments, the first fusion module can encode clinical context information into a clinical feature vector through a neural network or embedding layer; use an attention network to calculate the attention weights of the clinical feature vector and the common sub-components corresponding to each dimension to determine the weights corresponding to each dimension; and perform weighted fusion based on the common sub-components and weights corresponding to each dimension to determine the common components.
[0123] In some embodiments, the attention network is implemented using a multilayer perceptron (MLP). This MLP computes a set of attention weights (corresponding to inter-layer, modal, and temporal dimensions) conditioned on clinical feature vectors. The computation process can be represented as follows:
[0124]
[0125] The Softmax function ensures that the sum of all weights is 1, representing the proportion of resource allocation. The parameters of the MLP are learned during model training, enabling it to learn to allocate different importance levels according to different clinical scenarios. , , These are the attention weights corresponding to the inter-layer dimension, modal dimension, and temporal dimension, respectively. This is a clinical feature vector.
[0126] As an example only, hippocampal atrophy is a key diagnostic indicator for follow-up MRI in Alzheimer's disease. At this point, the first fusion module can: increase... This ensures that temporal variations (hippocampal volume changes) are accurately captured. Increase... This ensures that the image details in the hippocampus region are sufficiently clear. Reduce... Because for this disease, structural consistency may not be the primary concern. This dynamism allows the entire compression process to move beyond a "one-size-fits-all" approach, adaptively balancing compression rates with the specific diagnostic needs of different diseases, achieving "on-demand compression" and maximizing compression efficiency while maintaining diagnostic confidence.
[0127] Similarly, the second fusion module can encode clinical context information into a clinical feature vector through a neural network or embedding layer; it then uses an attention network to calculate attention weights for this clinical feature vector and the differential components corresponding to each dimension, determining the weights for each dimension; finally, it performs weighted fusion based on the differential components and weights for each dimension to determine the differential components. The second fusion module's method of weighted fusion of differential components corresponding to at least two dimensions using an attention mechanism is similar to the first fusion module's method of weighted fusion of common components corresponding to at least two dimensions using an attention mechanism; see the previous descriptions for further explanation.
[0128] In some embodiments of this specification, the weights of different dimensions are dynamically adjusted through clinical context information and attention mechanisms. This intelligently allocates limited bitrate resources (compressed bits) to critical diagnostic regions, achieving ultra-high definition fidelity in critical regions while maintaining an overall high compression ratio. This fundamentally resolves the contradiction between fidelity and compression ratio. Furthermore, the compressed data is no longer a cold binary stream but intelligent data tagged with clinical importance. This ensures that radiologists can clearly see the key features needed for diagnosis when reviewing medical image data, greatly reducing the risk of misdiagnosis due to compression and significantly improving the clinical usability of the compressed results.
[0129] It should be noted that raw medical image datasets may typically have the following problems:
[0130] (1) Spatial misalignment: Images acquired in different modalities or at different time points have inconsistent positions, angles and scales in the coordinate system due to factors such as patient position and scanning parameters.
[0131] (2) Interference information exists: such as motion artifacts caused by patient movement, breathing, heartbeat, etc., which will reduce image quality and introduce erroneous information.
[0132] If these issues are left unaddressed, they will severely interfere with the decoupling model's judgment of "common information" and "difference information." Therefore, data preprocessing can provide the decoupling model with high-quality, standardized input, thereby ensuring that it can accurately uncover the inherent correlations and differences in the data.
[0133] In some embodiments, the processor may also preprocess the medical image dataset before inputting it into the decoupled model.
[0134] In some embodiments, data preprocessing includes at least one of multimodal image registration, temporal data motion artifact correction, etc.
[0135] The embodiments in this specification do not impose any special limitations on the technical means of multimodal image registration and temporal data motion artifact correction; any operation familiar to those skilled in the art can be used.
[0136] Step 230: Input the common component into the first coding model to generate the common component coding stream; input the difference component into the second coding model to generate the difference component coding stream.
[0137] The first coding model is a model used to encode common components. For example, the first coding model could be a high-fidelity encoder used to compress common components. The common component coded stream refers to the coded stream after the common components are compressed.
[0138] Since the common component contains the underlying structure and critical diagnostic information shared by all data, a high-fidelity or even lossless coding strategy (such as lossless coding based on wavelet transform or lossy coding with extremely low loss) is required to ensure that this information is almost not lost during compression, thus prioritizing the integrity and reliability of the information. Even with a relatively low compression ratio, it is essential to ensure that this part of the data is lossless or nearly lossless.
[0139] The second coding model is used to encode the differential components. For example, the second coding model can be an efficient compression encoder used to compress the differential components. The differential component encoded stream refers to the encoded stream after the differential components are compressed. Since the differential components have a small data volume and allow for a certain degree of distortion, high-efficiency, high-compression-ratio lossy coding strategies (such as perceptual quantization and entropy coding) can be used to significantly reduce the data volume.
[0140] In some embodiments, the encoding mechanisms of the first encoding model and the second encoding model are different. The first encoding model generates a common component encoded stream through lossless encoding or high-fidelity lossy encoding; the second encoding model generates a differential component encoded stream through lossy compression encoding based on perceptual weighting.
[0141] Lossless coding refers to compression techniques that can completely and without distortion recover the original data after encoding. The decoded data is completely identical to the unencoded data, bit-to-bit. In some embodiments, lossless coding can be achieved through predictive coding or entropy coding.
[0142] The first coding model generates a common component coded stream through lossless coding, including the following steps S41-S43:
[0143] S41. Convert the common components into a serialization format suitable for subsequent encoding. Optional preprocessing includes data normalization, adjusting feature values to a specific numerical range to optimize the performance of the predictor in predictive coding.
[0144] S42. Utilizing the spatial or statistical correlation between data points within common components, a predictor estimates the value of the current data point to be encoded, and encodes the prediction residual between the predicted value (i.e., the estimated value of the current data point to be encoded) and the true value. Compression is achieved because the entropy of the prediction residual is much lower than that of the original data.
[0145] For each data point in the common component, the difference between its true value and the predicted value is calculated, resulting in a residual sequence. A linear predictor can be used for prediction, such as based on previously encoded adjacent data points.
[0146] S43. For the residual sequence obtained after predictive coding and decorrelation, entropy coding is applied to convert it into the final binary bitstream. The entire residual sequence is input into an arithmetic encoder. The arithmetic encoder maps the entire input sequence to a single real number interval between [0, 1) based on the probability of each symbol (residual value). Symbols with high occurrence probabilities (such as small residuals near 0) shorten the interval range more slowly, while symbols with low probabilities shorten the interval range significantly. After encoding, a sufficiently precise binary decimal within this interval is selected as the final compressed representation. This process approximates the theoretical limit of information entropy, achieving efficient compression.
[0147] When the common components are already very compact after decoupling, or when clinical applications require absolute zero distortion, lossless coding is used.
[0148] High-fidelity lossy coding refers to the encoding process that introduces extremely small distortions, imperceptible to the human eye or diagnostic software, but achieves a higher compression ratio than lossless coding. Its "high fidelity" is relative to the "lossy compression" of the second coding model.
[0149] The first coding model generates a common component coded stream through high-fidelity lossy coding, including the following steps S44-S46:
[0150] S44. Transform the common components from the spatial domain to the transform domain (such as the frequency domain or wavelet domain) to obtain multiple transformation coefficients.
[0151] After the transformation from the spatial domain to the transform domain, the common components are represented as multiple transform coefficients, where the low-frequency coefficients carry the main outline and structural information of the image, while the high-frequency coefficients represent finer details.
[0152] In some embodiments, transformation methods for converting common components from the spatial domain to the transform domain include discrete cosine transform (DCT), wavelet transform, etc.
[0153] S45. Define a quantization table with a very small quantization step size, and use the quantization table to quantize each transform coefficient to generate quantized coefficients.
[0154] In some embodiments, the principle for setting the quantization step size is to make it as small as possible while meeting the target compression ratio, so as to ensure that the quantization error (i.e., distortion) is negligible.
[0155] In some embodiments, using a quantization table The formula for quantizing each transformation coefficient is shown in the following formula (1):
[0156]
[0157] in, These are the quantized transform coefficients. For quantization tables, transformation coefficients This is the coefficient of variation.
[0158] S46. Apply arithmetic coding to the coefficient matrix formed by the quantized transform coefficients to obtain the common component coded stream.
[0159] After fine quantization, the coefficient matrix contains a large number of consecutive zero values (especially in high-frequency regions), and the non-zero values are mainly concentrated in low-frequency regions, making its statistical distribution highly predictable. Based on this statistical characteristic, the arithmetic encoder assigns extremely short codewords to symbols with high probability of occurrence (such as zero-value sequences), thereby efficiently compressing the data and generating the final common component encoded stream.
[0160] While ensuring visual and diagnostic non-destructive performance, in order to pursue the ultimate compression efficiency, high-fidelity lossy encoding of common components can be adopted in this way, which is "visually non-destructive" or "diagnostoriically non-destructive".
[0161] Perceptually weighted lossy compression coding prioritizes preserving features sensitive to the human visual system while discarding less sensitive features during compression. This coding method does not pursue pixel-level precision but rather high quality at the perceptual level.
[0162] The second encoding model generates the differential component encoded stream through perceptually weighted lossy compression encoding, including the following steps S51-S53:
[0163] S51. Transform the input difference components from the spatial domain to the transform domain (such as the frequency domain or wavelet domain) to obtain a set of transform coefficients, which include DC and low-frequency coefficients representing low-frequency information, and AC and high-frequency coefficients representing high-frequency details.
[0164] In some embodiments, the transformation method includes, but is not limited to, Discrete Cosine Transform (DCT), Wavelet Transform, etc.
[0165] S52. Define a basic quantization step size matrix and generate a perceptual weight matrix corresponding to the transform coefficients. For example, assign larger weight values (i.e., smaller effective quantization step sizes) to frequency regions that are visually sensitive or crucial for diagnosis (typically low to mid-frequency, corresponding to smooth regions and major edges). Assign smaller weight values (i.e., larger effective quantization step sizes) to insensitive regions (typically high-frequency, corresponding to subtle textures and noise). Then, quantize each transform coefficient, as shown in the formula:
[0166]
[0167] in, These are the quantized transform coefficients. W is the basic quantization step size matrix, and W is the perceptual weight matrix. Through this operation, optimal allocation of bit resources is achieved within a limited bit rate budget.
[0168] S53. Entropy encoding is performed on the quantized transform coefficients to eliminate statistical redundancy.
[0169] The second encoding module can perform entropy encoding on the quantized transform coefficients using an arithmetic encoder. The arithmetic encoder dynamically assigns codewords based on the probability of occurrence of each symbol (quantized coefficient value), assigning short codes to symbols with high probability of occurrence (such as zero-value coefficients) and long codes to symbols with low probability of occurrence, thereby generating the final difference component encoded stream.
[0170] In some embodiments of this specification, by fusing common and differential sub-components corresponding to different dimensions, effective information aggregation and unified representation are achieved. Common structural components become more refined and robust, and the management of differential components becomes more efficient. This provides a clear framework for subsequent differentiated coding strategies, further optimizing the compression process and ensuring the overall system architecture's synergy and efficiency. Simultaneously, it intelligently balances the trade-off between compression ratio and fidelity. For example, a high-fidelity strategy is adopted for common structural components to ensure the diagnostic "skeleton" remains intact; a high-efficiency strategy is used for differential components to significantly reduce data volume. This "divide and conquer" strategy is key to achieving high-fidelity decompression (also known as decoding) at a high compression ratio.
[0171] Step 240: Integrate the common component encoding stream and the differential component encoding stream to form a compressed data stream.
[0172] Compressed data streams refer to the final compressed files that are significantly smaller in size.
[0173] Integration refers to the operation of packaging two encoded data streams together.
[0174] In some embodiments, during integration, a description file header (also known as metadata) may be added. Metadata includes at least one of the following: stream identifier information, dimension information, quantization table, model parameter identifier, and clinical context label. Stream identifier information indicates whether it is a common component encoding stream or a differential component encoding stream. Dimension information indicates the size of the original image, the number of slices, etc. The quantization table indicates the quantization step size or weight information of the differential components, which the decoder uses for inverse quantization. The model parameter identifier indicates the model parameters used by the decoder (such as the model type of the decoded model). The clinical context label contains information such as examination type and body part. Further details about the decoded model can be found in the following description.
[0175] Decoding refers to the process of recovering the original medical image data from a compressed data stream. For example, decoding can refer to high-level semantic reconstruction using generative prior models, rather than the inverse transform operation in traditional codec standards.
[0176] Integration can be achieved in several ways. For example, a strict bitstream syntax structure can be defined, concatenating or overlaying the common component encoded stream and the difference component encoded stream. A typical bitstream syntax structure is as follows: [Header Information] + [Common Component Encoded Stream] + [Difference Component Encoded Stream] + [Optional: Termination Character]. The header information contains all the metadata needed to understand the compressed encoded stream.
[0177] In some embodiments, the processor can integrate the common component encoded stream and the differential component encoded stream using an entropy encoder to form a compressed data stream.
[0178] In some embodiments, the step of the processor integrating the common component coded stream and the difference component coded stream by the entropy encoder includes the following steps S61-S62:
[0179] S61. Define a strict codestream syntax structure. A typical syntax structure is as follows: [Header information] + [Common component encoded stream] + [Different component encoded stream] + [Optional: Termination character].
[0180] S62. The entropy encoder will encode different data segments sequentially according to the above bitstream syntax structure to form the final compressed data stream.
[0181] The entropy encoder first assigns short codes to frequently occurring header fields. Then, it entropy-encodes the data from the common component code stream and the differential component code stream separately. Because the common and differential component code streams have different statistical characteristics (e.g., the common component code stream may have a more concentrated numerical distribution, while the differential component code stream is very sparse), the entropy encoder may adaptively switch or use different probability models to achieve optimal compression for each. Finally, all these encoded binary bitstreams are sequentially written to a file, forming the final compressed data stream.
[0182] In some embodiments of this specification, the entropy encoder achieves secondary compression by utilizing the probability distribution of data, further reducing the final file size. Simultaneously, through a carefully designed bitstream syntax structure, multiple dispersed data components (common component encoded stream, differential component encoded stream, and metadata) are integrated into a self-contained, structured data packet, achieving structured storage. The decoding end only needs to parse according to predetermined rules to accurately reconstruct all information, ensuring data integrity and decodeability.
[0183] Figure 3 This is one of the exemplary schematic diagrams of an image data compression method according to some embodiments of this specification.
[0184] like Figure 3As shown, the processor can input a medical image dataset containing at least two of the following: first image data, second image data, and third image data, into the decoupling model. The inter-layer decoupling module in the decoupling model processes the first image data to generate a first common sub-component and a first differential sub-component. The modal decoupling module in the decoupling model processes the second image data to generate a second common sub-component and a second differential sub-component. The temporal decoupling module in the decoupling model processes the third image data to generate a third common sub-component and a third differential sub-component. A first fusion module fuses at least two of the first, second, and third common sub-components to generate a common component, and a second fusion module fuses at least two of the first, second, and third differential sub-components to generate a differential component. A first encoding model encodes the common component to obtain a common component encoded stream, and a second encoding model encodes the differential components to obtain a differential component encoded stream. Finally, the processor integrates the common component encoded stream and the differential component encoded stream to obtain a compressed encoded stream. See the preceding descriptions for more details.
[0185] This embodiment achieves deep redundancy mining of medical image data across multiple dimensions, including interlayer, modality, and temporal sequence, through the aforementioned steps. By decoupling to obtain common and differential components, and employing optimized encoding strategies for each, it ultimately achieves extremely high overall compression efficiency while ensuring high fidelity of diagnostic information.
[0186] Figure 4 This is a second exemplary schematic diagram of an image data compression method according to some embodiments of this specification.
[0187] like Figure 4As shown, the processor can input a medical image dataset containing at least two of the following: first image data, second image data, and third image data, into the decoupling model. The inter-layer decoupling module in the decoupling model processes the first image data to generate a first common sub-component and a first differential sub-component. The modal decoupling module in the decoupling model processes the second image data to generate a second common sub-component and a second differential sub-component. The temporal decoupling module in the decoupling model processes the third image data to generate a third common sub-component and a third differential sub-component. Guided by clinical context information, the first fusion module fuses at least two of the first, second, and third common sub-components to generate a common component. Guided by clinical context information, the second fusion module fuses at least two of the first, second, and third differential sub-components to generate a differential component. The first encoding model encodes the common component to obtain a common component encoded stream, and the second encoding model encodes the differential component to obtain a differential component encoded stream. Finally, the processor integrates the common component encoded stream and the differential component encoded stream to obtain a compressed encoded stream. For more details, please refer to the relevant description above.
[0188] This embodiment systematically achieves deep decoupling and feature fusion of medical image data across three dimensions: interlayer, modality, and temporal sequence, through the aforementioned steps, and intelligently optimizes the fusion weights using clinical context information. By employing optimized encoding strategies for the generated common and differential components, extremely high overall compression efficiency is achieved while ensuring high fidelity of diagnostic information.
[0189] Some embodiments in this specification include at least the following beneficial effects:
[0190] (1) By decoupling the data into common structural components and differential components in multiple dimensions and performing differential encoding, a paradigm shift in compression from the overall data level of the PACS system was achieved. This breaks through the efficiency bottleneck of traditional methods for independent compression of single images or single modalities, thereby fundamentally and significantly improving the overall compression ratio and systematically and efficiently solving the global redundancy problem of medical image data.
[0191] (2) Overcoming the traditional contradiction between fidelity and compression ratio: By decomposing the data into parts corresponding to "stable commonalities" and "sparse differences," optimal encoding strategies can be applied to each. High-fidelity encoding is applied to the "stable commonalities" part, while efficient compression is applied to the "sparse differences" part. This ensures that the anatomical structure and pathological change information required for diagnosis are completely preserved under high compression ratios, solving the problem that traditional high compression will lose key diagnostic features.
[0192] (3) Intelligent diagnosis provides support: The "common components" and "difference components" obtained after decoupling are themselves highly refined feature representations. For example, the "difference components" in the time dimension directly quantify the dynamic changes of lesions, which can be used as radiomics features to assist doctors in quantitative assessment and disease progression prediction, so that the compression process also has a certain feature extraction function.
[0193] It is understandable that the decoupling model described above (including the decoupling module, the first fusion module, and the second fusion module), the first encoding model, and the second encoding model together constitute a complex processing chain. However, during the model training phase, a prominent technical challenge lies in how to obtain high-quality, quantifiable training targets or "labels" for the intermediate stages of the model (such as the decoupling module, the first fusion module, the second fusion module, the first encoding model, and the second encoding model corresponding to different dimensions). Attempting to train the above models independently or in stages will face the following severe challenges:
[0194] The decoupling model suffers from a lack of training labels: The ideal output of a decoupling model is a "perfect" set of common and differential components. However, in the real world, a "perfect" decoupling result cannot be defined by objective, quantifiable standards. We cannot label a medical image with "standard" common structures and differential details, making it impossible to set a clear supervised learning objective for the decoupling model.
[0195] The training objective of the coding model is ambiguous: the output of the coding model is a binary bitstream after quantization and entropy encoding. Evaluating the quality of this bitstream in isolation from the decoding model is extremely difficult. The only valid criterion for measuring coding effectiveness is the quality of the final decompressed image, which depends on the complete decoding process.
[0196] In short, the core obstacle to training alone is that the ideal output of the intermediate module is itself unknown and not directly obtainable, and there is a lack of reliable labels for supervised learning.
[0197] To address the aforementioned challenge in tag acquisition, this specification proposes a key technical implementation: end-to-end joint training of the decoupled model, the first encoding model, the second encoding model, and the decoding model.
[0198] A decoding model is a model used to decompress a compressed data stream. Decoding models can be machine learning models, such as generative artificial intelligence models. For example, a decoding model could be a diffusion model or a conditional generative adversarial network.
[0199] In some embodiments, the input to the decoding model is a compressed data stream, and the output is decompressed medical image data (or restored medical image data).
[0200] In other words, the perfect "label" for the decoding model is the raw, uncompressed medical image data. This raw data is naturally available and easily accessible. By jointly training the decoding model with the decoupled model, the first encoding model, and the second encoding model, the label source problem can be perfectly solved.
[0201] For more information on joint training, please see below. Figure 5 Related explanations.
[0202] Figure 5 This is an exemplary schematic diagram illustrating joint training according to some embodiments of this specification.
[0203] In some embodiments, such as Figure 5 As shown, the decoupling model, the first encoding model, and the second encoding model are obtained through the following joint training process: Acquire a sample medical image dataset; input the sample medical image dataset into the initial decoupling model to obtain predicted common components and predicted difference components; input the predicted common components into the initial first encoding model to obtain a predicted common component encoding stream; input the predicted difference components into the initial second encoding model to obtain a predicted difference component encoding stream; input the predicted common component encoding stream and the predicted difference component encoding stream into the initial decoding model to obtain a predicted image dataset; with the objective of minimizing the difference between the predicted image dataset and the sample medical image dataset, the parameters of the initial decoupling model, the initial first encoding model, the initial second encoding model, and the initial decoding model are jointly optimized. The above training process can be executed by a processor.
[0204] The initial decoupling model, initial first encoding model, initial second encoding model, and initial decoding model described above can be randomly initialized neural networks, or they can be models pre-trained on relevant tasks and then fine-tuned in some embodiments of this specification. For example, the initial decoding model can use a diffusion model pre-trained on a publicly available medical image dataset as a starting point.
[0205] In some embodiments, the sample dataset contains a large amount of uncompressed raw medical image data. The sample dataset can serve as a "standard answer" or "supervisory signal" for the training process, and therefore can cover multiple modalities (CT, MRI, etc.), multiple anatomical sites, and different pathological conditions to ensure that the trained model has broad generalization ability.
[0206] In some embodiments, the difference between the predicted image dataset and the sample medical image dataset can be quantified by a loss function, which may include Mean Squared Error (MSE), Multi-Scale Structural Similarity (MS-SSIM), or Learned Perceptual Image Patch Similarity (LPIPS). MSE measures the average difference at the pixel level. MS-SSIM or LPIPS better reflect differences in perceived quality.
[0207] In some embodiments, the processor can utilize a backpropagation algorithm to propagate the calculated loss value (i.e., the difference) back through the entire network, starting from the decoding model. This means the error signal sequentially passes through the decoding model, the second encoding model, the first encoding model, and finally reaches the decoupled model. Next, the processor can apply an optimization algorithm (such as Adam) to simultaneously update the parameters of all four models based on this gradient information. By iteratively adjusting the model parameters, the ultimate goal is to minimize the loss function. When the loss is sufficiently small, it means that the four models have learned how to efficiently compress and decompress compressed data with high fidelity.
[0208] In some embodiments of this specification, the decoupling model, the first encoding model, the second encoding model, and the decoding model are jointly trained end-to-end. This helps to solve the problem of difficulty in obtaining labels when training the decoupling model, the first encoding model, and the second encoding model separately, i.e., it eliminates the need to manually define "what the perfect common component should look like" or "what standards the optimal encoded stream should meet." The training signal comes directly from the quality of the final and most intuitive decompression (also known as decoding), simplifying the training process. Under the "pressure" of joint training, the encoding model (including the first encoding model and the second encoding model) guides the decoupling model to produce more "compressible" components, while the needs of the decoding model simultaneously optimize the decoupling and encoding models, forming a virtuous cycle of co-evolution. At the same time, only through end-to-end joint training can the potential of the above architecture (i.e., the decoupling model, the first encoding model, and the second encoding model) be fully realized, achieving high-fidelity decoding at extremely high compression ratios. In addition, through joint training, the parameters of the decoupled model, the first encoding model, the second encoding model, and the decoding model are optimized simultaneously, enabling the decoupled model to learn to generate feature representations that are easier to encode and decode in subsequent processes, thereby improving the overall compression efficiency and decoding quality.
[0209] In the embodiments described in this specification, the order of the steps is interchangeable unless otherwise specified, and steps may be omitted. Other steps may also be included in the operation process.
[0210] The embodiments described in this specification, which depict the system and its modules, are for convenience only and should not be considered as limiting the scope of the embodiments. The modules may be combined in any way without departing from the system principle, or formed into subsystems connected to other modules.
[0211] The embodiments in this specification are merely illustrative and not intended to limit the scope of this specification. Various modifications and alterations that can be made by those skilled in the art under the guidance of this specification remain within its scope.
[0212] Certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.
[0213] Various aspects of this specification may be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The aforementioned hardware or software may be referred to as a "data block," "module," "engine," "unit," "component," or "system," etc. Furthermore, various aspects of this specification may be presented as a computer product located on one or more computer-readable media, including computer-readable program code.
[0214] Computer storage media can be any computer-readable medium that can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. Program code located on a computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above.
[0215] The computer program code required for the operation of each part of this manual can be written in any one or more programming languages. This program code can run entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer via any network, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).
[0216] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. An image data compression method, characterized by, The method comprises: obtaining a medical image data set of a target object, the medical image data set comprising multiple types of image data originating from different data sources; inputting the medical image data set into a decoupling model, and performing decoupling processing on the medical image data set in at least two different dimensions by the decoupling model to obtain a common component and a difference component; wherein the at least two dimensions comprise at least two of an inter-layer dimension, a modality dimension, and a time sequence dimension, the decoupling model comprises at least two decoupling modules corresponding to the at least two dimensions, a first fusion module, and a second fusion module, the at least two decoupling modules comprise at least two of an inter-layer decoupling module, a modality decoupling module, and a time sequence decoupling module, wherein the inter-layer decoupling module is configured to perform decoupling processing on first image data in the medical image data set in the inter-layer dimension to obtain a first common sub-component and a first difference sub-component, the first image data being three-dimensional image data corresponding to a same scan; the modality decoupling module is configured to perform decoupling processing on second image data corresponding to different imaging modalities in the medical image data set in the modality dimension to obtain a second common sub-component and a second difference sub-component; the time sequence decoupling module is configured to perform decoupling processing on third image data corresponding to different acquisition time points in the medical image data in the time sequence dimension to obtain a third common sub-component and a third difference sub-component; the decoupling processing on the medical image data set in at least two different dimensions by the decoupling model to obtain a common component and a difference component comprises: fusing at least two of the first common sub-component, the second common sub-component, and the third common sub-component by the first fusion module to generate the common component; fusing at least two of the first difference sub-component, the second difference sub-component, and the third difference sub-component by the second fusion module to generate the difference component; inputting the common component into a first encoding model to generate a common component encoding stream; inputting the difference component into a second encoding model to generate a difference component encoding stream; integrating the common component encoding stream and the difference component encoding stream to form a compressed data stream.
2. The method of claim 1, wherein, In the decoupling processing in the inter-layer dimension, the association between image data of different layers in the three-dimensional image data corresponding to the same scan in the medical image data set is analyzed. In the decoupling processing in the modality dimension, the association between image data under different imaging modalities is analyzed. In the decoupling processing in the time sequence dimension, the association between image data at different acquisition time points is analyzed.
3. The image data compression method of claim 2, wherein the imaging modalities are divided from the data generation stage level, and comprise at least two of raw data, imaging data, and target scan data. The imaging modalities are divided from the imaging principle level, and at least two of computer tomography data, X-ray photography data, magnetic resonance imaging data, positron emission tomography data, and ultrasonic imaging data are included.
4. The image data compression method as described in claim 1, characterized in that, The decoupling model is used to decouple the medical image data set in at least two different dimensions to obtain common components and difference components. The decoupling model is used to decouple the medical image data set in three different dimensions, including: The inter-layer decoupling module in the decoupling model is used to decouple the first image data in the inter-layer dimension to obtain the first common sub-component and the first difference sub-component. The modality decoupling module in the decoupling model is used to decouple the second image data in the modality dimension to obtain the second common sub-component and the second difference sub-component. The time sequence decoupling module in the decoupling model is used to decouple the third image data in the time sequence dimension to obtain the third common sub-component and the third difference sub-component.
5. The image data compression method as described in claim 1, characterized in that, The method further includes: Clinical context information related to the target object is obtained, and the clinical context information is input into the first fusion module and the second fusion module, respectively, wherein: The first fusion module is configured to use an attention mechanism to weight and fuse the common sub-components corresponding to the at least two dimensions based on the clinical context information; The second fusion module is configured to use an attention mechanism to weight and fuse the difference sub-components corresponding to the at least two dimensions based on the clinical context information.
6. The image data compression method as described in claim 1, characterized in that, The encoding mechanisms of the first encoding model and the second encoding model are different, The first encoding model generates the common component encoding stream through lossless encoding or high-fidelity lossy encoding; The second encoding model generates the difference component encoding stream through perceptual weight-based lossy compression encoding.
7. The image data compression method as described in claim 2, characterized in that, Wherein: The decoupling processing in the inter-layer dimension is implemented through a 3D convolutional neural network; The decoupling processing in the modality dimension is implemented through a shared-private encoder architecture; The decoupling processing in the time sequence dimension is implemented through a deformation registration network and a recurrent neural network.
8. The method of claim 1, wherein the step of compressing the image data is performed by a method comprising the steps of: Wherein: The first fusion module is configured to fuse at least two of the first common sub-component, the second common sub-component, and the third common sub-component in a feature-level fusion manner to generate the common component; The second fusion module is configured to fuse at least two of the first difference sub-component, the second difference sub-component, and the third difference sub-component in a splicing manner to generate the difference component.
9. An image data compression apparatus, characterized by comprising: The device includes at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is used to execute at least part of the computer instructions to implement the image data compression method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, and when a computer reads the computer instructions in the storage medium, the computer executes the image data compression method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-modal medical data classification method, model, equipment and storage medium
CN119322973A
Multi-modal sentiment analysis method based on tensor dynamic interaction and decoupling
CN119829985A