A Massive Medical Image Data Storage Device and Method
By extracting the standardized processing of global anatomical structure and local lesion characteristics, combined with neural radiation field model and sparse dictionary learning, the storage efficiency and fidelity problems of medical image data are solved, efficient data compression and reconstruction quality are achieved, and storage density and access efficiency are optimized.
Patent Information
- Application Number
- CN202510525165.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the existing medical image data storage methods, the coordinated compression efficiency of global structure and local features is low, and the storage space-time correlation is insufficient, resulting in high-frequency details loss and anatomical structure distortion during compression, and the parameters of NeRF models are highly redundant, and the reconstruction quality is unstable.
The pre-trained anatomical structure segmentation model extracts global anatomical structure and local lesion features, generates a standardized feature tensor, and enters a neural radiation field model with medical prior constraints. Combined with weight matrix sparseness and light field sparse dictionary learning, it is decomposed into the basic anatomical layer and the detailed lesion layer, and uses xenoor erasure coding and dynamic optimization of storage locations to construct a hierarchical index table for storage.
It improves the storage efficiency and data fidelity of medical images, reduces the amount of model parameters, ensures reconstruction quality, optimizes data block access efficiency, and improves storage density and diagnostic accuracy.
Smart Images

Figure CN120072168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and in particular to a storage device and method for a large amount of medical imaging data. Background Art
[0002] Currently, for medical imaging data, it mainly relies on a voxel-based 3D segmentation model (such as 3D U-Net) to extract anatomical structure features and combines with Neural Radiance Field (NeRF) to achieve data compression. For example, after locating the lesion area through the segmentation model, the original voxel data or low-dimensional features are directly stored, while NeRF compresses the 3D image into Multi-Layer Perceptron (MLP) parameters through implicit neural representation, significantly reducing storage redundancy.
[0003] However, the global features extracted by traditional 3D segmentation models lack a collaborative coding mechanism with local lesion features, resulting in the loss of high-frequency details or anatomical structure distortion during compression; especially, the existing NeRF models do not incorporate medical prior constraints, have a high parameter redundancy and unstable reconstruction quality, and are prone to introducing noise when processing multi-modal and multi-scale images; in addition, the hierarchical storage strategy does not fully consider spatio-temporal correlation, resulting in discretization of the physical distribution of data blocks and increasing access latency. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for storing a large amount of medical imaging data to solve the problems of low collaborative compression efficiency of the global structure and local features of medical imaging data and insufficient spatio-temporal correlation in storage.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for storing a large amount of medical image data, which includes obtaining medical image data, extracting global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly mapping them to a shared implicit feature space through a fully connected layer to generate a standardized feature tensor; inputting the standardized feature tensor into a neural radiance field model with medical prior constraints, generating an auxiliary supervision signal through a voxel density distribution constrained by the anatomical hierarchy, compressing the parameters of the multi-layer perceptron using a weight matrix sparsification method, performing bit-width compression processing based on integer quantization, and generating a lightweight storage parameter set including a dictionary encoding structure, an index table, and residual encoding; decomposing it into a basic anatomical layer and a detailed lesion layer using a light field sparse dictionary learning algorithm, constructing a data block association network based on three-dimensional spatial proximity, performing spatio-temporal feature modeling in combination with historical access time correlation, generating a hierarchical index table through a clustering analysis method, and storing the basic layer data block and the detailed layer data block adjacent to each other in adjacent sectors of the same storage node according to the physical position relationship of the storage medium; for the stored basic layer data block and detailed layer data block, generating redundant check blocks using an XOR erasure code algorithm and implementing cross-node distributed storage, and dynamically optimizing the storage location through historical access frequency.
[0008] As a preferred solution of the method for storing a large amount of medical image data according to the present invention, wherein: the lightweight storage parameter set includes,
[0009] Constraining the voxel density distribution based on anatomical hierarchy labels, compressing the model weights through dynamic pruning and low-rank decomposition, and performing dimensionality reduction compression on the weight matrix using a low-rank decomposition algorithm to generate sparsified parameters;
[0010] Performing INT8 integer quantization on the sparsified parameters, mapping them to a lightweight sparse dictionary, index table, and residual data based on dictionary encoding, and converting them into a binary format and writing them to the distributed node;
[0011] Reconstructing and validating the sparse dictionary, index table, and residual data through a decoder, and outputting a lightweight storage parameter set.
[0012] As a preferred solution of the method for storing a large amount of medical image data according to the present invention, wherein: the basic anatomical layer and the detailed lesion layer respectively contain orthogonal sparse dictionary atoms corresponding to low-frequency global anatomical structure features obtained by decomposing using a light field sparse dictionary learning algorithm, as well as sparse coefficients and residuals corresponding to high-frequency local lesion features;
[0013] The basic anatomical layer is used to represent the low-frequency features of the overall anatomical structure, and the detailed lesion layer is used to represent the high-frequency features of the local lesion area.
[0014] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: the hierarchical index table is generated based on the three-dimensional spatial proximity and historical access time correlation of the basic anatomical layer data blocks and the detailed lesion layer data blocks, and records the adjacent sector addresses of the basic anatomical layer and the detailed lesion layer data blocks stored in the same node;
[0015] The three-dimensional spatial proximity is determined by calculating the spatial distance between data blocks, and the historical access time correlation is determined by calculating the ratio of the number of times the data blocks are jointly accessed to the total number of accesses.
[0016] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: generating redundant check blocks by using the XOR erasure code algorithm means grouping the basic anatomical layer data blocks and the detailed lesion layer data blocks by adjacent sectors within the storage node, performing XOR operations on each group of data blocks to generate redundant check blocks, and storing them independently in another physical node.
[0017] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: the dynamic optimization of the storage location means migrating the frequently accessed detailed lesion layer data blocks to high-IOPS storage nodes, and migrating the infrequently accessed basic anatomical layer data blocks to high-capacity storage nodes;
[0018] The determination of frequent access and infrequent access is based on the access count statistics results within a preset time period.
[0019] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: the standardized feature tensor includes spatial coordinates, modality identifiers, and anatomical level labels;
[0020] The spatial coordinates are transformed into high-dimensional position vectors through position encoding, the modality identifiers and anatomical level labels are respectively transformed into feature vectors through embedding matrices, and cross-modal association is performed through feature fusion.
[0021] Second aspect, the present invention provides a mass medical image data storage device, including a feature extraction module, an image compression module, a data layering module, and a redundancy optimization module; the feature extraction module is used to obtain medical image data, extract global anatomical structure features and local lesion features by using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; the image compression module is used to input the standardized feature tensor into a medical prior-constrained neural radiance field model, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set; the data layering module is used to decompose it into a basic anatomical layer and a detailed lesion layer by using a light field sparse dictionary learning algorithm, generate a hierarchical index table based on spatio-temporal correlation, and store the basic layer data block and the detailed layer data block in adjacent sectors of the same storage node according to spatial proximity; the redundancy optimization module is used to generate redundant check blocks for the stored basic layer data blocks and detailed layer data blocks by using an XOR erasure code algorithm, and dynamically optimize the storage location based on historical access frequencies.
[0022] Third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the mass medical image data storage method described in the first aspect of the present invention is implemented.
[0023] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the mass medical image data storage method described in the first aspect of the present invention is implemented.
[0024] The beneficial effects of the present invention are as follows: By fusing global anatomical features and local lesion features, the present invention solves the problem of compression distortion caused by heterogeneous feature spaces, and improves the storage efficiency and data fidelity of medical images. At the same time, through a medical prior-constrained neural radiance field model, combined with an anatomy label-driven auxiliary loss function and low-rank dynamic pruning technology, the number of model parameters is significantly reduced, and the reconstruction quality is ensured, overcoming the deficiencies of parameter redundancy and medical semantic disconnection of traditional NeRF models. In addition, by using light field sparse dictionary learning and spatio-temporal correlation index optimization, the basic anatomical layer and the detailed lesion layer are allocated to adjacent sectors of the same storage node according to spatial proximity, and a hierarchical index table is constructed by using spectral clustering to optimize the data block access efficiency and improve the storage density, thus achieving a significant improvement in storage efficiency and medical diagnosis accuracy. Description of the Drawings
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a flowchart of a method for storing a large amount of medical image data.
[0027] Figure 2 It is a flowchart of feature extraction and standardization for a method for storing a large amount of medical image data.
[0028] Figure 3 It is a flowchart of neural radiance field compression for a method for storing a large amount of medical image data.
[0029] Figure 4 It is a flowchart of hierarchical storage and redundancy check for a method for storing a large amount of medical image data. Specific embodiments
[0030] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0031] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0032] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0033] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides a method for storing a large amount of medical image data, including the following steps:
[0034] S1. Obtain medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space through a fully connected layer to generate a standardized feature tensor.
[0035] Specifically, it includes the following steps:
[0036] Obtain three-dimensional medical image data in DICOM format (such as CT, MRI, etc.) from publicly available medical image datasets, and perform data cleaning, data normalization, and data augmentation.
[0037] Specifically, remove invalid or damaged image data through data cleaning, perform denoising on the image data, and use the non-local means denoising algorithm to reduce noise interference; use data normalization to uniformly resample the image data to the same resolution (such as 1mm³ voxel size), perform normalization processing on the image intensity values, and map them to the [0, 1] interval to eliminate device differences; finally, uniformly resample the image data to the same resolution (such as 1mm³ voxel size), perform normalization processing on the image intensity values, and map them to the [0, 1] interval to eliminate device differences.
[0038] Preferably, use a pre-trained 3D U-Net model as the anatomical structure segmentation model because it performs excellently in medical image segmentation tasks.
[0039] It should be noted that for the pre-trained 3D U-Net model, first, a large-scale medical image dataset needs to be prepared, such as BraTS or LiTS. These medical image datasets should contain high-quality three-dimensional medical images and their corresponding anatomical structure annotations; then, preprocess the data, including uniformly resampling the images to 1mm³ voxel size, performing normalization processing on the intensity values and mapping them to the [0, 1] interval, and performing data augmentation using geometric transformations such as random rotation, translation, and scaling; then, construct the 3D U-Net model architecture, where the encoder part uses 3D convolutional layers to extract features, the decoder part gradually restores the spatial resolution through deconvolutional layers, and fuses the features of the encoder and decoder in the skip connections; in the training stage, use the cross-entropy loss function to measure the difference between the predicted segmentation result and the true annotation, and use the Adam optimizer to update the parameters. The initial learning rate is set to 1e-4, and the learning rate is decayed to 0.1 of the original value every 10 epochs; during the training process, use batch normalization layers and Dropout layers to prevent overfitting, and monitor the model performance on the validation set. Select the model with the highest Dice coefficient on the validation set as the final pre-trained 3D U-Net model; and save the weights of the pre-trained 3D U-Net model for fine-tuning and feature extraction in subsequent tasks.
[0040] Input the preprocessed three-dimensional medical image data into the pre-trained 3D U-Net model to segment the main anatomical structures (such as organs, blood vessels, etc.);
[0041] Extract global features from the encoder part of the 3D U-Net model to generate a global feature tensor;
[0042] Locate the lesion area (such as tumors, diseased tissues, etc.) in the segmentation result, and use ROI (Region of Interest) extraction technology to crop local image patches from the lesion area;
[0043] Among them, the segmentation result is a complete segmentation map containing the main anatomical structures and the lesion area.
[0044] Input the local image patches into the encoder part of the 3D U-Net model to extract local features and generate local feature tensors;
[0045] Use a fully connected layer to map the global features and local features to a feature space of the same dimension to ensure consistent feature dimensions. Specifically, adopt the Adaptive Pooling technique to unify the spatial dimensions of the feature tensors into a fixed size (such as 32×32×32);
[0046] Concatenate the global features and local features to form a joint feature tensor, and use a Transformer encoder to further encode the joint feature tensor to capture the context relationship between the global and local features and generate an implicit feature representation;
[0047] Perform layer normalization on the implicit feature representation to eliminate the difference in feature distribution, and use the L2 normalization technique to map the feature values to the unit sphere to generate a standardized feature tensor containing spatial coordinates, modality identifiers, and anatomical level labels.
[0048] S2. Input the standardized feature tensor into the neural radiance field model with medical prior constraints, generate auxiliary supervision signals through the voxel density distribution constrained by the anatomical level, adopt the method of sparse weight matrix to compress the parameters of the multi-layer perceptron, and perform bit-width compression processing based on integer quantization to generate a lightweight storage parameter set containing a dictionary encoding structure, an index table, and residual encoding.
[0049] Specifically, it includes the following steps:
[0050] Encode the anatomical level labels in the standardized feature tensor into embedding vectors through a learnable embedding matrix;
[0051] Encode the modality identifiers into embedding vectors;
[0052] Map the spatial coordinates to high-dimensional position vectors through positional encoding (Positional Encoding), and concatenate them with the standardized feature tensor, anatomical level label embedding vectors, and modality identifier embedding vectors to form a joint feature representation ;
[0053] Use a lightweight MLP (Multi-Layer Perceptron) network that includes a medical prior constraint branch for the main branch and the prior constraint branch;
[0054] Specifically, the main branch refers to predicting voxel density and color , expressed as: ;
[0055] The prior constraint branch refers to ensuring that the density distribution is consistent with the anatomical labels through anatomical structure constraints and an auxiliary loss function. The prior constraint loss function is expressed as:
[0056] ;
[0057] In the formula, is the density value of the th voxel predicted, that is, the voxel density output by the neural radiance field model, is the preset density value based on the anatomical label, that is, the density value preset according to the anatomical structure (such as bone, soft tissue, etc.) to which the th voxel belongs. For example, the density of bone is usually higher than that of soft tissue, represents the voxel index, represents the L2 norm;
[0058] It should be noted that lightweight includes low-rank decomposition and dynamic parameter pruning. Specifically, it means decomposing the weight matrix of the MLP into a low-rank matrix to reduce the number of parameters; dynamically masking irrelevant neurons according to the anatomical hierarchy label. For example, masking the high-resolution detail branch for non-lesion areas.
[0059] The low-rank matrix , expressed as:
[0060]
[0061] In the formula, represents the low-rank decomposition matrix, is the transpose matrix of the low-rank decomposition matrix , represents the transpose of the matrix and represents one of the decomposition results of the low-rank matrix ;
[0062] Furthermore, for the parameters of the trained neural radiance field model (including MLP weights, embedding matrices, etc.), methods of parameter quantization and dictionary encoding are used for compression;
[0063] Specifically, parameter quantization: converting 32-bit floating-point parameters to 8-bit integers (INT8 quantization) to reduce storage occupancy. Dictionary encoding: constructing a sparse dictionary based on the similarity of parameter distributions , map the parameters to dictionary indices and residuals.
[0064] Sparse dictionary , expressed as:
[0065] ;
[0066] In the formula, is the real number field, is the number of atoms in the dictionary, that is, the number of column vectors contained in the sparse dictionary, represents the dimension of the dictionary atoms, that is, each atom is an 8-dimensional vector;
[0067] Among them, binary format storage is achieved through parameter quantization (INT8) and dictionary encoding.
[0068] Define the implicit neural representation as the joint encoding of the neural radiance field model model parameters and the sparse dictionary:
[0069] ;
[0070] In the formula, represents the lightweight storage parameter set, represents the index mapping table, represents the residual;
[0071] Lightweight verification: Restore the original parameters through the decoder to ensure the peak signal-to-noise ratio ;
[0072] It should be noted that the lightweight storage parameter set includes the following three parts:
[0073] Specifically, it includes the sparse dictionary : The low-dimensional representation of the storage parameters for efficient compression. Index mapping: Record the corresponding index of each parameter in the dictionary. Residual: Store the difference between the parameter and the nearest neighbor in the dictionary for high-precision reconstruction.
[0074] Preferably, through parameter quantization and dictionary encoding, the huge parameter set of the traditional NeRF is compressed into a lightweight form, significantly reducing the storage space.
[0075] It should be noted that after the generation of the implicit neural representation, in order to ensure that the compressed lightweight storage parameter set can reconstruct three-dimensional medical images with high precision, the neural radiance field model must be strictly verified and optimized. Restoring the original parameters through the decoder and verifying that the peak signal-to-noise ratio (PSNR) > 40dB can initially confirm the effectiveness of the compression. However, to further improve the representation ability and reconstruction accuracy of the neural radiance field model, it is necessary to use a multi-objective loss function in the training stage, including reconstruction loss, medical prior constraints, and sparse constraints, to optimize the parameters of the neural radiance field model 。
[0076] Among them, the reconstruction loss is one of the core objectives, which is used to constrain the consistency between the rendered image and the original image, ensuring that the model can accurately capture the detailed features of medical images. Specifically, the reconstruction loss is defined as follows:
[0077] To constrain the consistency between the rendered image and the original image, the reconstruction loss function , is expressed as:
[0078] ;
[0079] In the formula, is the ray sampling point, is the rendered color, is the true color;
[0080] By using L1 regularization to encourage parameter sparsity, the sparse constraint loss function , is expressed as:
[0081] ;
[0082] In the formula, represents the sparse weight coefficient, represents calculating the absolute value of each parameter, which is used for L1 regularization;
[0083] It should be noted that the value range of the sparse weight coefficient is to , and the specific value needs to be determined through grid search, cross-validation or empirical values. A suitable value can achieve a balance between sparsity and model performance, ensuring that the model can both reduce complexity and maintain high accuracy.
[0084] Specifically, the training strategy includes a pre-training stage and a joint training stage: for the pre-training stage, only the reconstruction loss function is used to train the base ; for the joint training stage, and are added to optimize the medical prior and the lightweight objective.
[0085] Finally, the cosine annealing strategy is adopted, with an initial learning rate of 5×10−4 and a minimum learning rate of 10−6.
[0086] Preferably, the lightweight storage parameter set provides a basis for subsequent light field sparse dictionary learning, reducing the complexity of subsequent storage optimization.
[0087] S3. Decompose it into the basic anatomical layer and the detailed lesion layer by using the light field sparse dictionary learning algorithm, construct a data block association network based on three-dimensional spatial proximity, perform spatio-temporal feature modeling by combining the correlation of historical access time, generate a hierarchical index table through the clustering analysis method, and store the basic layer data block and the detailed layer data block adjacently in the adjacent sectors of the same storage node according to the physical position relationship of the storage medium.
[0088] Specifically, it includes the following steps:
[0089] Decompose the lightweight storage parameter set into two sparse subspaces, namely the basic anatomical layer and the detailed lesion layer, and construct an orthogonal sparse dictionary to eliminate redundancy.
[0090] Specifically, the basic anatomical layer (low-resolution global feature) includes the global information that retains anatomical structures such as organs and blood vessels, with high sparsity and small data volume; the detailed lesion layer (high-resolution local feature) includes the high-frequency details that retain tumor and lesion areas, with low sparsity and large data volume.
[0091] Extract the parameters of the implicit neural representation (sparse dictionary D, index IndexMap, and residuals) from the lightweight storage parameter set for dictionary construction:
[0092] It should be noted that for the basic layer dictionary construction of the low-resolution global dictionary, it specifically covers the low-frequency features of the anatomical structure, which is expressed as:
[0093] ;
[0094] In the formula, represents the basic layer dictionary, represents the number of atoms of the basic layer dictionary, where each atom corresponds to a low-frequency feature pattern, represents the dimension of the dictionary atom, which is usually consistent with the feature dimension of the data block;
[0095] For the detailed layer dictionary construction of the high-resolution local dictionary, it specifically covers the high-frequency details of the lesion area, which is expressed as:
[0096] ;
[0097] In the formula, represents the detailed layer dictionary, represents the number of atoms of the detailed layer dictionary;
[0098] Furthermore, the basic layer dictionary and the detailed layer dictionary need to satisfy the orthogonality constraint and the sparsity constraint.
[0099] Specifically, the constraint conditions: is orthogonal to to avoid feature redundancy, which is expressed as: ;
[0100] Sparsity constraint: Sparse coefficients of the base layer with sparse weights , indicating that the sparsity of the base layer is forced to be higher.
[0101] The base layer dictionary and the detail layer dictionary are jointly decomposed into two-level cascaded sparse representations through the light field sparse model to achieve collaborative constraints in the parameter space, expressed as:
[0102] ;
[0103] In the formula, represents the sparse coefficients of the base layer dictionary, represents the sparse coefficients of the detail layer dictionary, is the weight coefficient of the base layer sparse term, indicating the intensity of controlling sparsity, represents the weight coefficient of the detail layer sparse term, controlling the intensity of sparsity, represents the set of non-zero indices of the sparse coefficients, ensuring that the non-zero regions of the two layers do not overlap;
[0104] It should be noted that since the base layer mainly contains low-frequency global anatomical structure features, it has higher sparsity and less data volume. In the sparse optimization problem, is usually used to control the intensity of the base layer sparsity, ensuring that the sparse coefficients of the base layer dictionary are as sparse as possible. Therefore, usually ranges from 0.1 to 1.0.
[0105] The detail layer mainly contains high-frequency local lesion features, with lower sparsity and larger data volume. is used to control the intensity of the detail layer sparsity, ensuring that the sparse coefficients of the detail layer dictionary are as sparse as possible while retaining high-frequency details. Therefore, usually ranges from 0.01 to 0.1.
[0106] Among them, the decomposition implementation: The optimization problem will be solved using the Alternating Direction Method of Multipliers (ADMM), and and will be iteratively updated, and the low-frequency and high-frequency components will be separated through Thresholding (low frequency → base layer, high frequency → detail layer).
[0107] Define the spatio-temporal correlation of data blocks (base layer data blocks and detail layer data blocks), including spatial proximity and access time correlation, generate a hierarchical index table, and achieve fast positioning and on-demand loading of data blocks within the storage node;
[0108] For spatial proximity, by calculating the three-dimensional spatial coordinates of the data block , define the adjacency relationship:
[0109] ;
[0110] In the formula, represents the adjacency threshold (such as voxel distance ≤ 3) represents the th data block, represents the th data block, represents the set of neighboring blocks of data block , represents the index of the current data block, represents the index of the neighboring data block;
[0111] Statistical historical access sequence , construct the time correlation matrix , define the time correlation, where the matrix elements are expressed as:
[0112] ;
[0113] In the formula, represents the proportion of and being co-accessed among all the times is accessed;
[0114] It should be noted that for the statistical historical access sequence , construct the time correlation matrix ;
[0115] Among them, represents the time point of the th access , represents the total number of all data blocks in the three-dimensional medical image;
[0116] Jointly optimize the index structure:
[0117] ;
[0118] In the formula, represents the minimization objective function, and the optimization variable is , represents the value of the minimization objective function on data block , represents the value of the minimization objective function on the set of neighboring blocks of data block , represents the Frobenius norm, is the regularization parameter, used to balance the weights of the two terms, represents the mask matrix, represents the element-wise multiplication (Hadamard product), represents minimizing the objective function in the data block of the high-frequency components the value on represents the norm;
[0119] It should be noted that the value range of the regularization parameter is to , and the specific value needs to be determined through grid search, cross-validation or empirical values. The appropriate value should balance between spatial proximity and temporal correlation to ensure that the objective function can optimize both constraints simultaneously.
[0120] Use spectral clustering to group the data blocks according to spatio-temporal correlation, and sort them in ascending order of spatial coordinates within each group to generate index table entries;
[0121] Prioritize allocating the base layer data blocks to the low-address sectors of the storage nodes (such as the starting sectors 0 - 1000); allocate the detail layer data blocks to the adjacent high-address sectors (such as 1001 - 2000) to ensure physical storage continuity.
[0122] Reserve 10% of the sector capacity for each storage node to dynamically migrate the detail layer data blocks with high-frequency access, and perform storage verification and performance testing through access latency and storage density;
[0123] Specifically, measure the read time difference between the base layer and detail layer data blocks within the same node (target ≤ 1ms) as the access latency; calculate the effective data occupancy ratio per unit sector (target ≥ 85%) to output the storage density.
[0124] S4. For the stored base layer data blocks and detail layer data blocks, use the XOR erasure code algorithm to generate redundant check blocks and implement cross-node distributed storage, and dynamically optimize the storage location through historical access frequencies.
[0125] Specifically, it includes the following steps:
[0126] Group the stored base layer and detail layer data blocks according to adjacent sectors within the same storage node.
[0127] Among them, each group contains a data block consisting of 4 consecutive sectors (for example, sectors 0 - 3 are the base layer block group, and sectors 1001 - 1004 are the detail layer block group). The size of each data block is fixed at 128KB, which is aligned with the sector capacity of the storage node to ensure read and write efficiency.
[0128] Perform an exclusive OR operation bit - by - bit on the 4 data blocks within each group to generate 1 redundant check block.
[0129] The generated check block is independently stored in another physical storage node. For example, the original data block group of node A corresponds to the storage location of the check block in node B. Record the mapping relationship between the check block and the original data block through the hierarchical index table in step 3 to ensure fast positioning.
[0130] Simulate the scenario of single data block loss (such as randomly deleting any 1 block within the group), triggering the check and recovery process. Locate the check block through the index table and recover the lost data through reverse exclusive OR operation. Use the SHA - 256 hash algorithm to verify the consistency between the recovered data and the original data, and require a hash matching rate of 100%.
[0131] Record the access log of each data block in real - time, including the access time, access source (such as clinical diagnosis, scientific research analysis), and access frequency. Assign a weight of 1 to the number of accesses in the past 7 days, a weight of 0.5 to the number of accesses between 7 - 30 days, and accesses over 30 days are not included in the statistics.
[0132] Calculate the heat value of each data block, expressed as:
[0133] ;
[0134] Adopt NVMe SSD hardware as the high - IOPS storage node, which supports low latency (<1ms) and high - concurrency access (≥100K IOPS), and is dedicated to storing high - frequency detail layer data blocks (heat value ≥1000 times / week).
[0135] Adopt large - capacity HDD or QLC SSD as the high - capacity storage node to store low - frequency base layer data blocks (heat value <100 times / week).
[0136] When the detail layer data blocks reach the high - frequency threshold for 3 consecutive days and the remaining capacity of the target high - IOPS node is ≥10%, trigger the migration; when the base layer data blocks have not been accessed for 30 consecutive days, migrate them to the high - capacity node.
[0137] Suspend the read and write operations on the data blocks to be migrated (base layer or detail layer) to prevent data conflicts during the migration process. Temporarily lock the target block through a distributed lock mechanism (such as Redis lock) to ensure that the data state remains unchanged during the migration.
[0138] Copy the locked data block from the original storage node to the target node (high-IOPS or high-capacity node) completely. During the copying process, update the hierarchical index table in real time, record the new storage location (such as node ID, sector address), and ensure that subsequent accesses can be accurately routed.
[0139] After the data copying is completed, delete the original data block in the original node to release the storage space. The deletion operation needs to be recorded through a transaction log (such as MySQL Binlog) to ensure that the operation is traceable. If the deletion fails, trigger an alarm and retain the original data block.
[0140] Furthermore, when the migration process is interrupted due to node downtime or network failure, first restore the original storage location in the index table according to the transaction log and clear part of the data in the target node, and then ensure that the rolled-back data is consistent with that before migration through SHA-256 hash comparison; after the rollback is completed, initiate a 4K random read test (queue depth 32) on the data migrated to the high-IOPS node using the FIO tool. If the average latency exceeds 1 ms or the IOPS is lower than 100,000, automatically trigger load balancing or hardware expansion; at the same time, monitor the storage utilization rate of the high-capacity node in real time. When it reaches 85%, automatically mount a new hard disk to expand the storage pool, and compress and archive the basic layer data that has not been accessed for 60 consecutive days to the tape library to release space, forming a closed-loop process from exception handling to performance guarantee and then to capacity optimization.
[0141] This embodiment also provides a mass medical image data storage device, including: a feature extraction module, an image compression module, a data layering module, and a redundancy optimization module; the feature extraction module is used to obtain medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; the image compression module is used to input the standardized feature tensor into a medical prior-constrained neural radiance field model, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set; the data layering module is used to decompose it into a basic anatomical layer and a detailed lesion layer using a light field sparse dictionary learning algorithm, and generate a hierarchical index table based on spatio-temporal correlation, and store the basic layer data block and the detailed layer data block in adjacent sectors of the same storage node according to spatial proximity; the redundancy optimization module is used to generate redundancy check blocks for the stored basic layer data blocks and detailed layer data blocks using an XOR erasure code algorithm, and dynamically optimize the storage location based on historical access frequencies.
[0142] This embodiment also provides a computer device, applicable to the case of the mass medical image data storage method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the mass medical image data storage method proposed in the above embodiment.
[0143] The computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0144] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for storing a large amount of medical image data as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviated as SRAM), electrically erasable programmable read-only memory (abbreviated as EEPROM), erasable programmable read-only memory (abbreviated as EPROM), programmable read-only memory (abbreviated as PROM), read-only memory (abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0145] In summary, by integrating global anatomical features and local lesion features, the present invention solves the problem of compression distortion caused by heterogeneous feature spaces, improving the storage efficiency and data fidelity of medical images. At the same time, through a neural radiance field model with medical prior constraints, combined with an anatomy-label-driven auxiliary loss function and low-rank dynamic pruning technology, the number of model parameters is significantly reduced, and the reconstruction quality is ensured, overcoming the deficiencies of redundant parameters and medical semantic disconnection in traditional NeRF models. In addition, by using light field sparse dictionary learning and spatio-temporal correlation index optimization, the basic anatomical layer and the detailed lesion layer are assigned to adjacent sectors of the same storage node according to spatial proximity, and a hierarchical index table is constructed using spectral clustering to optimize the data block access efficiency and increase the storage density, thus achieving a significant improvement in storage efficiency and medical diagnosis accuracy.
[0146] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for storing a large amount of medical image data, characterized in that: Including, Obtain medical image data, use a pre-trained anatomical structure segmentation model to extract global anatomical structure features and local lesion features, and uniformly map them to a shared implicit feature space through a fully connected layer to generate a standardized feature tensor; Input the standardized feature tensor into a medical prior-constrained neural radiance field model, generate an auxiliary supervision signal through the voxel density distribution constrained by the anatomical hierarchy, use a weight matrix sparsification method to compress the parameters of the multi-layer perceptron, and perform bit-width compression processing based on integer quantization to generate a lightweight storage parameter set including a dictionary encoding structure, an index table, and residual encoding. The specific steps are as follows: Use a lightweight MLP network including a medical prior-constrained branch for the main branch and the prior-constrained branch; The main branch refers to the predicted voxel density and color , expressed as: ; The prior constraint branch refers to ensuring that the density distribution is consistent with the anatomical label through anatomical structure constraints and an auxiliary loss function, and the prior constraint loss function is expressed as: ; In the formula, is the predicted density value of the th voxel, that is, the voxel density output by the neural radiance field model. is the preset density value based on the anatomical label, that is, the density value preset according to the th anatomical structure to which the voxel belongs. represents the voxel index. represents the L2 norm. Among them, lightweight includes low-rank decomposition and dynamic parameter pruning, specifically referring to decomposing the weight matrix of the MLP into a low-rank matrix , reducing the number of parameters; dynamically masking irrelevant neurons according to anatomical hierarchical labels; Low-rank matrix , expressed as: ; In the formula, represents the low-rank decomposition matrix, is the low-rank decomposition matrix of the transpose matrix, represents the transpose of the matrix, representing the low-rank matrix is one of the decomposition results; Decompose it into a basic anatomical layer and a detailed lesion layer using a light field sparse dictionary learning algorithm, construct a data block association network based on three-dimensional spatial proximity, perform spatio-temporal feature modeling in combination with historical access time correlation, generate a hierarchical index table through a clustering analysis method, and store the basic layer data block and the detailed layer data block adjacent to each other in the adjacent sectors of the same storage node according to the physical location relationship of the storage medium; For the stored basic layer data block and detailed layer data block, use an XOR erasure code algorithm to generate redundant check blocks and implement cross-node distributed storage, and dynamically optimize the storage location through historical access frequencies.
2. The method for storing a large amount of medical image data according to claim 1, wherein: The lightweight storage parameter set includes: Based on the anatomical hierarchy label to constrain the voxel density distribution, compress the model weights through dynamic pruning and low-rank decomposition, use the low-rank decomposition algorithm to reduce the dimension of the weight matrix, and generate sparse parameters; Perform INT8 integer quantization on the sparse parameters, map them to a lightweight sparse dictionary, index table, and residual data based on dictionary encoding, and convert them to binary format and write them to the distributed nodes; Reconstruct and verify the sparse dictionary, index table, and residual data through a decoder, and output a lightweight storage parameter set.
3. The method for storing massive medical image data according to claim 2, characterized in that: The basic anatomical layer and the detailed lesion layer respectively contain orthogonal sparse dictionary atoms corresponding to the low-frequency global anatomical structure features obtained by decomposing through the light field sparse dictionary learning algorithm, as well as sparse coefficients and residuals corresponding to the high-frequency local lesion features; The basic anatomical layer is used to represent the low-frequency features of the overall anatomical structure, and the detailed lesion layer is used to represent the high-frequency features of the local lesion area.
4. The method for storing massive medical image data according to claim 3, characterized in that: The hierarchical index table refers to being generated based on the three-dimensional spatial proximity and historical access time correlation of the basic anatomical layer data block and the detailed lesion layer data block, and records the storage addresses of the basic anatomical layer and the detailed lesion layer data blocks in the adjacent sectors of the same node; The three-dimensional spatial proximity is determined by calculating the spatial distance between data blocks, and the historical access time correlation is determined by calculating the ratio of the number of times the data blocks are jointly accessed to the total number of accesses.
5. The method for storing a large amount of medical image data according to claim 1, wherein: The generation of redundant check blocks using the XOR erasure code algorithm means grouping the basic anatomical layer data blocks and the detailed layer lesion data blocks according to the adjacent sectors within the storage node, performing an XOR operation on each group of data blocks to generate redundant check blocks, and independently storing them to another physical node.
6. The method for storing a large amount of medical image data according to claim 5, wherein: The dynamic optimized storage location refers to migrating the data blocks of the detailed lesion layer with high-frequency access to high-IOPS storage nodes, and migrating the data blocks of the basic anatomy layer with low-frequency access to high-capacity storage nodes; The determination of high-frequency access and low-frequency access is based on the statistical result of the access times within a preset time period.
7. The method for storing a large amount of medical image data according to claim 2, wherein: The standardized feature tensor includes spatial coordinates, modality identifiers, and anatomical level labels; The spatial coordinates are transformed into high-dimensional position vectors through position encoding, the modality identifiers and anatomical level labels are respectively transformed into feature vectors through embedding matrices, and cross-modal association is performed through feature fusion.
8. A mass medical image data storage device, based on the mass medical image data storage method according to any one of claims 1 to 7, characterized in that: It includes a feature extraction module, an image compression module, a data layering module, and a redundancy optimization module; The feature extraction module is used to obtain medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; The image compression module is used to input the standardized feature tensor into a neural radiance field model with medical prior constraints, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set. The specific steps are as follows, Use a lightweight MLP network including a medical prior constraint branch for the main branch and the prior constraint branch; The main branch refers to the predicted voxel density and color , expressed as: ; The prior constraint branch refers to ensuring that the density distribution is consistent with the anatomical label through anatomical structure constraints and an auxiliary loss function. The prior constraint loss function is expressed as: ; In the formula, is the predicted density value of the th voxel, that is, the voxel density output by the neural radiance field model, is the preset density value based on anatomical labels, that is, according to the th voxel belongs to the anatomical structure of the preset density value, represents the voxel index, represents the L2 norm; Among them, lightweight includes low-rank factorization and dynamic parameter pruning, specifically referring to decomposing the weight matrix of the MLP into a low-rank matrix , reducing the number of parameters; dynamically masking irrelevant neurons according to the anatomical hierarchical labels; Low-rank matrix , expressed as: ; In the formula, represents the low-rank decomposition matrix, is the low-rank decomposition matrix of the transposed matrix, represents the transpose of the matrix, representing the low-rank matrix is one of the decomposition results; The data layering module is used to decompose it into a basic anatomy layer and a detailed lesion layer using a light field sparse dictionary learning algorithm, generate a hierarchical index table based on spatio-temporal correlation, and store the basic layer data blocks and the detailed layer data blocks in adjacent sectors of the same storage node according to spatial proximity; The redundancy optimization module is used to generate redundancy check blocks for the stored basic layer data blocks and detailed layer data blocks using the XOR erasure code algorithm, and dynamically optimize the storage location based on the historical access frequency.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for storing massive medical image data according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method for storing massive medical image data according to any one of claims 1 to 7.
Citation Information
Patent Citations
Edible novel view synthesis method based on intrinsic nerve radiation field
CN115512036A
New view angle synthesis method based on spatial progressive neural radiation field
CN118298092A