Mass medical image data storage device and method

By extracting and unifying the mapping of global and local features of medical images, and combining the neural radiation field model with medical prior constraints and the light field sparse dictionary learning algorithm, the problems of low collaborative compression efficiency and insufficient storage space-time correlation in medical image data storage are solved, and efficient storage and fidelity reconstruction are achieved.

CN120072168AActive Publication Date: 2025-05-30THE THIRD MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510525165.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing medical image data storage methods are inefficient when the global structure and local features are compressed in a coordinated manner, and the storage space-time correlation is insufficient, resulting in high-frequency details loss and anatomical structure distortion.

Method used

By acquiring medical image data, the pre-trained anatomical segmentation model is used to extract global and local features and map them uniformly to a shared implicit feature space to generate a normalized feature tensor. These features are then input into the neural radiation field model of medical prior constraints to generate auxiliary supervision signals, and use weight matrix sparseness and bit width compression techniques to generate a lightweight storage parameter set. At the same time, the light field sparse dictionary learning algorithm is used to decompose the image as the basic anatomical layer and the detailed lesion layer, and a hierarchical index table is constructed based on the spatial and temporal correlation to optimize data block storage and access.

Benefits of technology

The coordinated compression of global anatomical features and local lesion features is realized, the storage efficiency and data fidelity of medical images are improved, the amount of model parameters is reduced, the reconstruction quality is ensured, the data block access efficiency is optimized, and the storage density is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072168A_ABST
    Figure CN120072168A_ABST
Patent Text Reader

Abstract

The invention discloses a massive medical image data storage device and method, and relates to the technical field of medical image.The method comprises the steps that medical image data are obtained, global anatomical structure features and local focus features are extracted through a pre-trained anatomical structure segmentation model and are uniformly mapped to a shared implicit feature space through a full-connection layer, and an implicit feature space is obtained; generating a standardized feature tensor; inputting the standardized feature tensor into a neural radiation field model of medical priori constraint, generating an auxiliary supervision signal through voxel density distribution of anatomical hierarchy constraint, compressing a multi-layer perceptron parameter by adopting a weight matrix sparse method, executing bit width compression processing based on integer quantization, and obtaining a multi-layer perceptron model; and generating a lightweight storage parameter set comprising a dictionary coding structure, an index table and residual coding. According to the method, the global anatomical features and the local lesion features are fused, so that the problem of compression distortion caused by feature space isomerism is solved, and the storage efficiency and the data fidelity of medical images are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging technology, and particularly to a storage device and method for a large amount of medical imaging data. Background Art

[0002] Currently, for medical imaging data, it mainly relies on a voxel-based three-dimensional segmentation model (such as 3D U-Net) to extract anatomical structure features and combines with Neural Radiance Field (NeRF) to achieve data compression. For example, after the lesion area is located by the segmentation model, the original voxel data or low-dimensional features are directly stored, and NeRF compresses the three-dimensional image into Multi-Layer Perceptron (MLP) parameters through implicit neural representation, significantly reducing storage redundancy.

[0003] However, the global features extracted by the traditional three-dimensional segmentation model and the local lesion features lack a collaborative coding mechanism, resulting in the loss of high-frequency details or anatomical structure distortion during compression; especially, the existing NeRF model does not combine medical prior constraints, has a high parameter redundancy and unstable reconstruction quality, and is prone to introducing noise when processing multi-modal and multi-scale images; in addition, the hierarchical storage strategy does not fully consider spatio-temporal correlation, resulting in the physical distribution of data blocks being discretized and increasing access latency. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method for storing a large amount of medical imaging data to solve the problems of low collaborative compression efficiency of the global structure and local features of medical imaging data and insufficient spatio-temporal correlation in storage.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for storing a large amount of medical image data, which includes obtaining medical image data, extracting global anatomical structure features and local lesion features by using a pre-trained anatomical structure segmentation model, and uniformly mapping them to a shared implicit feature space through a fully connected layer to generate a standardized feature tensor; inputting the standardized feature tensor into a medical prior-constrained neural radiance field model, generating an auxiliary supervision signal through a voxel density distribution constrained by anatomical levels, compressing the parameters of a multi-layer perceptron by using a weight matrix sparsification method, performing bit-width compression processing based on integer quantization, and generating a lightweight storage parameter set including a dictionary encoding structure, an index table, and residual encoding; decomposing it into a basic anatomical layer and a detailed lesion layer by using a light field sparse dictionary learning algorithm, constructing a data block association network based on three-dimensional spatial proximity, performing spatio-temporal feature modeling in combination with historical access time correlation, generating a hierarchical index table through a clustering analysis method, and storing the basic layer data block and the detailed layer data block adjacently in adjacent sectors of the same storage node according to the physical location relationship of the storage medium; for the stored basic layer data block and detailed layer data block, generating redundant check blocks by using an XOR erasure code algorithm and implementing cross-node distributed storage, and dynamically optimizing the storage location through historical access frequency.

[0007] As a preferred solution of the method for storing a large amount of medical image data according to the present invention, wherein: the lightweight storage parameter set includes Constraining the voxel density distribution based on anatomical level labels, compressing the model weights through dynamic pruning and low-rank decomposition, and performing dimensionality reduction compression on the weight matrix by using a low-rank decomposition algorithm to generate sparse parameters; Performing INT8 integer quantization on the sparse parameters, mapping them to a lightweight sparse dictionary, index table, and residual data based on dictionary encoding, and converting them into a binary format and writing them into a distributed node; Reconstructing and verifying the sparse dictionary, index table, and residual data through a decoder, and outputting a lightweight storage parameter set.

[0008] As a preferred solution of the method for storing a large amount of medical image data according to the present invention, wherein: the basic anatomical layer and the detailed lesion layer respectively include orthogonal sparse dictionary atoms corresponding to low-frequency global anatomical structure features obtained by decomposing through a light field sparse dictionary learning algorithm, as well as sparse coefficients and residuals corresponding to high-frequency local lesion features; The basic anatomical layer is used to represent the low-frequency features of the overall anatomical structure, and the detailed lesion layer is used to represent the high-frequency features of the local lesion area.

[0009] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: the hierarchical index table is generated based on the three-dimensional spatial proximity and historical access time correlation of the basic anatomical layer data blocks and the detailed lesion layer data blocks, and records the adjacent sector addresses of the basic anatomical layer and the detailed lesion layer data blocks stored in the same node; The three-dimensional spatial proximity is determined by calculating the spatial distance between data blocks, and the historical access time correlation is determined by calculating the ratio of the number of times the data blocks are jointly accessed to the total number of access times.

[0010] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: generating redundant check blocks by using the XOR erasure code algorithm means grouping the basic anatomical layer data blocks and the detailed lesion layer data blocks by adjacent sectors within the storage node, performing XOR operations on each group of data blocks to generate redundant check blocks, and independently storing them to another physical node.

[0011] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: dynamically optimizing the storage location means migrating the frequently accessed detailed lesion layer data blocks to a high-IOPS storage node, and migrating the infrequently accessed basic anatomical layer data blocks to a high-capacity storage node; The determination of frequent access and infrequent access is based on the access count statistics results within a preset time period.

[0012] As a preferred solution of the method for storing massive medical image data according to the present invention, wherein: the standardized feature tensor includes spatial coordinates, modality identifiers, and anatomical level labels; The spatial coordinates are transformed into high-dimensional position vectors through position encoding, the modality identifiers and anatomical level labels are respectively transformed into feature vectors through embedding matrices, and cross-modal association is performed through feature fusion.

[0013] Second aspect, the present invention provides a mass medical image data storage device, including a feature extraction module, an image compression module, a data layering module, and a redundancy optimization module; the feature extraction module is used to obtain medical image data, extract global anatomical structure features and local lesion features by using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; the image compression module is used to input the standardized feature tensor into a neural radiance field model with medical prior constraints, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set; the data layering module is used to decompose it into a basic anatomical layer and a detailed lesion layer by using a light field sparse dictionary learning algorithm, generate a hierarchical index table based on spatio-temporal correlation, and store the basic layer data block and the detailed layer data block in adjacent sectors of the same storage node according to spatial proximity; the redundancy optimization module is used to generate redundancy check blocks for the stored basic layer data blocks and detailed layer data blocks by using an XOR erasure code algorithm, and dynamically optimize the storage location based on historical access frequencies.

[0014] Third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, it implements any step of the mass medical image data storage method as described in the first aspect of the present invention.

[0015] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, it implements any step of the mass medical image data storage method as described in the first aspect of the present invention.

[0016] The beneficial effects of the present invention are as follows: By fusing global anatomical features and local lesion features, the present invention solves the problem of compression distortion caused by heterogeneous feature spaces, and improves the storage efficiency and data fidelity of medical images. At the same time, through a neural radiance field model with medical prior constraints, combined with an anatomy label-driven auxiliary loss function and low-rank dynamic pruning technology, the number of model parameters is significantly reduced, and the reconstruction quality is ensured, overcoming the deficiencies of redundant parameters and medical semantics disconnection in traditional NeRF models. In addition, by using light field sparse dictionary learning and spatio-temporal correlation index optimization, the basic anatomical layer and the detailed lesion layer are allocated to adjacent sectors of the same storage node according to spatial proximity, and a hierarchical index table is constructed by using spectral clustering, optimizing the data block access efficiency and increasing the storage density, thus achieving a significant improvement in storage efficiency and medical diagnosis accuracy. Description of the Drawings

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of a method for storing massive medical image data.

[0019] Figure 2 It is a flowchart of feature extraction and standardization for a method for storing massive medical image data.

[0020] Figure 3 It is a flowchart of neural radiance field compression for a method for storing massive medical image data.

[0021] Figure 4 It is a flowchart of hierarchical storage and redundancy check for a method for storing massive medical image data. Specific Embodiments

[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.

[0023] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.

[0025] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides a method for storing massive medical image data, including the following steps: S1. Obtain medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space through a fully connected layer to generate a standardized feature tensor.

[0026] Specifically, it includes the following steps: Obtain 3D medical image data in DICOM format (such as CT, MRI, etc.) from publicly available medical image datasets, and perform data cleaning, data normalization, and data augmentation.

[0027] Specifically, remove invalid or damaged image data through data cleaning, perform denoising on the image data, and use the non-local means denoising algorithm to reduce noise interference; use data normalization to uniformly resample the image data to the same resolution (such as 1mm³ voxel size), perform normalization processing on the image intensity values, and map them to the [0, 1] interval to eliminate equipment differences; finally, uniformly resample the image data to the same resolution (such as 1mm³ voxel size), perform normalization processing on the image intensity values, and map them to the [0, 1] interval to eliminate equipment differences.

[0028] Preferably, use a pre-trained 3D U-Net model as the anatomical structure segmentation model because it performs excellently in medical image segmentation tasks.

[0029] It should be noted that for the pre-trained 3D U-Net model, first, a large-scale medical image dataset needs to be prepared, such as BraTS or LiTS. These medical image datasets should contain high-quality 3D medical images and their corresponding anatomical structure annotations; then, preprocess the data, including uniformly resampling the images to 1mm³ voxel size, performing normalization processing on the intensity values and mapping them to the [0, 1] interval, and performing data augmentation using geometric transformations such as random rotation, translation, and scaling; then construct the 3D U-Net model architecture, where the encoder part uses 3D convolutional layers to extract features, the decoder part gradually restores the spatial resolution through transposed convolutional layers, and fuses the features of the encoder and decoder in the skip connections; in the training stage, use the cross-entropy loss function to measure the difference between the predicted segmentation result and the true annotation, and use the Adam optimizer to update the parameters, with the initial learning rate set to 1e-4, and the learning rate is decayed to 0.1 of the original value every 10 epochs; during the training process, use batch normalization layers and Dropout layers to prevent overfitting, and monitor the model performance on the validation set, and select the model with the highest Dice coefficient on the validation set as the final pre-trained 3D U-Net model; and save the weights of the pre-trained 3D U-Net model for fine-tuning and feature extraction in subsequent tasks.

[0030] Input the preprocessed 3D medical image data into the pre-trained 3D U-Net model to segment the main anatomical structures (such as organs, blood vessels, etc.); Extract global features from the encoder part of the 3D U-Net model to generate a global feature tensor; Locate the lesion area (such as tumors, diseased tissues, etc.) in the segmentation result, and use ROI (Region of Interest) extraction technology to crop local image patches from the lesion area; Among them, the segmentation result is a complete segmentation map containing the main anatomical structures and the lesion area.

[0031] Input the local image patches into the encoder part of the 3D U-Net model to extract local features and generate local feature tensors; Use a fully connected layer to map the global features and local features to a feature space of the same dimension to ensure consistent feature dimensions. Specifically, adopt the Adaptive Pooling technique to unify the spatial dimensions of the feature tensors into a fixed size (such as 32×32×32); Concatenate the global features and local features to form a joint feature tensor, and use a Transformer encoder to further encode the joint feature tensor to capture the context relationship between the global and local features and generate an implicit feature representation; Perform layer normalization on the implicit feature representation to eliminate the differences in feature distributions, and use the L2 normalization technique to map the feature values to the unit sphere to generate a normalized feature tensor containing spatial coordinates, modality identifiers, and anatomical hierarchy labels.

[0032] S2. Input the normalized feature tensor into the neural radiance field model with medical prior constraints, generate auxiliary supervision signals through the voxel density distribution constrained by the anatomical hierarchy, adopt the method of sparsifying the weight matrix to compress the parameters of the multi-layer perceptron, and perform bit-width compression processing based on integer quantization to generate a lightweight storage parameter set containing a dictionary encoding structure, an index table, and residual encoding.

[0033] Specifically, it includes the following steps: Encode the anatomical hierarchy labels in the normalized feature tensor into embedding vectors through a learnable embedding matrix; Encode the modality identifiers into embedding vectors; Map the spatial coordinates to high-dimensional position vectors through positional encoding and concatenate them with the normalized feature tensor, anatomical hierarchy label embedding vectors, and modality identifier embedding vectors to form a joint feature representation ; Use a lightweight MLP (multi-layer perceptron) network with medical prior constraints for the main branch and the prior constraint branch; Specifically, the main branch refers to predicting the voxel density and color , which is expressed as: ; The prior constraint branch refers to ensuring that the density distribution is consistent with the anatomical labels through anatomical structure constraints and an auxiliary loss function. The prior constraint loss function is expressed as: ; In the formula, is the density value of the th voxel predicted, that is, the voxel density output by the neural radiance field model, is the preset density value based on the anatomical label, that is, according to the th anatomical structure (such as bone, soft tissue, etc.) to which the voxel belongs. For example, the density of bone is usually higher than that of soft tissue, represents the voxel index, represents the L2 norm; It should be noted that lightweight includes low-rank decomposition and dynamic parameter pruning. Specifically, it means decomposing the weight matrix of the MLP into a low-rank matrix , reducing the number of parameters; dynamically masking irrelevant neurons according to the anatomical hierarchical label, for example, masking the high-resolution detail branch for non-lesion regions.

[0034] The low-rank matrix is expressed as:

[0035] In the formula, represents the low-rank decomposition matrix, is the transpose matrix of the low-rank decomposition matrix , represents the transpose of the matrix, indicating one of the decomposition results of the low-rank matrix ; Furthermore, for the parameters of the trained neural radiance field model (including MLP weights, embedding matrices, etc.), parameter quantization and dictionary encoding methods are used for compression; Specifically, parameter quantization: converting 32-bit floating-point parameters into 8-bit integers (INT8 quantization) to reduce storage occupancy. Dictionary encoding: Based on the similarity of parameter distributions, a sparse dictionary is constructed to map the parameters to dictionary indices and residuals.

[0036] The sparse dictionary is expressed as: ; In the formula, is the real number field, is the number of atoms in the dictionary, that is, the number of column vectors contained in the sparse dictionary, represents the dimension of the dictionary atom, that is, each atom is an 8-dimensional vector; Among them, binary format storage is achieved through parameter quantization (INT8) and dictionary encoding.

[0037] Define the implicit neural representation as the joint encoding of the neural radiance field model model parameters and the sparse dictionary: ; In the formula, represents the lightweight storage parameter set, represents the index mapping table, represents the residual; Lightweight verification: Restore the original parameters through the decoder to ensure the peak signal-to-noise ratio ; It should be noted that the lightweight storage parameter set includes the following three parts: Specifically, it includes the sparse dictionary : The low-dimensional representation of the storage parameters for efficient compression. Index mapping: Record the corresponding index of each parameter in the dictionary. Residual: Store the difference between the parameter and the nearest neighbor in the dictionary for high-precision reconstruction.

[0038] Preferably, through parameter quantization and dictionary encoding, the huge parameter set of the traditional NeRF is compressed into a lightweight form, significantly reducing the storage space.

[0039] It should be noted that after the generation of the implicit neural representation is completed, in order to ensure that the compressed lightweight storage parameter set can reconstruct three-dimensional medical images with high precision, the neural radiance field model must be strictly verified and optimized. Restoring the original parameters through the decoder and verifying that the peak signal-to-noise ratio (PSNR)>40dB can initially confirm the effectiveness of the compression. However, in order to further improve the representation ability and reconstruction accuracy of the neural radiance field model, a multi-objective loss function needs to be adopted in the training stage, including reconstruction loss, medical prior constraint, and sparse constraint, to optimize the parameters of the neural radiance field model 。

[0040] Among them, the reconstruction loss is one of the core objectives, used to constrain the consistency between the rendered image and the original image, ensuring that the model can accurately capture the detailed features of the medical image. Specifically, the reconstruction loss is defined as follows: Constraining the consistency between the rendered image and the original image, the reconstruction loss function , is expressed as: ; In the formula, is the ray sampling point, is the rendered color, is the true color; Encouraging parameter sparsity through L1 regularization, the sparse constraint loss function , is expressed as: ; In the formula, represents the sparse weight coefficient, represents calculating the absolute value of each parameter for L1 regularization; It should be noted that the sparse weight coefficient has a value range of to , and the specific value needs to be determined through grid search, cross-validation or empirical values. A suitable value can achieve a balance between sparsity and model performance, ensuring that the model can reduce complexity while maintaining high accuracy.

[0041] Specifically, the training strategy includes a pre-training stage and a joint training stage: for the pre-training stage, only the reconstruction loss function is used to train the base ; for the joint training stage, and are added to optimize the medical prior and the lightweight objective.

[0042] Finally, the cosine annealing strategy is adopted, with an initial learning rate of 5×10−4 and a minimum learning rate of 10−6.

[0043] Preferably, the lightweight storage parameter set provides a basis for subsequent light field sparse dictionary learning, reducing the complexity of subsequent storage optimization.

[0044] S3. Decompose using the light field sparse dictionary learning algorithm into a basic anatomical layer and a detailed lesion layer, construct a data block association network based on three-dimensional spatial proximity, perform spatio-temporal feature modeling by combining historical access time correlation, generate a hierarchical index table through clustering analysis, and store the basic layer data block and the detailed layer data block adjacent to each other in adjacent sectors of the same storage node according to the physical location relationship of the storage medium.

[0045] Specifically, it includes the following steps: Decompose the lightweight storage parameter set into two sparse subspaces, namely the basic anatomical layer and the detailed lesion layer, and construct an orthogonal sparse dictionary to eliminate redundancy.

[0046] Specifically, the basic anatomical layer (low-resolution global features) includes global information retaining anatomical structures such as organs and blood vessels, with high sparsity and small data volume; the detailed lesion layer (high-resolution local features) includes high-frequency details retaining tumor and lesion areas, with low sparsity and large data volume.

[0047] Extract the parameters of the implicit neural representation (sparse dictionary D, index IndexMap, and residuals) from the lightweight storage parameter set for dictionary construction: It should be noted that for the construction of the basic layer dictionary, a low-resolution global dictionary is constructed, which specifically covers the low-frequency features of the anatomical structure and is expressed as: ; In the formula, represents the basic layer dictionary, represents the number of atoms in the basic layer dictionary, where each atom corresponds to a low-frequency feature pattern, represents the dimension of the dictionary atom, which is usually consistent with the feature dimension of the data block; The detail layer dictionary constructs a high-resolution local dictionary, which specifically covers the high-frequency details of the lesion area and is expressed as: ; In the formula, represents the detail layer dictionary, represents the number of atoms in the detail layer dictionary; Furthermore, the basic layer dictionary and the detail layer dictionary need to satisfy the orthogonality constraint and the sparsity constraint.

[0048] Specifically, the constraint conditions: is orthogonal to to avoid feature redundancy and is expressed as: ; Sparsity constraint: The sparsity weight of the basic layer sparse coefficient , indicating that the basic layer is forced to be sparser.

[0049] The basic layer dictionary and the detail layer dictionary are jointly decomposed into two-level cascaded sparse representations through the light field sparse model to achieve the collaborative constraint of the parameter space, which is expressed as: ; In the formula, represents the sparse coefficient of the basic layer dictionary, represents the sparse coefficient of the detail layer dictionary, is the weight coefficient of the basic layer sparse term, indicating the strength of controlling the sparsity, represents the weight coefficient of the detail layer sparse term, controlling the strength of the sparsity, represents the set of non-zero indices of the sparse coefficient, ensuring that the non-zero regions of the two layers do not overlap; It should be noted that since the basic layer mainly contains low-frequency global anatomical structure features, it has a high sparsity and a small amount of data. In the sparse optimization problem, is usually used to control the strength of the basic layer sparsity to ensure that the sparse coefficient of the basic layer dictionary is as sparse as possible. Therefore, usually ranges from 0.1 to 1.0.

[0050] The detail layer mainly contains high-frequency local lesion features, with low sparsity and a large amount of data. To control the sparsity intensity of the detail layer and ensure the sparse coefficients of the detail layer dictionary are as sparse as possible while retaining high-frequency details. Therefore the value range of

[0051] is usually between 0.01 and 0.1. Among them, the decomposition implementation: The alternating direction multiplier method (ADMM) will be used to solve the optimization problem, and and will be iteratively updated, and the low-frequency and high-frequency components will be separated through thresholding (low-frequency → base layer, high-frequency → detail layer).

[0052] Define the spatio-temporal correlation of data blocks (base layer data blocks and detail layer data blocks), including spatial proximity and access time correlation, generate a hierarchical index table, and achieve fast positioning and on-demand loading of data blocks within storage nodes; For spatial proximity, by calculating the three-dimensional spatial coordinates of the data block define the adjacency relationship: ; In the formula, represents the adjacency threshold (such as voxel distance ≤ 3) represents the th data block, represents the th data block, represents the set of neighboring blocks of data block , represents the index of the current data block, represents the index of the neighboring data block; Statistical historical access sequences are used to construct a time correlation matrix , and the time correlation is defined, where the matrix elements are expressed as: ; In the formula, represents the proportion of times that and and are co-accessed among all the times that It should be noted that for statistical historical access sequences , a time correlation matrix is constructed; Among them, represents the time point of the th access , represents the total number of all data blocks in the three-dimensional medical image; Joint optimization of the index structure: ; In the formula, denotes minimizing the objective function, and the optimization variable is , denotes minimizing the objective function on the data block . denotes minimizing the objective function on the set of neighboring blocks of the data block . denotes the Frobenius norm, is the regularization parameter used to balance the weights of the two terms, denotes the mask matrix, denotes element-wise multiplication (Hadamard product), denotes minimizing the objective function on the high-frequency component of the data block . denotes norm; It should be noted that the value range of the regularization parameter is to , and the specific value needs to be determined through grid search, cross-validation or empirical values. The appropriate value should balance spatial proximity and temporal correlation to ensure that the objective function can optimize both constraints simultaneously.

[0053] Use spectral clustering to group data blocks according to spatio-temporal correlation, and arrange them in ascending order of spatial coordinates within each group to generate index table entries; Prioritize allocating the base layer data blocks to the low-address sectors of the storage node (such as the starting sectors 0 - 1000); allocate the detail layer data blocks to the adjacent high-address sectors (such as 1001 - 2000) to ensure physical storage continuity.

[0054] Reserve 10% of the sector capacity for each storage node for dynamically migrating the detail layer data blocks with high-frequency access, and perform storage verification and performance testing through access latency and storage density; Specifically, measure the read time difference between the base layer and detail layer data blocks within the same node (target ≤ 1ms) as the access latency; calculate the effective data occupancy ratio per unit sector (target ≥ 85%) to output the storage density.

[0055] S4. For the stored base-layer data blocks and detail-layer data blocks, use the XOR erasure code algorithm to generate redundant check blocks and implement cross-node distributed storage, and dynamically optimize the storage location according to the historical access frequency.

[0056] Specifically, it includes the following steps: Group the stored base-layer and detail-layer data blocks according to adjacent sectors within the same storage node.

[0057] Among them, each group contains data blocks of 4 consecutive sectors (for example, sectors 0 - 3 are the base-layer block group, and sectors 1001 - 1004 are the detail-layer block group). The size of each data block is fixed at 128KB, which is aligned with the storage node sector capacity to ensure read and write efficiency.

[0058] Perform an XOR operation on the 4 data blocks within each group bit by bit to generate 1 redundant check block.

[0059] The generated check block is independently stored in another physical storage node. For example, the original data block group of node A corresponds to the storage location of the check block of node B. Record the mapping relationship between the check block and the original data block through the hierarchical index table in step 3 to ensure fast positioning.

[0060] Simulate the scenario of single data block loss (such as randomly deleting any 1 block within the group), and trigger the check recovery process. Locate the check block through the index table, and perform reverse XOR operation to recover the lost data. Use the SHA-256 hash algorithm to verify the consistency between the recovered data and the original data, and the required hash matching rate is 100%.

[0061] Record the access logs of each data block in real time, including the access time, access source (such as clinical diagnosis, scientific research analysis), and access frequency. Assign a weight of 1 to the number of accesses in the past 7 days, a weight of 0.5 to the number of accesses between 7 - 30 days, and accesses over 30 days are not included in the statistics.

[0062] Calculate the heat value of each data block, expressed as: ; Use NVMe SSD hardware as the high-IOPS storage node, which supports low latency (<1ms) and high-concurrency access (≥100K IOPS), and is dedicated to storing high-frequency detail-layer data blocks (heat value ≥1000 times / week).

[0063] Use large-capacity HDD or QLC SSD as the high-capacity storage node to store low-frequency base-layer data blocks (heat value <100 times / week).

[0064] When the detail-layer data blocks reach the high-frequency threshold for 3 consecutive days and the remaining capacity of the target high-IOPS node is ≥10%, trigger migration; if the base-layer data blocks have not been accessed for 30 consecutive days, migrate them to the high-capacity node.

[0065] Suspend the read and write operations on the data blocks to be migrated (basic layer or detail layer) to prevent data conflicts during the migration process. Temporarily lock the target block through a distributed lock mechanism (such as Redis lock) to ensure that the data state remains unchanged during the migration.

[0066] Completely copy the locked data block from the original storage node to the target node (high-IOPS or high-capacity node). During the copying process, update the hierarchical index table in real time to record the new storage location (such as node ID, sector address) to ensure that subsequent accesses can be accurately routed.

[0067] After the data copying is completed, delete the original data block in the original node to free up storage space. The deletion operation needs to be recorded through a transaction log (such as MySQL Binlog) to ensure that the operation is traceable. If the deletion fails, trigger an alarm and retain the original data block.

[0068] Furthermore, when the migration process is interrupted due to node downtime or network failure, first restore the original storage location in the index table according to the transaction log and clear part of the data on the target node, and then ensure that the rolled-back data is consistent with that before migration through SHA-256 hash comparison; after the rollback is completed, initiate a 4K random read test (queue depth 32) on the data migrated to the high-IOPS node using the FIO tool. If the average latency exceeds 1 ms or the IOPS is lower than 100,000, automatically trigger load balancing or hardware expansion; at the same time, monitor the storage utilization rate of the high-capacity node in real time. When it reaches 85%, automatically mount a new hard disk to expand the storage pool, and compress the basic layer data that has not been accessed for 60 consecutive days and archive it to the tape library to free up space, forming a closed-loop process from exception handling to performance guarantee and then to capacity optimization.

[0069] This embodiment also provides a mass medical image data storage device, including: a feature extraction module, an image compression module, a data layering module, and a redundancy optimization module; the feature extraction module is used to obtain medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; the image compression module is used to input the standardized feature tensor into a medical prior-constrained neural radiance field model, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set; the data layering module is used to decompose it into a basic anatomical layer and a detail lesion layer using a light field sparse dictionary learning algorithm, and generate a hierarchical index table based on spatio-temporal correlation, and store the basic layer data block and the detail layer data block in adjacent sectors of the same storage node according to spatial proximity; the redundancy optimization module is used to generate redundancy check blocks for the stored basic layer data blocks and detail layer data blocks using the XOR erasure code algorithm, and dynamically optimize the storage location based on the historical access frequency.

[0070] This embodiment also provides a computer device, which is applicable to the case of a method for storing massive medical image data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for storing massive medical image data proposed in the above embodiment.

[0071] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0072] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for storing massive medical image data proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0073] In summary, by integrating global anatomical features and local lesion features, the present invention solves the compression distortion problem caused by heterogeneous feature spaces, improving the storage efficiency and data fidelity of medical images. At the same time, through a neural radiance field model with medical prior constraints, combined with an auxiliary loss function driven by anatomical labels and low-rank dynamic pruning technology, the number of model parameters is significantly reduced, and the reconstruction quality is ensured, overcoming the deficiencies of traditional NeRF models such as redundant parameters and disconnection from medical semantics. In addition, by using light field sparse dictionary learning and spatio-temporal correlation index optimization, the basic anatomical layer and the detailed lesion layer are assigned to adjacent sectors of the same storage node according to spatial proximity, and a hierarchical index table is constructed using spectral clustering to optimize the data block access efficiency and increase the storage density, thus achieving a significant improvement in storage efficiency and medical diagnosis accuracy.

[0074] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for storing massive medical image data, characterized in that: include, Obtain medical imaging data, use the pre-trained anatomical structure segmentation model to extract global anatomical structure features and local lesion features, and uniformly map them to the shared implicit feature space through a fully connected layer to generate a standardized feature tensor; The standardized feature tensor is input into the neural radiation field model with medical prior constraints, and the auxiliary supervision signal is generated through the voxel density distribution with anatomical level constraints. The weight matrix sparsification method is used to compress the multi-layer perceptron parameters, and the bit width compression processing based on integer quantization is performed to generate a lightweight storage parameter set including a dictionary coding structure, an index table and a residual coding. The light field sparse dictionary learning algorithm is used to decompose the data into the basic anatomical layer and the detailed lesion layer, and a data block association network is constructed based on the three-dimensional spatial proximity. The spatiotemporal feature modeling is performed in combination with the historical access time correlation. A hierarchical index table is generated through the clustering analysis method, and the basic layer data blocks and the detailed layer data blocks are stored adjacent to each other in adjacent sectors of the same storage node according to the physical position relationship of the storage medium. For the stored basic layer data blocks and detail layer data blocks, the XOR erasure code algorithm is used to generate redundant check blocks and implement cross-node distributed storage, and the storage location is dynamically optimized based on the historical access frequency.

2. The method for storing massive medical image data according to claim 1, characterized in that: The lightweight storage parameter set includes: Based on the anatomical level label constraint voxel density distribution, the model weights are compressed through dynamic pruning and low-rank decomposition, and the low-rank decomposition algorithm is used to reduce the dimension of the weight matrix to generate sparse parameters; The sparsification parameters are quantized into INT8 integers, mapped into lightweight sparse dictionaries, index tables, and residual data based on dictionary coding, and converted into binary format and written into distributed nodes; The sparse dictionary, index table and residual data are reconstructed and verified through a decoder, and a lightweight storage parameter set is output.

3. The method for storing massive medical image data according to claim 2, characterized in that: The basic anatomical layer and the detailed lesion layer respectively contain orthogonal sparse dictionary atoms corresponding to low-frequency global anatomical structure features decomposed by a light field sparse dictionary learning algorithm, and sparse coefficients and residuals corresponding to high-frequency local lesion features; The basic anatomical layer is used to characterize the low-frequency characteristics of the overall anatomical structure, and the detailed lesion layer is used to characterize the high-frequency characteristics of the local lesion area.

4. The method for storing massive medical image data according to claim 3, characterized in that: The hierarchical index table is generated based on the three-dimensional spatial proximity and historical access time correlation of the basic anatomical layer data block and the detailed lesion layer data block, and records the adjacent sector addresses of the basic anatomical layer and the detailed lesion layer data block stored in the same node; The three-dimensional spatial proximity is determined by calculating the spatial distance between data blocks, and the historical access time correlation is determined by calculating the ratio of the number of common accesses to the data blocks to the total number of accesses.

5. The method for storing massive medical image data according to claim 1, characterized in that: The use of XOR erasure code algorithm to generate redundant check blocks refers to grouping the basic anatomical layer data blocks and the detail layer lesion layer data blocks according to adjacent sectors in the storage node, performing XOR operation on each group of data blocks to generate redundant check blocks, and independently storing them in another physical node.

6. The method for storing massive medical image data according to claim 5, characterized in that: The dynamic optimization of storage location refers to migrating the frequently accessed detail lesion layer data blocks to the high IOPS storage nodes, and migrating the infrequently accessed basic anatomical layer data blocks to the high capacity storage nodes; The determination of high-frequency access and low-frequency access is based on the statistical result of the number of accesses within a preset time period.

7. The method for storing massive medical image data according to claim 2, characterized in that: The normalized feature tensor includes spatial coordinates, modality identifiers, and anatomical level labels; The spatial coordinates are converted into high-dimensional position vectors through position encoding, and the modality identifiers and anatomical level labels are converted into feature vectors through embedding matrices, respectively, and cross-modality association is performed through feature fusion.

8. A massive medical image data storage device, based on the massive medical image data storage method according to any one of claims 1 to 7, characterized in that: Including feature extraction module, image compression module, data stratification module and redundancy optimization module; The feature extraction module is used to acquire medical image data, extract global anatomical structure features and local lesion features using a pre-trained anatomical structure segmentation model, and uniformly map them to a shared implicit feature space to generate a standardized feature tensor; The image compression module is used to input the standardized feature tensor into the neural radiation field model of the medical prior constraints, compress the three-dimensional medical image into an implicit neural representation, and output a lightweight storage parameter set; The data stratification module is used to decompose the data into a basic anatomical layer and a detailed lesion layer using a light field sparse dictionary learning algorithm, and generate a hierarchical index table based on spatiotemporal correlation, and store the basic layer data blocks and the detailed layer data blocks in adjacent sectors of the same storage node according to spatial proximity; The redundancy optimization module is used to generate redundant check blocks for the stored basic layer data blocks and detail layer data blocks using an XOR erasure code algorithm, and dynamically optimize the storage location based on historical access frequencies.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for storing massive medical image data according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for storing massive medical image data according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Edible novel view synthesis method based on intrinsic nerve radiation field

    CN115512036A

  • New view angle synthesis method based on spatial progressive neural radiation field

    CN118298092A

  • Financial big data optimization storage method

    CN118363961A

  • Enterprise management method and system based on big data

    CN118446415A

  • Intelligent storage method for medical image data

    CN118626665A