Distributed medical image data storage method based on cloud platform

By implementing the distributed medical image data storage method on the cloud platform, the gradient direction conflict caused by data distribution differences in cross-site medical image joint training is solved, and higher medical image segmentation accuracy and data consistency are achieved.

CN120067353AActive Publication Date: 2025-05-30THE THIRD MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL

Patent Information

Application Number
CN202510549346.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art has failed to effectively solve the problem of gradient direction conflict caused by data distribution differences in joint cross-site medical imaging training.

Method used

By implementing distributed medical image data storage methods on the cloud platform, including encrypted transmission, two-factor verification, data calibration and metadata extraction, standardized medical image sequences are generated and bound to structured metadata. Then, independent logical shards are formed according to the site, metadata is analyzed to identify the modal categories of medical images, set a dedicated compression path and dynamically adjust the resource weights, calculate the activity entropy value and divide the hot and cold storage layers to form a compressed shard collection. Finally, the loss gradient is calculated through the initial segmentation model, and the directionality and arbitrary gradient are aligned with historical playback data to generate synthetic medical images and filter through reinforcement learning to form a structured playback data set.

Benefits of technology

It effectively reduces the average gradient angle of heterogeneous sites, improves the Dice coefficient of pulmonary nodule segmentation, reduces the generalization error of new sites, and improves the cosine similarity of gradient direction and anatomical overlap rate between synthetic data and real data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067353A_ABST
    Figure CN120067353A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed medical image data storage method based on a cloud platform, and relates to the technical field of medical image.The method comprises the steps that original medical images of all medical sites are subjected to encryption transmission, dual verification, data calibration and metadata extraction, and a standardized medical image sequence is generated and bound with structured metadata; dividing into independent logic fragments according to sites to form standardized sub-site data streams; inputting the compressed fragment set into an initial segmentation model to calculate a loss gradient, and performing directional and arbitrary gradient alignment in combination with historical playback data; based on the gradient alignment constraint of the joint loss function and the site identifier, generating a synthetic medical image through a diffusion model, and forming a structured playback data set; the gradient direction cosine similarity of site specific synthetic data generated by the diffusion model and real data is improved, the anatomical structure overlapping rate is improved, and the historical data playback efficiency is improved in combination with dynamic prompt vector optimization of reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging technology, and in particular to a distributed medical imaging data storage method based on a cloud platform. Background Art

[0002] The cross-site medical imaging joint training technology is a key path to break medical data silos and build a general medical AI model. The current mainstream methods are based on the federated learning framework, aggregating model updates of each site through a parameter server, or synthesizing cross-site data with the help of a generative adversarial network to alleviate distribution differences. Federated learning based on gradient sharing can achieve high detection accuracy in different site scenarios.

[0003] Existing technologies have tried to improve model convergence by gradient clipping or dynamic weight adjustment, but have not solved the problem that the gradient vectors generated by backpropagation of data at different sites have directional conflicts in the parameter space. Although the gradient averaging strategy of traditional federated learning can reduce the update variance, it cannot eliminate the inconsistency of gradient directions. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a distributed medical imaging data storage method based on a cloud platform to solve the problem of gradient direction conflicts caused by data distribution differences in cross-site medical imaging joint training.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a distributed medical imaging data storage method based on a cloud platform, which includes that the original medical images of each medical site are encrypted and transmitted, double-verified, data-calibrated and metadata-extracted, generating a standardized medical image sequence and binding it with structured metadata, and dividing it into independent logical shards by site to form a standardized sub-site data stream; Parse the metadata of the sub-site data stream to identify the medical image modality category, set a dedicated compression path according to the medical image modality category and dynamically adjust the resource weight, calculate the activity entropy value, divide the cold and hot storage layers according to the entropy value, and form a compressed shard set; Input the compressed shard set into an initial segmentation model to calculate the loss gradient, perform directional and arbitrary gradient alignment in combination with historical replay data, and form a joint loss function by combining the segmentation task loss; Based on the gradient alignment constraint of the joint loss function and the site identifier, generate synthetic medical images through a diffusion model, and bind site parameters after screening by reinforcement learning to form a structured replay data set.

[0007] As a preferred solution of the distributed medical image data storage method based on the cloud platform according to the present invention, wherein: the generation of the standardized medical image sequence includes the following steps, The original medical images of each medical site are uploaded to the cloud platform through an encrypted transmission channel, and double verification of the hash algorithm is performed; Perform in-depth parsing of the DICOM file header, verify the transfer syntax identifier, and reject error files; In view of the imaging differences of different devices, the key anatomical regions are located through a pre-trained anatomical structure detection model, the weights of the bilinear interpolation kernel function are dynamically adjusted, the size of the medical images is calibrated, the key metadata is extracted from the DICOM file header, and a standardized medical image sequence is generated.

[0008] As a preferred solution of the distributed medical image data storage method based on the cloud platform according to the present invention, wherein: the formation of the standardized sub-site data stream means extracting device information, imaging protocols, and anonymized patient identifiers from the metadata, encoding the extracted metadata into JSON format structured tags, and binding them to the standardized medical image sequence.

[0009] As a preferred solution of the distributed medical image data storage method based on the cloud platform according to the present invention, wherein: the formation of the compressed shard set includes the following steps, Parse the metadata of the standardized sub-site data stream and identify the modality categories of the medical images; For CT and MRI shards, a compression path combining a three-dimensional sparse convolutional layer and a deformable attention module is adopted, the spatial weights of the lesion regions are generated through Grad-CAM heat maps, and the deformation offsets of the convolutional kernels are dynamically adjusted to retain calcification points and vascular textures; Apply discrete wavelet transform to decompose ultrasound shards into low-frequency approximation components and high-frequency detail components, and perform Huffman coding on the high-frequency sub-bands to compress blood flow noise; Fuse the identified CT shards, MRI shards, and ultrasound shards, and output a compressed shard set.

[0010] As a preferred solution of the distributed medical image data storage method based on the cloud platform according to the present invention, wherein: the division of the hot and cold storage layers according to the entropy value includes the following steps, Extract the gradient amplitude of the feature map during compression and dynamically allocate the modality weight coefficients; Statistically calculate the activity entropy value of the L1 norm of the feature map from the compressed shard set, and set an entropy threshold to divide the compressed shards into hot and cold layers: The compressed shards with entropy values higher than the entropy threshold are allocated to high-bandwidth storage nodes, and the compressed shards with entropy values lower than the entropy threshold are secondarily compressed using the Zstandard algorithm and then stored in low-performance storage nodes.

[0011] As a preferred solution of the distributed medical image data storage method based on a cloud platform according to the present invention, wherein: the generation of the joint loss function includes the following steps, Input the compressed shard set with storage level labels into the initial segmentation model, calculate the loss value between the prediction result and the true label, and generate the original gradient vector; When there is historical synthetic replay data, calculate the cosine similarity between the compressed shard set and the historical data gradient, and maximize the inner product to constrain the directional alignment; Randomly split the current site data into a virtual training set and a test set, and enforce the same gradient direction for the training set and the test set for arbitrary alignment; Fuse the directional alignment loss, the arbitrary alignment loss, and the original segmentation loss to generate the joint loss function.

[0012] As a preferred solution of the distributed medical image data storage method based on a cloud platform according to the present invention, wherein: the generation of the synthetic medical image includes the following steps, Parse the site identifier of the compressed shard set, and extract the gradient alignment information from the joint loss function as the generation constraint; Load the pre-trained diffusion model, input the compressed shard feature map and the site identifier embedding vector, and control the generation data distribution through the learnable prompt vector; In the noise addition stage, gradually superimpose Gaussian noise on the compressed shards to generate an intermediate noise sequence, and in the denoising stage, combine the site prompt vector to restore the original distribution to generate the synthetic medical image.

[0013] As a preferred solution of the distributed medical image data storage method based on a cloud platform according to the present invention, wherein: the formation of the structured replay data set means setting a reinforcement learning reward function, calculating the gradient direction cosine similarity, the feature distribution entropy, and the segmentation overlap rate between the synthetic data and the real data, screening the synthetic data with anatomical structure consistency, and binding the screened synthetic data with the site identifier, the noise step number, and the prompt vector version parameter.

[0014] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the distributed medical image data storage method based on a cloud platform as described in the first aspect of the present invention is implemented.

[0015] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, any step of the distributed medical image data storage method based on a cloud platform as described in the first aspect of the present invention is implemented.

[0016] The beneficial effects of the present invention are as follows: The dynamic interpolation algorithm based on device metadata improves the sharpness of cross-device anatomical boundaries. The DICOM metadata parsing and linear normalization eliminate outliers, providing a high-consistency data basis for subsequent training. The three-dimensional sparse convolution combined with the deformation convolution kernel guided by the Grad-CAM heatmap preserves texture features in the liver tumor region and saves storage space. The hot and cold stratification strategy based on activity entropy reduces storage costs, and at the same time, the access latency of high-priority shards is shortened. The two-stage gradient alignment mechanism reduces the average gradient angle between heterogeneous sites, improves the Dice coefficient of lung nodule segmentation, and reduces the generalization error of new sites. The cosine similarity of the gradient directions between the site-specific synthetic data generated by the diffusion model and the real data is increased, and the anatomical structure overlap rate is increased. Combining the dynamic prompt vector optimization of reinforcement learning improves the historical data playback efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the distributed medical image data storage method based on the cloud platform in Embodiment 1.

[0019] Figure 2 It is a flowchart of data standardization and sharding in Embodiment 1.

[0020] Figure 3 It is a schematic diagram of the data compression and storage stratification mechanism in Embodiment 1.

[0021] Figure 4 It is a schematic diagram of gradient alignment and synthetic data generation in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings in the specification.

[0023] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0025] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides a distributed medical image data storage method based on a cloud platform, including the following steps: S1. The original medical images (such as MRI (Magnetic Resonance Imaging), CT (Computed Tomography)) of each medical site (such as Hospital A, Hospital B) are uploaded to the cloud platform through an encrypted transmission channel, and a dual-verification mechanism is executed. Based on the SHA-3 hash algorithm, it detects possible data corruption during the transmission process, and deeply analyzes the file header metadata to verify the compliance of core elements such as the transmission syntax identifier and the required tag group, rejecting non-medical images or files with incorrect formats. For the imaging differences of different devices (such as Siemens MRI (Siemens magnetic resonance imaging device), Philips CT (Philips CT device)), an integrated pre-trained anatomical structure detection model is fine-tuned on the datasets of different devices and then deployed to real-time locate key anatomical regions (such as the edge of the corpus callosum in brain MRI, the bronchial bifurcation point in CT lung window); generate a spatial attention mask according to the organ contour, dynamically adjust the weight distribution of the bilinear interpolation kernel function, and adopt a sharpening-guided weight matrix for the anatomical boundary region to enhance the sharpness of the tissue edge; maintain standard interpolation for homogeneous regions, and the calibrated medical images remove redundant blank regions to ensure that the anatomical structure is centered and displayed, and output a medical image sequence with standardized dimensions; construct a double-layer bidirectional LSTM neural network, input the device-specific parameters extracted from the DICOM (Digital Imaging and Communications in Medicine) file header (including features such as the peak kilovoltage of CT devices, the repetition time / echo time of MRI, etc.), and output the optimal truncation threshold range; perform dynamic range truncation according to the optimal truncation threshold to eliminate outliers outside the human tissue range (such as CT value < -1000 or > 3000); then linearly normalize the maximum and minimum valid values of the current medical image; finally, apply histogram stretching to enhance the contrast to ensure the visual consistency of medical images from different sites.

[0026] Extract key metadata from the DICOM file header, including device information, imaging protocol, and patient identification.

[0027] Among them, the device information includes the manufacturer name and device model; the imaging protocol refers to the sequence type, slice thickness, and magnetic field strength; the patient identifier includes the anonymized ID, with sensitive information removed; the extracted key metadata is encoded into structured tags in JSON format and bound to the medical image correspondingly to form an image-metadata paired dataset.

[0028] The paired dataset is divided into independent logical shards according to the site identifier, and each logical shard only contains the images and metadata of the current site; the logical shards enter the training queue in a preset order, and immediately trigger the deletion of the original medical image data and the release of memory after training is completed, forming a standardized site-by-site data stream.

[0029] S2. Analyze the device type and imaging protocol fields in the metadata in the standardized site-by-site data stream, automatically identify the medical image modality category, and output a shard set classified by medical image modality. For example, if the device model field in the metadata is judged to be a high-resolution modality, and the dynamic blood flow imaging label in the imaging protocol is identified as a low-resolution time-series modality.

[0030] For different modalities, design dedicated compression paths as follows: The CT / MRI compression path adopts a three-dimensional sparse convolutional layer and embeds a deformable attention module to dynamically adjust the geometric shape of the convolutional kernel; for lesion regions such as liver tumors, according to the spatial weights generated by the Grad-CAM heatmap, the convolutional kernel produces the maximum pixel deformation offset at the tumor boundary to accurately capture heterogeneous texture features; capture multi-scale anatomical structures (such as organ boundaries, lesion regions) through multi-level dilated convolutions, retain high-frequency details (such as calcification points, vascular textures) during compression, and constrain the reconstruction quality through peak signal-to-noise ratio; the ultrasound compression path combines wavelet transform and entropy coding to decompose the image into a low-frequency approximation component (retaining the tissue contour) and a high-frequency detail component (compressing blood flow noise).

[0031] Extract the gradient amplitude from the feature map of the compression model during the compression process to quantify the contribution of different modalities to the update of the compression model parameters; design a reward function to dynamically adjust the modality weight coefficients according to the compression quality (PSNR / SSIM (structural similarity index)) and computational overhead (such as GPU memory occupancy); for example, if the PSNR (peak signal-to-noise ratio) of the CT shard drops by more than the compression threshold after compression, then reduce the weight of the PSNR to reduce resource allocation.

[0032] Perform shard compression according to the dynamically adjusted modality weight coefficients to obtain a compressed shard set, as follows: The 3D sparse convolution model for CT / MRI is loaded in slices, and the divided and compressed slices are stitched together to form a complete ciphertext slice, and verify whether the PSNR meets the standard; for ultrasound slices, the discrete wavelet transform (DWT) is applied to decompose the image, and Huffman coding is performed on the high-frequency subbands, and the SSIM and compression ratio are verified.

[0033] Extract the activation features from the compressed slice set, calculate the L1 norm of the feature maps of each channel, calculate the probability distribution, and calculate the activity entropy value. The higher the entropy value, the stronger the feature diversity (such as the heterogeneous features in the tumor area), and high-frequency access is required.

[0034] Divide the storage strategy into cold layer and hot layer according to the activity entropy value; by analyzing the distribution of the activity entropy values of the historical compressed slices, take the top 20% quantile as the activity entropy value threshold (such as entropy value ≥ 5.7), and assign the slices with activity entropy values greater than the activity entropy value threshold (such as CT slices containing lesions) to high-bandwidth storage nodes to support low-latency access. For slices with activity entropy values less than or equal to the activity entropy value threshold (such as ultrasound slices of normal tissues), the Zstandard compression algorithm is used for further compression to save storage space. Integrate the activity entropy value list with the compressed slices to form a slice set with storage level labels.

[0035] S3. Input the slice set with storage level labels into the pre-trained initial segmentation model for forward inference, and calculate the loss value between the prediction result of each sample and the true label; based on the calculated loss value, backpropagate to generate the original gradient vector of the initial segmentation model parameters, and record the gradient distribution characteristics of the current site data slice set. At the same time, if there is historical synthetic replay data, synchronously calculate the gradient vector of the historical synthetic replay data for gradient feature extraction.

[0036] For the current site data slice set and the historical replay data, calculate the gradient direction similarity between the current site data slice set and the historical replay data for directional gradient alignment, including the initial stage processing and the iterative stage processing.

[0037] Further explanation, the initial stage processing means that if the historical replay data is empty (such as when training site A for the first time), skip the directional alignment and directly enter the arbitrary alignment; the iterative stage processing means that the replay buffer loads synthetic data (such as the generated images of site A), extracts the gradient vector of the synthetic data, calculates the cosine similarity with the gradient of the current site slice set, and maximizes the inner product of the two to constrain the gradient directions to be the same.

[0038] Randomly divide the current site data (such as the CT images of site B) into a virtual training set and a virtual test set, and simulate the arbitrary gradient alignment of the cross-site data distribution differences as follows: First, the current site data is randomly split into two subsets according to the security preset ratio, and the loss values are calculated by inputting them into the segmentation model respectively to generate corresponding gradient vectors. Then, the gradient directions of the two subsets are forced to be consistent, that is, to maximize the loss gradients of the segmentation model parameters on the virtual training set and the virtual test set, so that the segmentation model maintains a stable feature extraction ability on any divided data subset, thereby improving the generalization ability for unknown sites.

[0039] If there is historical replay data, add the directional alignment loss term and the arbitrary alignment loss; combine the original segmentation task loss with the task loss adding the directional alignment loss term and the arbitrary alignment loss to form a joint loss function.

[0040] S4. Parse the site identifier for the compressed shard set and metadata label, and at the same time extract the gradient alignment information from the joint loss function as the constraint condition for generating contract data; load the pre-trained diffusion model, the input channels include the compressed shard feature map and the embedded vector encoded by the site identifier. At the same time, initialize the learnable prompt vector, and each site corresponds to an independent prompt vector, which is used to control the distribution characteristics of the generated data; gradually add Gaussian noise to the compressed shard feature map to generate an intermediate noise image sequence. For example, for the CT shard of site B, a high-noise feature map is generated after 100 steps of noise accumulation; through the denoising network of the diffusion model, combined with the site prompt vector, gradually restore the original feature distribution. For example, input the noise feature map and the hospital prompt vector to generate a synthetic medical image that conforms to the imaging characteristics of site B.

[0041] Minimize the mean square error between the noise predicted by the denoising network and the real noise to ensure the fidelity of the generated medical image; limit the number of non-zero elements of the prompt vector through L1 regularization to reduce redundant information, and generate optimized diffusion model parameters and sparse prompt vectors; design a reinforcement learning reward function to dynamically adjust the prompt vector to balance the gradient difference and generation diversity between the synthetic data and the real data, specifically as follows: Calculate the cosine similarity of the gradient directions of the synthetic data and the real data in the segmentation model. The higher the similarity, the greater the reward value; evaluate the generation diversity through the feature distribution entropy of the synthetic data, and impose a penalty when the entropy value is lower than the entropy threshold; calculate the overlap rate of the segmentation results of the synthetic data and the real data through the pre-trained segmentation model to evaluate the anatomical structure consistency.

[0042] Based on the gradient alignment result, calculate the cosine similarity between the gradient of the synthetic data and the historical replay data; set a similarity score threshold, and the data with a similarity score lower than the similarity threshold is eliminated, and the high-score data is screened out; bind the screened synthetic data with the corresponding site identifier and generation parameters (such as noise steps, prompt vector version) to form a structured replay data set.

[0043] This embodiment also provides a computer device, which is applicable to the case of a distributed medical image data storage method based on a cloud platform, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the distributed medical image data storage method based on the cloud platform proposed in the above embodiment.

[0044] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0045] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the distributed medical image data storage method based on the cloud platform proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0046] In summary, the dynamic interpolation algorithm based on device metadata in the present invention improves the sharpness of cross-device anatomical boundaries. The DICOM metadata parsing and linear normalization eliminate outliers, providing a data basis with high consistency for subsequent training. The three-dimensional sparse convolution combined with the deformable convolution kernel guided by the Grad-CAM heatmap preserves texture features in the liver tumor region and saves storage space. The hot and cold stratification strategy based on activity entropy reduces storage costs, and at the same time shortens the access latency of high-priority shards. The two-stage gradient alignment mechanism reduces the average gradient angle between heterogeneous sites, improves the Dice coefficient of lung nodule segmentation, and reduces the generalization error of new sites. The cosine similarity of the gradient direction between the site-specific synthetic data generated by the diffusion model and the real data is increased, and the anatomical structure overlap rate is increased. Combining the dynamic prompt vector optimization of reinforcement learning improves the historical data replay efficiency.

[0047] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A distributed medical image data storage method based on a cloud platform, characterized in that: include, The original medical images of each medical site are encrypted for transmission, double-verified, data calibrated, and metadata extracted to generate standardized medical image sequences and bind them with structured metadata. They are divided into independent logical slices by site to form standardized sub-site data streams. Parsing the metadata of the subsite data stream to identify the medical imaging modality category, setting a dedicated compression path according to the medical imaging modality category and dynamically adjusting the resource weight, calculating the activity entropy value, dividing the cold and hot storage layers according to the entropy value, and forming a compressed shard set; Input the compressed shard set with storage level labels into the initial segmentation model, calculate the loss value between the predicted result and the true label and generate the original gradient vector; When there is historical synthetic playback data, the cosine similarity between the compressed shard set and the historical data gradient is calculated, and the inner product is maximized to constrain the directional alignment; The current site data is randomly split into virtual training sets and test sets, forcing the gradient directions of the training set and the test set to be consistent, and performing arbitrary alignment; Fusion of directional alignment loss, arbitrary alignment loss and original segmentation loss to generate a joint loss function; Based on the gradient alignment constraints and site identifiers of the joint loss function, synthetic medical images are generated through a diffusion model, and site parameters are bound after reinforcement learning screening to form a structured playback dataset.

2. The distributed medical image data storage method based on a cloud platform as claimed in claim 1, characterized in that: The generating of the standardized medical image sequence comprises the following steps: The original medical images of each medical site are uploaded to the cloud platform through an encrypted transmission channel and double verification by hash algorithm is performed; Deeply analyze the DICOM file header, verify the transfer syntax identifier, and reject incorrect files; In view of the imaging differences of different devices, the pre-trained anatomical structure detection model is used to locate key anatomical areas, dynamically adjust the weights of the bilinear interpolation kernel function, calibrate the medical image size, extract key metadata from the DICOM file header, and generate standardized medical image sequences.

3. The distributed medical image data storage method based on a cloud platform as claimed in claim 2, characterized in that: The forming of the standardized sub-site data stream refers to encoding the extracted key metadata into JSON format structured tags and binding them to the standardized medical image sequences.

4. The distributed medical image data storage method based on a cloud platform as claimed in claim 3, characterized in that: The forming of the compressed fragment set comprises the following steps: Parse metadata of standardized subsite data streams to identify modality categories of medical images; A compression path combining three-dimensional sparse convolutional layers and deformable attention modules is used for CT and MRI slices. The spatial weight of the lesion area is generated through the Grad-CAM heat map, and the convolution kernel deformation offset is dynamically adjusted to retain the calcification points and vascular texture. Discrete wavelet transform is applied to the ultrasound slices to decompose them into low-frequency approximate components and high-frequency detail components, and Huffman coding is performed on the high-frequency sub-bands to compress the blood flow noise. The identified CT slices, MRI slices and ultrasound slices are fused and a compressed slice set is output.

5. The distributed medical image data storage method based on a cloud platform as claimed in claim 4, characterized in that: The method of dividing the cold and hot storage layers according to the entropy value includes the following steps: Extract feature map gradient amplitude during compression and dynamically assign modal weight coefficients; The L1 norm of the feature graph is used to calculate the entropy value from the compressed shard set, and the entropy threshold is set to divide the compressed shards into hot and cold layers: The compressed shards with entropy values ​​higher than the entropy threshold are allocated to high-bandwidth storage nodes, and the compressed shards with entropy values ​​lower than the entropy threshold are compressed twice using the Zstandard algorithm and stored in low-performance storage nodes.

6. The distributed medical image data storage method based on a cloud platform as claimed in claim 5, characterized in that: The generating of the synthetic medical image comprises the following steps: Parse the site identifiers of the compressed shard set and extract the gradient alignment information from the joint loss function as a generation constraint; Load the pre-trained diffusion model, input the compressed slice feature map and the site identifier embedding vector, and control the generated data distribution through the learnable prompt vector; In the noise adding stage, Gaussian noise is gradually superimposed on the compressed slice feature map to generate an intermediate noise sequence. In the denoising stage, the site prompt vector is combined to restore the original distribution and generate a synthetic medical image.

7. The distributed medical image data storage method based on a cloud platform as claimed in claim 6, characterized in that: The forming of the structured playback data set refers to setting a reinforcement learning reward function, calculating the gradient direction cosine similarity, feature distribution entropy and segmentation overlap rate between the synthetic data and the real data, screening the synthetic data with anatomical structure consistency, and binding the screened synthetic data with the site identifier, noise step number, and prompt vector version parameters.

8. The distributed medical image data storage method based on a cloud platform as claimed in claim 2, characterized in that: The key metadata includes device information, imaging protocol, and patient identification.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the distributed medical image data storage method based on a cloud platform are implemented in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the distributed medical image data storage method based on a cloud platform are implemented.

Citation Information

Patent Citations

  • Medical image data management method and device, equipment and storage medium

    CN114242210A

  • Metadata and image feature collaborative perception semi-supervised longitudinal federated learning method

    CN117038053A

  • Medical image data optimization method for medical information system

    CN118841140A

  • Federated learning with training metadata

    US20230316090A1

  • Cognitive Communications, Collaboration, Consultation and Instruction with Multimodal Media and Augmented Generative Intelligence

    US20240266074A1

Cited By

  • A cloud platform-based distributed medical image data storage method

    CN122598976A