Distributed medical big data storage method
By adopting real-time traffic analysis, rack perception technology and data feature classification methods in distributed storage systems, combined with blockchain technology, the problems of low efficiency and insufficient data security in the existing technology are solved, and efficient and secure medical big data storage and management are achieved.
Patent Information
- Application Number
- CN202510026706.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
Existing distributed storage technology is difficult to efficiently manage large amounts of small files and optimize storage resource allocation, resulting in reduced storage efficiency and complex metadata management. The centralized storage model has the risk of data leakage and tampering, which cannot meet the high integrity and immutability of medical data.
Dynamic routing optimization based on real-time traffic analysis, intelligent distribution and migration based on rack-awareness, and distributed medical big data storage methods based on classification and merging of medical data characteristics are adopted. Data operations are recorded through blockchain technology to ensure the authenticity and immutability of data.
It significantly improves data reading and writing speed and transmission efficiency, reduces data reading delay, enhances the security and collaboration of data management, and supports efficient retrieval and distribution cooperation of medical data.
Smart Images

Figure CN119938791A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of medical information technology and computer technology, and in particular to a distributed medical big data storage method. Background Art
[0002] With the rapid advancement of the digitalization process in the medical industry, medical data has become massive, diverse, and heterogeneous. These data include structured data (such as basic patient information), semi-structured data (such as electronic medical records), and unstructured data (such as medical images). The core significance of distributed medical big data storage technology is to solve the problems of efficient storage, fast reading, and secure management of massive medical data, support data sharing and collaboration among medical institutions, and thus improve the efficiency of medical services and the level of scientific research.
[0003] Existing distributed storage technologies (such as Hadoop / HDFS, GoogleGFS, GlusterFS, Minio, etc.) have the following defects: There are a large number of small files in the medical field, and traditional systems find it difficult to efficiently manage these files, resulting in reduced storage efficiency and complex metadata management; in the face of medical data needs with strong real-time and sudden nature, existing technologies find it difficult to optimize storage resource allocation, often leading to hot spot bottlenecks and read and write delays; centralized storage models have the risk of data leakage and tampering, and cannot meet the requirements of high integrity and non-tamperability of medical data; existing technologies manage metadata in a relatively extensive manner and cannot effectively support efficient retrieval and distribution collaboration of medical data. Summary of the invention
[0004] In view of the many problems existing in the above-mentioned prior art, the present invention provides a distributed medical big data storage method, which is based on dynamic routing optimization of real-time traffic analysis, intelligent distribution and migration based on rack perception, and classification and merging based on medical data features. The technology of the present invention improves storage efficiency, reduces data reading latency, and enhances the security and collaboration of data management.
[0005] A distributed medical big data storage method comprises the following steps:
[0006] Collect medical data and pre-process the collected data to generate data suitable for distributed storage;
[0007] Based on a distributed storage network, the data transmission path is dynamically optimized by analyzing the real-time monitoring data of data flow characteristics, and the path adjustment operation is recorded using blockchain technology;
[0008] According to the storage requirements of medical data, based on the data type, access frequency and available resources of the storage nodes, the data is intelligently allocated to the target storage nodes in the distributed storage system; by real-time monitoring of the resource load of the storage nodes, dynamic adjustments are triggered when the resource utilization exceeds the preset threshold, and an optimized storage distribution plan is generated;
[0009] Classify and merge medical data. Based on the structural characteristics, relevance and access frequency of medical data, use clustering algorithms based on association rules for structured data, semantic analysis algorithms for semi-structured data, and image recognition and feature extraction technologies for unstructured data. Determine the optimal storage node based on the type of optimization algorithm, including multi-attribute decision-making models and utility maximization methods, and use blockchain technology to record merging and storage distribution operations.
[0010] In the data read request, the optimal read path is determined according to the optimized storage distribution scheme, and the required medical data is returned after verifying the data integrity.
[0011] Preferably, the medical data is preprocessed by cleaning the original data to remove duplicate records, interpolating missing values or completing them based on historical data, and using format conversion tools to unify data from different sources into a standardized structure format.
[0012] Preferably, the format conversion tool uses a hierarchical parsing technique to parse the labels and contents of the semi-structured data into a corresponding relationship between attributes and values, and converts the unstructured data into a searchable file format containing metadata descriptions.
[0013] Preferably, the dynamic optimization of the data transmission path is based on real-time analysis of traffic size and network delay, and the optimal path is selected by calculating the path cost value, which is calculated by weighted values of traffic load, delay time and number of transmission hops between nodes.
[0014] Preferably, in the calculation of the path cost value, when the network delay exceeds a preset threshold, the weight of the delay factor is increased, while the weight of the hop factor is reduced.
[0015] Preferably, the operation of intelligently allocating data to the target storage node is performed by dynamically evaluating the node's disk capacity, CPU utilization, and network bandwidth utilization, and performing predictive allocation based on the storage node's historical load change trend.
[0016] Preferably, the triggering conditions for dynamic adjustment include cross-node migration when disk utilization exceeds a preset threshold, optimizing the transmission path when network bandwidth utilization is higher than a preset threshold, or starting high-priority task allocation when CPU utilization is lower than a predetermined lower limit.
[0017] Preferably, the classification and merging of the medical data are performed through an analysis method based on data characteristics, wherein structured data are clustered according to patient identification and medical records, semi-structured data are classified by extracting key diagnostic terms through semantic analysis technology, and unstructured data are grouped and processed based on image type and patient association.
[0018] Preferably, the semantic analysis technology extracts diagnosis-related keywords by performing word segmentation processing on the diagnosis text in the semi-structured data, combining word frequency statistics and context semantic matching algorithms, and determines the classification group of the data accordingly.
[0019] Preferably, the data integrity is verified by matching the hash value of the blockchain record, and the specific steps of verification include calculating the hash value of the current data and comparing it one by one with the original hash value of the blockchain record.
[0020] Compared with the prior art, the advantages and beneficial effects of the present invention are:
[0021] The present invention realizes dynamic optimization of data transmission paths and network load balancing through the multi-factor storage acceleration technology of distributed storage network IP, significantly improving data reading and writing speed and transmission efficiency;
[0022] The present invention realizes dynamic data distribution and load balancing based on rack resources through distributed rack-aware storage optimization technology, solving storage performance bottlenecks and overload problems;
[0023] The present invention realizes efficient data classification and file distribution optimization through small file classification and merging technology based on medical data characteristics, thereby improving the management efficiency of small files;
[0024] The present invention ensures the authenticity and non-tamperability of data operation records through the introduction of blockchain technology, and enhances the security and credibility of the system;
[0025] The present invention realizes intelligent classification and rapid retrieval of various medical data through active metadata management and intelligent distribution technology, and promotes medical data sharing and collaboration. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the process of the present invention;
[0027] Figure 2 It is a schematic diagram of the data flow analysis and path optimization process in the present invention;
[0028] Figure 3 It is a schematic diagram of the dynamic data distribution and load balancing process in the present invention;
[0029] Figure 4It is a schematic diagram of the small file classification and merging process of the present invention. DETAILED DESCRIPTION
[0030] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0031] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0032] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0033] Traditional network routing usually adopts fixed strategies, which are difficult to flexibly adjust according to real-time traffic changes and specific data access patterns, resulting in congestion during network peak hours and increased data transmission delays, affecting the normal development of medical services, such as image data transmission jams in remote medical diagnosis and slow response of electronic medical record systems.
[0034] The lack of effective network operation records and verification mechanisms makes it difficult to accurately trace the root cause and responsibility of the problem once a network failure or data loss occurs, and the integrity and reliability of the data cannot be guaranteed. For medical data, which has extremely high requirements for accuracy and security, there are great risks. Figure 1-Figure 2 As shown, the present invention provides the following technical contents:
[0035] Assume that there are n network devices (nodes) in the network topology, and use the set V = {v 1, v2,......,v n.,} means. For any two nodes v i and v j The link between them defines the weight function W(v i, v j )=(b ij ,d ij ,l ij ,cij , r ij , h ij , e ij ), The input parameters are introduced as shown in the following table:
[0036] Table 1 Parameters for constructing the full graph model
[0037]
[0038]
[0039] Shortest path algorithm variant (improved based on Dijkstra's algorithm): Define the path cost function where p is a path from the source node to the target node, and α and β are weight adjustment factors (which can be adjusted according to the network's emphasis on bandwidth, latency, and the resource situation of each node, etc.).
[0040] That is: α × (load / bandwidth + I / O bandwidth occupancy rate) + β × (latency + (1 - CPU idle rate) / 2 + number of routing hops - read / write efficiency)
[0041] The algorithm steps include:
[0042] Initialization: Mark the distance of the source node as 0, and mark the distances of other nodes as infinity. Create a set S = {} of visited nodes and a set Q = V of unvisited nodes.
[0043] Loop through nodes: When Q is not empty, select the node u with the minimum distance from Q and add it to S. For each adjacent node v of u, calculate the new path cost C new (v) = C(p u ) + C((u, v)) (where p u is the current optimal path to node u). If C new (v) < C(v), then update the distance of v to C new (v) and record its predecessor node as u.
[0044] End condition: When the target node is added to S, the algorithm ends, and the optimal path is obtained by backtracking the predecessor nodes.
[0045] Input parameters, network topology information: including network devices (nodes) and their connection relationships, the initial bandwidth b ij , latency d ij information of each link. At the same time, it is also necessary to obtain the CPU idle rate c ij , I / O bandwidth occupancy rate r ij , number of routing hops h ij to other relevant nodes, and the nearest read / write efficiency e ijAnd other characteristic data.
[0046] Real-time traffic data, real-time traffic load of each link ij , as well as the source and target information of data traffic requests between different IP network segments (such as traffic requests from the IP network segment of the outpatient area to the IP network segment of the imaging diagnosis area).
[0047] The weight adjustment factors α and β are used to adjust the weights of bandwidth, delay, and node resource-related factors in the path cost function in path selection according to network requirements.
[0048] Output parameter, optimal data transmission path p (the smaller the output value, the better): the sequence of network device nodes from the source IP segment to the target IP segment, along which data traffic should be transmitted to optimize network performance. At the same time, the future availability model built based on the resource characteristics of each node can also output the expected availability of each node in the future (depending on the specific forecast period) (such as the predicted value of each resource indicator, etc.), to assist network managers in making planning work such as resource allocation in advance.
[0049] The weight adjustment factors and adjustments include:
[0050] Reference medical services: For high real-time requirements (such as intraoperative image transmission), increase α (such as 0.7) and reduce β (0.3) to ensure real-time requirements; for large-scale backup, increase β (0.6) and reduce α (0.4) to ensure bandwidth requirements.
[0051] According to network congestion: if the latency is high and the bandwidth is sufficient, the latency will be increased by α (0.6-0.7); if the bandwidth bottleneck is high, the bandwidth will be increased by β (0.7-0.8).
[0052] The adjustment of node feature related parameters includes:
[0053] CPU idle rate: low idle (<30%) increases the penalty factor (1.2-1.5); high idle reduces the factor (0.3-0.4).
[0054] I / O bandwidth ratio: if it is too high (>80%), the penalty will be increased; if it is too low, the coefficient will be reduced (0.3).
[0055] Routing hop count and read / write efficiency: Large-scale networks increase the hop count weight (0.2-0.3); efficient node link plus reward projects are given priority.
[0056] Based on the distributed storage IP address pool, that is, the hadoopd data node IP address pool, a traffic monitoring module is established to track the traffic conditions of different network segments in real time, intelligently adjust the routing strategy according to the medical data access mode, and use the blockchain to ensure network stability and data integrity, thereby improving data reading and writing speeds.
[0057] like Figure 3As shown in the figure, the distributed rack-aware data access acceleration includes:
[0058] Resource information integration and evaluation: When the system is initialized, it collects and integrates resource information such as rack disks, networks, CPUs, and memory, and builds a database after evaluation and marking, laying the foundation for subsequent storage decisions, making resource utilization more reasonable and efficient, and improving the overall planning and stability of the storage system.
[0059] Intelligent data distribution: When new medical data enters, AI identifies the type and estimates the access frequency. Based on the data distribution scoring mechanism and rack resources, small files are stored in racks with low latency, fast read and write speeds and sufficient CPU, and large imaging data are stored in racks with high bandwidth and large disks, ensuring efficient data access.
[0060] Real-time load monitoring and balancing: The system monitors the CPU, network, disk I / O, and memory loads of the rack in real time during operation. When the threshold is exceeded, AI migrates part of the data of the overloaded rack to the lightly loaded rack based on data migration logic and multiple factors, maintaining system stability and efficiency and avoiding performance impacts due to overload.
[0061] Blockchain ensures data credibility: When data is distributed and migrated, blockchain records operation details, including data identification, time, rack information, etc., to provide a basis for auditing, enhance system credibility and security, meet the medical industry's requirements for data integrity and privacy, and ensure that data is traceable throughout the process.
[0062] The present invention solves the problem that traditional distributed storage mostly adopts static layout, does not consider the differences in rack resources and data access modes, causes high-frequency access data to be stored in poor racks, and leads to performance bottlenecks during peak periods, such as remote diagnosis delays and medical record system freezes, which affects medical services and reduces efficiency and service quality.
[0063] The present invention solves the problem that traditional storage encounters uneven rack loads, lacks an automatic adjustment mechanism, and cannot intelligently migrate data when overloaded, affecting system performance and even making data inaccessible for storage, increasing the risk of loss and failure. This is extremely harmful to the integrity and availability of medical data because it is related to the life and health of patients.
[0064] The present invention solves the problems of traditional distributed storage in data operations, lack of reliable recording and verification mechanisms, difficulty in tracing the root causes and responsible parties when problems occur, and inability to meet the needs of medical regulatory audits. It is easy to lead to erroneous conclusions and decisions in data sharing and scientific research, endangering patient safety.
[0065] The data distribution scoring function of the present invention includes:
[0066] Data distribution scoring function: Let the data type be T, the access frequency be F, the disk read and write speed of rack R be S, the network connection bandwidth be B, and the disk capacity be D c , the CPU usage is Cr , the memory usage is M r , the data distribution score DS(T,F,R) can be defined as:
[0067]
[0068] where w 1, w 2, w 3, w 4, w5 is the weight coefficient, which is adjusted according to the different emphasis of the data on each factor and the actual situation of the system. cavg It is the average disk capacity of all racks and is used for normalization so that the capacity factors of different racks are comparable.
[0069] Load evaluation function, for rack R, let its CPU usage be C r , the network bandwidth utilization is B r , the disk I / O load is IO r , the memory usage is M r , the load evaluation value L(R) can be expressed as:
[0070] L(R)=w c ×C r +w b ×B r +w i ×IO r +w m ×M r
[0071] Among them, w c ,w b ,w i , w m It is a weight coefficient, which is used to adjust the influence of each load factor on the overall load assessment according to the actual situation.
[0072] Data migration decision function, when the load evaluation value L(R) of rack R is greater than the preset overload threshold TH, data migration is triggered. Let the amount of data to be migrated be MD, and migrate from the overloaded rack R to the target rack R t , the target rack is selected based on minimizing the load increase after the target rack is migrated. Let the target rack R before and after migration t The load evaluation values are L(R t ) and L new (R t ),but
[0073]
[0074] Among them, D crt is the remaining disk capacity of the target rack, Mrt is the remaining memory capacity of the target rack, λ and u are coefficients related to the impact of data migration on disk and memory load, and L is selected so that new (R t )The smallest target rack for data migration.
[0075] When the system is initialized, it evaluates and marks the rack storage resources. When storing data, it intelligently distributes data based on data type, access frequency, and the status of multi-dimensional rack resources (network, CPU, memory, disk), monitors the load in real time, and migrates as needed to achieve load balancing. At the same time, it uses blockchain to record data distribution and migration information to ensure its reliability and traceability, thereby improving system performance and credibility.
[0076] Intelligently sense rack resources and medical data characteristics to achieve precise adaptive storage; real-time monitoring and automatic adjustment of load balancing; blockchain ensures the authenticity, integrity and traceability of data throughout its life cycle, solving traditional storage problems.
[0077] File merging based on medical data features: Using dedicated algorithms, medical files are accurately classified and merged based on data relevance, keyword semantics, and the relationship between images and patients. Structured, semi-structured, and unstructured files have their own methods. At the same time, the merged information is recorded on the blockchain to improve efficiency, ensure credibility, promote the circulation and application of medical data, and enhance the efficiency of management and utilization.
[0078] Optimization of file distribution in distributed storage: With the help of multi-factor intelligent distribution algorithm, intelligent and comprehensive evaluation of storage node resources, location and access mode, intelligent distribution of merged medical files, local files stored nearby, multi-node backup of important files, blockchain-secured operations, reduced latency, improved availability, promoted medical services and research, and ensured data quality.
[0079] The present invention solves the problem that traditional medical data storage is not integrated according to features, and structured, semi-structured and unstructured data are scattered, resulting in chaotic management, inefficient query, and difficulty for doctors to quickly obtain complete patient information. For example, when making a diagnosis, data needs to be pieced together, which is time-consuming and affects efficiency. File merging can accurately classify data, make data clear and easy to manage and query, and solve this problem.
[0080] The present invention solves the problem that the previous distributed medical storage did not coordinately optimize the node resources, location and access mode, resulting in unbalanced node load, high latency and poor availability, such as slow transmission and data loss in remote consultation. File distribution optimization intelligently allocates files, balances loads, reduces latency, increases fault tolerance, ensures efficient storage access, and alleviates the above problems.
[0081] This invention solves the problem that in medical data sharing and research, traditional methods cannot guarantee the legality of data, the authenticity and integrity of operations, there is a risk of tampering, and it is difficult to comply with regulations and research requirements. Blockchain technology solves this trust crisis, provides a verification and traceability mechanism with tamper-proof records, ensures data security and reliability, and builds a solid defense line for its legal application and sharing.
[0082] For structured medical data, a clustering model based on data association rules is constructed, and key identifiers such as patient ID are used as the association basis to cluster related data files together.
[0083] For semi-structured medical data, keyword extraction and semantic analysis algorithms in natural language processing (NLP) technology are used to identify characteristic vocabulary and semantic information related to disease departments, and file merging is performed based on this.
[0084] For unstructured medical imaging data, image recognition and patient information matching algorithms are used to determine their merging relationships based on image types (such as X-ray, CT, MRI, etc.) and patient identity information.
[0085] Using the multi-attribute decision-making theory, the resource attributes of storage nodes (such as storage capacity, read / write speed, CPU performance, etc.), geographical location attributes (such as network distance and latency between nodes, etc.), and data access mode attributes (such as access frequency and access time distribution, etc.) are used as decision factors to construct a file distribution model based on utility maximization. Through a comprehensive evaluation of these factors, the storage utility value of each storage node for different medical files is calculated to determine the optimal file distribution plan.
[0086] The comprehensive score and utility function calculation expression is:
[0087]
[0088] Table 2 Comprehensive score and utility function parameter table
[0089]
[0090]
[0091]
[0092]
[0093] This function is used to comprehensively evaluate the priority of medical files in the entire processing flow, including the priority in the merge operation and the priority in selecting storage nodes in the distributed storage system. By reasonably setting the weight coefficient, the emphasis on different factors can be flexibly adjusted according to actual needs, thereby achieving optimized decisions on medical data storage and management.
[0094] P t The larger the value, the higher the priority of distributed storage node selection and merging.
[0095] like Figure 4 As shown, the steps for merging files based on medical data features are:
[0096] Data collection and preprocessing: Collect various medical documents, perform data cleaning and format standardization on structured documents; perform text parsing and preprocessing on semi-structured documents; perform image enhancement and feature extraction on unstructured image files for subsequent analysis and processing.
[0097] Document classification: Using the classification algorithm mentioned above, medical documents are classified into three categories: structured, semi-structured, and unstructured.
[0098] Merge decision: Calculate the merge priority of each file based on the file merge scoring formula. For files with high priority, determine their merge method and target file group.
[0099] File merging execution: According to the merging decision, relevant files are merged and the merged file is generated.
[0100] Blockchain records: During the file merging process, blockchain technology is used to record detailed information about the merging operation, including the source of the merged files, the basis for the merging, the time of the merging, etc.
[0101] Steps to optimize file distribution in distributed storage:
[0102] System initialization: Collect resource information (storage capacity, read / write speed, CPU performance, etc.), geographic location information (network topology, inter-node latency, etc.) and historical data access pattern information (access frequency, access time distribution, etc.) of each node in the distributed storage system, and initialize the weight coefficient in the file distribution utility formula.
[0103] File allocation decision: When a new merged medical file needs to be stored, the storage utility value of each storage node for the file is calculated according to the file distribution utility formula.
[0104] Node selection and file storage: Select the storage node with the highest utility value as the storage location of the file, store the file on the node, and update the resource usage and historical access information of the node.
[0105] Blockchain verification and synchronization: After the file is stored in the node, the file distribution operation is managed and verified using blockchain technology. Each node confirms the correctness of the file distribution operation through the blockchain consensus mechanism to ensure the consistency and integrity of medical data in the distributed storage system.
[0106] The innovative medical document classification and merging algorithm achieves precise merging based on the characteristics of different data structures, improving data management efficiency and availability.
[0107] The multi-factor intelligent file distribution algorithm comprehensively considers storage node resources, geographical location and access patterns to optimize distributed storage performance and ensure data security and efficient access.
[0108] The introduction of blockchain technology ensures data credibility throughout the process, records file merging and distribution operations, ensures the legality, authenticity, consistency and integrity of data, meets the strict security and regulatory requirements for medical data, promotes the circulation and application of medical data, and promotes the digital development of the medical industry.
[0109] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware.
[0110] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A distributed medical big data storage method, characterized in that: The following steps are involved: Collect medical data and pre-process the collected data to generate data suitable for distributed storage; Based on a distributed storage network, the data transmission path is dynamically optimized by analyzing the real-time monitoring data of data flow characteristics, and the path adjustment operation is recorded using blockchain technology; According to the storage requirements of medical data, intelligently allocate data to the target storage nodes in the distributed storage system based on the data type, access frequency and available resources of the storage nodes; By monitoring the resource load of storage nodes in real time, dynamic adjustments are triggered when resource utilization exceeds the preset threshold to generate an optimized storage distribution plan; Classify and merge medical data. Based on the structural characteristics, relevance and access frequency of medical data, clustering algorithms based on association rules are used for structured data, semantic analysis algorithms are used for semi-structured data, and image recognition and feature extraction technologies are used for unstructured data. Determine the optimal storage nodes based on the type of optimization algorithm, including multi-attribute decision-making models and utility maximization methods, and use blockchain technology to record merging and storage distribution operations; In the data read request, the optimal read path is determined according to the optimized storage distribution scheme, and the required medical data is returned after verifying the data integrity.
2. The distributed medical big data storage method according to claim 1, characterized in that: The medical data preprocessing includes cleaning the original data to remove duplicate records, interpolating missing values or completing them based on historical data, and using format conversion tools to unify data from different sources into a standardized structure format.
3. The distributed medical big data storage method according to claim 2, characterized in that: The format conversion tool adopts hierarchical parsing technology to parse the labels and contents of semi-structured data into corresponding relationships between attributes and values, and converts unstructured data into a searchable file format containing metadata descriptions.
4. The distributed medical big data storage method according to claim 1, characterized in that: The dynamic optimization of the data transmission path is based on real-time analysis of traffic size and network delay, and the optimal path is selected by calculating the path cost value, which is calculated by weighted values of traffic load, delay time and number of transmission hops between nodes.
5. The distributed medical big data storage method according to claim 4, characterized in that: In the calculation of the path cost value, when the network delay exceeds a preset threshold, the weight of the delay factor is increased, and the weight of the hop factor is reduced.
6. The distributed medical big data storage method according to claim 1, characterized in that: The operation of intelligently allocating data to the target storage node dynamically evaluates the node's disk capacity, CPU utilization, and network bandwidth utilization, and performs predictive allocation based on the storage node's historical load change trend.
7. The distributed medical big data storage method according to claim 6, characterized in that: The trigger conditions for dynamic adjustment include cross-node migration when disk utilization exceeds a preset threshold, optimizing the transmission path when network bandwidth utilization is above a preset threshold, or initiating high-priority task allocation when CPU utilization is below a predetermined lower limit.
8. The distributed medical big data storage method according to claim 1, characterized in that: The classification and merging of the medical data are performed through an analysis method based on data characteristics, wherein structured data are clustered based on patient identification and medical records, semi-structured data are classified by extracting key diagnostic terms through semantic analysis technology, and unstructured data are grouped and processed based on image type and patient association.
9. The distributed medical big data storage method according to claim 8, characterized in that: The semantic analysis technology extracts diagnosis-related keywords by performing word segmentation on the diagnosis text in the semi-structured data, combining word frequency statistics and context semantic matching algorithms, and determines the classification group of the data accordingly.
10. The distributed medical big data storage method according to claim 1, characterized in that: The data integrity is verified by matching the hash value of the blockchain record. The specific steps of the verification include calculating the hash value of the current data and comparing it one by one with the original hash value of the blockchain record.
Citation Information
Cited By
Medical image data storage method and device and storage medium
CN120823939A
Intelligent medical record data management system and method
CN120977469A
Clinical laboratory data rapid processing and sharing method based on big data
CN121011300A
Medical data informatization safety communication method and system based on big data
CN122419824A