Medical data efficient storage and reading method

By organizing medical data as a basic unit and using specific technical means such as correlation identification, modular storage, distributed storage and multi-level indexing, the problem of slow redundant storage and processing speed in existing medical data storage and reading solutions is solved, and efficient data storage and fast data query are achieved.

CN120199401APending Publication Date: 2025-06-24INSPUR FINANCIAL INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510354468.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing medical data storage and reading solutions have problems such as redundant storage, large amount of space and slow processing speed, especially inefficient in complex queries, affecting medical diagnosis and research.

Method used

The method of organizing medical data according to the patient as the basic unit is adopted, and the data is divided into basic information, medical record, image data and genetic testing data, and connected through specific correlation identifiers. Image data adopts modular storage, metadata is standardized storage, distributed storage and redundant strategies for data backup, basic information is highly compressed and encoded, multi-level index is established, and pre-match and data cache and prefetch technologies are used for data reading.

Benefits of technology

It effectively reduces the storage redundancy of medical data, improves the utilization rate of storage space, significantly improves the reading speed of complex multimodal data queries, can quickly and accurately meet the data query needs of medical personnel, and demonstrates good stability and reliability in long-term storage and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199401A_ABST
    Figure CN120199401A_ABST
Patent Text Reader

Abstract

The invention discloses a method for efficiently storing and reading medical data. The method comprises the following steps: organizing the medical data by taking a patient as a basic unit; dividing the medical data corresponding to each patient into basic information, medical record, image data and gene detection data; the divided data are connected through a specific association identifier, and the association identifier is a unique code with an index function, so that associated module data can be quickly positioned; according to the method, through innovative data structure design and storage architecture optimization, the storage redundancy of medical data can be effectively reduced, and the utilization rate of the storage space is remarkably improved; meanwhile, in the aspect of reading, the reading speed for complex multi-modal data query is greatly increased by constructing an efficient index mechanism and adopting an intelligent data reading algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for efficiently storing and retrieving medical data. Background Art

[0002] In the modern medical field, with the wide application of various medical detection devices and electronic medical record systems, the amount of medical data generated has shown an explosive growth. These medical data include different types such as patients' basic information, medical record records, diagnostic images (such as X-ray films, CT scans, MRI images, etc.), and gene detection data. The data scale is huge and the formats are complex and diverse. Existing medical data storage and retrieval solutions have many problems. On the one hand, when storing, a simple and general storage architecture is often adopted, without fully considering the special structure and relationship of medical data, resulting in a high redundancy of data storage and occupying a large amount of storage space. On the other hand, when retrieving data, especially for some complex queries, such as multi-modal data association queries (for example, associating a patient's medical record information with the results of multiple imaging examinations), the processing speed of traditional solutions is slow and the efficiency is low, seriously affecting the efficiency of medical diagnosis and research, and also being unfavorable for the long-term preservation and management of medical data. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for efficiently storing and retrieving medical data for the above problems in the prior art, and thus solve all or one of the above problems existing in the prior art.

[0004] To solve the above technical problems, the specific technical solution of the present invention is as follows: The present invention provides a method for efficiently storing and retrieving medical data, including: Organizing medical data with patients as the basic unit; The medical data corresponding to each patient is divided into basic information, medical record records, imaging data, and gene detection data; Connecting the divided data through specific association identifiers, and the association identifiers are unique codes with an indexing function, so that relevant module data can be quickly located.

[0005] As an improved solution, the imaging data is stored using an imaging data module; In the imaging data module, there are sub-modules for separately storing different types of imaging data. Inside each sub-module, the data is pre-sorted and stored in chronological order and the order of image acquisition, and detailed metadata is marked for each imaging data item; The metadata includes the collection time, the model number of the collection device, and the collection parameter information. The metadata is stored in a metadata file closely associated with the image data in a standardized format for quick query and management.

[0006] As an improved solution, it further includes: Using distributed storage to disperse the medical data with the patient as the basic unit among multiple nodes, and data backup and recovery management are carried out between nodes through a data redundancy strategy; The data redundancy strategy dynamically adjusts the storage location of replicas according to the set number of replicas and the data importance weight to ensure that data can still be accessed completely and quickly in case of partial node failures.

[0007] As an improved solution, the basic information module is stored in the basic information module; The basic information in the basic information module includes key information such as the patient's demographic information, allergy history, family medical history, etc. The basic information is stored in a highly compressed coding format, and the coding format is pre-designed based on the specific statistical distribution law of medical data, and can store the basic information in a very small amount of data without losing key information.

[0008] As an improved solution, it further includes: establishing the top-level index based on the patient's unique identifier, establishing the second-level index based on the patient's disease type or main symptoms, and establishing the third-level index based on specific examination items or test results. Through the comprehensive application of multi-level indexes, the specific medical data of the target patient can be quickly located.

[0009] As an improved solution, when reading medical data, a technology based on query condition pre-matching and data cache prefetching is adopted; When receiving a medical data query request, first perform a pre-matching query in the index structure to identify the nodes and ranges where the medical data that may meet the query requirements is located, and then prefetch the relevant cache data; The cache data includes the recently queried data and the data with a high degree of association with the query conditions, so as to improve the hit rate of data reading and reduce the response time of data reading.

[0010] As an improved solution, it further includes: When performing data query, decompose and optimize the complex query conditions, and convert the query conditions including logical operators and restrictions such as time range and data type into simplified expressions suitable for searching in the index structure, so as to improve the efficiency and accuracy of pre-matching.

[0011] As an improved solution, the gene detection data is stored in the gene detection data module; The gene detection data module encrypts and stores gene detection data using homomorphic encryption technology, enabling basic statistical analysis operations without decrypting the data. At the same time, access control technology is adopted to restrict access to gene detection data according to the roles and permissions of medical users, ensuring the security of gene detection data and the protection of patient privacy.

[0012] As an improved solution, it further includes: When medical data is updated, the associated identifiers, index structures, and backup copy data in distributed storage are automatically adjusted according to the updated content, ensuring data consistency and integrity, while reducing the impact on storage and reading performance caused by data updates.

[0013] As an improved solution, it further includes: Regularly check the integrity of the stored medical data. By verifying the correspondence between the associated identifiers, index entries, and the actual stored data, ensure that the data is always in a consistent and correct state throughout its life cycle; If data inconsistency is found, it can automatically trigger a data repair process, which includes but is not limited to operations such as re - establishing associations, updating indexes, and redownloading copy data.

[0014] The beneficial effects of the technical solution of the present invention are: Through innovative data structure design and storage architecture optimization, it can effectively reduce the storage redundancy of medical data and significantly improve the utilization rate of storage space; At the same time, in terms of reading, by constructing an efficient index mechanism and adopting an intelligent data reading algorithm, the reading speed for complex multi - modal data queries is greatly improved, and it can quickly and accurately meet the query needs of medical staff for patients' all - round data, and also shows good stability and reliability in the long - term preservation and management of medical data. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a flowchart of the method for efficient storage and reading of medical data according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the scope of protection of the present invention more clearly defined.

[0018] In the description of the present invention, it should be noted that the embodiments described herein are some, but not all, of the embodiments of the present invention; all other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0019] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this article are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment. Embodiment

[0020] This embodiment provides a method for efficient storage and retrieval of medical data, as Figure 1 shown, including: Organize medical data with patients as the basic unit; The medical data corresponding to each patient is divided into basic information, medical records, image data, and genetic test data; Connect the divided data through specific association identifiers, and the association identifiers are unique and have an indexing function, enabling quick positioning to related module data.

[0021] As an improved solution, the image data is stored using an image data module; In the image data module, there are sub-modules for separately storing different types of image data. Inside each sub-module, the data is pre-sorted and stored in chronological order and the order of image acquisition, and detailed metadata is marked for each image data item; The metadata includes acquisition time, acquisition device model, and acquisition parameter information. The metadata is stored in a standardized format in a metadata file closely associated with the image data for easy query and management.

[0022] As an improved solution, it further includes: The medical data with the patient as the basic unit is stored dispersedly on multiple nodes by using distributed storage, and data backup and recovery management are carried out between nodes through a data redundancy strategy; The data redundancy strategy dynamically adjusts the storage locations of replicas according to the set number of replicas and the data importance weights, ensuring that data can still be accessed completely and quickly in case of partial node failures.

[0023] As an improved solution, the basic information module is stored in the basic information module; The basic information in the basic information module includes key information such as the patient's demographic information, allergy history, family medical history, etc. The basic information is stored in a highly compressed coding format, and the coding format is pre-designed based on the specific statistical distribution law of medical data. Without losing key information, the basic information can be stored with extremely small data volume.

[0024] As an improved solution, it also includes: establishing the top-level index based on the patient's unique identifier, establishing the second-level index based on the patient's disease type or main symptoms, and establishing the third-level index based on specific examination items or test results. By comprehensively using the multi-level index, the specific medical data of the target patient can be quickly located.

[0025] As an improved solution, when reading medical data, a technology based on query condition pre-matching and data cache prefetching is adopted; When receiving a medical data query request, first perform a pre-matching query in the index structure to identify the nodes and ranges where the medical data that may meet the query requirements is located, and then prefetch the relevant cache data; The cache data includes the recently queried data and the data with a high degree of correlation with the query conditions, thereby improving the hit rate of data reading and reducing the response time of data reading.

[0026] As an improved solution, it also includes: When performing a data query, decompose and optimize the conversion of complex query conditions, and convert the query conditions containing logical operators, time ranges, data types, etc. into simplified expressions suitable for searching in the index structure, so as to improve the efficiency and accuracy of pre-matching.

[0027] As an improved solution, the gene detection data is stored in the gene detection data module; The gene detection data module encrypts and stores the gene detection data by using homomorphic encryption technology, enabling basic statistical analysis operations without decrypting the data. At the same time, access control technology is adopted to restrict the access of medical users to the gene detection data according to their roles and permissions, ensuring the security of gene detection data and the protection of patient privacy.

[0028] As an improved solution, it further includes: When the medical data is updated, the associated identifier, index structure, and backup copy data in the distributed storage are automatically adjusted according to the updated content to ensure the consistency and integrity of the data, while reducing the impact on storage and reading performance caused by data updates.

[0029] As an improved solution, it further includes: Regularly perform integrity checks on the stored medical data. By verifying the correspondence between the associated identifier, index entries, and the actual stored data, ensure that the data is always in a consistent and correct state throughout its life cycle; If data inconsistency is found, it can automatically trigger a data repair process, and the repair process includes, but is not limited to, operations such as re - establishing associations, updating indexes, and redownloading copy data.

[0030] It should be noted that the examples here are only for explaining the present invention and should not limit the protection scope of the present invention accordingly.

[0031] It should be understood that in various embodiments of this article, the magnitudes of the serial numbers of the above - mentioned processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this article.

[0032] It should also be understood that in the embodiments of this article, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front - and - back associated objects.

[0033] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this article.

[0034] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific logical processes of the methods described above can refer to the corresponding working processes of the systems, devices, and units in the foregoing method embodiments, and will not be elaborated here.

[0035] In several embodiments provided in this document, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed coupling, direct coupling, or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.

[0036] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments in this document.

[0037] In addition, in each embodiment of this document, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0038] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this document. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0039] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. All equivalent structural or equivalent process transformations made by using the specifications and drawings of the present invention, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A method for efficiently storing and reading medical data, characterized in that: include: Organize medical data based on patients as the basic unit; The medical data corresponding to each patient is divided into basic information, medical records, imaging data and genetic testing data; The divided data are connected by a specific association identifier, which is a unique code with an index function, so that the associated module data can be quickly located.

2. The method for efficiently storing and reading medical data according to claim 1, characterized in that: The image data is stored using an image data module; The image data module includes submodules for storing different types of image data separately. Each submodule pre-sorts and stores the data in chronological order and in the order of image acquisition, and annotates each image data item with detailed metadata. The metadata includes acquisition time, acquisition equipment model, and acquisition parameter information. The metadata is stored in a metadata file closely associated with the image data in a standardized format to facilitate rapid query and management.

3. The method for efficiently storing and reading medical data according to claim 1, characterized in that: Also includes: Distributed storage is used to store the medical data of the patient as the basic unit in multiple nodes, and data backup and recovery management are performed between nodes through data redundancy strategies; The data redundancy strategy dynamically adjusts the storage location of the replicas according to the set number of replicas and the data importance weight, ensuring that the data can still be fully and quickly accessed when some nodes fail.

4. The method for efficiently storing and reading medical data according to claim 1, characterized in that: The basic information module is stored in the basic information module; The basic information in the basic information module includes the patient's demographic information, allergy history, family medical history and other key information. The basic information is stored in a highly compressed coding format, which is pre-designed based on the specific statistical distribution rules of medical data. Without losing key information, the basic information can be stored in a very small amount of data.

5. The method for efficiently storing and reading medical data according to claim 1, characterized in that: Also includes: The top-level index is established based on the patient's unique identifier, the second-level index is established based on the patient's disease type or main symptoms, and the third-level index is established based on specific examination items or test results. Through the comprehensive use of multi-level indexes, the specific medical data of the target patient can be quickly located.

6. The method for efficiently storing and reading medical data according to claim 1, characterized in that: When reading medical data, the technology based on query condition pre-matching and data cache pre-fetching is adopted; When a medical data query request is received, a pre-match query is first performed in the index structure to identify the nodes and ranges where medical data that may meet the query requirements are located, and then the relevant cache data is pre-fetched; The cached data includes recently queried data and data that is highly correlated with the query conditions, thereby improving the hit rate of data reading and reducing the response time of data reading.

7. The method for efficiently storing and reading medical data according to claim 6, characterized in that: Also includes: When performing data queries, complex query conditions are decomposed and optimized, and query conditions containing logical operators and restrictions such as time range and data type are converted into simplified expressions suitable for searching in the index structure, thereby improving the efficiency and accuracy of pre-matching.

8. The method for efficiently storing and reading medical data according to claim 1, characterized in that: The gene detection data is stored in a gene detection data module; The genetic testing data module uses homomorphic encryption technology to encrypt and store genetic testing data, so that basic statistical analysis operations can be performed without decrypting the data. At the same time, access control technology is used to limit the access of medical users to genetic testing data according to their roles and permissions, ensuring the security of genetic testing data and the protection of patient privacy.

9. The method for efficiently storing and reading medical data according to claim 1, characterized in that: Also includes: When medical data is updated, the data's associated identifiers, index structure, and backup copy data in distributed storage are automatically adjusted based on the updated content to ensure data consistency and integrity while reducing the impact on storage and reading performance caused by data updates.

10. The method for efficiently storing and reading medical data according to claim 1, characterized in that: Also includes: Regularly perform integrity checks on stored medical data to ensure that the data is always consistent and correct during its life cycle by verifying the correspondence between the data's associated identifiers, index entries, and the actual stored data; If data inconsistency is found, the data repair process can be automatically triggered, including but not limited to re-establishing associations, updating indexes, and re-downloading copy data.