Medical data processing system and method

By performing format conversion, standardization, and intelligent algorithm deduplication on medical data, combined with distributed storage and deep learning analysis, the problems of low efficiency and low quality in existing technologies have been solved, achieving efficient and intelligent medical data processing.

CN120913730APending Publication Date: 2025-11-07SUZHOU GUOKE MEDICAL TECH DEV CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510765192.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies are inefficient in medical data processing, have low data quality, and lack automation, while also incurring high costs for manual intervention.

Method used

By collecting medical data from multiple sources and performing format conversion and standardization, intelligent algorithms are used to remove duplicate data and outliers. Distributed storage and processing technologies are employed, combined with statistical analysis, machine learning, and deep learning for in-depth analysis.

Benefits of technology

It significantly improves data processing efficiency and quality, reduces the cost of manual intervention, ensures data security and reliability, and achieves efficient and intelligent medical data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913730A_ABST
    Figure CN120913730A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data processing system and method, and the method comprises the following steps: collecting medical data from an electronic medical record, a hospital information system, a medical image and a biosensor, and carrying out the format conversion and standardization processing, so as to guarantee the uniformity and comparability of the data; an intelligent algorithm is used for automatically detecting and removing repeated data, missing values and abnormal values are processed, and the problem of data inconsistency is solved; the availability of the data is improved through data transformation, normalization and feature selection; the beneficial effects of the invention are that through the systematic and automatic medical data processing method, the data processing efficiency and quality are significantly improved, and the manual intervention cost is reduced; the compliance use of patient information is ensured by enhancing data privacy protection and security measures; efficient and intelligent processing and analysis of medical data are realized through an integrated medical data processing system; not only are the data processing capability and efficiency improved, but also the manual intervention cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical data processing, and particularly relates to a medical data processing system and method. BACKGROUND

[0002] With the rapid development of medical technology and information technology, a large amount of medical data has been accumulated in the medical field; these data contain valuable information that can be used to improve medical quality, optimize medical processes, and promote the progress of medical research; medical data contain a large amount of patient health information, medical history records, examination results, etc., which are crucial for doctors' diagnosis and treatment decisions; through effective data processing, doctors can more accurately understand the patient's condition, and thus develop more accurate treatment plans; processing of medical data helps medical institutions better understand resource usage, such as bed occupancy rate, equipment utilization rate, etc.; based on these data, medical institutions can optimize resource allocation, improve resource utilization efficiency, and reduce operating costs; medical data are an important basis for medical research, and through data analysis, the occurrence regularity of diseases, treatment effects, etc. can be found. These information has important significance for promoting medical innovation, developing new drugs and treatment methods, however, processing and analyzing large-scale medical data is a complex and massive task.

[0003] The patent CN112799317A discloses a medical data processing system and method, which includes collecting first medical data; verifying the first medical data, storing the first medical data that passes the verification; sending the first medical data to the cloud platform; monitoring whether there is second medical data in the preset storage module; if so, sending the second medical data to the cloud platform.

[0004] Although the prior art has achieved processing of medical data to some extent, there is still room for improvement in terms of processing efficiency, data quality and automation degree. SUMMARY

[0005] The purpose of the present application is to provide a medical data processing system and method that significantly improves data processing efficiency and quality, and reduces the cost of manual intervention.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution: a medical data processing method, comprising the following steps:

[0007] Collecting medical data from electronic medical records, hospital information systems, medical images, and biological sensors, and performing format conversion and standardization processing to ensure the uniformity and comparability of the data;

[0008] The intelligent algorithm is used for automatically detecting and removing repeated data, processing missing values and abnormal values, and solving the inconsistency of data; through data transformation, normalization and feature selection, the availability of data is improved;

[0009] The distributed storage and processing technology is adopted to realize efficient storage, fast retrieval and parallel calculation of medical data; the data backup and recovery mechanism is established to ensure the safety and reliability of data;

[0010] Statistical analysis, machine learning, deep learning and data mining are used to analyze the processed medical data in depth.

[0011] As a preferred technical solution of the present application, medical data is collected from multiple sources and is subjected to format conversion and standardization processing, and the specific implementation method is as follows:

[0012] Data is collected from electronic medical records, hospital information systems, medical imaging systems and biological sensing multiple medical data sources, and these data exist in different formats and structures;

[0013] Data conversion tools or conversion scripts are used to convert data in different formats into a unified format;

[0014] Data cleaning: including verifying the accuracy of data, filling missing data, removing repeated data;

[0015] Data conversion: unifying the format, structure and unit of data;

[0016] Data unification: using unified data coding, naming and data model to ensure the comparability of data.

[0017] As a preferred technical solution of the present application, the implementation method of automatically detecting and removing repeated data, processing missing values and abnormal values by using intelligent algorithm is as follows:

[0018] Hash algorithm or similarity measurement is used to detect repeated data;

[0019] For the detected repeated data, one copy can be selected to be retained or combined;

[0020] According to the distribution and characteristics of missing values, a filling method is selected; for missing values that cannot be filled, the related records or features are deleted;

[0021] Statistical methods or machine learning algorithms are used to detect abnormal values; for the detected abnormal values, deletion, correction or replacement with reasonable values is selected.

[0022] As a preferred technical solution of the present application, the similarity measurement includes cosine similarity and Jaccard similarity; the filling method includes mean filling, median filling, mode filling and interpolation filling.

[0023] As a preferred technical solution of the present application, the statistical method is standard deviation method; and the machine learning algorithm is isolated forest.

[0024] As a preferred technical solution of the present application, the distributed storage technology: uses a distributed file system or a distributed database to store medical data; divides the data into multiple small blocks and stores them on different nodes, improving storage efficiency and scalability.

[0025] As a preferred technical solution of the present application, fast retrieval and parallel computing: using a distributed data processing framework to process data in parallel, reducing computing time; using index technology to speed up data retrieval speed.

[0026] As a preferred technical solution of the present application, the index technology includes inverted index and B-tree index.

[0027] The present application also discloses a medical data processing system, comprising

[0028] Data collection and integration module: responsible for collecting data from multiple medical data sources, and performing format conversion and standardization processing; this module uses intelligent recognition technology to automatically recognize and integrate medical data of different formats, ensuring the uniformity and comparability of data;

[0029] Data cleaning and preprocessing module: using algorithms and models, automatically detecting and removing duplicate data, handling missing values and outliers, solving data inconsistency problems; providing data transformation, normalization and feature selection functions to improve data quality and usability.

[0030] Data storage and management module: using distributed storage and processing technology, realizing efficient storage, fast retrieval and parallel computing of medical data; through intelligent indexing and classification technology, improving data retrieval efficiency;

[0031] Data analysis and mining module: integrating statistical analysis, machine learning, deep learning and data mining technology, for in-depth analysis of processed medical data.

[0032] Compared with the prior art, the present application has the following advantages:

[0033] The present application significantly improves the data processing efficiency and quality through systematic and automated medical data processing method, and reduces the cost of manual intervention;

[0034] Through strengthening the data privacy protection and security measures, the compliant use of patient information is ensured;

[0035] Through the integrated medical data processing system, efficient and intelligent processing and analysis of medical data are realized, which not only improves the data processing capacity and efficiency, but also reduces the cost of manual intervention. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 A medical data processing method flowchart of the present application;

[0037] Figure 2 A medical data processing system block diagram of the present application. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0039] Embodiment 1

[0040] Please refer to Figure 1 , the first embodiment of the present application, the embodiment provides a medical data processing method, comprising the following steps:

[0041] Collect medical data from electronic medical records, hospital information systems, medical images, and biological sensors, and perform format conversion and standardization processing to ensure data uniformity and comparability;

[0042] Use intelligent algorithms to automatically detect and remove duplicate data, process missing values and outliers, and solve data inconsistency problems; through data transformation, normalization and feature selection, improve data usability;

[0043] Adopt distributed storage and processing technology to realize efficient storage, fast retrieval and parallel computing of medical data; establish data backup and recovery mechanism to ensure data security and reliability;

[0044] Use statistical analysis, machine learning, deep learning and data mining to analyze the processed medical data in depth.

[0045] In this embodiment, preferably, medical data is collected from multiple sources, and format conversion and standardization processing is performed, the specific implementation method is as follows:

[0046] Collect data from electronic medical records, hospital information systems, medical image systems, and biological sensors, which exist in different formats and structures;

[0047] Convert data of different formats into a unified format using data conversion tools or writing conversion scripts;

[0048] Data cleaning: including verifying the accuracy of data, filling in missing data, removing duplicate data;

[0049] Data conversion: unify the format, structure and unit of data;

[0050] Data unification: use unified data coding, naming and data model to ensure data comparability.

[0051] Different formats and structures include CSV, XML, JSON, database tables, image files.

[0052] In this embodiment, the preferred method of automatically detecting and removing duplicate data, handling missing values and outliers using intelligent algorithms is as follows:

[0053] Use hash algorithm or similarity measure (such as cosine similarity, Jaccard similarity) to detect duplicate data;

[0054] For detected duplicate data, you can choose to keep one or merge them;

[0055] According to the distribution and characteristics of missing values, select filling method (such as mean filling, median filling, mode filling, interpolation method);

[0056] For missing values that cannot be filled, delete the relevant records or features;

[0057] Use statistical methods (such as standard deviation method) or machine learning algorithms (such as isolation forest) to detect outliers;

[0058] For detected outliers, choose to delete, correct or replace with reasonable values.

[0059] In this embodiment, the preferred method is to use distributed storage and processing technology to realize efficient storage, fast retrieval and parallel computing of medical data; Establish data backup and recovery mechanism, the method is as follows:

[0060] Distributed storage technology: use distributed file system (such as Hadoop HDFS) or distributed database (such as NoSQL database) to store medical data;

[0061] Divide the data into multiple small blocks and store them on different nodes to improve storage efficiency and scalability;

[0062] Fast retrieval and parallel computing: use distributed data processing framework (such as Apache Hadoop, Spark) to process data in parallel, significantly reducing computing time;

[0063] Use index technology (such as inverted index, B-tree index) to accelerate data retrieval speed;

[0064] Data backup and recovery mechanism: develop a reasonable backup strategy (such as full backup, incremental backup, off-site backup); verify the backup data regularly to ensure the effectiveness and recoverability of the backup; when data is lost or damaged, use backup data for recovery to ensure data integrity and availability.

[0065] In this embodiment, preferably, statistical analysis, machine learning, deep learning and data mining are used to analyze the processed medical data in depth, and the specific implementation methods include

[0066] Statistical analysis: use descriptive statistics (such as mean, median, standard deviation) to summarize the basic characteristics of the data; use inferential statistics (such as hypothesis testing, confidence interval) to infer the overall characteristics or compare the differences between different groups;

[0067] Machine learning: use supervised learning algorithms (such as decision tree, random forest, support vector machine) to train data with known labels and predict data with unknown labels; use unsupervised learning algorithms (such as clustering algorithm, association rule mining) to automatically classify and associate analyze data without labels;

[0068] Deep learning: use deep learning models (such as convolutional neural network CNN, recurrent neural network RNN, generative adversarial network GAN) to process and analyze complex data such as images and time series; apply deep learning technology in medical image analysis, disease prediction, drug response prediction, etc.

[0069] Data mining: use data mining techniques (such as association rule mining, clustering analysis, time series analysis) to discover interesting patterns and rules in data; apply data mining technology in disease and drug association analysis, patient classification and personalized treatment, etc.

[0070] Embodiment 2

[0071] Please refer to Figure 2 For the second embodiment of the present application, the embodiment provides a medical data processing system, comprising

[0072] Data collection and integration module: responsible for collecting data from multiple medical data sources, and performing format conversion and standardization processing; this module uses intelligent recognition technology, which can automatically recognize and integrate medical data of different formats, ensuring the uniformity and comparability of data;

[0073] Data cleaning and preprocessing module: using algorithms and models, automatically detect and remove duplicate data, handle missing values and outliers, solve data inconsistency problems; provide data transformation, normalization and feature selection functions to improve data quality and availability.

[0074] Data storage and management module: using distributed storage and processing technology, realize efficient storage, fast retrieval and parallel computing of medical data; through intelligent indexing and classification technology, improve the retrieval efficiency of data;

[0075] Data analysis and mining module: integrate statistical analysis, machine learning, deep learning and data mining technology to analyze the processed medical data in depth.

[0076] Although the embodiments of the present application have been shown and described in detail, as described above, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A medical data processing method, characterized by: The method comprises the following steps: Collecting medical data from multiple sources such as electronic medical records, hospital information systems, medical images, and biosensors, and performing format conversion and standardization processing to ensure data consistency and comparability; Using intelligent algorithms to automatically detect and remove duplicate data, handle missing values and outliers, and solve data inconsistency problems; through data transformation, normalization and feature selection, improve data usability; Using distributed storage and processing technology, realize efficient storage, fast retrieval and parallel computing of medical data; establish data backup and recovery mechanism, ensure data security and reliability; Using statistical analysis, machine learning, deep learning and data mining, analyze the processed medical data in depth.

2. The medical data processing method of claim 1, wherein: Collecting medical data from multiple sources and performing format conversion and standardization processing, the specific implementation method is as follows: Collect data from multiple medical data sources such as electronic medical records, hospital information systems, medical image systems, and biosensors, which exist in different formats and structures; Use data conversion tools or write conversion scripts to convert data in different formats to a unified format; Data cleaning: including verifying the accuracy of data, filling missing data, and removing duplicate data; Data conversion: unify the format, structure and unit of data; Data unification: use unified data coding, naming and data model to ensure data comparability.

3. The medical data processing method of claim 2, wherein: Different formats and structures include CSV, XML, JSON, database tables, and image files.

4. The medical data processing method of claim 1, wherein: The implementation method of using intelligent algorithms to automatically detect and remove duplicate data, handle missing values and outliers is as follows: Use hash algorithm or similarity measure to detect duplicate data; For duplicate data detected, you can choose to keep one or merge them; According to the distribution and characteristics of missing values, choose filling method; for missing values that cannot be filled, delete related records or features; Use statistical methods or machine learning algorithms to detect outliers; for detected outliers, choose to delete, correct or replace with reasonable values.

5. The medical data processing method of claim 4, wherein: The similarity measure includes cosine similarity and Jaccard similarity; the filling method includes mean filling, median filling, mode filling and interpolation filling.

6. The medical data processing method of claim 4, wherein: The statistical method is standard deviation method; the machine learning algorithm is Isolation Forest.

7. The medical data processing method of claim 1, wherein: Distributed storage technology: use distributed file system or distributed database to store medical data; divide data into multiple small blocks and store them on different nodes to improve storage efficiency and scalability.

8. The medical data processing method of claim 1, wherein: Fast retrieval and parallel computing: use distributed data processing framework to process data in parallel to reduce computing time; use index technology to speed up data retrieval speed.

9. The medical data processing method of claim 8, wherein: The index technology includes inverted index and B-tree index.

10. A medical data processing system, characterized by: Including Data collection and integration module: responsible for collecting data from multiple medical data sources and performing format conversion and standardization processing; this module uses intelligent recognition technology to automatically recognize and integrate medical data in different formats, ensuring data consistency and comparability; Data cleaning and preprocessing module: using algorithms and models, automatically detect and remove duplicate data, handle missing values and outliers, solve data inconsistency problems; provide data transformation, normalization and feature selection functions to improve data quality and usability. Data storage and management module: using distributed storage and processing technology, realize efficient storage, fast retrieval and parallel computing of medical data; through intelligent indexing and classification technology, improve the retrieval efficiency of data; Data analysis and mining module: integrate statistical analysis, machine learning, deep learning and data mining technology, and conduct in-depth analysis on the processed medical data.

Citation Information

Patent Citations

  • Medical data processing system and method

    CN112799317A