Medical image information co-processing system and method based on big data analysis
By building a collaborative processing system for medical image information analysis by big data and integrating medical health data, the problems of data dispersion and inconsistent format are solved, and unified storage and efficient diagnosis and treatment decisions of electronic health files are realized.
Patent Information
- Application Number
- CN202510665158.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to the differences in technical architecture and inconsistent data standards in the current medical information system, the fragmented storage of patient health data has increased the communication cost and risk of repeated examinations for cross-institutional medical treatment, affecting the efficiency and quality of diagnosis and treatment.
Build a medical image information collaborative processing system based on big data analysis, collect data through multiple data adapters, perform format conversion and abnormal analysis, and integrate index structures to form electronic health files to ensure unified storage and reliability of data.
It has achieved comprehensive and continuous integration of medical and health data, improved data usage efficiency and accuracy of diagnosis and treatment decisions, and solved the problems of data dispersion and inconsistent format.
Smart Images

Figure CN120473064A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical systems, and more specifically, to a system and method for collaborative processing of medical image information based on big data analysis. Background Art
[0002] In the current process of medical informatization, medical institutions at all levels have generally built independent business systems to support core functions such as registration, payment, and electronic medical records. However, due to differences in technical architecture and inconsistent data standards, these systems have formed information silos. This results in fragmented storage of patient health data across different institutions, making it difficult to form a continuous and complete health record. This fragmented state not only increases communication costs and the risk of duplicate examinations for patients seeking medical treatment across institutions, but also hinders doctors' comprehensive understanding of medical history, affecting the efficiency and quality of diagnosis and treatment.
[0003] To address data interoperability challenges, a unified data integration platform must be built. Key challenges include: 1) adapting to heterogeneous data sources, requiring compatibility with diverse system interfaces and unstructured medical data; 2) standardizing the processing of large volumes of data, such as medical images, requiring format conversion and efficient storage; 3) data quality control, requiring automated detection and labeling of abnormal information; and 4) linking cross-source data, requiring the establishment of logical links based on unique patient identifiers to integrate fragmented data into structured archival views.
[0004] There is currently no effective technical solution to the above problems. Summary of the Invention
[0005] This application aims to provide a medical image information collaborative processing system and method based on big data analysis to effectively integrate medical and health data, address data fragmentation and inconsistent formats, and automatically tag abnormal data, providing users with comprehensive and continuous health information, thereby improving data utilization efficiency and the accuracy of diagnosis and treatment decisions.
[0006] In a first aspect, the present application provides a medical image information collaborative processing system based on big data analysis for building electronic health records, the system comprising: An acquisition module is used to collect medical and health data from system interfaces or data export methods of different medical institutions through multiple data adapters. The medical and health data includes medical image files and structured data. A data conversion module, configured to convert the format of medical image files using the distributed file storage module to form a unified data format within the platform, generate an image identifier for each medical image file, perform content parsing and format conversion on structured data to form data metadata, and store the converted medical image files, image identifiers, and metadata, including a unique patient identifier; The anomaly analysis module is used to use the big data identification module to analyze whether there is any anomaly in the meta-information stored in the distributed file storage module, and if so, generate an anomaly mark and add it to the meta-information; The data display module is used to establish an index structure based on patient information after the user identity authentication is passed, and then use the metadata management module to match and associate data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure about the patient's electronic health record, so as to integrate it into a standardized file view for display. The index structure includes the patient's unique identifier and the image identifier.
[0007] The medical image information collaborative processing system based on big data analysis of this application aggregates scattered data into electronic health records through the steps of collection, conversion, detection, and integration. It can effectively integrate medical and health data, solve the problems of data dispersion and inconsistent formats, provide users with comprehensive and continuous health information, and improve the efficiency of data use and the accuracy of diagnosis and treatment decisions.
[0008] In a second aspect, the present application further provides a method for collaboratively processing medical image information based on big data analysis for constructing an electronic health record, the method comprising the following steps: S1. Collect medical and health data from system interfaces or data exports of different medical institutions through multiple data adapters, where the medical and health data includes medical image files and structured data; S2. Utilizing a distributed file storage module to convert the format of medical image files to form a unified data format within the platform, generating an image identifier for each medical image file, performing content parsing and format conversion on structured data to form data metadata, and storing the converted medical image files, image identifiers, and metadata, including a unique patient identifier; S3. Analyze the metadata stored in the distributed file storage module using the big data identification module to see if there is any anomaly. If so, generate an anomaly mark and add it to the metadata. S4. After the user identity authentication is passed, an index structure is established based on the patient information. Then, the metadata management module is used to match and associate the data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure of the patient's electronic health record, which is integrated into a standardized file view for display. The index structure includes the patient's unique identifier and the image identifier.
[0009] The collaborative processing method of medical image information based on big data analysis in this application aggregates scattered data into electronic health records through the steps of collection, conversion, detection, and integration. It can effectively integrate medical and health data, solve the problems of data dispersion and inconsistent formats, provide users with comprehensive and continuous health information, and improve the efficiency of data use and the accuracy of diagnosis and treatment decisions.
[0010] The method for collaborative processing of medical image information based on big data analysis, wherein the big data identification module is configured with a meta-information anomaly identification rule library; Step S3 includes: S31, using the metadata anomaly recognition rule base of the big data recognition module to scan the metadata stored in the distributed file storage module in parallel, and determine whether each metadata satisfies the anomaly recognition rule in the metadata anomaly recognition rule base; S32. If the meta information satisfies any of the exception identification rules, it is determined to be abnormal data, an abnormal tag including the abnormal type, abnormal field and abnormal value is generated, and the generated abnormal tag is added to the corresponding meta information.
[0011] In this example, data quality issues in the original metadata are clearly identified, marked, and recorded, solving the problem of insufficient abnormal marking information, providing a detailed basis for subsequent data cleaning, correction, or manual review, and improving the sophistication of data quality management.
[0012] The medical image information collaborative processing method based on big data analysis, wherein the metadata anomaly identification rule base includes data format verification rules, numerical range verification rules and data consistency verification rules.
[0013] The method for collaborative processing of medical image information based on big data analysis, wherein the data format verification rule is used to check whether the data type and length of the metadata conform to the preset format, the numerical range verification rule is used to determine whether the numerical metadata exceeds the preset reasonable range, and the data consistency verification rule is used to compare the metadata of the same patient at different time points to check whether there are obvious contradictions or conflicts.
[0014] The method for collaborative processing of medical image information based on big data analysis, wherein step S1 comprises: S11. Selecting a secure socket layer protocol or a virtual private network protocol based on the data transmission channel of a pre-configured data adapter according to the system interface type of different medical institutions, and establishing an encrypted data transmission channel; S12. Start the data collection task to collect medical and health data from different medical institutions’ system interfaces or data export methods; S13. Monitor the data transmission status in real time. If a data transmission interruption or error is detected, automatically re-initiate the data collection task according to the retry strategy.
[0015] The method for collaborative processing of medical image information based on big data analysis, wherein the distributed file storage module includes an image format conversion module, a desensitization module and a storage management module, and the storage management module is used to store the converted medical image files, image identifiers and meta-information; The steps of converting the format of the medical image files using the distributed file storage module to form a unified data format within the platform and generating an image identifier for each medical image file include: S21. Using the image format conversion module, for medical image files of different formats, a multi-threaded parallel processing mechanism is used to schedule image conversion tasks to perform image format conversion; S22. During the image format conversion process, based on the desensitization rules predefined in the desensitization module, identify and process the patient information in the medical image file, and record the desensitization operation log; S23. Use the UUID algorithm to generate a globally unique image identifier for each medical image file, and establish a mapping relationship between the image identifier and the original medical image file and storage path.
[0016] In the method for collaborative processing of medical image information based on big data analysis, the step of performing content parsing and format conversion on structured data to form data metadata includes: S24. Based on a preset data cleaning rule library, a parallel processing framework is used to clean the structured data. S25. Use natural language processing technology to extract medical terms from the cleaned structured data, standardize the medical terms based on the medical knowledge graph, and construct a knowledge representation that includes medical terms, attributes, and relationships; S26. According to a predefined data meta-model, the cleaned structured data and knowledge representation are converted into data meta-information.
[0017] The method for collaborative processing of medical image information based on big data analysis, wherein the data cleaning rule base includes missing value filling strategies, outlier detection rules and data format standardization schemes.
[0018] The method for collaborative processing of medical image information based on big data analysis, wherein step S4 includes: S41. When a user initiates an electronic health record access request, the user's identity is verified through the identity authentication module, the user role is queried based on the user identity, and the data access rights corresponding to the role are obtained from the user role permission rule library; S42. Build a dynamic data desensitization strategy based on the data access rights corresponding to the user role; S43: Establish a multi-level index structure based on the patient information corresponding to the access request, wherein the index structure includes a patient unique identifier, a medical institution identifier, an examination item identifier, and an image identifier; S44, using the metadata management module to retrieve patient-related data metadata and medical image files from the distributed file storage module according to the multi-level index structure, and desensitizing the retrieved data metadata according to the dynamic data desensitization strategy; S45. Integrate the desensitized data metadata and medical image files to form a metadata structure of the patient's electronic health record, and display the metadata structure under a standardized record view.
[0019] As can be seen from the above, the present application provides a medical image information collaborative processing system and method based on big data analysis, wherein the medical image information collaborative processing system based on big data analysis of the present application collects data from different medical institutions through multiple data adapters, ensuring the breadth and compatibility of data sources. Then, the collected data is standardized, including the format conversion of medical image files and the parsing and conversion of structured data, so that heterogeneous data can be stored and managed in a unified manner. Then, the metadata is analyzed for abnormalities through the big data identification module, which improves the quality and reliability of the data. Finally, based on user identity authentication, the index structure is used to match and associate data from different institutions to form a complete electronic health record view. The system aggregates scattered data into electronic health records through the steps of collection, conversion, detection, and integration. It can effectively integrate medical and health data, solve the problems of data dispersion and inconsistent formats, provide users with comprehensive and continuous health information, and improve the efficiency of data use and the accuracy of diagnosis and treatment decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic structural diagram of a medical image information collaborative processing system based on big data analysis provided in some embodiments of the present application.
[0021] Figure 2 A structural diagram of a medical image information collaborative processing system based on big data analysis provided in some other embodiments of the present application.
[0022] Figure 3 This is a structural diagram of the distributed file storage module.
[0023] Figure 4 A flowchart of a method for collaborative processing of medical image information based on big data analysis provided in an embodiment of the present application.
[0024] Reference numerals: 101, acquisition module; 102, data conversion module; 103, abnormality analysis module; 104, data display module. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0027] First, please refer to Figure 1 Some embodiments of the present application provide a medical image information collaborative processing system based on big data analysis for building electronic health records, the system comprising: An acquisition module 101 is configured to collect medical and health data from system interfaces or data exports of different medical institutions through multiple data adapters, wherein the medical and health data includes medical image files and structured data; The data conversion module 102 is used to convert the format of medical image files using the distributed file storage module to form a unified data format within the platform, generate an image identifier for each medical image file, perform content analysis and format conversion on structured data to form data metadata, and store the converted medical image files, image identifiers, and metadata, including the patient's unique identifier; The abnormality analysis module 103 is used to analyze whether the meta information stored in the distributed file storage module contains abnormalities using the big data identification module, and if so, to generate an abnormality mark and add it to the meta information; The data display module 104 is used to establish an index structure based on the patient information after the user identity authentication is passed, and then use the metadata management module to match and associate the data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure about the patient's electronic health record, so as to integrate it into the standardized file view for display. The index structure includes the patient's unique identifier and image identifier.
[0028] Specifically, the acquisition module 101 is used for data acquisition. Multiple data adapters are configured to connect to system interfaces at different medical institutions, such as HIS, PACS, and LIS systems, or to process data export files, such as burning CDs or transferring files via FTP. The data adapters are responsible for establishing connections with source systems and reading or receiving medical and health data according to pre-set rules. The acquired data streams include medical image files, such as DICOM format files, as well as structured data such as basic patient information, examination application forms, and examination report text. This enables the acquisition of raw data from dispersed medical institutions, overcoming the problem of heterogeneous data sources.
[0029] More specifically, the data conversion module 102 is used to perform data standardization and preliminary processing. The distributed file storage module receives the collected data. For medical image files, the image format conversion component performs a format conversion operation to convert them into a unified data format within the platform. During the conversion process, the image identifier generation component generates a globally unique image identifier for each converted image file, for example, using a UUID algorithm. For structured data, the content parsing component parses the data and extracts key information, and the format conversion component converts the parsed data into the data metadata format specified within the platform. The metadata contains a unique patient identifier for subsequent data association. The converted image files, image identifiers, and metadata are stored in the distributed file storage system. In this way, standardized processing and preliminary organization of the data are achieved, laying the foundation for subsequent steps.
[0030] More specifically, the anomaly analysis module 103 is used to perform data quality checks. The big data identification module accesses the metadata stored in the distributed file storage module. This module performs data analysis tasks, checking the metadata for anomalies. If an anomaly is detected, the big data identification module generates an anomaly flag. This flag is added to the corresponding metadata record. This improves data reliability, provides quality assurance for subsequent data use, and enables users to intuitively identify data anomalies during subsequent data presentations.
[0031] More specifically, the data display module 104 is used to integrate and display data. After the user passes the identity authentication, the system determines the user's access rights based on the user's identity (such as patient, medical staff or regulatory department). The system establishes an index structure based on the patient information contained in the user request, such as identity information. The index structure contains the patient's unique identifier and the image identifier related to the patient. The metadata management module uses the index structure to search and match the meta-information and medical image files belonging to the patient in the distributed file storage module. The matched and associated data is organized into a metadata structure about the patient's electronic health record. Finally, the metadata structure is integrated and presented to the user in a unified, standardized archive view. In this way, data scattered across different institutions are linked together to form a complete patient health record view, which is convenient for users to access and manage.
[0032] The medical image information collaborative processing system based on big data analysis in the embodiment of the present application collects data from different medical institutions through multiple data adapters, ensuring the breadth and compatibility of data sources. Then, the collected data is standardized, including the format conversion of medical image files and the parsing and conversion of structured data, so that heterogeneous data can be uniformly stored and managed. Then, the metadata is analyzed for abnormalities through the big data identification module, which improves the quality and reliability of the data. Finally, based on user identity authentication, the index structure is used to match and associate data from different institutions to form a complete electronic health record view. The system aggregates scattered data into electronic health records through the steps of collection, conversion, detection, and integration. It can effectively integrate medical and health data, solve the problems of data dispersion and inconsistent formats, provide users with comprehensive and continuous health information, and improve the efficiency of data use and the accuracy of diagnosis and treatment decisions.
[0033] Second, please refer to Figure 2-Figure 4 Some embodiments of the present application also provide a method for collaboratively processing medical image information based on big data analysis, for building an electronic health record, the method comprising the following steps: S1. Collect medical and health data from system interfaces or data exports of different medical institutions through multiple data adapters, where the medical and health data includes medical image files and structured data; S2. Utilize the distributed file storage module to convert the format of medical image files to form a unified data format within the platform, generate an image identifier for each medical image file, perform content parsing and format conversion on structured data to form data metadata, and store the converted medical image files, image identifiers, and metadata, including the patient's unique identifier. S3. Analyze the metadata stored in the distributed file storage module using the big data identification module to see if there is any anomaly. If so, generate an anomaly mark and add it to the metadata. S4. After the user identity authentication is passed, an index structure is established based on the patient information. Then, the metadata management module is used to match and associate the data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure of the patient's electronic health record, which is integrated into the standardized file view for display. The index structure includes the patient's unique identifier and image identifier.
[0034] The collaborative processing method of medical image information based on big data analysis in the embodiment of the present application collects data from different medical institutions through multiple data adapters, ensuring the breadth and compatibility of data sources. Then, the collected data is standardized, including the format conversion of medical image files and the parsing and conversion of structured data, so that heterogeneous data can be uniformly stored and managed. Then, the metadata is analyzed for abnormalities through the big data identification module, which improves the quality and reliability of the data. Finally, based on user identity authentication, the index structure is used to match and associate data from different institutions to form a complete electronic health record view. This method aggregates scattered data into electronic health records through the steps of collection, conversion, detection, and integration. It can effectively integrate medical and health data, solve the problems of data dispersion and inconsistent formats, provide users with comprehensive and continuous health information, and improve the efficiency of data use and the accuracy of diagnosis and treatment decisions.
[0035] In some preferred embodiments, the big data identification module is configured with a meta-information anomaly identification rule base; Step S3 includes: S31, using the metadata anomaly recognition rule base of the big data recognition module to scan the metadata stored in the distributed file storage module in parallel, and determine whether each metadata satisfies the anomaly recognition rule in the metadata anomaly recognition rule base; S32. If the meta information satisfies any of the exception identification rules, it is determined to be abnormal data, an abnormal tag including the abnormal type, abnormal field and abnormal value is generated, and the generated abnormal tag is added to the corresponding meta information.
[0036] Specifically, the meta-information anomaly identification rule base provides a standard for meta-information anomaly judgment. The rule base can be implemented as a structured storage, such as a database table or a configuration file set, in which multiple anomaly identification rules are defined.
[0037] More specifically, step S31 describes the execution process of identifying anomalies in meta-information. The big data identification module is configured to scan the meta-information stored in the distributed file storage module according to the rules in the rule base. To improve processing efficiency, this scanning process is scheduled and executed using a parallel processing mechanism.
[0038] More specifically, step S32 describes the judgment logic of abnormal data and the generation and attachment of abnormal tags. When a certain meta-information is found to meet any of the abnormal identification rules defined in the rule base during the scanning process, the meta-information is determined to be abnormal data. Subsequently, the method of the present application will generate an abnormal tag, which is designed to contain specific information about the abnormality, such as the type of abnormality (such as format error, numerical value out of bounds, data inconsistency), the name of the field where the abnormality occurred, and the specific abnormal value of the field. The generated abnormal tag is then attached to the corresponding meta-information record, thereby directly associating the abnormal information with the problem data. As a result, the data quality problems existing in the original meta-information are clearly identified, marked and recorded, solving the problem of insufficient abnormal tag information, providing a detailed basis for subsequent data cleaning, correction or manual review, and improving the precision of data quality management.
[0039] In some preferred implementations, the meta-information anomaly identification rule base includes data format verification rules, value range verification rules, and data consistency verification rules.
[0040] In some preferred embodiments, the data format verification rules are used to check whether the data type and length of the metadata conform to the preset format, the numerical range verification rules are used to determine whether the numerical metadata exceeds the preset reasonable range, and the data consistency verification rules are used to compare the metadata of the same patient at different time points to check whether there are obvious contradictions or conflicts.
[0041] Specifically, by including data format verification rules, the method of the present application can identify data type errors or length inconsistencies, and ensure the structural correctness of the data. By including numerical range verification rules, the method of the present application can filter out extreme values that do not conform to the actual situation and identify anomalies in numerical data. By including data consistency verification rules, the method of the present application can discover logical contradictions between different records of the same patient and identify correlation errors between data. As a result, the big data identification module can detect data quality problems in different dimensions, thereby improving the accuracy and comprehensiveness of metadata anomaly identification. The data format verification rules ensure the standardization of the data, the numerical range verification rules filter out extreme or impossible values, and the data consistency verification rules discover logical contradictions between data. These rules work together to enable the metadata anomaly identification process to more effectively identify and mark anomalies in metadata, laying the foundation for building high-quality electronic health records.
[0042] More specifically, the implementation of data format validation rules can include defining the expected data type (e.g., string, integer, floating point number, date) and the maximum / minimum length allowed for each metadata field. The implementation of numerical range validation rules can include setting reasonable upper and lower limits for specific numerical metadata fields. The implementation of data consistency validation rules can include establishing associations between records of the same patient at different time points and defining comparison logic, such as checking the rationality of changes in weight, height, age, etc. over time, or whether diagnostic information is inconsistent.
[0043] In some preferred embodiments, step S1 includes: S11. Selecting a secure socket layer protocol or a virtual private network protocol based on the data transmission channel of a pre-configured data adapter according to the system interface type of different medical institutions, and establishing an encrypted data transmission channel; S12. Start the data collection task to collect medical and health data from different medical institutions’ system interfaces or data export methods; S13. Monitor the data transmission status in real time. If a data transmission interruption or error is detected, automatically re-initiate the data collection task according to the retry strategy.
[0044] Specifically, in step S11, the method of this application dynamically selects an appropriate encrypted transmission protocol based on the interface characteristics of the target connection (medical institution system) (for example, whether it is a web-based API or requires a point-to-point connection). The Secure Sockets Layer (SSL) protocol is suitable for scenarios where data is exchanged over public networks, providing end-to-end encryption. The Virtual Private Network (VPN) protocol is suitable for scenarios where a secure tunnel is required over an insecure network to connect a remote network to a local network. By selecting and establishing these encrypted channels, the confidentiality and integrity of data during transmission from the medical institution to the platform are ensured.
[0045] More specifically, the data transmission status is monitored in real time. If a data transmission interruption or error is detected, the data collection task is automatically re-initiated according to the retry strategy. The monitoring of the data transmission status can be achieved in a variety of ways, such as checking the activity of the network connection, verifying the integrity of the received data (such as checksum), monitoring the transmission rate or timeout. Once an anomaly is detected, the method of the present application triggers the preset retry logic. The retry strategy may include waiting for a certain time interval before retrying, gradually increasing the waiting time (exponential backoff), or setting a maximum number of retries. Automatically re-initiating the collection task improves the robustness of data collection and reduces data loss due to transient network problems or system failures.
[0046] More specifically, by establishing an encrypted channel, data is protected during transmission, preventing unauthorized access and tampering, thus resolving data transmission security issues. A monitoring and retry mechanism addresses the issue of data collection task failures due to transmission interruptions or errors, improving data collection success rates and data integrity.
[0047] In some preferred embodiments, the distributed file storage module includes an image format conversion module, a desensitization module, and a storage management module, wherein the storage management module is used to store the converted medical image files, image identifiers, and meta-information; The steps of using the distributed file storage module to convert the format of medical image files to form a unified data format within the platform and generate an image identifier for each medical image file include: S21. Using the image format conversion module, for medical image files of different formats, a multi-threaded parallel processing mechanism is used to schedule image conversion tasks to perform image format conversion; S22. During the image format conversion process, based on the desensitization rules predefined in the desensitization module, identify and process the patient information in the medical image file, and record the desensitization operation log; S23. Use the UUID algorithm to generate a globally unique image identifier for each medical image file, and establish a mapping relationship between the image identifier and the original medical image file and storage path.
[0048] Specifically, the image format conversion module is configured to receive medical image files of different formats and schedule image conversion tasks using a multi-threaded parallel processing mechanism, thereby achieving image processing. The desensitization module is configured to store predefined desensitization rules and is called during the image format conversion process to identify and process patient information in medical image files, while also recording desensitization operation logs to protect patient privacy. The storage management module is configured to store the converted medical image files, the generated image identifiers, and related metadata. The image identifiers are generated using a UUID algorithm and a mapping relationship is established with the original medical image files and storage paths.
[0049] More specifically, this solution uses a distributed file storage module to process collected medical image files. When medical image files of different formats are received, the image format conversion module is activated and uses a multi-threaded parallel processing mechanism to simultaneously handle multiple image conversion tasks, converting these images into a unified data format within the platform, thereby achieving processing. During the image format conversion process, the desensitization module identifies and processes patient information in the image according to predefined desensitization rules, such as blurring or removing sensitive areas, while also recording a detailed log of the desensitization operation, thereby ensuring patient privacy. Each converted medical image file will generate a globally unique image identifier using the UUID algorithm, and a mapping relationship will be established between this identifier and the original file path and storage path, thereby achieving uniqueness of the massive image identifiers and facilitating subsequent management and retrieval. Finally, the converted image files, the generated image identifiers, and related metadata are stored in the storage management module, forming a standardized and traceable medical image data asset for subsequent access and management.
[0050] In some preferred embodiments, the steps of performing content parsing and format conversion on structured data to form data metadata include: S24. Based on a preset data cleaning rule library, a parallel processing framework is used to clean the structured data. S25. Use natural language processing technology to extract medical terms from the cleaned structured data, standardize the medical terms based on the medical knowledge graph, and construct a knowledge representation that includes medical terms, attributes, and relationships; S26. According to a predefined data meta-model, the cleaned structured data and knowledge representation are converted into data meta-information.
[0051] Specifically, simple data parsing and format conversion may not be able to effectively deal with quality issues in the data, such as missing, abnormal or inconsistent formats. It is also difficult to extract and standardize medical concepts from unstructured or semi-structured texts, resulting in inaccurate or incomplete generated data metadata, affecting subsequent archive construction and utilization. The method of the present application solves these problems through steps S24, S25 and S26.
[0052] More specifically, step S24 uses a data cleaning rule library and a parallel processing framework to clean the structured data, and the parallel processing framework improves the cleaning efficiency.
[0053] More specifically, in step S25, natural language processing technology is used to extract medical terminology from structured data (particularly text content). A medical knowledge graph is used to standardize the extracted medical terminology, eliminating synonyms and variant forms, and construct a knowledge representation encompassing terminology, attributes, and relationships. This step transforms unstructured or semi-structured medical text into a structured knowledge representation, enhancing the semantic value and usability of the data.
[0054] More specifically, step S26 converts the cleaned structured data and the knowledge representation generated in step S25 according to a predefined data metamodel. The data metamodel defines the structure, format, and meaning of data elements. By converting according to the model, structured data and extracted knowledge from different sources and formats are unified into standardized data metainformation. This step ensures that the resulting data metainformation conforms to unified specifications, facilitating subsequent storage, management, and application, and supporting the construction of standardized electronic health records.
[0055] More specifically, the method of the present application systematically addresses the data quality and standardization challenges faced during the processing of structured medical data, generates high-quality, standardized data metadata, and lays the foundation for building accurate and complete electronic health records.
[0056] In some preferred embodiments, the data cleaning rule base includes missing value filling strategies, outlier detection rules, and data format standardization schemes.
[0057] Specifically, in response to the problems of missing, abnormal and inconsistent formats in structured data, step S24 is filled, identified and converted according to the rules, ensuring the quality and consistency of the structured data and providing a reliable foundation for subsequent processing. Among them, the establishment of a data cleaning rule base provides specific guidance for the cleaning process of structured data. The missing value filling strategy is used to handle blanks or invalid items in the data to ensure the integrity of the data. The outlier detection rule is used to identify values in the data that do not conform to the preset pattern or range to avoid erroneous data affecting subsequent processing. The data format standardization scheme is used to unify the data formats of different sources or representations and improve the consistency of the data. These strategies and rules work together to standardize the data cleaning process. As a result, the quality of structured data is improved, and the generated data metadata is more accurate and standardized, laying the foundation for subsequent data matching and association, and improving the availability of data metadata.
[0058] In some preferred embodiments, step S4 includes: S41. When a user initiates an electronic health record access request, the user's identity is verified through the identity authentication module, the user role is queried based on the user identity, and the data access rights corresponding to the role are obtained from the user role permission rule library; S42. Build a dynamic data desensitization strategy based on the data access rights corresponding to the user role; S43. Establish a multi-level index structure based on the patient information corresponding to the access request, where the index structure includes a patient unique identifier, a medical institution identifier, an examination item identifier, and an image identifier; S44, using the metadata management module to retrieve patient-related data metadata and medical image files from the distributed file storage module according to the multi-level index structure, and desensitizing the retrieved data metadata according to the dynamic data desensitization strategy; S45. Integrate the desensitized data metadata and medical image files to form a metadata structure of the patient's electronic health record, and display the metadata structure under a standardized archive view.
[0059] Specifically, the above-described approach addresses the issue of secure access control and data privacy protection during the display of integrated electronic health records. First, through identity authentication and permission acquisition, the user's identity, role, and the scope of data they are permitted to access are determined, laying the foundation for secure access. Then, a data desensitization policy is dynamically constructed based on user roles, setting different sensitive information processing rules for different user types to ensure data privacy is protected during the display phase. Next, a multi-level index structure is established, encompassing the patient's unique identifier, medical institution identifier, examination item identifier, and image identifier. This provides a more granular indexing approach than simply including the patient's unique identifier and image identifier, facilitating more accurate and efficient location and retrieval of medical data related to specific patients, institutions, and examination items. Subsequently, the multi-level index structure is used to retrieve relevant data from the storage module, and the constructed dynamic desensitization policy is immediately applied to the retrieved data metadata, ensuring that only data that meets the user's permission and has been desensitized is used for subsequent display. Finally, the desensitized data metadata is integrated with the medical image files to form the final electronic health record metadata structure for display, which is presented to the user in a unified view. Through the above steps, this technical solution realizes refined permission control and dynamic data desensitization for access to electronic health records, ensuring the security and privacy of patient data, while supporting users of different roles to view the required information according to their responsibilities and permissions.
[0060] More specifically, for the dynamic data desensitization strategy, for the patient role, sensitive personal information is desensitized; for the medical staff role, non-diagnosis and treatment-related patient information is desensitized; for the regulatory department role, patient information that exceeds the scope of compliance review is desensitized.
[0061] In some preferred embodiments, the structured data includes one or more of patient information, examination application form, and examination report text.
[0062] Specifically, patient information provides basic data for identifying patients and is a core element in building electronic health records. The examination application form provides background information for medical image examinations, including examination items, applying doctors, etc., which helps to understand the source and purpose of the images. The examination report text provides clinical interpretation and diagnostic conclusions of medical images and is an important manifestation of the value of medical images. By clearly collecting these types of structured data, the integrity of the collected data is ensured, providing a data foundation for the subsequent generation of comprehensive data metadata, effective abnormality identification, and the construction of electronic health record metadata structures containing necessary contextual information. Each type of structured data contributes an indispensable information dimension to electronic health records.
[0063] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0064] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0065] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.
[0066] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A medical image information collaborative processing system based on big data analysis, used to build electronic health records, characterized by: The system comprises: An acquisition module is used to collect medical and health data from system interfaces or data export methods of different medical institutions through multiple data adapters. The medical and health data includes medical image files and structured data. A data conversion module, configured to convert the format of medical image files using the distributed file storage module to form a unified data format within the platform, generate an image identifier for each medical image file, perform content parsing and format conversion on structured data to form data metadata, and store the converted medical image files, image identifiers, and metadata, including a unique patient identifier; The anomaly analysis module is used to use the big data identification module to analyze whether there is any anomaly in the meta-information stored in the distributed file storage module, and if so, generate an anomaly mark and add it to the meta-information; The data display module is used to establish an index structure based on patient information after the user identity authentication is passed, and then use the metadata management module to match and associate data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure about the patient's electronic health record, so as to integrate it into a standardized file view for display. The index structure includes the patient's unique identifier and the image identifier.
2. A collaborative processing method for medical image information based on big data analysis, used to build electronic health records, characterized in that: The method comprises the following steps: S1. Collect medical and health data from system interfaces or data exports of different medical institutions through multiple data adapters, where the medical and health data includes medical image files and structured data; S2. Utilizing a distributed file storage module to convert the format of medical image files to form a unified data format within the platform, generating an image identifier for each medical image file, performing content parsing and format conversion on structured data to form data metadata, and storing the converted medical image files, image identifiers, and metadata, including a unique patient identifier; S3. Analyze the metadata stored in the distributed file storage module using the big data identification module to see if there is any anomaly. If so, generate an anomaly mark and add it to the metadata. S4. After the user identity authentication is passed, an index structure is established based on the patient information. Then, the metadata management module is used to match and associate the data metadata and medical image files from different medical institutions according to the index structure to form a metadata structure of the patient's electronic health record, which is integrated into a standardized file view for display. The index structure includes the patient's unique identifier and the image identifier.
3. The method for collaborative processing of medical image information based on big data analysis according to claim 2, characterized in that: The big data identification module is configured with a meta-information anomaly identification rule base; Step S3 includes: S31, using the metadata anomaly recognition rule base of the big data recognition module to scan the metadata stored in the distributed file storage module in parallel, and determine whether each metadata satisfies the anomaly recognition rule in the metadata anomaly recognition rule base; S32. If the meta information satisfies any of the exception identification rules, it is determined to be abnormal data, an abnormal tag including the abnormal type, abnormal field and abnormal value is generated, and the generated abnormal tag is added to the corresponding meta information.
4. The method for collaborative processing of medical image information based on big data analysis according to claim 3, characterized in that: The meta-information anomaly identification rule base includes data format verification rules, value range verification rules and data consistency verification rules.
5. The method for collaborative processing of medical image information based on big data analysis according to claim 4, characterized in that: The data format verification rule is used to check whether the data type and length of the metadata conform to the preset format, the numerical range verification rule is used to determine whether the numerical metadata exceeds the preset reasonable range, and the data consistency verification rule is used to compare the metadata of the same patient at different time points to check whether there are obvious contradictions or conflicts.
6. The method for collaborative processing of medical image information based on big data analysis according to claim 2, characterized in that: Step S1 includes: S11. Selecting a secure socket layer protocol or a virtual private network protocol based on the data transmission channel of a pre-configured data adapter according to the system interface type of different medical institutions, and establishing an encrypted data transmission channel; S12. Start the data collection task to collect medical and health data from different medical institutions’ system interfaces or data export methods; S13. Monitor the data transmission status in real time. If a data transmission interruption or error is detected, automatically re-initiate the data collection task according to the retry strategy.
7. The method for collaborative processing of medical image information based on big data analysis according to claim 2, characterized in that: The distributed file storage module includes an image format conversion module, a desensitization module and a storage management module, and the storage management module is used to store the converted medical image files, image identifiers and meta information; The steps of converting the format of the medical image files using the distributed file storage module to form a unified data format within the platform and generating an image identifier for each medical image file include: S21. Using the image format conversion module, for medical image files of different formats, a multi-threaded parallel processing mechanism is used to schedule image conversion tasks to perform image format conversion; S22. During the image format conversion process, based on the desensitization rules predefined in the desensitization module, identify and process the patient information in the medical image file, and record the desensitization operation log; S23. Use the UUID algorithm to generate a globally unique image identifier for each medical image file, and establish a mapping relationship between the image identifier and the original medical image file and storage path.
8. The method for collaborative processing of medical image information based on big data analysis according to claim 2, characterized in that: The steps of performing content parsing and format conversion on structured data to form data metadata include: S24. Based on a preset data cleaning rule library, a parallel processing framework is used to clean the structured data. S25. Use natural language processing technology to extract medical terms from the cleaned structured data, standardize the medical terms based on the medical knowledge graph, and construct a knowledge representation that includes medical terms, attributes, and relationships; S26. According to a predefined data meta-model, the cleaned structured data and knowledge representation are converted into data meta-information.
9. The method for collaborative processing of medical image information based on big data analysis according to claim 8, characterized in that: The data cleaning rule base includes missing value filling strategies, outlier detection rules and data format standardization solutions.
10. The method for collaborative processing of medical image information based on big data analysis according to claim 2, characterized in that: Step S4 includes: S41. When a user initiates an electronic health record access request, the user's identity is verified through the identity authentication module, the user role is queried based on the user identity, and the data access rights corresponding to the role are obtained from the user role permission rule library; S42. Build a dynamic data desensitization strategy based on the data access rights corresponding to the user role; S43: Establish a multi-level index structure based on the patient information corresponding to the access request, wherein the index structure includes a patient unique identifier, a medical institution identifier, an examination item identifier, and an image identifier; S44, using the metadata management module to retrieve patient-related data metadata and medical image files from the distributed file storage module according to the multi-level index structure, and desensitizing the retrieved data metadata according to the dynamic data desensitization strategy; S45. Integrate the desensitized data metadata and medical image files to form a metadata structure of the patient's electronic health record, and display the metadata structure under a standardized record view.