A standardized synchronization method, apparatus, equipment, and medium for multi-source heterogeneous DICOM data.
By combining hierarchical cleaning rules and a three-dimensional collaborative mapping model, the problems of unclear data cleaning and mapping dependence on a single field in the standardization and synchronization of multi-source heterogeneous DICOM data are solved, and accurate and reliable standardized data processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MEDICAL IMAGE INSIGHTS INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-26
AI Technical Summary
In the process of standardizing and synchronizing multi-source heterogeneous DICOM data, existing technologies suffer from unclear data cleaning and data mapping that relies on a single field correspondence, making it difficult to achieve high-quality data standardization and synchronization.
Data cleaning is performed by setting hierarchical cleaning rules, and data mapping is performed using a three-dimensional collaborative mapping model, which includes mapping rules for three dimensions: data structure, value range, and standard data logic model, to ensure the accuracy of data cleaning and the precision of mapping.
It improves the accuracy and reliability of standardized processing of multi-source heterogeneous DICOM data, avoids mapping errors in single-field correspondence, and ensures the consistency of data structure and semantics.
Smart Images

Figure CN122086880A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a standardized synchronization method, apparatus, device, and medium for multi-source heterogeneous DICOM data. Background Technology
[0002] When medical images are acquired in medical institutions, DICOM data is generated. The data formats of DICOM data generated by different equipment models or types of imaging equipment vary greatly. Therefore, standardization processing of this multi-source, heterogeneous DICOM data is necessary during data management, querying, and sharing.
[0003] Currently, the standardization and synchronization of DICOM data is achieved through data cleaning and data mapping. However, the existing process lacks clear data cleaning methods, and the data mapping relies solely on the correspondence of a single field, making it difficult to support high-quality data standardization and synchronization services. Summary of the Invention
[0004] This disclosure provides a standardized synchronization method, apparatus, device, and medium for multi-source heterogeneous DICOM data. By setting explicit data cleaning rules, the reliability and accuracy of data cleaning are improved. Data mapping is performed through a three-dimensional collaborative mapping model, thereby improving processing accuracy in cross-modal and cross-mechanism scenarios.
[0005] According to one aspect of this disclosure, a standardized synchronization method for multi-source heterogeneous DICOM data is provided, comprising: Acquire DICOM data to be synchronized from at least one data source; Based on pre-built hierarchical cleaning rules, the data to be cleaned in the DICOM data to be synchronized is determined and data cleaning is performed to obtain the cleaned DICOM data; The cleaned DICOM data is mapped using a three-dimensional collaborative mapping model based on a pre-built standard data dictionary to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
[0006] According to another aspect of this disclosure, a standardized synchronization device for multi-source heterogeneous DICOM data is provided, comprising: The data acquisition module is used to acquire DICOM data to be synchronized from at least one data source. The data cleaning module is used to determine the data to be cleaned in the DICOM data to be synchronized based on pre-built hierarchical cleaning rules and perform data cleaning to obtain cleaned DICOM data. The data mapping module is used to map the cleaned DICOM data based on a pre-built standard data dictionary using a three-dimensional collaborative mapping model to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
[0007] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the standardized synchronization method for multi-source heterogeneous DICOM data as described in any embodiment of this disclosure.
[0008] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the standardized synchronization method for multi-source heterogeneous DICOM data as described in any embodiment of this disclosure.
[0009] According to another aspect of this disclosure, a computer program product is provided that, when executed by a processor, implements a standardized synchronization method for multi-source heterogeneous DICOM data as described in any of the embodiments of this disclosure.
[0010] The technical solution provided in this disclosure formulates hierarchical and explicit cleaning rules through pre-built hierarchical cleaning rules, enabling accurate and reliable data cleaning of multi-source heterogeneous DICOM data, and providing a reliable data foundation for data standardization mapping. Through a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped from three dimensions: data structure, value range, and standard data logic model, avoiding mapping based on single-field correspondences and improving the accuracy of standardization processing.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure; Figure 2 This is a flowchart of a standardized synchronization method for multi-source heterogeneous DICOM data in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of a standardized synchronization device for multi-source heterogeneous DICOM data in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] Figure 1This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure. The standardized synchronization device in this embodiment can be understood as an electronic device capable of performing standardized synchronization processing on multi-source heterogeneous DICOM data. This standardized synchronization device can be an electronic device such as a server or computer device. Figure 1 The number of institutions and the imaging equipment included in each institution are merely examples.
[0018] This standardized synchronization device can connect to imaging equipment in different medical institutions. Each imaging device in a medical institution can serve as a data source, capable of acquiring medical images and generating DICOM data. The same medical institution may include different types of imaging equipment, such as, but not limited to, MRI and CT scanners. The same type of imaging equipment may correspond to different manufacturers. Different types of imaging equipment can serve as different data sources, and the same type of imaging equipment from different manufacturers can serve as different data sources. DICOM data generated from different data sources will differ in at least one of the following: label naming, field format, and value range, resulting in differences in the data structure, semantics, and value range of the DICOM data.
[0019] To achieve standardized management, querying, and sharing of DICOM data, a standardized synchronization device is used as a data platform to enable data transmission with multiple heterogeneous data sources and to perform standardized synchronization processing on multi-source heterogeneous DICOM data to obtain standardized synchronized data.
[0020] To address the technical problems existing in the standardized synchronization process, this disclosure provides a standardized synchronization method for multi-source heterogeneous DICOM data, see [link to relevant documentation]. Figure 2 , Figure 2 This flowchart illustrates a standardized synchronization method for multi-source heterogeneous DICOM data provided in this embodiment. This embodiment is applicable to scenarios involving the aggregation and synchronization of medical imaging data (i.e., DICOM data) across devices, manufacturers, institutions, and systems, where multi-source heterogeneous DICOM data undergoes standardized synchronization processing. This method can be executed by a standardized synchronization device for multi-source heterogeneous DICOM data in this embodiment, which can be implemented using software and / or hardware, such as... Figure 2 As shown, the method specifically includes the following steps: S110, acquire DICOM data to be synchronized from at least one data source.
[0021] S120, based on pre-built hierarchical cleaning rules, determine the data to be cleaned in the DICOM data to be synchronized and perform data cleaning to obtain the cleaned DICOM data.
[0022] S130, through a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped based on a pre-built standard data dictionary to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
[0023] In application scenarios such as regional imaging centers, medical consortia, or remote diagnostic platforms, standardized synchronous processing of multi-source heterogeneous DICOM data is required. Examples include scenarios involving the aggregation of medical data from different departments within the same medical institution, or from different medical institutions within the same region. These medical data aggregation scenarios enable cross-device, cross-vendor, cross-institutional, and cross-system data synchronization, management, and sharing.
[0024] like Figure 1 As shown, the standardized synchronization device establishes communication with at least one imaging device from at least one organization. In this embodiment, the data source can be any imaging device that establishes communication with the standardized synchronization device. When each data source acquires image data, it encapsulates the image data and the corresponding data tags into a DICOM file and transmits the DICOM file to the standardized synchronization device to achieve standardized synchronization processing across devices, manufacturers, organizations, and systems.
[0025] The standardized synchronization device receives DICOM data to be synchronized, parses the data, cleans the parsed data, and removes or corrects non-standard, abnormal, redundant, and semantically contradictory data. This ensures that the data entering the mapping stage has a complete structure, valid fields, and valid values, preventing erroneous data from being fixed to the target library after standardized mapping. This improves the accuracy, reliability, and usability of subsequent standardized synchronization data, ensuring that the final standardized data can be directly used for data sharing, clinical analysis, and AI model training.
[0026] To avoid issues such as missing or ambiguous data cleaning mechanisms, this disclosure establishes tiered cleaning rules to achieve tiered cleaning of DICOM data to be synchronized. Tiered cleaning rules can be understood as a set of rules divided into different levels and with different constraint strengths based on at least one of the following: the importance of the DICOM data, the type of anomaly, and the magnitude of its impact on subsequent business. Optionally, the tiered cleaning rules include mandatory cleaning rules, optional cleaning rules, and prohibited cleaning rules. Mandatory cleaning rules can be understood as mandatory processing rules set for anomaly data that affects the legality, integrity, and security of the data. For example, data anomaly types in mandatory cleaning rules may include, but are not limited to, missing labels, invalid fields, mismatched value ranges, redundancy, and semantic contradictions. Optional cleaning rules can be understood as flexible processing rules set for anomaly data that does not affect the legality of the data subject but affects standardization, aesthetics, or user experience. Data anomaly types in optional cleaning rules may include format errors such as inconsistent capitalization, extra spaces, empty non-critical fields, and inconsistent formatting. Prohibited cleaning rules can be understood as rules that prohibit modification of original image data, diagnostic key information, and non-reproducible original information. The data tags included in the prohibited cleaning rules may include, but are not limited to, image pixel data, lesion descriptions, original diagnostic information, and original parameters of key equipment.
[0027] The mandatory cleaning rules, the optional cleaning rules, and the prohibited cleaning rules each correspond to a data label. The mandatory cleaning rules and the optional cleaning rules include data anomaly types, which include at least one of the following: missing label, mismatched value range, incorrect format, redundancy, and semantic contradiction.
[0028] Accordingly, based on the pre-built hierarchical cleaning rules, the data to be cleaned in the DICOM data to be synchronized is determined, including: determining the range of data to be cleaned in the DICOM data to be synchronized based on the data tags in the mandatory cleaning rules and the optional cleaning rules; and performing anomaly detection on the data within the range of data to be cleaned based on the data anomaly types in the mandatory cleaning rules and the optional cleaning rules to determine the data to be cleaned, wherein the data to be cleaned includes at least the abnormal data corresponding to the data tags in the mandatory cleaning rules and the data anomaly types.
[0029] Optionally, the optional cleaning rules may include at least one of a cleaning business identifier and a non-cleaning business identifier. For example, the business of DICOM data may include at least one of AI training, data visualization, archiving and storage, and data sharing. The same DICOM data may correspond to different cleaning rules for different businesses. For example, in the AI training business, abnormal data corresponding to the optional cleaning rules may be left uncleaned to improve the diversity and generalization of the sample data; for example, in the archiving and storage and data sharing businesses, abnormal data corresponding to the optional cleaning rules may be cleaned to improve data accuracy and reliability.
[0030] By setting at least one of a cleaning service identifier and a non-cleaning service identifier in the optional cleaning rules, the actual service identifier of the DICOM data is matched in the optional cleaning rules. If the actual service identifier of the DICOM data is a cleaning service identifier, the data to be cleaned in the DICOM data to be synchronized is determined according to the mandatory cleaning rules and the optional cleaning rules. If the actual service identifier of the DICOM data is a non-cleaning service identifier, the data to be cleaned in the DICOM data to be synchronized is determined according to the mandatory cleaning rules.
[0031] Optionally, the data tags in the prohibited cleaning rules are matched in the DICOM data to be synchronized to determine the prohibited cleaning data tags. Based on the data tags in the DICOM data to be synchronized other than the prohibited cleaning data tags, the range of cleaned data is determined. Data anomaly detection is performed in the cleaned data range. If the anomaly type of the data content in the range of data to be cleaned meets the data anomaly type in the mandatory cleaning rules and / or optional cleaning rules, it is determined as data to be cleaned.
[0032] Based on the data cleaning method corresponding to each type of data anomaly, data cleaning is performed on the abnormal data in the DICOM data to be synchronized to obtain the cleaned DICOM data.
[0033] In this embodiment, by setting hierarchical cleaning rules and clearly defining boundaries, an adaptive dynamic cleaning strategy is achieved, avoiding a one-size-fits-all cleaning approach.
[0034] The cleaned DICOM data is standardized and mapped to obtain standardized synchronized data. Specifically, the cleaned DICOM data is standardized and mapped based on a pre-built standard data dictionary. The pre-built standard data dictionary can be understood as a pre-constructed dataset containing unified standard values and semantic correspondences, used to provide a unified standard basis and legal value range for data cleaning and data mapping. This standard data dictionary can include standard data elements or standard value ranges. Data standardization is achieved by converting data tags in the cleaned DICOM data into standard data elements in the standard data dictionary, and by converting the data content of data tags in the cleaned DICOM data into standard values in the standard data dictionary. For example, for a certain data source, the data tag for patient gender can be "name" or "name". The data content of the patient gender data tag can be represented by 1 for male and 2 for female. In the standard data dictionary, the standard value range for this data tag includes: male and female. Accordingly, data standardization can be achieved by mapping "name" or "name" to "patient name", mapping 1 to male, and mapping 2 to female.
[0035] In some embodiments, the pre-built standard data dictionary includes at least one of the following: a general value domain dictionary, a vendor-adapted value domain dictionary, a modality-adapted standard dictionary, and a business scenario-adapted standard dictionary. The general value domain dictionary can be understood as a standard dictionary applicable to all scenarios, without distinguishing between vendors, devices, or modalities. For example, the general value domain dictionary may include standard data elements and standard value ranges for data tags such as gender and date. The vendor-adapted value domain dictionary can be understood as being used to address inconsistencies in the descriptions of equipment produced by different vendors, such as vendor name, device model, vendor-specific tags, and vendor-specific parameters. Different vendors can configure a vendor-adapted value domain dictionary. The modality-adapted standard dictionary can be understood as being used to provide standard data elements and standard value ranges for different imaging modalities. For example, imaging modalities may include, but are not limited to, CT, MR, DR, ultrasound, and PET-CT. For example, slice thickness in CT, sequence parameters in MR, and probe information in ultrasound are professional data tags for different imaging modalities. Different imaging modalities can configure a modality-adapted standard dictionary. A business scenario adaptation standard dictionary can be understood as a dictionary set up for different data uses, such as using standard value ranges with different precision and field sets for different scenarios such as clinical review, scientific research analysis, AI training, and data reporting.
[0036] Optionally, the attribute information of the cleaned DICOM data is identified, including at least one of the following: equipment manufacturer, equipment model, image modality, affiliated organization, and business scenario; a matching dictionary is determined in the pre-built standard data dictionary based on the attribute information; a mapping rule group is determined based on the matching dictionary; and the cleaned DICOM data is mapped based on the mapping rule group to obtain standardized synchronization data.
[0037] For example, determining the matching dictionary based on the attribute information in the pre-built standard data dictionary includes at least one of the following: matching the device manufacturer's corresponding manufacturer-matching value range dictionary in the pre-built standard data dictionary; matching the image modality's corresponding modality-matching standard dictionary in the pre-built standard data dictionary; and matching the business scenario-matching standard dictionary corresponding to the business scenario in the pre-built standard data dictionary. Furthermore, the matching dictionary includes at least a general value range dictionary. Determining the matching dictionary through attribute information allows for targeted mapping processing, avoiding interference from other irrelevant dictionaries. In some embodiments, the pre-built standard data dictionary may also include an institution-matching standard dictionary, an equipment model-matching standard dictionary, etc., to be applicable to different medical institutions and equipment models.
[0038] Mapping rule groups are determined based on the matching dictionary, wherein the matching dictionary can be a local dictionary in a pre-built standard data dictionary, each dictionary in the matching dictionary can correspond to a mapping rule, the matching dictionary includes at least one dictionary, and the corresponding mapping rule group includes at least one mapping rule. For example, the mapping rule group can be a rule group of "Manufacturer A - CT Equipment - Hospital A - AI Training Scenario".
[0039] Determining a mapping rule group based on the matching dictionary includes: calling the mapping rules corresponding to the matching dictionaries in the mapping rule library to form the mapping rule group, wherein the mapping rule library includes mapping rules corresponding to each dictionary; wherein the mapping rule corresponding to each dictionary is used to map the data tags in the DICOM data to be mapped to the standard data elements in the dictionary, and to map the data content in the DICOM data to be mapped to the standard values in the dictionary.
[0040] Each of the aforementioned dictionaries corresponds to a mapping rule specifically designed to achieve a precise mapping between the DICOM data to be mapped and the dictionary standard. The mapping rule is established by: for each level of data tags extracted from the DICOM data to be mapped (such as patient information tags, examination parameter tags, and device information tags), determining the standard data element corresponding to the native DICOM data tag in the dictionary. For example, a manufacturer's proprietary tag "device model_GE" corresponds to the standard data element "device model" in the manufacturer's adaptation value domain dictionary. A mapping rule corresponding to the manufacturer's adaptation value domain dictionary is established based on this proprietary tag and the standard data element in the dictionary. Similarly, for the "slice thickness tag" of a CT device, the standard data element in the modality adaptation standard dictionary is "CT slice thickness." A mapping rule corresponding to the modality adaptation standard dictionary is established based on the original data identifier and the standard data element in the modality adaptation standard dictionary. Likewise, for each data content in the DICOM data to be mapped (i.e., the specific value corresponding to the data tag), mapping rules are established based on the correspondence between the original data content and the standard value range in the dictionary.
[0041] Based on the aforementioned mapping rule set, the cleaned DICOM data is mapped to obtain standardized synchronized data. Specifically, for the data tags extracted from each level of the DICOM data to be mapped, the native DICOM data tags are accurately mapped to the preset standard data elements in the corresponding dictionary using the mapping rules of the corresponding dictionary. For example, the mapping rules corresponding to the manufacturer-adaptive value domain dictionary map a manufacturer's private tag "device model_GE" to the standard data element "device model" in the dictionary; the mapping rules corresponding to the modality-adaptive standard dictionary map the "slice thickness tag" of the CT device to the standard data element "CT slice thickness" in the dictionary. Through the mapping rules of the corresponding dictionary, combined with the preset standard values and semantic relationships in the dictionary, the native data content is accurately mapped to the standard values in the dictionary. For example, the mapping rules corresponding to the general value domain dictionary map the data content "male 1" to the standard value "male" in the dictionary; the mapping rules corresponding to the business scenario-adaptive standard dictionary map "slice thickness 2.0mm" in the AI training scenario to the preset standard value "2.0mm" in the dictionary (adapting to the accuracy requirements of AI training).
[0042] In some embodiments, to avoid data mapping based on a single field correspondence, a three-dimensional collaborative mapping model is used to achieve data mapping. The three-dimensional collaborative mapping model can be understood as a set of mutually collaborative and indivisible standardized mapping models built around three core dimensions: data structure, value range, and standard data logical model, based on a pre-built standard data dictionary. Its core function is to achieve accurate mapping of cleaned, multi-source heterogeneous DICOM data to standard data. The three dimensions work synchronously and mutually verify each other, unlike traditional single-field mapping. It adapts to the heterogeneous characteristics of DICOM data across multiple vendors, modalities, and scenarios, while also supporting efficient invocation and execution of mapping rule groups.
[0043] Specifically, through a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped based on a pre-built standard data dictionary to obtain standardized synchronization data. This includes: parsing the original hierarchical structure of the cleaned DICOM data and extracting data tags and data content at each level of the original hierarchical structure; mapping the data tags at each level of the original hierarchical structure sequentially based on the pre-built standard data dictionary to obtain standard data elements, wherein the mapping of related tags in different levels is consistent, and the related tags are data tags of the same object and the same inspection type at each level; mapping the data content of the data content to standard values based on the pre-built standard data dictionary and semantic relevance to obtain value range mapping data; and performing correlation verification on the mapped standard data elements and the value range mapping data based on the standard data logic model to obtain standardized synchronization data that passes the verification.
[0044] The original hierarchical structure of the cleaned DICOM data can include patient layer, examination layer, sequence layer, and image layer. The DICOM tags and data content corresponding to each layer are extracted.
[0045] Mapping rules are determined based on a pre-built standard data dictionary. For example, a matching dictionary is determined based on the attribute information of the cleaned DICOM data within the pre-built standard data dictionary. Mapping rule groups are then determined based on this matching dictionary. Following a hierarchical linkage logic, the native DICOM tags at each level are mapped one by one to the corresponding DICOM-specific standard data elements: first, the mapping of core tags at the top level (e.g., patient level) (such as patient ID and name) is completed; then, the mapping of associated tags at the downstream levels (examination level, sequence level, imaging level) (such as examination ID, sequence ID, and pixel parameter tags) is completed in a linked manner. The mapping of associated tags within the same level is consistent, meaning the mapped data structure is unified and the associations are consistent. The associated tags are data tags for the same object and the same examination type at each level. Here, the object can be a patient. During the mapping process, the consistency between the data structure at each level and the target standard data structure is simultaneously verified to ensure no structural missingness or field misalignment. If a discrepancy is found, rule adaptation verification is triggered (to confirm whether it is a rule configuration deviation). If no error is found, dynamic mapping of value domain semantics is performed.
[0046] Dynamic semantic mapping of value domains must ensure semantic consistency: 1. For each standard data element whose structure mapping has been completed, extract its cleaned original value; call the hierarchical standard dictionary, first match the manufacturer's personalized expression corresponding to the original value through the manufacturer-adapted value domain dictionary, and then map it to the standard value in the core value domain dictionary based on the preset semantic association in the dictionary (e.g., the original value "computed tomography scan" is mapped to the standard value "CT" in the core dictionary through the manufacturer-adapted dictionary matching); for semantic ambiguity and newly added manufacturer-adapted values, start the dynamic semantic similarity matching logic of the 3D model, compare the semantic association between the original value and the standard value in the core value domain dictionary, automatically recommend the most suitable standard value, and trigger manual review (after the review is passed, the manufacturer-adapted dictionary is updated synchronously to optimize subsequent mapping); after the value domain mapping is completed, verify the binding constraints between the value and the standard data element (ensure that the value is within the legal range of the dictionary).
[0047] Logical Model Collaborative Verification: The system calls the target standard data logical model (DICOM specific), and substitutes the completed structure mapping and value range mapping results into the logical model for correlation verification. Verification covers three main logical correlations: 1) Data logic between levels (e.g., patient IDs are consistent between the patient layer and the examination layer); 2) Semantic logic between fields (e.g., the matching of the image modality "CT" with the examination site "head," avoiding logical contradictions); 3) Adaptation logic between data and business scenarios (e.g., data from AI training scenarios, where mapping accuracy meets preset thresholds). If verification fails, the system identifies the contradiction type (structural contradiction / value range contradiction / logical contradiction), returns to remapping, and continues until verification passes.
[0048] In the data acquisition process described above, the standardized synchronous data obtained after mapping undergoes secondary verification and correction: A full secondary verification is performed on the standardized synchronous data after 3D collaborative mapping: First, the data structure, field naming, and value range are verified to ensure they fully comply with the hierarchical standard dictionary and DICOM-specific standard data element requirements; second, the data logic is verified to ensure it fully matches the target standard data logic model; third, the mapping accuracy is sampled and verified (ensuring no mapping deviation). For minor deviations found during the secondary verification (such as inaccurate value range mapping or missing structural fields, etc., type 1 deviations), the mapping rules and standard dictionary are automatically invoked for correction; for serious deviations (such as logical contradictions that cannot be automatically corrected, etc., type 2 deviations), anomalies are marked and manual intervention is triggered. After correction, the secondary verification is re-executed until all standards are met.
[0049] After the secondary verification is passed, the mapping is confirmed to be complete, and DICOM data conforming to the unified standard is output and pushed to the next step of data entry and synchronization. A complete mapping traceability log is retained, recording core information: the original attributes of the data to be mapped (manufacturer, device, modality), the matching mapping rule group, the structural mapping correspondence, the values before and after the value range mapping (associated with standard dictionary entries), the logical verification results, the mapping time and the operator, to ensure that the mapping process is traceable and auditable. Simultaneously, rule optimization points found during the mapping process (such as adding personalized values, semantic matching deviations) are pushed to the rule management module to provide data support for the iterative optimization of the 3D collaborative mapping model and standard dictionary.
[0050] Based on the above embodiments, the standardized synchronized data is configured as an asset, which includes setting at least one of asset type, asset entity information and lifecycle parameters; and setting the sharing organization scope and access interface configuration information of the standardized synchronized data to respond to access requests for the standardized synchronized data.
[0051] The asset type can include image data, text reports, etc., the asset subject information can include owner information and manager information, and the lifecycle parameters can include creation time and validity period.
[0052] The scope of organizations sharing standardized synchronized data can be understood as the range of structures that can share the aforementioned standardized synchronized data. Access interface configuration information may include, but is not limited to, API request methods, return formats, paths, and encryption algorithms.
[0053] By configuring assets, standardized synchronized data can be managed as an asset, clarifying data ownership and attributes, and ensuring compliant and efficient management throughout the data lifecycle. By setting the scope of sharing institutions and access interface configuration information for the standardized synchronized data, secure and controllable sharing of the standardized synchronized data can be achieved, adapting to cross-institutional and multi-scenario needs, and supporting the efficient reuse and value mining of standardized synchronized data.
[0054] The technical solution in this embodiment formulates hierarchical and clearly defined cleaning rules through pre-built hierarchical cleaning rules, enabling accurate and reliable data cleaning of multi-source heterogeneous DICOM data, and providing a reliable data foundation for data standardization mapping. Through a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped from three dimensions: data structure, value range, and standard data logic model, avoiding mapping based on single-field correspondences and improving the accuracy of standardization processing.
[0055] Figure 3 This is a schematic diagram of a standardized synchronization device for multi-source heterogeneous DICOM data provided in an embodiment of this disclosure. The device can be implemented in software and / or hardware, and specifically includes: a data acquisition module 210, a data cleaning module 220, and a data mapping module 230.
[0056] The data acquisition module 210 is used to acquire DICOM data to be synchronized transmitted from at least one data source. The data cleaning module 220 is used to determine the data to be cleaned in the DICOM data to be synchronized based on pre-built hierarchical cleaning rules and perform data cleaning to obtain cleaned DICOM data. The data mapping module 230 is used to perform data mapping on the cleaned DICOM data based on a pre-built standard data dictionary through a three-dimensional collaborative mapping model to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
[0057] The technical solution in this embodiment formulates hierarchical and clearly defined cleaning rules through pre-built hierarchical cleaning rules, enabling accurate and reliable data cleaning of multi-source heterogeneous DICOM data, and providing a reliable data foundation for data standardization mapping. Through a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped from three dimensions: data structure, value range, and standard data logic model, avoiding mapping based on single-field correspondences and improving the accuracy of standardization processing.
[0058] Based on the above embodiments, optionally, the hierarchical cleaning rules include mandatory cleaning rules, optional cleaning rules, and prohibited cleaning rules. The mandatory cleaning rules, optional cleaning rules, and prohibited cleaning rules correspond to data labels respectively. The mandatory cleaning rules and optional cleaning rules include data anomaly types, which include at least one of the following: missing label, mismatched value range, format error, redundancy and duplication, and semantic contradiction.
[0059] Optionally, the data cleaning module 220 is used to determine the range of data to be cleaned in the DICOM data to be synchronized based on the data tags in the mandatory cleaning rules and the optional cleaning rules; and to perform anomaly detection on the data within the range of data to be cleaned based on the data anomaly types in the mandatory cleaning rules and the optional cleaning rules, and to determine the data to be cleaned, wherein the data to be cleaned includes at least the abnormal data corresponding to the data tags in the mandatory cleaning rules and the data anomaly types.
[0060] Based on the above embodiments, optionally, the data mapping module 230 is used to parse the original hierarchical structure of the cleaned DICOM data, extract the data tags and data content of each level in the original hierarchical structure; based on the pre-built standard data dictionary, it sequentially maps the data tags of each level in the original hierarchical structure to obtain standard data elements, wherein the mapping of related tags in different levels is consistent, and the related tags are data tags of the same object and the same inspection type in each level; based on the pre-built standard data dictionary and semantic association, it maps the data content of the data content to standard values to obtain value range mapping data; based on the standard data logic model, it performs association verification on the mapped standard data elements and the value range mapping data to obtain standardized synchronization data that passes the verification.
[0061] Based on the above embodiments, optionally, the pre-built standard data dictionary includes at least one of: a general value domain dictionary, a vendor-adapted value domain dictionary, a modality-adapted standard dictionary, and a business scenario-adapted standard dictionary; The data mapping module 230 is also used to identify the attribute information of the cleaned DICOM data, the attribute information including at least one of the following: equipment manufacturer, equipment model, image modality, affiliated organization, and business scenario; determine a matching dictionary in the pre-built standard data dictionary based on the attribute information; determine a mapping rule group based on the matching dictionary; and perform data mapping on the cleaned DICOM data based on the mapping rule group to obtain standardized synchronized data.
[0062] The data mapping module 230 is further configured to call the mapping rules corresponding to the respective matching dictionaries in the mapping rule library to form the mapping rule group, wherein the mapping rule library includes the mapping rules corresponding to each dictionary; wherein the mapping rule corresponding to each dictionary is used to map the data tags in the DICOM data to be mapped to the standard data elements in the dictionary, and to map the data content in the DICOM data to be mapped to the standard values in the dictionary.
[0063] Optionally, based on the above embodiments, the device further includes a data management module for configuring the standardized synchronized data as an asset. The asset configuration includes setting at least one of asset type, asset entity information, and lifecycle parameters; and setting the sharing organization scope and access interface configuration information of the standardized synchronized data to respond to access requests for the standardized synchronized data.
[0064] The above-described products can perform the methods provided in any embodiment of this disclosure, and have the corresponding functional modules and beneficial effects for performing the methods.
[0065] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0066] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0067] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0068] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as standardized synchronization methods for multi-source heterogeneous DICOM data.
[0069] In some embodiments, the method for standardizing and synchronizing multi-source heterogeneous DICOM data can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for standardizing and synchronizing multi-source heterogeneous DICOM data described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for standardizing and synchronizing multi-source heterogeneous DICOM data by any other suitable means (e.g., by means of firmware).
[0070] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0071] Computer programs used to implement the methods of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0072] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0073] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0074] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0075] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0076] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0077] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements a standardized synchronization method for multi-source heterogeneous DICOM data according to any embodiment of this disclosure.
[0078] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0079] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A standardized synchronization method for multi-source heterogeneous DICOM data, characterized in that, include: Acquire DICOM data to be synchronized from at least one data source; Based on pre-built hierarchical cleaning rules, the data to be cleaned in the DICOM data to be synchronized is determined and data cleaning is performed to obtain the cleaned DICOM data; The cleaned DICOM data is mapped using a three-dimensional collaborative mapping model based on a pre-built standard data dictionary to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
2. The method according to claim 1, characterized in that, The hierarchical cleaning rules include mandatory cleaning rules, optional cleaning rules, and prohibited cleaning rules. The mandatory cleaning rules, optional cleaning rules, and prohibited cleaning rules correspond to data labels respectively. The mandatory cleaning rules and optional cleaning rules include data anomaly types, which include at least one of the following: missing label, mismatched value range, incorrect format, redundancy and duplication, and semantic contradiction.
3. The method according to claim 2, characterized in that, Based on pre-built hierarchical cleaning rules, the data to be cleaned in the DICOM data to be synchronized is determined, including: Based on the data tags in the mandatory cleaning rules and the optional cleaning rules, the range of cleaning data is determined in the DICOM data to be synchronized. Based on the data anomaly types in the mandatory cleaning rules and the optional cleaning rules, anomaly detection is performed on the data within the cleaning data range to determine the data to be cleaned. The data to be cleaned includes at least the abnormal data corresponding to the data tags in the mandatory cleaning rules and the data anomaly types.
4. The method according to claim 1, characterized in that, Using a three-dimensional collaborative mapping model, the cleaned DICOM data is mapped based on a pre-built standard data dictionary to obtain standardized synchronization data, including: The original hierarchical structure of the cleaned DICOM data is analyzed, and the data tags and data content of each level in the original hierarchical structure are extracted. Based on the pre-built standard data dictionary, data labels of each level in the original hierarchical structure are mapped sequentially to obtain standard data elements. Among them, the mapping of related labels in different levels is consistent. The related labels are data labels of the same object and the same inspection type in each level. Based on the pre-built standard data dictionary and semantic association, the data content is mapped to standard values to obtain value range mapping data; Based on the standard data logic model, the mapped standard data elements and the value range mapping data are subjected to correlation verification to obtain standardized synchronization data that passes the verification.
5. The method according to claim 4, characterized in that, The pre-built standard data dictionary includes at least one of the following: a general value domain dictionary, a vendor-adapted value domain dictionary, a modal adaptation standard dictionary, and a business scenario adaptation standard dictionary; The method further includes: Identify the attribute information of the cleaned DICOM data, the attribute information including at least one of the following: equipment manufacturer, equipment model, image modality, affiliated organization, and business scenario; Based on the attribute information, a matching dictionary is determined from the pre-built standard data dictionary; Based on the matching dictionary, a mapping rule group is determined, and the cleaned DICOM data is mapped based on the mapping rule group to obtain standardized synchronization data.
6. The method according to claim 5, characterized in that, Determining mapping rule groups based on the matching dictionary includes: The mapping rules corresponding to the respective matching dictionaries are called from the mapping rule base to form the mapping rule group, wherein the mapping rule base includes the mapping rules corresponding to each dictionary; The mapping rules corresponding to each dictionary are used to map the data tags in the DICOM data to be mapped to the standard data elements in the dictionary, and to map the data content in the DICOM data to be mapped to the standard values in the dictionary.
7. The method according to claim 1, characterized in that, The method further includes: The standardized synchronized data is used for asset configuration, which includes setting at least one of the following: asset type, asset entity information, and lifecycle parameters. Configure the sharing organization scope and access interface information of the standardized synchronization data to respond to access requests for the standardized synchronization data.
8. A standardized synchronization device for multi-source heterogeneous DICOM data, characterized in that, include: The data acquisition module is used to acquire DICOM data to be synchronized from at least one data source. The data cleaning module is used to determine the data to be cleaned in the DICOM data to be synchronized based on pre-built hierarchical cleaning rules and perform data cleaning to obtain cleaned DICOM data. The data mapping module is used to perform data mapping on the cleaned DICOM data based on a pre-built standard data dictionary through a three-dimensional collaborative mapping model to obtain standardized synchronized data. The three-dimensional collaborative mapping model includes mapping rules for three dimensions: data structure, value range, and standard data logic model.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform a standardized synchronization method for multi-source heterogeneous DICOM data as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute and implement the standardized synchronization method for multi-source heterogeneous DICOM data as described in any one of claims 1-7.