Cross-standard data exchange method and system based on BRIDG model
Through the cross-standard data exchange method based on the BRIDG model, dynamic mapping rules are constructed using semantic label mapping and dynamic bridge ontology, which solves the problem of high maintenance costs of mapping rules and difficult to adapt to standard changes in traditional methods, and achieves efficient and flexible cross-standard data exchange.
Patent Information
- Application Number
- CN202510258039.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional cross-standard data exchange methods have the problem of high maintenance costs of mapping rules and difficult to adapt to standard changes.
A cross-standard data exchange method based on the BRIDG model is adopted to construct a rich semantic map by semantic label mapping and fusion of source system data; then, dynamic mapping rules are automatically generated through cross-map concept anchoring and dynamic bridging ontology construction.
It reduces the maintenance cost of mapping rules, improves the flexibility and adaptability of the system, and realizes efficient cross-standard data exchange.
Smart Images

Figure CN120179719A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross - standard data exchange, and particularly to a cross - standard data exchange method and system based on the BRIDG model. Background Art
[0002] Cross - standard data exchange refers to the process of data conversion and exchange between different data standards. Nowadays, various data standards emerge in an endless stream. For example, HL7 and DICOM in the medical field, ISO 20022 in the financial field, STEP in the manufacturing field, etc. Different industries, different organizations, and even different departments within the same organization adopt different data standards, which leads to the phenomenon of data islands and hinders information sharing and business collaboration. Therefore, cross - standard data exchange technology has emerged.
[0003] However, traditional cross - standard data exchange methods have the defect of high mapping rule maintenance costs. They usually rely on predefined and hard - coded mapping rules. When the source standard or target standard changes, or when new data standards need to be supported, a large number of mapping rules need to be manually modified or added, resulting in high maintenance costs, low efficiency, and easy errors.
[0004] Traditional cross - standard data exchange methods have the defect of being difficult to adapt to standard changes. Their predefined mapping rules lack flexibility and are difficult to adapt to the evolution of data standards. When the standard changes, a large number of rules need to be rewritten or modified, and even the entire data exchange system needs to be redesigned. Summary of the Invention
[0005] Based on this, it is necessary to provide a cross - standard data exchange method and system based on the BRIDG model to solve at least one of the above - mentioned technical problems.
[0006] To achieve the above object, a cross - standard data exchange method based on the BRIDG model includes the following steps:
[0007] Step S1: Collect data from the source system to obtain the source system data to be converted; perform semantic label mapping on the source system data to be converted to obtain a preliminary semantic mapping result; perform semantic label fusion based on the relationship graph on the preliminary semantic mapping result to obtain a rich semantic graph;
[0008] Step S2: Obtain the specification document of the target data standard; perform target standard structured parsing on the specification document of the target data standard to obtain a target standard structure model; construct a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph; perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; extract bridging concepts according to the initial anchoring candidate set and construct a dynamic bridging ontology to obtain a refined bridging ontology;
[0009] Step S3: Refine the bridging paths according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a refined bridging path set; model the transformation operation for the refined bridging path set to obtain a structured transformation logic; perform rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set;
[0010] Step S4: Apply the mapping rules to the source system data to be transformed according to the dynamic mapping rule set to obtain transformed data fragments; assemble the target structure for the transformed data fragments to obtain preliminary target instances; perform data quality verification on the preliminary target instances to obtain target standard data instances;
[0011] Step S5: Monitor the transformation results for the target standard data instances and perform root cause analysis of errors to obtain a set of error cause tags; evolve the bridging ontology according to the set of error cause tags to obtain a refined mapping knowledge base to achieve the cross-standard data exchange task.
[0012] Preferably, step S1 includes the following steps:
[0013] Step S11: Collect structured data from the source system to obtain the source system data to be transformed;
[0014] Step S12: Perform preliminary data parsing on the source system data to be transformed to obtain parsed data;
[0015] Step S13: Perform element-level semantic label mapping on the parsed data to obtain a preliminary semantic mapping result;
[0016] Step S14: Mine relational context for the parsed data according to the preliminary semantic mapping result and construct a relational graph to obtain a preliminary relational graph;
[0017] Step S15: Perform semantic label fusion according to the preliminary relational graph and the preliminary semantic mapping result to obtain a rich semantic graph.
[0018] Preferably, step S2 includes the following steps:
[0019] Step S21: Obtain the specification document of the target data standard; perform target standard structured parsing according to the specification document of the target data standard to obtain the target standard structure model;
[0020] Step S22: Construct the target standard semantic graph for the target standard structure model to obtain the target standard semantic graph;
[0021] Step S23: Perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set;
[0022] Step S24: Perform stable matching optimization based on the initial anchoring candidate set and the cross-graph node embedding data to obtain a preliminary concept anchoring set;
[0023] Step S25: Perform structured neighborhood alignment on the target standard semantic graph and the rich semantic graph according to the preliminary concept anchoring set to obtain a high-confidence concept mapping;
[0024] Step S26: Explore the bridging path for the high-confidence concept mapping to obtain bridging path exploration data; extract potential bridging paths from the bridging path exploration data to obtain potential bridging paths;
[0025] Step S27: Assemble the dynamic bridging ontology for the potential bridging paths to obtain the dynamic bridging ontology; optimize and refine the dynamic bridging ontology to obtain the refined bridging ontology.
[0026] Preferably, step S23 includes the following steps:
[0027] Step S231: Enhance the graph node features of the target standard semantic graph and the rich semantic graph to obtain enhanced node features;
[0028] Step S232: Perform cross-graph Transformer embedding based on the enhanced node features to obtain cross-graph node embedding data;
[0029] Step S233: Construct cross-graph node pairs for the cross-graph node embedding data to obtain cross-graph node pairs;
[0030] Step S234: Calculate the similarity of the embedding vectors for the cross-graph node pairs to obtain the node pair similarity;
[0031] Step S235: Define the similarity threshold for the node pair similarity to obtain the similarity threshold;
[0032] Step S236: Screen the initial candidate anchor points for the node pair similarity according to the similarity threshold to obtain the initial anchoring candidate set.
[0033] Preferably, step S3 includes the following steps:
[0034] Step S31: Refine the bridging paths based on the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a refined bridging path set;
[0035] Step S32: Model the transformation operation for the refined bridging path set to obtain a structured transformation logic;
[0036] Step S33: Generate mapping rule patterns for the structured transformation logic according to the predefined mapping rule pattern library to obtain a candidate mapping rule set;
[0037] Step S34: Extract sample source data from the source system data to be transformed and set the transformation expectation target to obtain sample source data and expected target data;
[0038] Step S35: Verify and optimize the candidate mapping rule set according to the sample source data and the expected target data to obtain a dynamic mapping rule set.
[0039] Preferably, step S31 includes the following steps:
[0040] Step S311: Identify the target elements to be mapped in the target standard semantic graph to obtain a set of targets to be searched;
[0041] Step S312: Traverse the target elements in the set of targets to be searched in a loop, and use the refined bridging ontology to conduct a preliminary path exploration on the rich semantic graph to obtain a preliminary path discovery result;
[0042] Step S313: Quantify and evaluate the path features of the preliminary path discovery result to obtain a set of paths with evaluation scores;
[0043] Step S314: Screen and sort the set of paths with evaluation scores to optimize and obtain a candidate bridging path set;
[0044] Step S315: Extract path elements from the candidate bridging path set to obtain a refined bridging path set.
[0045] Preferably, step S32 includes the following steps:
[0046] Step S321: Extract the relational semantic mapping rules for the refined bridging path set to obtain a set of paths with operation identifiers;
[0047] Step S322: Analyze the data difference features of the set of paths with operation identifiers to obtain a set of paths with difference features;
[0048] Step S323: Correlate and enhance the business rules according to the set of paths with difference features to obtain a set of paths to be formalized;
[0049] Step S324: Perform a formal language element mapping on the formalized path set to obtain a preliminary formalized path set;
[0050] Step S325: Construct a conversion logic expression based on the preliminary formalized path set to obtain a structured conversion logic.
[0051] Preferably, step S4 includes the following steps:
[0052] Step S41: Extract source data instances from the source system data to be converted to obtain data units to be converted;
[0053] Step S42: Perform semantic rule matching and execution on the data units to be converted according to the dynamic mapping rule set to obtain converted data fragments;
[0054] Step S43: Assemble the target structure for the converted data fragments according to the target standard structure model to obtain a preliminary target instance;
[0055] Step S44: Perform semantic constraint verification and enhancement on the preliminary target instance according to the target standard semantic graph to obtain a verified target instance;
[0056] Step S45: Serialize the target format of the verified target instance to obtain a target standard data instance.
[0057] Preferably, step S5 includes the following steps:
[0058] Step S51: Conduct a conversion performance audit on the target standard data instance to obtain a conversion performance report;
[0059] Step S52: Collect user problem feedback on the target standard data to obtain a user feedback list;
[0060] Step S53: Trace the errors in the dynamic mapping rule set according to the conversion performance report and the user feedback list to obtain an error cause marking set;
[0061] Step S54: Perform intelligent optimization of the mapping rules in the dynamic mapping rule set according to the error cause marking set to obtain an optimized mapping rule set;
[0062] Step S55: Iteratively update the bridging knowledge base according to the error cause marking set, the optimized mapping rule set, and the refined bridging ontology to obtain a refined mapping knowledge base.
[0063] Preferably, the present invention also provides a cross-standard data exchange system based on the BRIDG model for executing the cross-standard data exchange method based on the BRIDG model as described above. The cross-standard data exchange system based on the BRIDG model includes:
[0064] A semantic topology module for collecting data from the source system to obtain the source system data to be converted; performing semantic label mapping on the source system data to be converted to obtain a preliminary semantic mapping result; performing semantic label fusion based on the relationship graph on the preliminary semantic mapping result to obtain a rich semantic graph;
[0065] A bridging ontology construction module for obtaining the specification document of the target data standard; performing target standard structured parsing on the specification document of the target data standard to obtain a target standard structure model; constructing a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph; performing cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; extracting bridging concepts according to the initial anchoring candidate set and performing dynamic bridging ontology construction to obtain a refined bridging ontology;
[0066] A mapping paradigm generation module for refining the bridging path according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a refined bridging path set; performing transformation operation modeling on the refined bridging path set to obtain structured transformation logic; performing rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set;
[0067] A semantic data shaping module for applying mapping rules to the source system data to be converted according to the dynamic mapping rule set to obtain converted data fragments; assembling the target structure for the converted data fragments to obtain a preliminary target instance; performing data quality verification on the preliminary target instance to obtain a target standard data instance;
[0068] A transformation rule refinement module for monitoring the transformation result of the target standard data instance and performing root cause analysis of errors to obtain an error cause marking set; evolving the bridging ontology according to the error cause marking set to obtain a refined mapping knowledge base to achieve the cross-standard data exchange task.
[0069] Through structured data collection, the present invention efficiently obtains the source system data to be converted, laying a foundation for subsequent processing. Through preliminary parsing and element-level semantic tag mapping, using word vector models and rule-based methods, the meaning of data elements is initially understood, providing a semantic foundation for cross-standard conversion. Through semantic tag fusion based on a relational graph, semantic ambiguity is effectively eliminated, the accuracy of data understanding is improved, and the constructed rich semantic graph provides high-quality semantic knowledge for the construction of subsequent bridging ontologies and the generation of mapping rules. Through structured parsing of the target data standard specification document, a target standard structure model is quickly constructed, providing a structured target blueprint for subsequent data conversion. Through the construction of a target standard semantic graph, the semantic information of the target standard is deeply understood, laying a foundation for cross-standard concept matching and mapping. Through cross-graph concept anchoring and stable matching optimization, potential corresponding relationships between source data and target data are effectively identified, providing an accurate starting point for the exploration of bridging paths. Through the construction and refinement of a dynamic bridging ontology, according to specific source and target standards, bridging knowledge is automatically constructed without the need for a large number of manually predefined mapping rules, greatly improving flexibility and efficiency. Through the refinement of bridging paths, the best conversion paths from source data elements to target data elements are effectively found, providing a basis for modeling subsequent conversion operations. Through the modeling of conversion operations, semantic relationships, data differences, and business rules are transformed into structured conversion logic, laying a foundation for the automatic generation of mapping rules. Through rule pattern matching and generation, mapping rules are automatically generated, greatly reducing the cost of manually writing and maintaining mapping rules, and improving the efficiency and accuracy of rule generation. Through the use of sample data for rule verification and optimization, it is ensured that the generated mapping rules can correctly convert source data into target data, improving the quality and reliability of data conversion. Through the extraction of source data instances, specific operation objects are provided for subsequent data conversion. Through the matching and execution of semantic rules, source data can be accurately converted into data segments that conform to the target standard according to the dynamically generated mapping rules, realizing semantic-driven intelligent conversion. Through the assembly of the target structure, it is ensured that the converted data can meet the structural requirements of the target standard, laying a foundation for the effective utilization of the target system. Through the verification and enhancement of semantic constraints, the quality of the converted data is further improved, ensuring that data types, value ranges, and logical consistency conform to the target standard, and improving the usability and reliability of the data. Through the serialization of the target format, the converted data can be seamlessly integrated into the target system. Through the auditing of conversion performance, the quality and efficiency of data conversion can be comprehensively monitored, and potential problems can be discovered in a timely manner. Through the collection of user problem feedback, valuable information directly from end-users is provided for improving data conversion. Through error tracing and diagnosis, errors occurring in the data conversion process and their root causes can be effectively located, providing a clear direction for subsequent rule optimization. Through the intelligent optimization of mapping rules, according to error feedback and analysis results, mapping rules can be automatically adjusted and optimized, continuously improving the accuracy and efficiency of data conversion and reducing the manual maintenance cost.The iterative update of the bridging knowledge base enables the system to continuously learn and evolve, better adapt to changes in data standards, and in the long run, continuously improve the intelligence level and adaptability of cross-standard data exchange. Therefore, the present invention provides a cross-standard data exchange method based on the BRIDG model, which effectively solves the drawbacks of high mapping rule maintenance costs and difficulty in adapting to standard changes by dynamically constructing a bridging ontology, reasoning using a semantic graph, automatically generating mapping rules, and introducing a feedback mechanism, and realizes more flexible, efficient, and intelligent cross-standard data exchange. Description of the Drawings
[0070] Figure 1 It is a schematic diagram of the step flow of a cross-standard data exchange method based on the BRIDG model;
[0071] Figure 2 It is a schematic diagram of the detailed implementation steps of step S1 in the present invention;
[0072] Figure 3 It is a schematic diagram of the detailed implementation steps of step S3 in the present invention.
[0073] The realization of the object, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0074] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative work belong to the scope of protection of the present invention.
[0075] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0076] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0077] To achieve the above object, please refer to Figures 1 to 3 , a cross-standard data exchange method based on the BRIDG model, comprising the following steps:
[0078] Step S1: Collect data from the source system to obtain the source system data to be converted; perform semantic tag mapping on the source system data to be converted to obtain a preliminary semantic mapping result; perform semantic tag fusion based on the relationship graph on the preliminary semantic mapping result to obtain a rich semantic graph;
[0079] Step S2: Obtain the specification document of the target data standard; perform target standard structured parsing on the specification document of the target data standard to obtain a target standard structure model; construct a target standard semantic graph on the target standard structure model to obtain a target standard semantic graph; perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; extract bridging concepts according to the initial anchoring candidate set and perform dynamic bridging ontology construction to obtain a refined bridging ontology;
[0080] Step S3: Refine the bridging path according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a refined bridging path set; perform transformation operation modeling on the refined bridging path set to obtain a structured transformation logic; perform rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set;
[0081] Step S4: Apply the mapping rule set to the source system data to be converted according to the dynamic mapping rule set to obtain a converted data segment; perform target structure assembly on the converted data segment to obtain a preliminary target instance; perform data quality verification on the preliminary target instance to obtain a target standard data instance;
[0082] Step S5: Monitor the conversion result of the target standard data instance and perform error root cause analysis to obtain an error cause marking set; evolve the bridging ontology according to the error cause marking set to obtain a refined mapping knowledge base to implement the cross-standard data exchange task.
[0083] In the embodiment of the present invention, refer to Figure 1As shown in the figure, it is a schematic diagram of the step flow of the cross-standard data exchange method based on the BRIDG model of the present invention. In this example, the cross-standard data exchange method based on the BRIDG model includes the following steps:
[0084] Step S1: Collect data from the source system to obtain the source system data to be converted; perform semantic tag mapping on the source system data to be converted to obtain a preliminary semantic mapping result; perform semantic tag fusion based on the relationship graph on the preliminary semantic mapping result to obtain a rich semantic graph;
[0085] In the embodiment of the present invention, first, patient information is collected from the source system electronic medical record database through a JDBC connector and stored as a data frame. Subsequently, the date of birth is format-checked and converted using regular expressions, and the parsed data is stored in dictionary form. Then, a pre-trained Word2Vec model and a rule-based method are used to map data fields to medical ontology concepts and BRIDG model attributes. For example, "patient_name" is mapped to "patient name". Then, a relationship graph of patient information is constructed based on data structure and association rule mining. For example, the "patient name" and "date of birth" nodes are connected. Finally, disambiguation of multiple mappings is performed according to the semantic tags of adjacent nodes, and the finally determined ontology concepts are added to the relationship graph to form a rich semantic graph containing rich semantic information.
[0086] Step S2: Obtain the specification document of the target data standard; perform target standard structured parsing on the specification document of the target data standard to obtain a target standard structure model; construct a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph; perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; extract bridging concepts according to the initial anchoring candidate set and perform dynamic bridging ontology construction to obtain a refined bridging ontology;
[0087] In the embodiments of the present invention, a StructureDefinition document of the HL7 FHIR R4 version is loaded, and an XML parser is used to extract data element definitions to construct a target standard structure model. Then, the Sentence-BERT model is used to map the target standard data elements to medical ontology concepts, and a target standard semantic graph is constructed. Next, the GraphSAGE model is used to learn the node embeddings of the two graphs, and preliminary concept anchoring is performed by calculating the cosine similarity. Subsequently, the Gale-Shapley algorithm is used to optimize the anchoring results, and high-confidence concept mappings are obtained by comparing the semantic similarities of neighbor nodes. Then, bidirectional breadth-first search is performed to extract potential bridging paths connecting the source and target concepts. Finally, the potential bridging concepts are instantiated into ontology classes, and an ontology reasoner is used for consistency checking and redundant merging to obtain a refined bridging ontology.
[0088] Step S3: According to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph, refine the bridging paths to obtain a refined bridging path set; model the transformation operation on the refined bridging path set to obtain structured transformation logic; perform rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set;
[0089] In the embodiments of the present invention, using the refined bridging ontology, the rich semantic graph, and the target standard semantic graph, an improved breadth-first search algorithm is used to refine the bridging paths. Then, according to the semantic relationships on the paths and predefined mapping rules, transformation operation identifiers are extracted, and the data type, structure, and constraint differences between the source and target data are analyzed. Next, relevant business rules are retrieved and applied, and the path elements are mapped to Lambda expressions. Finally, the Lambda expressions are assembled according to the operation order to construct structured transformation logic. For the structured transformation logic, matching patterns are found in the mapping rule pattern library and parameters are filled to generate candidate mapping rules. Sample source data is extracted and the expected target data is set by domain experts to evaluate the accuracy, recall rate, and F1 value of the candidate mapping rules, and the rule parameters are automatically repaired or adjusted according to the evaluation results to finally obtain a dynamic mapping rule set.
[0090] Step S4: Apply the mapping rules to the source system data to be transformed according to the dynamic mapping rule set to obtain transformed data segments; assemble the target structures for the transformed data segments to obtain preliminary target instances; perform data quality verification on the preliminary target instances to obtain target standard data instances;
[0091] In an embodiment of the present invention, patient data is extracted from a laboratory system through a JDBC connector and encapsulated into data units to be converted. Then, the rule matching engine looks up and executes applicable rules in the dynamic mapping rule set according to the element names and semantic tags of the data units. For example, the height is converted from centimeters to meters. Next, an FHIRPatient object is created according to the target standard structure model, and the converted data is filled into the corresponding attributes. Then, according to the constraints defined in the target standard semantic graph, data type, value range, and logical consistency checks are performed on the target instance, and data enhancement is carried out. Finally, the verified FHIR Patient object is serialized into a JSON string using the JsonParser of HAPI FHIR to generate a target standard data instance.
[0092] Step S5: Monitor the conversion results of the target standard data instance, perform root cause analysis of errors, and obtain an error cause tag set; perform bridging ontology evolution according to the error cause tag set to obtain a refined mapping knowledge base to achieve cross-standard data exchange tasks;
[0093] In an embodiment of the present invention, the conversion status and time consumption of each target standard data instance are recorded, the conversion success rate, rule hit rate, and exception types are statistically analyzed, and a conversion performance report is generated. Collect user feedback on data errors, omissions, and format issues in the target system user interface, and associate the feedback information with the target data instance. Then, perform correlation analysis on the user feedback and the conversion performance report, trace the source data and mapping rules corresponding to the problem data, and mark the error causes, such as mapping rule errors or source data quality problems. According to the error causes, adjust the mapping rules automatically or manually, such as adding new code mapping relationships or adjusting unit conversion formulas, and evaluate the effectiveness of the new rules through backtesting. Finally, according to the error causes and optimized mapping rules, adjust or supplement the concept definitions and relationships in the bridging ontology, such as refining the concept of contact information, and perform version management and consistency checks to obtain a refined mapping knowledge base.
[0094] Preferably, step S1 includes the following steps:
[0095] Step S11: Perform structured data collection on the source system to obtain source system data to be converted;
[0096] Step S12: Perform preliminary data parsing on the source system data to be converted to obtain parsed data;
[0097] Step S13: Perform element-level semantic tag mapping on the parsed data to obtain a preliminary semantic mapping result;
[0098] Step S14: Perform relational context mining on the parsed data according to the preliminary semantic mapping result, and perform relational graph construction to obtain a preliminary relational graph;
[0099] Step S15: Perform semantic tag fusion based on the preliminary relationship graph and the preliminary semantic mapping result to obtain a rich semantic graph.
[0100] As an example of the present invention, referring to Figure 2 as shown, in this example, step S1 includes:
[0101] Step S11: Collect structured data from the source system to obtain the source system data to be converted;
[0102] In the embodiment of the present invention, by configuring a data connection component, specifying the JDBC connection string, username, and password of a relational database management system (for example, MySQL 8.0), a connection with the electronic medical record database of the source system is established. Execute a pre-written SQL query statement, such as "SELECT patient_id,patient_name,gender,birthday FROM patients WHERE update_time>'2023-01-01'", to extract structured data such as patient identifiers, patient names, genders, and birth dates within a specified time range from the patient information table. The collected data is stored in a data frame structure in memory in tabular form, with each row of the data frame representing a patient record and each column corresponding to a data field.
[0103] Step S12: Perform preliminary data parsing on the source system data to be converted to obtain the parsed data;
[0104] In the embodiment of the present invention, for the data frame obtained in step S11, each row of patient records is traversed. Using a predefined regular expression pattern, such as the "YYYY-MM-DD" pattern for the date field, the format of the string-type birth date field is verified. If the date format conforms to the predefined pattern, it is converted into a date object within the program; if the format does not conform, the identifier of the record and its date format error information are recorded in the error log. For the patient name and gender fields, their string values are directly extracted. The parsed data is stored in dictionary form, with the keys of the dictionary being the field names and the values being the parsed data values or the original string values.
[0105] Step S13: Perform element-level semantic tag mapping on the parsed data to obtain a preliminary semantic mapping result;
[0106] In the embodiments of the present invention, for each parsed patient record in step S12, its field name is analyzed one by one. Using a pre-trained word vector model (for example, the Word2Vec model trained based on the Wikipedia corpus), the semantic similarity between the field name "patient_name" and the concept names in a predefined medical ontology (for example, SNOMED CT) is calculated. If the similarity with the concept of "patient name" exceeds the set threshold of 0.8, then this field is mapped to the concept of "patient name". Using a rule-based mapping method, if the field name is "gender" and its value is "male" or "female", they are respectively mapped to the enumerated values "male" or "female" of the "Patient.gender" attribute in the BRIDG model. The mapping results are stored in the form of a list, and each element of the list contains the field name, the original value, and the ontology concept or BRIDG model attribute mapped to.
[0107] Step S14: Perform relational context mining on the parsed data according to the preliminary semantic mapping results, and construct a relational graph to obtain a preliminary relational graph;
[0108] In the embodiments of the present invention, based on the preliminary semantic mapping results of step S13 and the structural information of the parsed data in step S12, a relational graph of patient information is constructed. Each data element in the graph (for example, patient identifier, patient name, gender, date of birth) is used as a node. If there is a direct structural relationship between two pieces of data (for example, belonging to the same patient record), then an undirected edge is created between the corresponding nodes. Using an association rule mining algorithm (for example, the Apriori algorithm), analyze the frequently occurring field value combination patterns in the historical patient data. If it is found that "patient name" and "date of birth" often appear simultaneously, then an edge of "has personal attribute" is added between their corresponding nodes and a higher weight is assigned. The preliminary relational graph is stored in an adjacency list structure, and each node maintains a list of the nodes directly connected to it and their edge attributes.
[0109] Step S15: Perform semantic label fusion according to the preliminary relational graph and the preliminary semantic mapping results to obtain a rich semantic graph.
[0110] In the embodiment of the present invention, for each node in the preliminary relationship graph constructed in step S14, if there are multiple candidate ontology concept mappings in the preliminary semantic mapping result of step S13, disambiguation is performed according to the semantic labels of its adjacent nodes. For example, if the "gender" field may be mapped to the concept of "social gender" in addition to the concept of "gender", the semantic label of its adjacent node "date of birth" is analyzed. Since the date of birth belongs to biological characteristics, "gender" is selected as the final semantic label. The finally determined ontology concept or BRIDG model attribute is added to the relationship graph as the attribute of the node. Each node of the rich semantic graph includes a data value, an original field name, a final semantic label, and connection relationships and relationship types with other nodes.
[0111] Preferably, step S2 includes the following steps:
[0112] Step S21: Obtain the specification document of the target data standard; perform target standard structured parsing according to the specification document of the target data standard to obtain a target standard structure model;
[0113] Step S22: Construct a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph;
[0114] Step S23: Perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set;
[0115] Step S24: Perform stable matching optimization according to the initial anchoring candidate set and cross-graph node embedding data to obtain a preliminary concept anchoring set;
[0116] Step S25: Perform structured neighborhood alignment on the target standard semantic graph and the rich semantic graph according to the preliminary concept anchoring set to obtain high-confidence concept mappings;
[0117] Step S26: Explore bridging paths for the high-confidence concept mappings to obtain bridging path exploration data; extract potential bridging paths from the bridging path exploration data;
[0118] Step S27: Assemble a dynamic bridging ontology for the potential bridging paths to obtain a dynamic bridging ontology; optimize and refine the dynamic bridging ontology to obtain a refined bridging ontology.
[0119] In the embodiments of the present invention, a StructureDefinition document of the target data standard HL7 FHIR R4 version (for example, Patient.StructureDefinition.xml) is loaded from a pre-configured file storage path. The XML document is parsed using an XML parser (for example, Apache Xerces) to extract the data element definitions that make up the patient resource, including the element path (for example, Patient.name.given), data type (for example, string), cardinality (for example, 0..*), and value set binding information (for example, ValueSet:http: / / hl7.org / fhir / ValueSet / administrative-gender). The parsing result is stored in a tree structure, where each node of the tree represents a data element, and the attributes of the node include the element path, data type, cardinality, and value set reference. Traverse the target standard structure model tree generated in step S21. For each data element node, extract its element path and description information. Using a pre-trained Transformer model (for example, Sentence-BERT), encode the element path "Patient.name.given" and the description "Given names for use to identify the patient" into vector representations. Calculate the cosine similarity between this vector and the vector representation of the concept name in a predefined medical ontology (for example, SNOMED CT). If the similarity with the "first name" concept is the highest, then link this data element node to the "first name" concept. The target standard semantic graph is stored using an attribute graph database (for example, Neo4j), where each data element is a node, and the attributes of the node include the element path, data type, and the linked ontology concept. The parent-child relationship between the elements defined in the structure model is represented as a directed edge of the "contains" type in the semantic graph. For all nodes in the target standard semantic graph constructed in step S22 and the rich semantic graph constructed in step S15, respectively extract the graph embedding vectors learned through a Transformer model (for example, GraphSAGE). Calculate the cosine similarity between the embedding vectors of each node in the rich semantic graph and all nodes in the target standard semantic graph. If the cosine similarity between the embedding vector of the node representing "patient name" in the rich semantic graph and the embedding vector of the node representing "Patient.name.text" in the target standard semantic graph is greater than the set threshold of 0.9, then add these two nodes as a preliminary concept anchor point to the anchor set. The preliminary concept anchor set is stored in a list form, and each element of the list contains the identifiers of the corresponding nodes from the two graphs.Based on the preliminary concept anchoring set generated in step S23 and the cross-graph node embedding vectors calculated in step S23, the Gale-Shapley algorithm is used for stable matching. For each node in the rich semantic graph, a preference list is constructed according to the similarity of its embedding vector to the nodes in the target standard semantic graph. Similarly, a preference list of the rich semantic graph nodes is constructed for each node in the target standard semantic graph. Run the Gale-Shapley algorithm to find a stable matching result through multiple rounds of proposal and acceptance processes. For example, if the "patient name" node most prefers the "Patient.name.text" node, and the "Patient.name.text" node also most prefers the "patient name" node, then this pair of anchorings is retained. The optimized preliminary concept anchoring set replaces the result of step S23. For each pair of anchor points in the preliminary concept anchoring set obtained in step S24, extract their first-order neighbor nodes in their respective graphs. Calculate the semantic similarity of the neighbor nodes of the two anchor points. For example, use the Jaccard similarity to calculate the proportion of shared neighbor nodes. If the neighbor nodes of the anchor point "patient name" contain "date of birth", and the neighbor nodes of the anchor point "Patient.name.text" also contain "Patient.birthDate", and "date of birth" and "Patient.birthDate" are also an anchor point, then increase the confidence of the pair of anchorings "patient name" and "Patient.name.text". Add the pairs of anchor points with confidence higher than the set threshold to the high-confidence concept mapping set. For each pair of mapping relationships in the high-confidence concept mapping obtained in step S25, such as "patient name" mapping to "Patient.name.text", perform bidirectional breadth-first search between the rich semantic graph and the target standard semantic graph. The search starts from the two mapped nodes and searches for paths connecting the two graphs within the set maximum path length (for example, 3 hops). The intermediate nodes on the path form potential bridging concepts. For example, if the path "patient name" - "has attribute" - "name" - "maps to" - "Patient.name.text" is found, then "name" is the potential bridging concept. All the found potential bridging paths and the bridging concept information they contain are stored in the bridging path exploration data. Based on the potential bridging paths extracted in step S26, construct a dynamic bridging ontology for the current data exchange task. Instantiate each potential bridging concept as a class in the ontology, and instantiate the relationships between the source data elements, target data elements, and bridging concepts (such as "has attribute", "maps to") as attributes or relationships in the ontology. For example, create the class "name", and establish the relationships "patient name" "has attribute" "name", "name" "maps to" "Patient.name.text".Use an ontology reasoner (e.g., Apache Jena Reasoner) to perform consistency checking on the dynamic bridging ontology and identify redundant or equivalent concepts. Merge bridging concepts with high semantic similarity. For example, merge "patient name" and "patient's name" into one concept. Next, refine the dynamic bridging ontology. The refinement process includes: ① Concept merging: Use the ontology reasoner to identify and merge semantically similar or equivalent bridging concepts. For example, if the semantic similarity between "patient's name" and "patient name" exceeds a preset threshold (e.g., 0.95), merge them into a more general concept, such as "name". ② Relationship simplification: Remove redundant or inconsistent relationships. For example, if there are "patient name" "belongs to" "patient information", "patient's name" "belongs to" "patient information", and "patient name" and "patient's name" have been merged, remove the redundant "belongs to" relationship. ③ Attribute supplementation: Add necessary attributes to the bridging concepts, such as data type, value range, etc. For example, add the attribute "data type: string" to the "name" concept. ④ Consistency checking: Use the reasoner to check whether there are logical conflicts in the ontology. For example, check whether there are circular definitions or concept exclusions. ⑤ Ontology debugging and repair: According to the results of consistency checking and the feedback of domain experts, repair the errors or inconsistencies in the ontology. For example, modify incorrect class definitions or relationships. After the refinement through the above steps, the final refined bridging ontology is obtained. This ontology is stored in OWL format and contains more concise, consistent, and rich concepts, relationships, and attributes.
[0120] Preferably, step S23 includes the following steps:
[0121] Step S231: Enhance the graph node features of the target standard semantic graph and the rich semantic graph to obtain enhanced node features;
[0122] Step S232: Perform cross-graph Transformer embedding based on the enhanced node features to obtain cross-graph node embedding data;
[0123] Step S233: Construct cross-graph node pairs from the cross-graph node embedding data to obtain cross-graph node pairs;
[0124] Step S234: Calculate the similarity of the embedding vectors of the cross-graph node pairs to obtain the node pair similarity;
[0125] Step S235: Define a similarity threshold for the node pair similarity to obtain the similarity threshold;
[0126] Step S236: Screen the initial candidate anchor points based on the similarity threshold for the node pair similarity to obtain the initial anchor candidate set.
[0127] In an embodiment of the present invention, for each node in the target standard semantic graph constructed in step S22 and the rich semantic graph constructed in step S15, its text description information, including the node name and description, is extracted. Using the pre-trained BERT model, the text description "Given names foruse to identify the patient" of the "Patient.name.given" node in the target standard semantic graph and the text description "Patient's name" representing the "patient's name" node in the rich semantic graph are encoded into text embedding vectors with a dimension of 768. At the same time, using the GraphSAGE algorithm, the structural embedding vectors of each node in the two graphs are learned respectively, and the vector captures the neighbor information of the node in the graph structure. The text embedding vector and the structural embedding vector are concatenated to obtain an enhanced feature vector for each node, which contains both semantic information and structural information. The enhanced node feature vectors of the two graphs obtained in step S231 are used as input to train a cross-graph Transformer model. The Transformer model includes a multi-layer self-attention mechanism and a feedforward neural network for learning the association between cross-graph nodes. The training goal is to minimize the distance between the embedding vectors of aligned node pairs and maximize the distance between the embedding vectors of unaligned node pairs. After the training is completed, the enhanced feature vectors of each node in the target standard semantic graph and the rich semantic graph are mapped to a shared low-dimensional embedding space using the trained Transformer model to obtain the cross-graph node embedding vector of each node. For the cross-graph node embedding data obtained in step S232, each node in the rich semantic graph is paired with each node in the target standard semantic graph to generate all possible cross-graph node pairs. For example, the node representing "patient name" in the rich semantic graph is paired with all nodes representing "Patient.name.text", "Patient.gender" and so on in the target standard semantic graph. Each node pair contains the identifiers of the two nodes and their respective cross-graph node embedding vectors. The cross-graph node pairs are stored in a list form. For each cross-graph node pair constructed in step S233, the cross-graph node embedding vectors of the two nodes are obtained respectively. The cosine similarity between the two embedding vectors is calculated. The value range of cosine similarity is -1 to 1. The closer the value is to 1, the more similar the two vectors are in direction, that is, the more similar the two nodes are in semantics and structure. For example, calculate the cosine similarity between the embedding vector of the node representing "patient name" and the embedding vector of the node representing "Patient.name.text". The calculated similarity value is used as the similarity score of the node pair. The node pair similarity is stored in the form of a list, and each element of the list contains a node pair and a corresponding similarity score. Analyze the distribution of similarity scores of all node pairs calculated in step S234.Calculate the average value and standard deviation of the similarity scores. Using the statistical threshold method, set the similarity threshold to the average value plus one standard deviation. For example, if the average value of the similarity scores is 0.6 and the standard deviation is 0.2, then set the similarity threshold to 0.8. The similarity threshold is used to screen potential anchor points. Traverse the node pair similarity list calculated in step S234. For each node pair, compare its similarity score with the similarity threshold set in step S235. If the similarity score of the node pair is greater than or equal to the threshold, then determine this node pair as an initial candidate anchor point and add it to the initial candidate anchor set. For example, if the similarity score between the "patient name" node and the "Patient.name.text" node is 0.85, which is greater than the threshold 0.8, then add this pair of nodes to the initial candidate anchor set. The initial candidate anchor set is stored in the form of a list, and each element of the list contains the identifiers of a pair of nodes from the two graphs.
[0128] Preferably, step S3 includes the following steps:
[0129] Step S31: Refine the bridging paths according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a set of refined bridging paths;
[0130] Step S32: Model the transformation operations on the set of refined bridging paths to obtain structured transformation logic;
[0131] Step S33: Generate mapping rule patterns for the structured transformation logic according to the predefined mapping rule pattern library to obtain a set of candidate mapping rules;
[0132] Step S34: Extract sample source data from the source system data to be transformed and set the transformation expected target to obtain sample source data and expected target data;
[0133] Step S35: Verify and optimize the set of candidate mapping rules according to the sample source data and the expected target data to obtain a set of dynamic mapping rules.
[0134] As an example of the present invention, referring to Figure 3 as shown, in this example, step S3 includes:
[0135] Step S31: Refine the bridging paths according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a set of refined bridging paths;
[0136] In the embodiments of the present invention, for the refined bridging ontology generated in step S27, the rich semantic graph generated in step S15, and the target standard semantic graph generated in step S22, an improved breadth-first search algorithm is used to refine the bridging paths. For each pair of source data element nodes and target standard data element nodes anchored in the high-confidence concept mapping, starting from these two nodes, a two-way search is performed on the combined graph that integrates the information of the three graphs, and the search depth is limited to 3 hops. During the search process, paths passing through the bridging concepts and relationships defined in the refined bridging ontology are preferentially explored. For example, for the "patient name" node and the "Patient.name.text" node, if there is a path of "patient name" - "rdfs:subClassOf" - "patient's name" in the refined bridging ontology, and there is a path of "patient name" - "has attribute" - "name" in the rich semantic graph, then the path passing through the "patient's name" concept is preferentially explored. All the valid paths found are stored as refined bridging paths, and each path record contains an ordered list of nodes and edges.
[0137] Step S32: Model the transformation operations for the refined bridging path set to obtain structured transformation logic;
[0138] In the embodiments of the present invention, each refined bridging path obtained in step S31 is traversed, and the semantic relationships between adjacent nodes on the path are analyzed. Using the mapping rules from predefined semantic relationships to transformation operations, for example, the "rdfs:subClassOf" relationship is mapped to a "direct copy" operation, and the "hasUnit" relationship is mapped to a "unit conversion" operation. If there is a "hasCode" relationship connecting the code value and the code system on the path, and the corresponding element in the target standard is also bound to the code system, it is mapped to a "code value mapping" operation. For the "unit conversion" operation, the predefined unit conversion rule library is further searched to determine the source unit and the target unit, and the corresponding conversion formula is extracted. The results of the transformation operation modeling are stored in a structured form, with each path corresponding to a sequence of transformation operations, and each operation includes an operation type and operation parameters.
[0139] Step S33: Generate candidate mapping rule sets according to the predefined mapping rule pattern library for the structured transformation logic;
[0140] In an embodiment of the present invention, for each structured transformation logic obtained in step S32, a matching pattern is searched for in a predefined mapping rule pattern library. The mapping rule pattern library stores common transformation patterns, such as direct mapping pattern, type conversion pattern, unit conversion pattern, code value mapping pattern, etc. Each pattern defines input, output, and transformation logic. For example, if the structured transformation logic is "direct copy", the direct mapping pattern is matched; if it is "unit conversion, source unit: cm, target unit: m, formula: value / 100", the unit conversion pattern is matched. If a matching pattern is found, the pattern parameters are filled according to the current structured transformation logic to generate a specific mapping rule. For example, a rule "Map the value of the source data element 'patient height' from centimeters to meters and map it to the target data element 'Patient.height'" is generated. The generated mapping rule is stored in XML format and contains information such as source data elements, target data elements, transformation patterns, and pattern parameters.
[0141] Step S34: Extract sample source data from the source system data to be transformed, and set the expected target for transformation to obtain sample source data and expected target data;
[0142] In an embodiment of the present invention, 100 patient records are randomly selected from the source system data to be transformed collected in step S11 as sample source data. For each sample source data, the domain expert manually determines the expected target data after its transformation according to the target data standard specification. For example, for a patient record with a height of "170 cm", the value of the "Patient.height" field in its expected target data should be "1.7 m". The sample source data and the expected target data are stored in the form of a CSV file, including the original value of the source data and the expected value of the target data, and are used for subsequent rule verification and optimization.
[0143] Step S35: Verify and optimize the candidate mapping rule set according to the sample source data and the expected target data to obtain a dynamic mapping rule set;
[0144] In the embodiments of the present invention, for each rule in the candidate mapping rule set generated in step S33, the sample source data extracted in step S34 is used for testing. Each piece of sample source data is substituted into the mapping rule for conversion, and the conversion result is compared with the corresponding expected target data. Evaluation metrics such as the accuracy rate, recall rate, and F1 value of each rule are calculated. If the accuracy rate of a certain rule on the sample data is lower than the set threshold of 0.95, it is considered that the rule needs to be optimized. For the incorrect rules, analyze their conversion logic and try to perform automatic repair according to the error type. For example, for the rules with unit conversion errors, try to adjust the coefficients in the conversion formula. Replace the original rules with the optimized rules. Finally, a dynamic mapping rule set is obtained, which contains the mapping rules that have been verified and optimized.
[0145] Preferably, step S31 includes the following steps:
[0146] Step S311: Identify the target elements to be mapped in the target standard semantic graph to obtain the target set to be searched;
[0147] Step S312: Traverse the target elements in the target set to be searched in a loop, and use the refined bridging ontology to perform a preliminary path exploration on the rich semantic graph to obtain the preliminary path discovery result;
[0148] Step S313: Quantify and evaluate the path features of the preliminary path discovery result to obtain a path set with evaluation scores;
[0149] Step S314: Screen and sort and optimize the path set with evaluation scores to obtain a candidate bridging path set;
[0150] Step S315: Extract path elements from the candidate bridging path set to obtain a refined bridging path set.
[0151] In the embodiments of the present invention, all nodes in the target standard semantic graph constructed in traversal step S22 are traversed. Nodes with a data type of basic data type (e.g., string, integer, boolean) and a non-zero cardinality in the target standard specification are identified. These nodes represent target data elements that need to be mapped from the source data. For example, nodes such as "Patient.name.given", "Patient.gender", "Patient.birthDate" are identified. Structural nodes or nodes that are only containers are excluded. The identifiers of the identified target data element nodes are stored in the to-be-searched target set. For each target data element in the to-be-searched target set obtained in step S311, such as "Patient.name.given", in the rich semantic graph constructed in step S15, starting from all source data element nodes, path exploration is performed using the bridging concepts and relationships defined in the refined bridging ontology generated in step S27. The path exploration adopts a breadth-first search algorithm with a limited depth, and the search depth is limited to 3 hops. The search target is to start from the source data element node, pass through the bridging concepts in the bridging ontology, and finally reach the node semantically related to the current target data element. For example, if there is a mapping from "patient name" to "patient's name" in the refined bridging ontology, then explore the path from the "patient name" node in the rich semantic graph, pass through the "patient's name" concept, and finally reach the "Patient.name.given" node in the target standard semantic graph. All explored paths, including the sequence of nodes and edges on the path, are stored in the preliminary path discovery result. For each path found in step S312, calculate its path length, that is, the number of edges included on the path. Evaluate the confidence of the bridging concepts on the path. The confidence is based on the frequency of the bridging concepts being referenced in the refined bridging ontology and the quality score of manual annotation. Evaluate the clarity of the relationships on the path. For example, the clarity of the "rdfs:subClassOf" relationship is higher than the "association" relationship connected through multiple intermediate concepts. According to the predefined scoring function, comprehensively consider factors such as path length, bridging concept confidence, and relationship clarity, and calculate the score of each path. For example, the path length is given a negative weight, and the bridging concept confidence and relationship clarity are given positive weights. Store the path and its corresponding score in the path set with evaluation scores. For the path set with evaluation scores obtained in step S313, for paths connecting the same source data element and target data element, remove duplicate paths and retain the path with the highest score. Set a scoring threshold, such as 0.8, and only retain paths with a score higher than the threshold. For each target data element, if there are multiple paths with a score higher than the threshold, then sort them from high to low according to the score, and select the two paths with the highest scores as candidate bridging paths.For example, if there are multiple paths mapping "patient name" to "Patient.name.given", then retain the top two with the highest scores. Store the filtered and sorted candidate bridging paths in the candidate bridging path set. For each candidate bridging path obtained in step S314, extract all the nodes and edges on the path and arrange them in the path order to form the node sequence and edge sequence of the path. Clearly identify the bridging concept nodes on the path. For example, for the path "patient name" - "rdfs:subClassOf" - "patient name" - "mapped to" - "Patient.name.text", the extracted node sequence is ["patient name", "patient name", "Patient.name.text"], the edge sequence is ["rdfs:subClassOf", "mapped to"], and the bridging concept is identified as "patient name". Store the extracted path elements in the refined bridging path set in a structured form, and each path record includes its node sequence, edge sequence, and bridging concept information.
[0152] Preferably, step S32 includes the following steps:
[0153] Step S321: Extract the relationship semantic mapping rules from the refined bridging path set to obtain the path set with operation identifiers;
[0154] Step S322: Analyze the data difference characteristics of the path set with operation identifiers to obtain the path set with difference characteristics;
[0155] Step S323: Associate and enhance the business rules according to the path set with difference characteristics to obtain the path set to be formalized;
[0156] Step S324: Map the language elements of the path set to be formalized to obtain the preliminary formalized path set;
[0157] Step S325: Construct the conversion logic expression according to the preliminary formalized path set to obtain the structured conversion logic.
[0158] In the embodiments of the present invention, each path in the refined bridging path set obtained by traversing step S31 is traversed to identify the semantic relationship type between two adjacent nodes on the path. Query the predefined relationship-to-operation mapping rule library, which stores rules such as "rdfs:subClassOf" mapped to "direct copy", "hasUnit" mapped to "unit conversion", "mapsToCode" mapped to "code value mapping", etc. If there is a "hasUnit" relationship on the path connecting the "patient height" and "centimeter" concepts, the "unit conversion" operation identifier is extracted. For complex relationships, such as the relationship connecting the components of an address, it is decomposed into multiple basic operations. Each path is appended with the extracted operation identifier and stored in the path set with operation identifiers. For each path with operation identifiers in step S321, compare the data types of the source data element and the target data element it connects. For example, if the data type of the source data element "patient date of birth" is a string and the data type of the target data element "Patient.birthDate" is a date, mark the data type difference. Compare the structures of the source data and the target data. For example, if the source data is a single field containing the full name and the target data requires two independent fields for the first name and the last name, mark the structural difference as "field splitting". Analyze the constraint condition differences. For example, if the target data element has an enumeration value constraint while the source data element does not, mark the constraint difference. Each path is appended with the analyzed difference features and stored in the path set with difference features. For each path with difference features in step S322, retrieve the predefined business rule library. If the source data element involved in the path is "patient age" and the target data element is "drug dosage", retrieve the dosage calculation rules related to age. If there is a rule stating that "for patients under 12 years old, the dosage is halved", associate this business rule with the path. For the code value mapping operation, if the business rule specifies a specific code mapping relationship, add this mapping relationship to the conversion operation of the path. Each path is fused with the associated business rule information and stored in the path set to be formalized. For each path to be formalized in step S323, select a formal language based on Lambda expressions for mapping. Map the source data element to a variable in the Lambda expression. For example, "patient height" is mapped to the variable source.patientHeight. Map the target data element to the target variable. For example, "Patient.height" is mapped to target.Patient.height. Map the bridging concept to an intermediate variable. Map the relationship operation to a function or operator in the Lambda expression. For example, "unit conversion" is mapped to UnitConverter.convert(source.patientHeight,"cm","m").The elements of each path are mapped to preliminary formal language expressions and stored in the preliminary formal path set. For each preliminarily formalized path in step S324, according to the operation order on the path, the mapped Lambda expression elements are combined into a complete transformation logic expression. For example, if the path is "patient height" -> "unit conversion (cm -> m)" -> "Patient.height", then the expression target.Patient.height = UnitConverter.convert(source.patientHeight, "cm", "m") is constructed. For paths containing business rules, the conditional judgment logic of the business rules is embedded in the expression. For example, if (source.patientAge < 12) target.Medication.dose = calculateDose(source.patientWeight) / 2 else target.Medication.dose = calculateDose(source.patientWeight). The constructed transformation logic expression is stored in string form and associated with the corresponding bridging path and stored in the structured transformation logic.
[0159] Preferably, step S4 includes the following steps:
[0160] Step S41: Extract source data instances from the source system data to be transformed to obtain data units to be transformed;
[0161] Step S42: Perform semantic rule matching and execution on the data units to be transformed according to the dynamic mapping rule set to obtain transformed data segments;
[0162] Step S43: Assemble the target structure for the transformed data segments according to the target standard structure model to obtain preliminary target instances;
[0163] Step S44: Perform semantic constraint verification and enhancement on the preliminary target instances according to the target standard semantic graph to obtain verified target instances;
[0164] Step S45: Serialize the format of the verified target instances to obtain target standard data instances.
[0165] In an embodiment of the present invention, the pre-configured data source connection information is read, and a connection to the electronic medical record database of the source system is established. The SQL query statement "SELECT patient_id, patient_name, gender, birthday, body_height FROM patients WHERE data_source = 'LIS'" is executed to extract the patient identifier, patient name, gender, date of birth, and height information of patients whose data source is the laboratory system. Each row of data in the query result is encapsulated into a data unit, where the attribute names of the data unit correspond to the field names of the query result, and the attribute values are the field values of the query result. For example, a data unit may contain attributes patient_id: "123", patient_name: "Zhang San", gender: "male", birthday: "1990-01-01", body_height: "175". All the extracted data units are stored in a list as data units to be converted. For each data unit to be converted obtained in step S41, such as the unit containing the information of patient "Zhang San", the dynamic mapping rule set generated in step S35 is traversed. The rule matching engine searches for applicable mapping rules based on the data element names and semantic tags in the data unit. For example, the rule "Multiply the value of the source data element 'body_height' by 0.01 and map it to the target data element 'Patient.height'" is found. The conversion logic defined in the matched rule is executed, that is, multiply the height attribute value "175" by 0.01 to obtain the converted height value "1.75". The converted data value is stored as a converted data fragment, for example, stored in the form of a dictionary {"Patient.height": "1.75"}. For the converted data fragment obtained in step S42 and the target standard structure model generated in step S21, a data instance that conforms to the target standard is constructed. According to the target standard structure model, an FHIR Patient object representing the patient resource is created. The data element values in the converted data fragment are traversed and filled into the corresponding attributes of the Patient object. For example, the converted height value "1.75" is filled into the height attribute of the Patient object. If there is a one-to-many mapping relationship, the corresponding sub-resources or extension elements are created. For fields that are missing in the source data but are required in the target standard, default values are filled or marked as missing. The assembled preliminary target instance is an FHIR Patient object that conforms to the target standard structure. The preliminary target instance generated in step S43 is verified according to the semantic constraints defined in the target standard semantic graph constructed in step S22. The data type constraints are verified. For example, the value of the Patient.birthDate attribute must be of the date type.Verify value range constraints. For example, the value of the Patient.gender attribute must be within a predefined set of gender code. Use an ontology reasoner to check logical consistency. For example, if Patient.gender is "female", then the value of Patient.pregnancyStatus should not be "true". If the verification fails, record the fields and reasons for the verification failure. If permitted by the target standard, perform data enhancement based on ontology information. For example, if only the disease name is provided, complete the disease code according to the ontology. The verified target instance after verification and enhancement is the verified target instance. For the verified target instance obtained in step S44, serialize it into a JSON-formatted string according to the specifications of the HL7 FHIR R4 version of the target data standard. Select a suitable FHIR serialization tool (e.g., JsonParser of HAPI FHIR). Configure the parameters of the serializer, such as setting the format of dates and times, whether to include null attributes, etc. Call the encoding method of the serializer to convert the verified FHIR Patient object into a JSON string that complies with the FHIR specification. For example, serialize a Patient object containing the patient's name, gender, and height into a JSON string. The serialized JSON string is used as the target standard data instance and can be used for further processing in the target system.
[0166] Preferably, step S5 includes the following steps:
[0167] Step S51: Conduct a conversion performance audit on the target standard data instance to obtain a conversion performance report;
[0168] Step S52: Collect user problem feedback on the target standard data to obtain a user feedback list;
[0169] Step S53: Trace the errors in the dynamic mapping rule set based on the conversion performance report and the user feedback list to obtain an error cause marking set;
[0170] Step S54: Perform intelligent optimization of the mapping rules in the dynamic mapping rule set based on the error cause marking set to obtain an optimized mapping rule set;
[0171] Step S55: Iteratively update the bridging knowledge base according to the error cause marking set, the optimized mapping rule set, and the refined bridging ontology to obtain a refined mapping knowledge base.
[0172] In the embodiments of the present invention, for each target standard data instance generated in step S45, its conversion status (success or failure) and conversion time consumption are recorded. The number of successfully converted data instances and the number of failures are statistically counted within the past week, and the conversion success rate is calculated. For the data instances that failed to be converted, their failure reasons are recorded, such as data type verification failure, value range constraint verification failure, etc. The number of times each mapping rule is successfully applied and the number of times of matching failure are statistically counted. The exception logs generated during the conversion process are analyzed, such as recording data type conversion exceptions, null pointer exceptions, etc. The above statistical information is summarized to generate a conversion performance report, which includes the conversion success rate, failure reason distribution, rule hit rate, exception type statistics, etc., and is presented in HTML format. A feedback module is deployed on the user interface of the target system, allowing users to submit feedback for specific target standard data instances. The user selects the problem type (such as data error, data missing, format problem) through the feedback module, and fills in the problem description and expected value. The feedback information submitted by the user, such as "The gender of patient Zhang San is displayed as unknown, and the expected value is male", is automatically associated with the unique identifier of the corresponding target standard data instance. All the collected user feedback information, including the feedback time, user identifier, problem type, description, expected value, and the associated data instance identifier, is stored in the user feedback list and sorted according to the feedback time. For each piece of user feedback collected in step S52, such as "The gender of patient Zhang San is displayed as unknown", the conversion record related to this patient in the conversion performance report generated in step S51 is searched. The mapping rules hit during the conversion process of this record are analyzed. For example, if the rule that maps the source system gender code to the target system gender code is hit, then check whether the mapping relationship of this rule is complete or there are errors. If there is an exception log related to this data instance in the conversion performance report, such as code value mapping failure, then associate this exception information with the user feedback. According to the analysis results, the root cause of the error is marked, such as "Missing code value mapping in the mapping rule". All the analyzed errors and their root cause marks are stored in the error cause mark set. For the error causes marked in step S53, corresponding rule optimization strategies are adopted. If the error cause is "Missing code value mapping in the mapping rule", then according to the expected value provided by the user feedback, a new code mapping relationship is automatically added to the mapping rule library. If the error cause is "Unit conversion formula error", then the gradient descent algorithm is used to adjust the parameters in the unit conversion formula according to the historical conversion data and user feedback. After applying the new mapping rule or modifying the existing rule, backtesting is performed using the historical conversion data to evaluate the effectiveness of the new rule. All the optimized mapping rules replace the original dynamic mapping rule set to form an optimized mapping rule set.Analyze the error cause tags in step S53. If it is found that multiple data mapping errors are related to the unclear definition of a certain bridging concept. For example, the concept of "patient contact information" includes both phone numbers and email addresses, making it difficult to accurately distinguish mapping rules. Then refine the concept of "patient contact information" in the refined bridging ontology generated in step S27, splitting it into two concepts: "patient phone number" and "patient email address". According to the optimized mapping rules in step S54, update the relationships between concepts in the bridging ontology. For example, if a new rule that maps the "patient mobile phone number" in the source system to "Patient.telecom" in the target system is added, then add the corresponding mapping relationship in the bridging ontology. The updated mapping rules and the bridging ontology together constitute the refined mapping knowledge base.
[0173] Preferably, the present invention also provides a cross-standard data exchange system based on the BRIDG model for performing the cross-standard data exchange method based on the BRIDG model as described above. The cross-standard data exchange system based on the BRIDG model includes:
[0174] A semantic topology module for collecting data from the source system to obtain the source system data to be converted; performing semantic label mapping on the source system data to be converted to obtain a preliminary semantic mapping result; performing semantic label fusion based on the relational graph on the preliminary semantic mapping result to obtain a rich semantic graph;
[0175] A bridging ontology construction module for obtaining the specification document of the target data standard; performing target standard structured parsing on the specification document of the target data standard to obtain a target standard structure model; constructing a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph; performing cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; extracting bridging concepts according to the initial anchoring candidate set and performing dynamic bridging ontology construction to obtain a refined bridging ontology;
[0176] A mapping paradigm generation module for refining the bridging path according to the refined bridging ontology, the rich semantic graph, and the target standard semantic graph to obtain a refined bridging path set; performing transformation operation modeling on the refined bridging path set to obtain a structured transformation logic; performing rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set;
[0177] A semantic data shaping module for applying the mapping rules to the source system data to be converted according to the dynamic mapping rule set to obtain converted data segments; assembling the target structure for the converted data segments to obtain a preliminary target instance; performing data quality verification on the preliminary target instance to obtain a target standard data instance;
[0178] The conversion rule refinement module is used to monitor the conversion results of target standard data instances, perform root cause analysis of errors, and obtain an error cause tag set; bridge ontology evolution is performed according to the error cause tag set to obtain a refined mapping knowledge base, so as to implement cross-standard data exchange tasks.
[0179] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be embraced within the present invention.
[0180] The above description is only a specific implementation manner of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A cross-standard data exchange method based on the BRIDG model, characterized in that: The following steps are involved: Step S1: collect data from the source system to obtain source system data to be converted; perform semantic label mapping on the source system data to be converted to obtain preliminary semantic mapping results; perform semantic label fusion based on the relationship graph on the preliminary semantic mapping results to obtain a rich semantic graph; Step S2: Obtain the specification document of the target data standard; perform target standard structured analysis on the specification document of the target data standard to obtain the target standard structure model; construct the target standard semantic graph on the target standard structure model to obtain the target standard semantic graph; perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; perform bridge concept extraction based on the initial anchoring candidate set, and perform dynamic bridge ontology construction to obtain a refined bridge ontology; Step S3: Refine the bridge path according to the refined bridge ontology, the rich semantic graph and the target standard semantic graph to obtain a refined bridge path set; perform conversion operation modeling on the refined bridge path set to obtain a structured conversion logic; Perform rule pattern matching and generation on the structured transformation logic to obtain a dynamic mapping rule set; Step S4: applying mapping rules to the source system data to be converted according to the dynamic mapping rule set to obtain converted data segments; assembling the converted data segments into target structures to obtain preliminary target instances; Perform data quality verification on the preliminary target instance to obtain the target standard data instance; Step S5: Monitor the conversion results of the target standard data instance and perform error root cause analysis to obtain an error cause tag set; perform bridge ontology evolution based on the error cause tag set to obtain a refined mapping knowledge base to achieve cross-standard data exchange tasks.
2. The cross-standard data exchange method based on the BRIDG model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting structured data from the source system to obtain source system data to be converted; Step S12: Perform preliminary data analysis on the source system data to be converted to obtain analyzed data; Step S13: performing element-level semantic label mapping on the parsed data to obtain a preliminary semantic mapping result; Step S14: performing relational context mining on the parsed data according to the preliminary semantic mapping results, and constructing a relational graph to obtain a preliminary relational graph; Step S15: Perform semantic label fusion based on the preliminary relationship graph and the preliminary semantic mapping results to obtain a rich semantic graph.
3. The cross-standard data exchange method based on the BRIDG model according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Obtain a specification document of a target data standard; perform a structured analysis of the target standard according to the specification document of the target data standard to obtain a target standard structure model; Step S22: constructing a target standard semantic graph for the target standard structure model to obtain a target standard semantic graph; Step S23: performing cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; Step S24: Perform stable matching optimization based on the initial anchor candidate set and cross-graph node embedding data to obtain a preliminary concept anchor set; Step S25: performing structured neighborhood alignment on the target standard semantic graph and the rich semantic graph according to the preliminary concept anchor set to obtain a high-confidence concept mapping; Step S26: performing bridge path exploration on the high confidence concept mapping to obtain bridge path exploration data; performing potential bridge path extraction on the bridge path exploration data to obtain potential bridge paths; Step S27: assembling the potential bridging paths into a dynamic bridging ontology to obtain a dynamic bridging ontology; optimizing and refining the dynamic bridging ontology to obtain a refined bridging ontology.
4. The cross-standard data exchange method based on the BRIDG model according to claim 3, characterized in that: Step S23 includes the following steps: Step S231: enhancing graph node features of the target standard semantic graph and the rich semantic graph to obtain enhanced node features; Step S232: Perform cross-graph Transformer embedding according to the enhanced node features to obtain cross-graph node embedding data; Step S233: construct cross-graph node pairs for the cross-graph node embedding data to obtain cross-graph node pairs; Step S234: Calculate the embedding vector similarity of the cross-graph node pairs to obtain the node pair similarity; Step S235: Determine a similarity threshold for the node pair similarity to obtain a similarity threshold; Step S236: Screening initial candidate anchor points based on the similarity of the node pairs according to the similarity threshold to obtain an initial anchor candidate set.
5. The cross-standard data exchange method based on the BRIDG model according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Refining the bridge path according to the refined bridge ontology, the rich semantic graph and the target standard semantic graph to obtain a refined bridge path set; Step S32: Modeling the conversion operation of the refined bridge path set to obtain a structured conversion logic; Step S33: Generate a mapping rule pattern for the structured conversion logic according to a predefined mapping rule pattern library to obtain a candidate mapping rule set; Step S34: extracting sample source data from the source system data to be converted, and setting the expected conversion target to obtain sample source data and expected target data; Step S35: verifying and optimizing the candidate mapping rule set according to the sample source data and the expected target data to obtain a dynamic mapping rule set.
6. The cross-standard data exchange method based on the BRIDG model according to claim 5, characterized in that: Step S31 includes the following steps: Step S311: Identify the target elements to be mapped on the target standard semantic graph to obtain a target set to be searched; Step S312: loop through the target elements of the target set to be searched, and use the refined bridging ontology to perform preliminary path exploration on the rich semantic graph to obtain preliminary path discovery results; Step S313: quantifying and evaluating the path characteristics of the preliminary path discovery results to obtain a path set with evaluation scores; Step S314: Filter and sort the paths with evaluation scores to obtain a set of candidate bridge paths; Step S315: extracting path elements from the candidate bridging path set to obtain a refined bridging path set.
7. The cross-standard data exchange method based on the BRIDG model according to claim 5, characterized in that: Step S32 includes the following steps: Step S321: extracting relational semantic mapping rules from the refined bridge path set to obtain a path set with operation identifiers; Step S322: performing data difference feature analysis on the path set with operation identification to obtain a path set with difference features; Step S323: associate and enhance business rules according to the path set with difference characteristics to obtain a path set to be formalized; Step S324: performing formal language element mapping on the formalized path set to obtain a preliminary formalized path set; Step S325: construct a transformation logic expression according to the preliminary formalized path set to obtain a structured transformation logic.
8. The cross-standard data exchange method based on the BRIDG model according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: extracting source data instances from source system data to be converted to obtain data units to be converted; Step S42: matching and executing semantic rules on the data unit to be converted according to the dynamic mapping rule set to obtain a converted data segment; Step S43: assembling the converted data segments into target structures according to the target standard structure model to obtain a preliminary target instance; Step S44: performing semantic constraint verification and enhancement on the preliminary target instance according to the target standard semantic graph to obtain a verified target instance; Step S45: Serialize the verified target instance into a target format to obtain a target standard data instance.
9. The cross-standard data exchange method based on the BRIDG model according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: Perform a conversion performance audit on the target standard data instance to obtain a conversion performance report; Step S52: collecting user problem feedback on the target standard data to obtain a user feedback list; Step S53: Tracing the error source of the dynamic mapping rule set according to the conversion performance report and the user feedback list to obtain an error cause tag set; Step S54: intelligently optimizing the mapping rules of the dynamic mapping rule set according to the error cause mark set to obtain an optimized mapping rule set; Step S55: iteratively update the bridging knowledge base according to the error cause marking set, the optimized mapping rule set and the refined bridging ontology to obtain a refined mapping knowledge base.
10. A cross-standard data exchange system based on the BRIDG model, characterized in that: The cross-standard data exchange method based on the BRIDG model according to claim 1 is used to execute the cross-standard data exchange method based on the BRIDG model, and the cross-standard data exchange system based on the BRIDG model comprises: The semantic topology module is used to collect data from the source system to obtain the source system data to be converted; perform semantic label mapping on the source system data to be converted to obtain preliminary semantic mapping results; perform semantic label fusion based on the relationship graph on the preliminary semantic mapping results to obtain a rich semantic graph; The bridging ontology construction module is used to obtain the specification documents of the target data standard; perform target standard structured analysis on the specification documents of the target data standard to obtain the target standard structure model; construct the target standard semantic graph on the target standard structure model to obtain the target standard semantic graph; perform cross-graph concept anchoring on the target standard semantic graph and the rich semantic graph to obtain a preliminary concept anchoring set; perform bridging concept extraction based on the initial anchoring candidate set, and perform dynamic bridging ontology construction to obtain a refined bridging ontology; The mapping paradigm generation module is used to refine the bridge path according to the refined bridge ontology, the rich semantic graph and the target standard semantic graph to obtain a refined bridge path set; perform conversion operation modeling on the refined bridge path set to obtain a structured conversion logic; perform rule pattern matching and generation on the structured conversion logic to obtain a dynamic mapping rule set; The semantic data shaping module is used to apply mapping rules to the source system data to be converted according to the dynamic mapping rule set to obtain converted data fragments; assemble the converted data fragments into target structures to obtain preliminary target instances; perform data quality verification on the preliminary target instances to obtain target standard data instances; The conversion rule refining module is used to monitor the conversion results of the target standard data instance and perform error root cause analysis to obtain an error cause tag set; the bridge ontology is evolved based on the error cause tag set to obtain a refined mapping knowledge base to achieve cross-standard data exchange tasks.
Citation Information
Cited By
General data exchange system based on configurable label structure
CN121579581A
Automatic collection and intelligent verification system and method for text data
CN122263882A