Heterogeneous data fusion method and device, equipment and storage medium
By using a unified semantic middleware to perform structured parsing and mapping of heterogeneous data from medical terminal devices, a multi-dimensional spatiotemporal knowledge graph is constructed, solving the data integration problem caused by the diversity of communication protocols of medical terminal devices, and realizing standardized data management and effective support for upper-layer services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-14
AI Technical Summary
The diversity of medical terminal devices leads to a variety of communication protocols and inconsistent standards, making data integration difficult and hindering data management and use.
Data processing is performed using a unified semantic middleware. The original message data is uniformly structured and parsed through the target protocol parsing plugin. Key fields of the message are extracted and mapped to target standard semantic nodes to construct a multi-dimensional spatiotemporal knowledge graph for heterogeneous data fusion.
It enables standardized processing of heterogeneous data, simplifies data management and maintenance, and supports effective data access at the upper-level medical service layer.
Smart Images

Figure CN121858657A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data processing technology and is applied to the semantic fusion processing of data from multiple heterogeneous sources. It relates to a heterogeneous data fusion method, apparatus, device and storage medium. Background Technology
[0002] With the deep integration of cloud computing, edge computing, and smart terminal devices in the healthcare field, the "cloud-edge-device collaboration" architecture is gradually becoming a core technological support for scenarios such as smart hospitals, remote diagnosis and treatment, chronic disease management, and emergency rescue. This architecture achieves a balance between low latency response, high bandwidth utilization, and privacy and security by rationally distributing data processing tasks among the cloud computing layer (centralized high-performance computing), the edge computing layer (local near-source processing), and medical terminal devices (such as monitors, wearable sensors, and imaging equipment).
[0003] However, the diversity of medical terminal devices and the different suppliers have led to a variety of communication protocols with inconsistent standards, making data integration difficult and severely restricting the interoperability and scalability of cloud-edge-device systems. Due to the diversity of communication protocols, which may be international standard protocols, proprietary protocols of manufacturers, or custom serial / network protocols of hospitals, the output data varies greatly in terms of data format, encoding method, transmission mechanism, and semantic definition, which is not conducive to the management and use of health and medical testing data by the upper-level medical service layer. Summary of the Invention
[0004] The purpose of this application is to propose a heterogeneous data fusion method, apparatus, device, and storage medium to solve the technical problems of difficulty in data integration and data management and use in existing health and medical testing data.
[0005] Firstly, embodiments of this application provide a heterogeneous data fusion method, which employs the following technical solution: A heterogeneous data fusion method includes the following steps: Obtain the raw message data output by the data output interfaces of various heterogeneous protocols; The original message data is transmitted to a preset unified semantic middleware; The target protocol parsing plugin in the unified semantic middleware is used to perform unified structured parsing on the corresponding original message data, and key fields of the message are extracted based on the unified structured parsing results; Map all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields; The metadata information corresponding to the standardized semantic fields is annotated; All standardized semantic fields are summarized by timestamp; Construct a multidimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after aggregation and processing as graph nodes, and metadata information as the basis for edge relationships, to complete the fusion of heterogeneous data.
[0006] Secondly, this application also provides a heterogeneous data fusion device, which adopts the following technical solution: A heterogeneous data fusion device, comprising: The raw message data acquisition module is used to acquire the raw message data output by the data output interfaces of various heterogeneous protocols. The raw message data transmission module is used to transmit the raw message data to a preset unified semantic middleware; The message key field extraction module is used to perform unified structured parsing of the corresponding original message data using the target protocol parsing plugin in the unified semantic middleware, and extract message key fields based on the unified structured parsing results; The standardized semantic field acquisition module is used to map all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields. The metadata information annotation module is used to annotate the metadata information corresponding to the standardized semantic fields; The semantic field aggregation and processing module is used to aggregate all standardized semantic fields by timestamp. The heterogeneous data fusion module is used to construct a multidimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after aggregation and processing as graph nodes, and metadata information as the basis for edge relationships, thereby completing the heterogeneous data fusion.
[0007] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the heterogeneous data fusion method described above.
[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the heterogeneous data fusion method described above.
[0009] Compared with the prior art, the embodiments of this application have the following main advantages: The heterogeneous data fusion method described in this application acquires raw message data output from various heterogeneous protocol data output interfaces; transmits the raw message data to a pre-defined unified semantic middleware; utilizes the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; maps all message key fields to target standard semantic nodes to obtain standardized semantic fields corresponding to all message key fields; annotates the metadata information corresponding to the standardized semantic fields; summarizes all standardized semantic fields by timestamp; and constructs a multi-dimensional spatiotemporal knowledge graph with a pre-defined target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing the heterogeneous data fusion. By using the pre-defined unified semantic middleware to perform standardized semantic processing on the output data of various heterogeneous protocols, the heterogeneous data fusion is ultimately achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call the underlying data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of a heterogeneous data fusion method according to this application; Figure 3 This is a flowchart of a specific embodiment of the heterogeneous data fusion method described in this application, which integrates the protocol parsing plugin. Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 203 shown; Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown; Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown; Figure 7 yes Figure 2A flowchart of a specific embodiment of step 206 shown; Figure 8 yes Figure 7 A flowchart of a specific embodiment of step 703 shown; Figure 9 yes Figure 2 A flowchart of a specific embodiment of step 207 shown; Figure 10 This is a schematic diagram of one embodiment of a heterogeneous data fusion device according to this application; Figure 11 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0019] It should be noted that the heterogeneous data fusion method provided in this application embodiment is generally executed by a server, and correspondingly, a heterogeneous data fusion device is generally set in the server.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a heterogeneous data fusion method according to this application. The heterogeneous data fusion method includes the following steps: Step 201: Obtain the raw message data output by the data output interfaces of various heterogeneous protocols.
[0022] In this embodiment, the heterogeneous data fusion method can be applied in the field of health and medical technology. Specifically, in the scenario of medical data acquisition and integration, taking a routine physical examination as an example, it may involve multiple testing items such as blood tests, urine tests, breath tests, and electrocardiograms. Due to the large number of medical testing devices involved, and the complexity of the manufacturers and models of the medical testing devices used in different testing items, the data output interfaces of different medical testing devices use different data communication protocols, resulting in the output data being multi-heterogeneous protocol output data.
[0023] The data communication protocol used here may be an international standard protocol, a vendor-specific protocol, or a hospital-defined serial / network protocol. This diversity of protocols results in significant differences in output data format, encoding methods, transmission mechanisms, and semantic definitions.
[0024] Specifically, the acquisition of raw message data output by data output interfaces of various heterogeneous protocols is, for example, by connecting the data output interfaces of medical testing equipment used for different testing items, the output data of various heterogeneous protocols is acquired, i.e., the raw message data.
[0025] Of course, if the fintech business also involves data output interfaces using different data communication protocols, resulting in the output data being output data from multiple heterogeneous protocols, the heterogeneous data fusion method described above can also be used to complete the fusion of heterogeneous data, ensuring that upper-layer financial services can use the fused standardized financial data knowledge when using the data.
[0026] Step 202: Transmit the original message data to a preset unified semantic middleware.
[0027] In this embodiment, the problem that the diversity of protocols leads to huge differences in output data in terms of data format, encoding method, transmission mechanism, and semantic definition, which is not conducive to heterogeneous data fusion, is solved by constructing a unified semantic middleware.
[0028] Specifically, by designing a lightweight, embeddable unified semantic middleware that can be embedded in the edge computing layer, and setting it as an abstraction layer component between all medical testing devices and the upper medical service layer, the upper medical service layer will no longer obtain the output data of multiple heterogeneous protocols detected by all medical testing devices when acquiring test data. Instead, it will obtain the fused data after being integrated and organized by the unified semantic middleware in the abstraction layer. This facilitates unified management of health and medical testing data and better supports the upper medical service layer in using it for medical services.
[0029] The original message data is transmitted to a preset unified semantic middleware so that the preset unified semantic middleware can subsequently perform unified semantic processing on the original message data.
[0030] Step 203: Use the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing on the corresponding original message data, and extract key fields of the message based on the unified structured parsing results.
[0031] In this embodiment, the unified semantic middleware includes at least three processing parts: a raw message data parsing part, a unified semantic mapping processing part, and a data fusion implementation part. Specifically, in the raw message data parsing part, a protocol parsing plugin library is constructed by pre-integrating all protocol parsing plugins and added to the unified semantic middleware in a containerized manner to facilitate raw message data parsing during subsequent actual use. In the unified semantic mapping processing part, a standardized semantic mapping model is first generated based on the health and medical business scenario. Then, during actual processing, it is used to uniformly convert the data content indicators output by various medical testing devices according to a standardized format, avoiding the situation where one piece of data corresponds to multiple representations due to the complexity of medical testing devices, which is detrimental to management and maintenance. Finally, the data fusion implementation part performs data fusion processing on the output data after unified semantic mapping processing.
[0032] Specifically, by utilizing the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing of the corresponding original message data, and extracting key fields of the message based on the unified structured parsing results, the initial unified structured parsing of original message data of various heterogeneous protocols is achieved.
[0033] Step 204: Map all message key fields to the target standard semantic node to obtain the standardized semantic fields corresponding to all message key fields.
[0034] Specifically, in this embodiment, the unified semantic mapping processing part can be understood as a standardized processing model for medical testing services, pre-constructed based on medical equipment and testing indicators. This model encompasses semantic elements such as equipment category, measurement method, anatomical location, standard physiological parameter terminology (using LOINC encoding), and unit of measurement (UCUM). After receiving the target protocol parsing plugin in the unified semantic middleware and performing unified structured parsing on the corresponding original message data, the model maps different message key fields to a correspondence between standardized semantic fields and standard semantic nodes based on semantic elements. This facilitates subsequent fusion of heterogeneous data.
[0035] Step 205: Label the metadata information corresponding to the standardized semantic fields.
[0036] Specifically, the metadata information includes the source information of the original message data, such as the specific detection equipment and detection indicators, as well as the data accuracy level information, that is, the accuracy level of the detection results of this type of detection equipment in the industry, and the recommended use scenarios of the detection data, that is, the detection data of low accuracy or current detection methods can be used in certain specific medical assistance scenarios, while the detection data of relatively high accuracy or higher difficulty detection methods can be used in more accurate medical assistance scenarios.
[0037] By annotating the metadata information corresponding to the standardized semantic fields, the detection data can be selectively filtered and called in subsequent use.
[0038] Step 206: Summarize all standardized semantic fields by timestamp.
[0039] In this embodiment, since the results detected by medical testing equipment have certain time differences in the detection time sequence, and the data changes in this time difference often have a significant impact on medical auxiliary decision-making, all standardized semantic fields are first summarized and processed according to timestamps.
[0040] Step 207: Construct a multi-dimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after aggregation and processing as graph nodes, and metadata information as the basis for edge relationships, to complete the heterogeneous data fusion.
[0041] Ultimately, heterogeneous data fusion was achieved by constructing a multi-dimensional spatiotemporal knowledge graph. This facilitates subsequent heterogeneous data management and avoids excessive storage pressure caused by storing large amounts of heterogeneous data.
[0042] In this embodiment, raw message data output from various heterogeneous protocol data output interfaces is acquired; the raw message data is transmitted to a preset unified semantic middleware; the target protocol parsing plugin in the unified semantic middleware performs unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; all key message fields are mapped to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; metadata information corresponding to the standardized semantic fields is labeled; all standardized semantic fields are summarized by timestamp; a multi-dimensional spatiotemporal knowledge graph is constructed with a preset target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing heterogeneous data fusion. By using the preset unified semantic middleware to perform standardized semantic processing on the output data of various heterogeneous protocols, heterogeneous data fusion is finally achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call the underlying data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0043] Continue to refer to Figure 3 Before step 202, there is also a step of integrating the protocol parsing plugin. Figure 3 This is a flowchart of a specific embodiment of the heterogeneous data fusion method described in this application, which integrates the protocol parsing plugin, including: Step 301: Obtain the protocol parsing plugins corresponding to the data output interfaces of all target heterogeneous protocols; In this embodiment, the protocol parsing plugins used by all detection devices are first obtained; Step 302: Integrate the protocol parsing plugins to build a protocol adaptation plugin library; Specifically, the protocol parsing plugins used by all the aforementioned testing devices are integrated to build a protocol adaptation plugin library; Step 303: Containerize the protocol adapter plugin library and deploy the containerized protocol adapter plugin library into the unified semantic middleware in a plug-and-play edge deployment manner.
[0044] Specifically, using Docker container encapsulation, all protocol parsing plugins in the protocol adaptation plugin library are encapsulated into a processing component. Then, the containerized protocol adaptation plugin library is deployed to the unified semantic middleware in a plug-and-play edge deployment manner, which facilitates subsequent calls from the unified semantic middleware to the corresponding protocol parsing plugins to parse and process the original message data.
[0045] Continue to refer to Figure 4 , Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 203 shown includes: Step 401: Initially parse the original message data to obtain the data communication protocol corresponding to each of the original message data; Specifically, the original message data, for example, the HTTP protocol, originally contains the content "HTTP: / / ". Therefore, by initially parsing the original message data, the data communication protocol corresponding to each of the original message data can be obtained.
[0046] Step 402: Select the corresponding protocol parsing plugin from the containerized protocol adaptation plugin library according to the data communication protocol; In this embodiment, since the containerized protocol adaptation plugin library contains data parsing plugins corresponding to different data communication protocols, the corresponding protocol parsing plugin can be selected from the containerized protocol adaptation plugin library according to the data communication protocol.
[0047] Step 403: Use the selected corresponding protocol parsing plugin to perform JSON format parsing on all raw message data to obtain JSON format parsing results; Specifically, the parsed message data is standardized and organized into JSON format to facilitate the extraction of key fields later.
[0048] Step 404: Extract key fields of the message from the JSON formatted parsing result.
[0049] Specifically, since the JSON formatted parsing result has been obtained in step 403, the structured JSON parsing result can be directly extracted to obtain the key fields of the message.
[0050] In this embodiment, before executing the step of mapping all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields, the method further includes: constructing a domain description model of the target business domain using a web ontology language. The domain description model encompasses all standard semantic nodes included in the target business domain, as well as the semantic analysis basis and standardized semantic fields corresponding to all standard semantic nodes. Specifically, the web ontology language, namely OWL (Web Ontology Language), is a semantic web language designed to represent rich and complex knowledge about things, groups of things, and relationships between things. It describes the semantic relationships between different things in a specific domain. For example, if a domain description model for the health and medical business domain is constructed using a web ontology language, then this domain description model describes the semantic relationships between different things in the health and medical business domain. A standard semantic node can correspond to only one standardized semantic field, or a standard semantic node can correspond to multiple standardized semantic fields, or a standard semantic field can hit multiple standard semantic nodes. This is set according to the actual business processing scenario.
[0051] Continue to refer to Figure 5 , Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown includes: Step 501: Input all the key fields of the message into the domain description model; Specifically, since the domain description model is built based on the web ontology language, it covers all standard semantic nodes contained in the target business domain, as well as the semantic analysis basis and standardized semantic fields corresponding to all standard semantic nodes. Therefore, all the key fields of the message are input into the domain description model for semantic parsing processing.
[0052] Step 502: Perform semantic analysis on all message key fields according to the semantic analysis method corresponding to all standard semantic nodes, and identify the standard semantic nodes corresponding to each message key field. Specifically, when performing semantic parsing on all the key fields of the message, semantic analysis is performed on all the key fields of the message according to the semantic analysis criteria corresponding to all the standard semantic nodes, and the standard semantic nodes corresponding to each key field of the message are identified.
[0053] Step 503: Determine the standardized semantic fields corresponding to each of the key fields of the message by using the standard semantic nodes corresponding to each key field of the message.
[0054] Finally, by using the standard semantic nodes corresponding to each of the key fields in the message, the standardized semantic fields corresponding to each of the key fields in the message are determined.
[0055] Steps 501 to 503 above can be regarded as the second processing part in the unified semantic middleware, namely the processing steps of the unified semantic mapping processing part, thereby realizing the standardized semantic field conversion of all key fields in the original message data.
[0056] Continue to refer to Figure 6 , Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown includes: Step 601: Based on the source of the original message data, obtain the source information, data accuracy information, and specific application scenario information of all key fields of the message. Step 602: Based on the standardized semantic fields corresponding to all key fields of the message, mark the source information, data precision information, and specific business scenario information of all key fields of the message as metadata information to the corresponding standardized semantic fields.
[0057] In this embodiment, the source information, data accuracy information, and specific business scenario information of all key fields of the message are labeled as metadata information to the corresponding standardized semantic fields. This enables the source information of different medical testing equipment, testing indicators, etc., the testing accuracy level of the equipment, etc., and the data that can assist the application of the business scenario to be labeled as metadata information to the corresponding standardized semantic fields, so as to facilitate subsequent data management and use.
[0058] Continue to refer to Figure 7 , Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 206 shown includes: Step 701: Identify the generation time of the original message data corresponding to all standardized semantic fields; Step 702: Set the generation time to the timestamp of the corresponding standardized semantic field; Step 703: Organize all standardized semantic fields into a matrix based on the timestamp.
[0059] In this embodiment, all standardized semantic fields are summarized and organized. In essence, two aspects of summarization and organization are carried out. First, the standardized semantic fields corresponding to different key fields in the same original message data are summarized and processed. Second, the original message data output by different protocol transmission interfaces, i.e., different medical testing devices, are summarized and processed.
[0060] Continue to refer to Figure 8 , Figure 8 yes Figure 7 A flowchart of a specific embodiment of step 703 shown includes: Step 801: Using a cyclic filtering method, randomly select one standardized semantic field from all the standardized semantic fields as the current standardized semantic field; Step 802: Compare the timestamp of the current standardized semantic field with the timestamps of other standardized semantic fields in turn; Step 803: If the timestamps of the two are consistent, set the two standardized semantic fields to be compared as fields in the same column of the matrix; Specifically, there are two situations where the timestamps of the two are consistent. The first situation is that two standardized semantic fields corresponding to different key fields in the same original message data have the same generation time and timestamp because they are from the same original message data. The second situation is that two standardized semantic fields corresponding to different key fields in different original message data may also have the same timestamp because the two original message data coincidentally correspond to the same generation time.
[0061] Step 804: If the timestamps of the two are consistent, perform time serialization processing on the two standardized semantic fields to be compared according to the order of the timestamps, and generate the matrix row field according to the time serialization processing result; Specifically, for standardized semantic fields that have a chronological order, a matrix row field is generated according to the chronological order.
[0062] Step 805: Organize the columns of the matrix and the rows of the matrix to obtain a multidimensional field matrix containing all standardized semantic fields.
[0063] Specifically, the fields in the same column and the fields in the same row of the matrix are organized. Due to the different timestamps, a single original message may contain multiple key fields, thus forming a multi-dimensional matrix-like field array.
[0064] Continue to refer to Figure 9 , Figure 9 yes Figure 2 A flowchart of a specific embodiment of step 207 shown includes: Step 901: Obtain the multidimensional field matrix containing all standardized semantic fields; Step 902: Use all standardized semantic fields in the multidimensional field matrix as graph nodes; Step 903: Based on the metadata information corresponding to all standardized semantic fields, determine whether there is a target relationship between any two standardized semantic fields, wherein the target relationship includes business causal relationship, execution sequence relationship, and processing time sequence relationship; Step 904: If there is at least one target relationship between the two current standardized semantic fields, then the two current standardized semantic fields are connected in the multidimensional field matrix using the preset connection edge corresponding to the target relationship to generate a graph edge representation; Essentially, this involves constructing a graph edge representation from the multidimensional field matrix containing all standardized semantic fields, ultimately generating a multidimensional spatiotemporal knowledge graph composed of interconnected nodes, edge representations, and nodes. Specifically, for different target relationships, different colors or styles of line segments can be used to construct the graph edge relationship representation.
[0065] Step 905, until all standardized semantic fields have undergone target relationship judgment and graph edge representation generation, the multidimensional spatiotemporal knowledge graph is obtained; Step 906: Set the target object as the graph identifier of the multidimensional spatiotemporal knowledge graph, wherein the target object includes the entity object name.
[0066] Specifically, assuming that the multidimensional spatiotemporal knowledge graph is generated by heterogeneous data fusion of output data from multiple heterogeneous protocols in medical testing equipment, the business project name is Medical Heterogeneous Data Integration Project, and the business project name can be directly set as the graph identifier of the multidimensional spatiotemporal knowledge graph.
[0067] In this embodiment, raw message data output from various heterogeneous protocol data output interfaces is acquired; the raw message data is transmitted to a preset unified semantic middleware; the target protocol parsing plugin in the unified semantic middleware performs unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; all key message fields are mapped to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; metadata information corresponding to the standardized semantic fields is labeled; all standardized semantic fields are summarized by timestamp; a multi-dimensional spatiotemporal knowledge graph is constructed with a preset target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing heterogeneous data fusion. By using the preset unified semantic middleware to perform standardized semantic processing on the output data of various heterogeneous protocols, heterogeneous data fusion is finally achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call the underlying data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0068] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0069] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0070] In this embodiment, raw message data output from various heterogeneous protocol data output interfaces is acquired; the raw message data is transmitted to a preset unified semantic middleware; the target protocol parsing plugin in the unified semantic middleware performs unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; all key message fields are mapped to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; metadata information corresponding to the standardized semantic fields is labeled; all standardized semantic fields are summarized by timestamp; a multi-dimensional spatiotemporal knowledge graph is constructed with a preset target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing heterogeneous data fusion. By using the preset unified semantic middleware to perform standardized semantic processing on the output data of various heterogeneous protocols, heterogeneous data fusion is finally achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call the underlying data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0071] Further reference Figure 10 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a heterogeneous data fusion device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0072] like Figure 10 As shown, the heterogeneous data fusion device 10 described in this embodiment includes: a raw message data acquisition module 10a, a raw message data transmission module 10b, a message key field extraction module 10c, a standardized semantic field acquisition module 10d, a metadata information annotation module 10e, a semantic field summary processing module 10f, and a heterogeneous data fusion module 10g. Wherein: The raw message data acquisition module 10a is used to acquire raw message data output by data output interfaces of various heterogeneous protocols. The original message data transmission module 10b is used to transmit the original message data to a preset unified semantic middleware; The message key field extraction module 10c is used to perform unified structured parsing of the corresponding original message data using the target protocol parsing plugin in the unified semantic middleware, and extract message key fields based on the unified structured parsing results. The standardized semantic field acquisition module 10d is used to map all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields. Metadata information annotation module 10e is used to annotate the metadata information corresponding to the standardized semantic fields; The semantic field aggregation and processing module 10f is used to aggregate and process all standardized semantic fields by timestamp; The heterogeneous data fusion module 10g is used to construct a multidimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after summary processing as graph nodes, and metadata information as the basis for edge relationships, thereby completing the heterogeneous data fusion.
[0073] This application achieves heterogeneous data fusion by acquiring raw message data output from various heterogeneous protocol data output interfaces; transmitting the raw message data to a pre-defined unified semantic middleware; utilizing the target protocol parsing plugin within the unified semantic middleware to perform unified structured parsing on the corresponding raw message data, and extracting key message fields based on the unified structured parsing results; mapping all key message fields to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; annotating the metadata information corresponding to the standardized semantic fields; summarizing all standardized semantic fields by timestamp; and constructing a multi-dimensional spatiotemporal knowledge graph with a pre-defined target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis. By using the pre-defined unified semantic middleware to perform standardized semantic processing on the output data of various heterogeneous protocols, heterogeneous data fusion is ultimately achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call lower-layer data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0074] In this embodiment, the heterogeneous data fusion device 10 further includes a protocol parsing plugin acquisition module, a protocol parsing plugin integration module, and a containerization processing module. Wherein: The protocol parsing plugin acquisition module is used to acquire the protocol parsing plugins corresponding to the data output interfaces of all target heterogeneous protocols; The protocol parsing plugin integration module is used to integrate the protocol parsing plugins and build a protocol adaptation plugin library; The containerization module is used to containerize the protocol adapter plugin library and deploy the containerized protocol adapter plugin library into the unified semantic middleware in a plug-and-play edge deployment manner.
[0075] In this embodiment, the message key field extraction module 10c includes a message initial parsing unit, a protocol parsing plugin filtering unit, a JSON formatting parsing unit, and a message key field extraction unit. Wherein: The initial parsing unit is used to initially parse the original message data and obtain the data communication protocol corresponding to each of the original message data. The protocol parsing plugin filtering unit is used to filter out the corresponding protocol parsing plugins from the containerized protocol adaptation plugin library according to the data communication protocol. The JSON formatting and parsing unit is used to perform JSON formatting and parsing on all raw message data using the selected corresponding protocol parsing plugins, and obtain JSON formatted parsing results; The message key field extraction unit is used to extract message key fields from the JSON formatted parsing result.
[0076] In this embodiment, the heterogeneous data fusion device 10 further includes a domain description model construction module. The domain description model construction module is used to construct a domain description model of the target business domain through the web ontology language. The domain description model covers all standard semantic nodes contained in the target business domain, as well as the semantic analysis basis and standardized semantic fields corresponding to all standard semantic nodes.
[0077] In this embodiment, the standardized semantic field acquisition module 10d includes a message key field input unit, a standard semantic node identification unit, and a standardized semantic field determination unit. Wherein: The message key field input unit is used to input all the message key fields into the domain description model; The standard semantic node identification unit is used to perform semantic analysis on all message key fields according to the semantic analysis criteria corresponding to all standard semantic nodes, and identify the standard semantic nodes corresponding to each message key field. The standardized semantic field determination unit is used to determine the standardized semantic fields corresponding to each of the key fields of the message by using the standard semantic nodes corresponding to each key field of the message.
[0078] In this embodiment, the metadata information annotation module 10e includes a metadata information acquisition unit and a metadata information annotation unit. Wherein: The metadata information acquisition unit is used to acquire the source information, data accuracy information, and specific business scenario information of all key fields of the original message based on the source of the original message data. The metadata information annotation unit is used to annotate the source information, data precision information, and specific business scenario information of all message key fields as metadata information to the corresponding standardized semantic fields, based on the standardized semantic fields corresponding to all message key fields.
[0079] In this embodiment, the semantic field aggregation and processing module 10f includes a message generation time identification unit, a timestamp setting unit, and a matrix processing unit. Wherein: The message generation time identification unit is used to identify the generation time of the original message data corresponding to all standardized semantic fields; A timestamp setting unit is used to set the generation time as the timestamp of the corresponding standardized semantic field. The matrix organization unit is used to organize all standardized semantic fields into a matrix based on the timestamp.
[0080] In this embodiment, the matrix organization unit is specifically used to select one standardized semantic field from all the standardized semantic fields as the current standardized semantic field using a cyclic filtering method; it is also used to compare the timestamp of the current standardized semantic field with the timestamps of other standardized semantic fields in turn; if the timestamps of the two are the same, it sets the two standardized semantic fields being compared as fields in the same column of the matrix; if the timestamps of the two are the same, it performs time-series processing on the two standardized semantic fields being compared according to the order of the timestamps, and generates a field in the same row of the matrix based on the time-series processing result; it is also used to organize the fields in the same column of the matrix and the fields in the same row of the matrix to obtain a multi-dimensional field matrix containing all standardized semantic fields.
[0081] In this embodiment, the heterogeneous data fusion module 10g includes a multi-dimensional field matrix acquisition unit, a graph node setting unit, a graph edge relationship judgment unit, a graph edge representation generation unit, a graph acquisition unit, and a graph identifier setting unit. Wherein: A multidimensional field matrix acquisition unit is used to acquire the multidimensional field matrix containing all standardized semantic fields; The graph node setting unit is used to set all standardized semantic fields in the multidimensional field matrix as graph nodes. The graph edge relationship judgment unit is used to determine whether there is a target relationship between any two standardized semantic fields based on the metadata information corresponding to all standardized semantic fields. The target relationship includes business causal relationship, execution order relationship, and processing time sequence relationship. The graph edge representation generation unit is used to connect the two standardized semantic fields in the multidimensional field matrix by using the preset connection edge corresponding to the target relationship if there is at least one target relationship between the two standardized semantic fields, thereby generating a graph edge representation. The graph acquisition unit is used to obtain the multidimensional spatiotemporal knowledge graph until target relationship judgment and graph edge representation generation have been performed between all standardized semantic fields. The graph identifier setting unit is used to set the target object as the graph identifier of the multidimensional spatiotemporal knowledge graph, wherein the target object includes the entity object name.
[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0083] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0084] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed] for details. Figure 11 , Figure 11 This is a basic structural block diagram of the computer device in this embodiment.
[0085] The computer device 11 includes a memory 11a, a processor 11b, and a network interface 11c, which are interconnected via a system bus. It should be noted that... Figure 11Only a computer device 11 with component memory 11a, processor 11b, and network interface 11c is shown in the illustration. However, it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0086] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0087] The memory 11a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11a may be an internal storage unit of the computer device 11, such as the hard disk or memory of the computer device 11. In other embodiments, the memory 11a may also be an external storage device of the computer device 11, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 11a may include both internal storage units and external storage devices of the computer device 11. In this embodiment, the memory 11a is typically used to store the operating system and various application software installed on the computer device 11, such as computer-readable instructions for a heterogeneous data fusion method. In addition, the memory 11a can also be used to temporarily store various types of data that have been output or will be output.
[0088] In some embodiments, the processor 11b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 11b is typically used to control the overall operation of the computer device 11. In this embodiment, the processor 11b is used to execute computer-readable instructions stored in the memory 11a or to process data, for example, to execute computer-readable instructions of the heterogeneous data fusion method described above.
[0089] The network interface 11c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 11 and other electronic devices.
[0090] The computer device proposed in this embodiment belongs to the field of big data processing technology and is applied to the semantic fusion processing of data from multiple heterogeneous sources. This application acquires raw message data output from various heterogeneous protocol data output interfaces; transmits the raw message data to a preset unified semantic middleware; uses the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; maps all key message fields to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; annotates the metadata information corresponding to the standardized semantic fields; summarizes all standardized semantic fields by timestamp; and constructs a multi-dimensional spatiotemporal knowledge graph with a preset target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing the heterogeneous data fusion. By using the preset unified semantic middleware to perform standardized semantic processing on the output data of multiple heterogeneous protocols, the heterogeneous data fusion is finally completed, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call lower-layer data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0091] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the heterogeneous data fusion method described above.
[0092] The computer-readable storage medium proposed in this embodiment belongs to the field of big data processing technology and is applied to semantic fusion processing of data from multiple heterogeneous sources. This application acquires raw message data output from various heterogeneous protocol data output interfaces; transmits the raw message data to a preset unified semantic middleware; uses the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing on the corresponding raw message data, and extracts key message fields based on the unified structured parsing results; maps all key message fields to target standard semantic nodes to obtain standardized semantic fields corresponding to all key message fields; annotates the metadata information corresponding to the standardized semantic fields; summarizes all standardized semantic fields by timestamp; and constructs a multi-dimensional spatiotemporal knowledge graph with a preset target object as the graph identifier, all summarized standardized semantic fields as graph nodes, and metadata information as the edge relationship basis, thus completing heterogeneous data fusion. By using the preset unified semantic middleware to perform standardized semantic processing on the output data of multiple heterogeneous protocols, heterogeneous data fusion is ultimately achieved, which is beneficial for data management and maintenance, as well as for upper-layer business applications to call lower-layer data. This method can be applied in the field of health and medical technology, using the unified semantic middleware between different medical testing devices and the upper medical service layer, which facilitates unified management of health and medical testing data and can better support the upper medical service layer in using medical services.
[0093] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0094] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
Claims
1. A heterogeneous data fusion method, characterized in that, Includes the following steps: Obtain the raw message data output by the data output interfaces of various heterogeneous protocols; The original message data is transmitted to a preset unified semantic middleware; The target protocol parsing plugin in the unified semantic middleware is used to perform unified structured parsing on the corresponding original message data, and key fields of the message are extracted based on the unified structured parsing results; Map all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields; The metadata information corresponding to the standardized semantic fields is annotated; All standardized semantic fields are summarized by timestamp; Construct a multidimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after aggregation and processing as graph nodes, and metadata information as the basis for edge relationships, to complete the fusion of heterogeneous data.
2. The heterogeneous data fusion method according to claim 1, characterized in that, Before performing the step of transmitting the original message data to a preset unified semantic middleware, the method further includes: Obtain the protocol parsing plugins corresponding to the data output interfaces of all target heterogeneous protocols; The aforementioned protocol parsing plugins are integrated to build a protocol adaptation plugin library; The protocol adapter plugin library is containerized, and the containerized protocol adapter plugin library is deployed to the unified semantic middleware in a plug-and-play edge deployment manner.
3. The heterogeneous data fusion method according to claim 1 or 2, characterized in that, The step of using the target protocol parsing plugin in the unified semantic middleware to perform unified structured parsing of the corresponding original message data, and extracting key fields of the message based on the unified structured parsing results, specifically includes: The original message data is initially parsed to obtain the data communication protocol corresponding to each original message data. According to the data communication protocol, the corresponding protocol parsing plugins are selected from the containerized protocol adaptation plugin library; The selected protocol parsing plugins are used to perform JSON format parsing on all raw message data to obtain JSON formatted parsing results; Extract key fields from the parsed JSON format.
4. The heterogeneous data fusion method according to claim 1, characterized in that, Before performing the step of mapping all message key fields to the target standard semantic node to obtain the standardized semantic fields corresponding to all message key fields, the method further includes: A domain description model for the target business domain is constructed using a web ontology language. The domain description model includes all standard semantic nodes contained in the target business domain, as well as the semantic analysis basis and standardized semantic fields corresponding to all standard semantic nodes. The step of mapping all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields specifically includes: Input all the key fields of the message into the domain description model; Based on the semantic analysis criteria corresponding to all standard semantic nodes, semantic analysis is performed on all key fields of the message to identify the standard semantic nodes corresponding to each key field of the message. By identifying the standard semantic nodes corresponding to each of the key fields in the message, the standardized semantic fields corresponding to each of the key fields in the message are determined.
5. The heterogeneous data fusion method according to claim 1, characterized in that, The step of annotating the metadata information corresponding to the standardized semantic fields specifically includes: Based on the source of the original message data, obtain the source information, data accuracy information, and specific business scenario information of all key fields of the message; Based on the standardized semantic fields corresponding to all key fields of the message, the source information, data precision information, and specific business scenario information of all key fields of the message are labeled as metadata information to the corresponding standardized semantic fields.
6. The heterogeneous data fusion method according to claim 1, characterized in that, The step of summarizing all standardized semantic fields by timestamp specifically includes: Identify the generation time of the original message data corresponding to all standardized semantic fields; Set the generation time to the timestamp of the corresponding standardized semantic field; All standardized semantic fields are matrixed based on the timestamp.
7. The heterogeneous data fusion method according to claim 6, characterized in that, The step of matrixing all standardized semantic fields based on the timestamp specifically includes: A cyclic filtering method is used to randomly select one standardized semantic field from all the standardized semantic fields as the current standardized semantic field. The timestamp of the current standardized semantic field is compared with the timestamps of other standardized semantic fields in turn; If the timestamps of the two are the same, the two standardized semantic fields to be compared will be set as fields in the same column of the matrix; If the timestamps of the two are the same, the two standardized semantic fields to be compared are time-series processed according to the order of the timestamps, and the matrix row fields are generated according to the time-series processing results. Organize the fields in the same column and the fields in the same row of the matrix to obtain a multidimensional field matrix containing all standardized semantic fields.
8. A heterogeneous data fusion device, characterized in that, include: The raw message data acquisition module is used to acquire the raw message data output by the data output interfaces of various heterogeneous protocols. The raw message data transmission module is used to transmit the raw message data to a preset unified semantic middleware; The message key field extraction module is used to perform unified structured parsing of the corresponding original message data using the target protocol parsing plugin in the unified semantic middleware, and extract message key fields based on the unified structured parsing results; The standardized semantic field acquisition module is used to map all message key fields to target standard semantic nodes to obtain the standardized semantic fields corresponding to all message key fields. The metadata information annotation module is used to annotate the metadata information corresponding to the standardized semantic fields; The semantic field aggregation and processing module is used to aggregate all standardized semantic fields by timestamp. The heterogeneous data fusion module is used to construct a multidimensional spatiotemporal knowledge graph with preset target objects as graph identifiers, all standardized semantic fields after aggregation and processing as graph nodes, and metadata information as the basis for edge relationships, thereby completing the heterogeneous data fusion.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the heterogeneous data fusion method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the heterogeneous data fusion method as described in any one of claims 1 to 7.