Medical equipment communication protocol standardization method and system based on large language model

Through a method based on a large language model, automatic parsing and standardization of medical device communication protocols, the problems of complex device access processes and inconsistent data management in existing technologies are solved, and fast and accurate protocol parsing and unified management of device data are achieved.

CN120676056APending Publication Date: 2025-09-19RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510760837.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately parse and standardize medical device communication protocols, resulting in complex device access processes, long development cycles, high maintenance costs, and the inability to achieve unified management of device data.

Method used

This approach, based on a large language model, achieves automated understanding of medical device communication protocols through pre-training of the large language model and the design of prompt words. The specific steps include data collection, pre-processing, parsing data messages using the large language model to obtain structured semantic labels, and mapping these labels to the medical institution's unified device data model.

Benefits of technology

It achieves fast and accurate analysis and standardization of medical device communication protocols, significantly shortens the device access cycle, reduces the burden of manual development and maintenance, and supports convenient and efficient cross-device management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676056A_ABST
    Figure CN120676056A_ABST
Patent Text Reader

Abstract

The invention relates to the field of communication protocol analysis, and provides a medical equipment communication protocol standardization method and system based on a large language model, and the method comprises the following steps: S1, collecting the monitoring data of each medical equipment in real time, and transmitting the monitoring data to a data integration platform in the form of a data message through an original communication protocol of the medical equipment; s2, the data integration platform inputs the preprocessed data message into a large language model to obtain a structured semantic tag; S2.1, the large language model is trained; s2.2, guiding the large language model to analyze and output a structured semantic tag by designing a cue word; and S3, the data integration platform maps the structured semantic tag to a unified equipment data model of the medical structure, and automatically accesses an equipment management system of the medical structure. The invention aims to solve the problem that a medical equipment communication protocol cannot be quickly and accurately analyzed and standardized in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication protocol standardization, and in particular to a method and system for standardizing medical device communication protocols based on a large language model. Background Art

[0002] With the popularization of intelligent medical equipment, there are a large number of equipment of different brands and models in hospitals, such as ventilators, ECG monitors, infusion pumps, etc. In order to uniformly manage various medical equipment and multi-source medical data, these devices need to be connected to the hospital's equipment management system.

[0003] These medical devices use different communication protocols, including standard protocols such as HL7 and DICOM, as well as a large number of vendor-defined proprietary protocols. Compared to ordinary text protocols, medical device protocols are more structured (e.g., IEEE 11073, HL7, and DICOM standards specify segmentation and field formats, have high real-time requirements, strict naming specifications, and strict control of numerical precision). For example, HL7 messages use a clearly delimited paragraph structure (each paragraph is separated by a carriage return, and fields are separated by "|"); the DICOM imaging protocol uses a tag-element approach to organize data; and the IEEE 11073 series achieves consistent data representation through a glossary and object encoding.

[0004] Directly using existing large models to organize medical device communication protocols fails to meet the requirements for accurate understanding of protocol structure, domain terminology, and precise numerical values, nor can it automatically integrate into a hospital's internal device management system. Existing medical device adaptation methods primarily rely on manually developed parsers or hard-coded field mappings. Existing technologies suffer from the following issues: long development cycles: adding new devices requires detailed analysis of protocol content and manual adaptation logic; poor flexibility: protocol changes require redevelopment, preventing rapid adaptation; high maintenance costs: frequent updates require continuous manpower investment; and difficulty in standardization: inconsistent protocol descriptions across devices complicate device data integration.

[0005] Traditional manual adaptation methods can no longer meet the requirements of intelligent medical equipment management systems for high scalability, high real-time performance, and high standardization. Summary of the Invention

[0006] The technical problem to be solved by the present invention is that, in response to the above-mentioned defects of the prior art, a medical device communication protocol standardization method and system based on a large language model is intended to solve the problem in the prior art that medical device communication protocols cannot be quickly and accurately parsed and standardized.

[0007] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0008] In a first aspect, a method for standardizing a medical device communication protocol based on a large language model is provided, comprising the following steps:

[0009] S1, real-time collection of monitoring data from each medical device, and reporting the monitoring data in the form of data packets to the data integration platform through the original communication protocol of the medical device;

[0010] S2: The data integration platform inputs the pre-processed data message into the large language model to obtain structured semantic labels:

[0011] S2.1, training the large language model, including designing a supervised learning task based on inputting original communication protocol messages and outputting standardized field names and semantic labels; and adding domain terminology explanations for the original communication protocol fields to the training set;

[0012] S2.2, guiding the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol as context, domain terminology explanations of the data message fields, and the required output structure format;

[0013] S3, the data integration platform maps the structured semantic tags to the unified equipment data model of the medical institution and automatically connects to the equipment management system of the medical institution.

[0014] In one embodiment, the preprocessing includes character set conversion, escape character processing and segmentation, and time calibration to unify the timestamp formats of the various medical devices.

[0015] In one embodiment, in steps S2.1 and S2.2, the domain terminology interpretation of the original communication protocol fields and the domain terminology interpretation of the fields involved in the data message both adopt the same terminology standard, including LOINC and SNOMED CT.

[0016] In one embodiment, in step S2.1, contrastive learning technology is also used to generate positive samples using LOINC and SNOMED CT, and non-related terms are randomly sampled as negative samples, so that the large language model learns the relative distance of samples in the feature space, thereby improving the large language model's ability to distinguish medical terms with similar semantics but different expressions.

[0017] In one embodiment, in step S2.2, if the data message contains a numerical field, the output numerical format and precision need to be specified.

[0018] In one embodiment, in step S2.2, the large language model not only outputs structured semantic tags but also outputs a confidence score for each inference. When the confidence score is lower than a preset standard, manual review is triggered.

[0019] In one embodiment, step S2 also includes step S2.3. After the large language model outputs the structured semantic tag, the data integration platform also corrects the structured semantic tag, and the correction includes verifying the information format, unit, and outliers.

[0020] In one implementation, the manual review result and the correction result are fed back to the large language model to update the large language model.

[0021] In one embodiment, in step S3, the step of mapping the structured semantic tag to a unified equipment data model of a medical institution includes: performing field mapping according to a preset mapping rule table, performing unit conversion according to a preset unit conversion rule library, and integrating the field structure in the structured text according to a preset data structure template.

[0022] In a second aspect, a medical device communication protocol standardization system based on a large language model is provided, comprising:

[0023] The data acquisition module is used to collect monitoring data of each medical device in real time and report the monitoring data to the data integration platform in the form of data messages through the original communication protocol of the medical device;

[0024] The protocol parsing module is used in the data integration platform to input pre-processed data packets into the large language model to obtain structured semantic labels:

[0025] A large model training unit trains the large language model, including designing supervised learning tasks based on inputting original communication protocol messages and outputting standardized field names and semantic labels; and adding domain terminology explanations of the original communication protocol fields to the training set;

[0026] a prompt word design unit, which guides the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol, domain terminology explanations of the data message fields, and output structure formats;

[0027] The standardized mapping module and the data integration platform map the structured semantic tags to the unified equipment data model of the medical institution and automatically connect to the equipment management system of the medical institution.

[0028] The beneficial effects brought by the present invention are:

[0029] 1. This invention achieves automated understanding of medical device communication protocols through pre-training of large models and special design of prompt words, using advanced large model reasoning technology. This invention breaks the current dilemma of medical equipment relying on manual development, and can quickly and accurately identify the meaning represented by each field. On this basis, the new device access process is greatly simplified. The new device protocol analysis and adaptation work that previously took weeks or even months can now be greatly shortened with the help of this invention, significantly improving the efficiency of equipment online and helping medical institutions quickly realize the information management of equipment.

[0030] 2. This invention makes cross-device management more convenient and efficient by standardizing the data generated by all medical devices. Managers can centrally monitor and analyze data from different devices on the same platform; at the same time, it also lays a solid foundation for collaborative control between devices.

[0031] 3. Through the method of the present invention, there is no need to invest a lot of resources to rebuild a complex large language model. Only by using the existing mature large language model can the accurate analysis and standardization of medical device communication protocols be achieved, which significantly reduces the burden of manual development and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The present invention will be further described below with reference to the accompanying drawings.

[0033] Figure 1 It is a flowchart of a method for standardizing a medical device communication protocol based on a large language model according to an embodiment of the present invention.

[0034] Figure 2 This is a framework diagram of a medical device communication protocol standardization system based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The present invention relates to the following technical terms and constraints:

[0036] Medical device communication protocol: It is the rules and standards followed during data transmission and interaction between medical devices and systems, used to ensure accurate exchange of information such as device status and monitoring data.

[0037] Data message: A structured data unit organized in a specific format for transmission during the communication process.

[0038] Domain terminology: It is a proprietary word or phrase used in a specific professional field. It has strict definitions and precise semantics and is used to accurately describe the concepts, phenomena or operations in that field.

[0039] Prompt words: These are instructions or text descriptions that users input into a large language model to guide the model to generate specific content, perform specified tasks, or follow specific format requirements.

[0040] Context: refers to historical conversations, previous content, or background information that large language models refer to when processing current input. It is used to understand semantics, maintain conversation coherence, and accurately generate responses.

[0041] LOINC: Logical Observation Identifiers Names and Code is a medical standard terminology system jointly developed by the College of Laboratory Medicine (CLSI) and the U.S. National Library of Medicine (NLM). It aims to provide globally unified standardized identifiers and names for clinical test items and observation indicators.

[0042] SNOMED CT: Systematized Nomenclature of Medicine—Clinical Terms is a clinical terminology standard maintained by the International Health Terminology Standards Development Organization (IHTSDO). It is constructed using Description Logic and accurately describes all entities in the clinical field (such as diseases, symptoms, procedures, drugs, anatomical structures, etc.) through a structured combination of concepts, relationships, and attributes.

[0043] A unified equipment data model for medical institutions: This refers to the modeling of heterogeneous data (such as protocol formats and differences in field meanings) of different brands and models of equipment by medical institutions through unified field definitions, data structures, and semantic specifications in order to achieve standardized management of equipment data. This forms a standardized data template for the entire hospital and ensures the consistency and interoperability of various types of equipment data at the collection, storage, transmission, and application levels.

[0044] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] like Figure 1 As shown, taking the Nihon Kohden BMS-6000 series ECG monitor as an example (the device uses a proprietary communication protocol, and the protocol fields are not publicly standardized), an embodiment of the present invention provides a medical device communication protocol standardization method based on a large language model, including the following steps:

[0046] S1, real-time collection of monitoring data from each medical device, and reporting the monitoring data in the form of data packets to the data integration platform through the original communication protocol of the medical device;

[0047] Specifically, first connect the device to the local area network. Using a packet capture tool (Wireshark in this example), real-time monitoring data, such as electrocardiogram (ECG), heart rate, blood pressure, and blood oxygen levels, is captured. The data packets are then pushed to the data integration platform via the device gateway using the HL7 or IEEE 11073 protocol. If multiple medical devices are involved, follow these steps separately, with the specific communication protocol depending on the specific device. The gateway program receives the raw binary / ASCII packets and decompresses them according to the protocol to obtain the original field string.

[0048] In one implementation, the original message undergoes preliminary cleaning, including character set conversion, escape character processing, and segmentation. For example, an HL7 message is segmented by "\r\n" and parsed into a list of message segments; a blood pressure value of "120 / 80 mmHg" is split into "120 mmHg" and "80 mmHg" to facilitate subsequent parsing. Time calibration is also performed to unify the timestamp format across devices. There are many ways to clean messages. In addition to the methods mentioned above, other methods can also include removing redundant whitespace characters and verifying data integrity.

[0049] S2: The data integration platform inputs the pre-processed data message into the large language model to obtain structured semantic labels:

[0050] S2.1, training the large language model, including designing a supervised learning task based on inputting original communication protocol messages and outputting standardized field names and semantic labels; and adding domain terminology explanations for the original communication protocol fields to the training set;

[0051] Specifically, this example utilizes transfer learning to fine-tune a large, general-purpose model using annotated data related to medical device protocol semantics, including communication logs from actual devices, publicly available medical protocol examples, and synthetic data; as well as real HL7 or IEEE 11073 messages. This example also retains a validation set during training to prevent overfitting, and employs cross-validation to ensure the stability of the large language model.

[0052] In one embodiment, the domain terminology interpretation of the original communication protocol field adopts the terminology standards of LOINC and SNOMED CT. LOINC and SNOMED CT are internationally recognized standardized descriptions that ensure that the large language model learns the standard concepts involved in the protocol text. Of course, in addition to LOINC and SNOMED CT, other standardized descriptions can be introduced, such as ICD-10 (International Classification of Diseases, 10th Revision).

[0053] In one embodiment, by using contrastive learning technology, LOINC and SNOMED CT are used to generate positive samples, and non-related terms are randomly sampled as negative samples, so that the large language model learns the relative distance of samples in the feature space, thereby improving the large language model's ability to distinguish medical terms with similar semantics but different expressions.

[0054] Specifically, positive samples refer to different terms for the same concept, such as "hypertension" and "hypertension," or the LOINC code "55224-3" and the SNOMED CT code "42782006." Negative samples refer to terms for different concepts, such as diabetes and pneumonia. The large language model learns the relative distances of samples in feature space, meaning similar terms are close in vector space, while dissimilar terms are far apart. This forces the model to learn the semantic essence of medical terms, rather than superficial differences in characters or encodings. In addition to contrastive learning, label smoothing techniques can also be used to improve the large language model's generalization ability for different medical terms.

[0055] S2.2, guiding the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol as context, domain terminology explanations of the data message fields, and the required output structure format;

[0056] The prompt design incorporates protocol structure and domain knowledge to guide the large model to perform high-quality semantic parsing of the field:

[0057] First, the prompts clearly provide protocol background information and format requirements, such as example protocol message fragments and field meanings, as artifacts to reduce the misjudgment rate of large language models.

[0058] Second, the output structure format must be clearly defined, requiring the model to return field names, meanings, and units, and to correspond to the LOINC / SNOMED standard as much as possible to facilitate subsequent mapping to the unified device data model of medical institutions and access to the unified medical device management system;

[0059] Third, domain terms and examples are introduced into the prompt words. For example, for the field description "heart rate (HR)", the prompt word can include its LOINC code and name; for blood oxygen (SpO2), its standard definition and unit can be provided to enhance the large language model's recognition of medical field concepts.

[0060] The prompt words guide the large language model through the above three means, significantly improving the large language model's ability to parse medical device communication protocols.

[0061] In one embodiment, the domain terminology interpretation of the fields involved in the data message adopts LOINC and SNOMEDCT, which are the same as those of the original communication protocol fields, so as to facilitate the large language model to accurately understand the data message.

[0062] In one embodiment, when a data message contains numeric fields, the output format and precision of the numeric value must be clearly specified. For example, a statement such as "If the value contains decimals, please round it to two decimal places" or a regular expression specification pattern may be used. Medical device communication protocols require high numeric precision control, so this needs to be clearly stated in the prompt.

[0063] In this embodiment, the DeepSeek large model is taken as an example:

[0064] Input the cleaned data message text (or key fields) into the DeepSeek large model. The prompt words include the background information and format requirements of the protocol. For example, the sample protocol message fragment and field meaning are as follows:

[0065] [{"field_name":"SPO2_VAL","data_type":"integer","range":"0-100"},

[0066] {"field_name":"HR_ALARM_STATE","data_type":"boolean"},

[0067] {"field_name":"RESP_RATE_VAL","data_type":"integer","range":"0-60"}]

[0068] "SPO2_VAL" (blood oxygen value raw field)

[0069] "HR_ALARM_STATE" (heart rate alarm state raw field)

[0070] "RESP_RATE_VAL" (respiratory rate raw field)

[0071] Parse the following HL7 OBX line: OBX|1|NM|9279-1^Blood oxygen saturation^LN|...|98.6|%. These are fields in a medical device communication. Infer the medical indicator or device status corresponding to each field and standardize the naming and units (e.g., "Blood oxygen saturation," "98.6," "%").

[0072] Example of DeepSeek output:

[0073] [{"semantic_label":"Blood oxygen saturation","unit":"%"}.

[0074] Example of DeepSeek output results for inputting heart rate and respiratory rate related data packets:

[0075] {"semantic_label":"Heart rate alarm status","unit":"N / A"},

[0076] {"semantic_label":"Respiratory rate","unit":"breaths / minute"}]

[0077] In one embodiment, the large language model outputs not only structured semantic labels but also a confidence score for each inference. When the confidence score falls below a preset standard, manual review is triggered. Requiring the output of a confidence score makes the model output interpretable and reduces the risk of "erroneous inferences," but this is optional and not a necessary step in achieving communication protocol standardization. This embodiment sets the preset standard to 0.7. The large language model is required to output in JSON format, thereby obtaining structured semantic labels for each field.

[0078] In one embodiment, after the large language model outputs the structured semantic tags, the data integration platform further corrects the structured semantic tags, and the correction includes verifying the information format, unit, and outliers.

[0079] For example, it verifies whether the blood oxygen value format is digital and whether the unit is correct. If the model misidentifies, it will be handed over to the decision tree model to determine whether this value is more likely to be pulse or temperature, and correct it based on historical data models or simple threshold rules. At the same time, domain rules are applied: if the heart rate value is greater than 250 or less than 20, an abnormal value warning is triggered and marked for manual review. The rule engine can also automatically add missing unit information (such as automatically supplementing when the model output lacks units) or split unclear segmentation fields (such as solving compound alarm states into separate parameters).

[0080] In one embodiment, the manual review results and the correction results are fed back to the large language model to update the large language model, thereby enabling the large language model to continuously optimize itself during use and improve the parsing accuracy and adaptability of the large language model.

[0081] S3, the data integration platform maps the structured semantic tags to the unified equipment data model of the medical institution and automatically connects to the equipment management system of the medical institution.

[0082] In one embodiment, the step of mapping the structured semantic tags to a unified equipment data model of a medical institution includes: performing field mapping according to a preset mapping rule table, performing unit conversion according to a preset unit conversion rule library, and integrating the field structure in the structured text according to a preset data structure template.

[0083] Field Mapping: Assign corresponding internal model attributes to each parsed field. For example, if the model parses "HR" as heart rate (HeartRate), the internal system might store it as the vital_signs.heartRate field. In this case, you need to maintain a mapping rule table to map the output semantic labels to internal attributes. This process can be automated using a term library (e.g., mapping LOINC / SNOMED codes to internal standard terms).

[0084] Unit conversion: Medical data often involves conversions between different units. For example, a device might output a blood pressure value of "120 / 80 mmHg," but the internal model might store it in SI units (Pascals). In this case, the conversion must be performed using a dimensional conversion formula: for example, 1 mmHg ≈ 133.322 Pa. The code can call a unit conversion library to automatically convert values ​​using predefined conversion factors. For composite values ​​(such as blood pressure containing systolic and diastolic pressures), they must be split into two values ​​before conversion.

[0085] Data structure conversion: Integrate the parsed fields into internally defined objects or database records. For example, multiple fields such as heart rate, blood pressure, blood oxygen, etc. are combined into a "vital sign" object or associated with clinical observation records at the same time point. If the original protocol has complex nesting (such as multi-frame data packets), key frames must be extracted in the preprocessing stage. The entire mapping process is completed through configuration files and codes: field mapping tables, unit conversion rule libraries, data structure templates, etc. can be used as configuration maintenance to ensure that new protocol types can also be quickly accessed. Through the above technical process, the model parsing results can be seamlessly integrated into the standard equipment model of the system, realizing the automatic docking of the protocol to the application model.

[0086] In this embodiment, the standardized data example is completed according to the above steps:

[0087] json

[0088] {

[0089] "device_id":"NK_BMS6000_001",

[0090] "device_type":"ECG monitor",

[0091] "parameters":{

[0092] "heart rate":{"value":80,"unit":"bpm"},

[0093] "Blood oxygen saturation":{"value":96,"unit":"%"},

[0094] "Respiratory rate":{"value":20,"unit":"breaths / minute"}

[0095] },

[0096] "alarm_status":{

[0097] "Heart rate alarm": true

[0098] },

[0099] "timestamp":1714296300000}

[0100] like Figure 2 As shown, an embodiment of the present invention further provides a medical device communication protocol standardization system based on a large language model, including:

[0101] The data acquisition module is used to collect monitoring data of each medical device in real time and report the monitoring data to the data integration platform in the form of data messages through the original communication protocol of the medical device;

[0102] The protocol parsing module is used in the data integration platform to input pre-processed data packets into the large language model to obtain structured semantic labels:

[0103] A large model training unit trains the large language model, including designing supervised learning tasks based on inputting original communication protocol messages and outputting standardized field names and semantic labels; and adding domain terminology explanations of the original communication protocol fields to the training set;

[0104] a prompt word design unit, which guides the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol, domain terminology explanations of the data message fields, and output structure formats;

[0105] The standardized mapping module and the data integration platform map the structured semantic tags to the unified equipment data model of the medical institution and automatically connect to the equipment management system of the medical institution.

[0106] The advantages of the present invention are:

[0107] 1. This invention leverages advanced large-scale model inference technology, pre-training the large model, and specially designed prompt words to achieve automated understanding of medical device communication protocols. This invention overcomes the current dilemma of medical equipment relying on manual development, quickly and accurately identifying the meaning of each field. On this basis, the new device access process is greatly simplified. Previously, the new device protocol analysis and adaptation work required weeks or even months. With this invention, the access cycle can now be significantly shortened, significantly improving device online efficiency and helping medical institutions quickly implement information management of their equipment.

[0108] 2. This invention makes cross-device management more convenient and efficient by standardizing the data generated by all medical devices. Managers can centrally monitor and analyze data from different devices on the same platform; at the same time, it also lays a solid foundation for collaborative control between devices.

[0109] 3. Through the method of the present invention, there is no need to invest a lot of resources to rebuild a complex large language model. Only by using the existing mature large language model can the accurate analysis and standardization of medical device communication protocols be achieved, which significantly reduces the burden of manual development and maintenance.

[0110] The present invention is not limited to the specific technical solutions described in the above embodiments. In addition to the above embodiments, the present invention may also have other implementation methods. For those skilled in the art, any technical solutions formed by modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for standardizing a medical device communication protocol based on a large language model, comprising the following steps: S1, real-time collection of monitoring data from each medical device, and reporting the monitoring data in the form of data packets to the data integration platform through the original communication protocol of the medical device; S2: The data integration platform inputs the pre-processed data message into the large language model to obtain structured semantic labels: S2.1, training the large language model, including designing a supervised learning task based on inputting original communication protocol messages and outputting standardized field names and semantic labels; and adding domain terminology explanations for the original communication protocol fields to the training set; S2.2, guiding the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol as context, domain terminology explanations of the data message fields, and the required output structure format; S3, the data integration platform maps the structured semantic tags to the unified equipment data model of the medical institution and automatically connects to the equipment management system of the medical institution.

2. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: The preprocessing includes character set conversion, escape character processing and segmentation, and time calibration is performed to unify the timestamp formats of the various medical devices.

3. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: In steps S2.1 and S2.2, the domain terminology interpretations of the original communication protocol fields and the domain terminology interpretations of the fields involved in the data message both adopt the same terminology standards, including LOINC and SNOMED CT.

4. The method for standardizing medical device communication protocols based on a large language model according to claim 3, wherein: In step S2.1, contrastive learning technology is also used to generate positive samples using LOINC and SNOMED CT, and non-related terms are randomly sampled as negative samples, so that the large language model can learn the relative distance between samples in the feature space, thereby improving the large language model's ability to distinguish medical terms with similar semantics but different expressions.

5. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: In step S2.2, if the data message contains a numerical field, the output numerical format and precision must be specified.

6. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: In step S2.2, in addition to outputting structured semantic tags, the large language model also outputs a confidence score for each inference. When the confidence score is lower than a preset standard, manual review is triggered.

7. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: Step S2 also includes step S2.

3. After the large language model outputs the structured semantic tags, the data integration platform also corrects the structured semantic tags. The correction includes verifying the information format, unit, and abnormal values.

8. The method for standardizing a medical device communication protocol based on a large language model according to any one of claims 6 to 7, wherein: The manual review result and the correction result are fed back to the large language model to achieve an update of the large language model.

9. The method for standardizing medical device communication protocols based on a large language model according to claim 1, wherein: In step S3, the step of mapping the structured semantic tags to the unified equipment data model of the medical institution includes: performing field mapping according to a preset mapping rule table, performing unit conversion according to a preset unit conversion rule library, and integrating the field structure in the structured text according to a preset data structure template.

10. A medical device communication protocol standardization system based on a large language model, comprising: The data acquisition module is used to collect monitoring data of each medical device in real time and report the monitoring data to the data integration platform in the form of data messages through the original communication protocol of the medical device; The protocol parsing module is used in the data integration platform to input pre-processed data packets into the large language model to obtain structured semantic labels: A large model training unit, which trains the large language model, including designing supervised learning tasks based on inputting original communication protocol messages and outputting standardized field names and semantic labels; Adding domain terminology explanations about the original communication protocol fields to the training set; a prompt word design unit, which guides the large language model to parse and output structured semantic tags by designing prompt words; the prompt words at least include background information and field meaning descriptions of the original communication protocol, domain terminology explanations of the data message fields, and output structure formats; The standardized mapping module and the data integration platform map the structured semantic tags to the unified equipment data model of the medical institution and automatically connect to the equipment management system of the medical institution.

Citation Information

Cited By

  • Communication protocol analysis method and system, computer equipment and computer program product

    CN121486486A

  • Communication protocol analysis method and system, computer device and computer program product

    CN121486486B