Medical internet of things protocol adaptation method based on deep transfer learning

Through deep transfer learning and conditional adversarial adaptive algorithms, combined with a dynamic gradient reversal mechanism, multimodal data adaptation of the medical Internet of Things protocol is achieved, which solves the problems of information loss and poor real-time performance in existing technologies and improves the flexibility and applicability of protocol adaptation.

CN120711083APending Publication Date: 2025-09-26PROACTIVE MEDICAL DEVICES LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510881335.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing medical Internet of Things protocol adaptation methods lack the ability to jointly model and adaptively integrate multimodal data, resulting in information loss and poor adaptation effects. In addition, traditional methods have high computational overhead and poor real-time performance, making it difficult to meet the complex and changing needs of medical scenarios.

Method used

A multi-channel multimodal feature extraction network based on deep transfer learning and a conditional adversarial adaptive algorithm are used, combined with a dynamic adaptive gradient inversion mechanism, to achieve deep feature fusion and adaptive alignment of protocol distribution of multimodal medical Internet of Things data. Target protocol data is generated through multimodal feature encoding, internal and external modal fusion, conditional adversarial training and protocol mapping.

Benefits of technology

It achieves efficient and robust multi-protocol device interconnection, improves the flexibility and real-time performance of feature alignment, adapts to multi-protocol, multi-modal, and dynamically changing scenarios in the medical Internet of Things environment, and has efficient interconnection and intelligent data interaction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711083A_ABST
    Figure CN120711083A_ABST
Patent Text Reader

Abstract

The invention discloses a medical internet of things protocol adaptation method based on deep transfer learning, and the method comprises the following steps: S1, collecting source protocol data and target protocol data, and obtaining multi-modal input data; s2, inputting the multi-modal input data into a multi-channel multi-modal adaptive network, and extracting multi-modal features; s3, performing intra-modal fusion and inter-modal fusion on the multi-modal features to generate fusion feature representation; s4, inputting the fusion feature representation into a conditional adversarial adaptive algorithm, generating an adversarial sample and executing a protocol domain discrimination operation; s5, performing protocol domain adversarial training by using gradient inversion, updating network parameters through a back propagation algorithm, and outputting updated fusion feature representation; and S6, performing protocol mapping and data format conversion, generating target protocol data, and transmitting the target protocol data to the medical Internet of Things gateway. According to the invention, multi-modal feature deep fusion, protocol distribution dynamic alignment and efficient adaptation of the target protocol in a complex multi-protocol scene of the medical internet of things are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep transfer learning and medical Internet of Things protocol adaptation, and in particular to a medical Internet of Things protocol adaptation method based on deep transfer learning. Background Art

[0002] With the rapid development of medical IoT technology and the widespread adoption of diverse intelligent devices in healthcare scenarios, data interoperability between medical devices and platforms has become a crucial technical enabler for improving the quality and efficiency of healthcare services. Within the medical IoT environment, various medical devices and systems typically exchange data using different protocols, such as HL7, FHIR, DICOM, and MQTT. These protocols exhibit significant differences in data structure, format, and semantics, making seamless integration difficult. Existing medical IoT protocol adaptation methods primarily rely on rule-based protocol conversion or simple field mapping techniques, typically using static mapping tables or pre-set rules to perform format conversion and field adaptation for different protocol data. These methods can achieve basic data compatibility to a certain extent for protocol data with a fixed format and a high degree of standardization. However, in real-world medical IoT applications, with a wide variety of devices, frequent protocol updates, and diverse data structures, static rule-based adaptation approaches lack flexibility and scalability, making them difficult to adapt to complex and ever-changing protocol adaptation requirements.

[0003] At the same time, existing protocol adaptation methods have obvious limitations when processing multimodal data. Data in the medical Internet of Things environment usually includes multimodal information such as text, numbers, and images, and these data modalities have different feature expressions and distribution patterns. Most traditional protocol adaptation methods lack the ability to jointly model and adaptively fuse multimodal data, resulting in information loss and poor adaptation effects in the protocol conversion of multimodal data. In addition, the protocol adaptation process in the medical Internet of Things environment not only needs to ensure consistency in format, but also needs to adaptively align the data distribution between protocols at the feature level to achieve efficient interaction and in-depth understanding of the data. Existing methods generally lack multimodal feature extraction and protocol distribution alignment mechanisms based on deep models, and are unable to fully explore the potential relationships and complementary advantages of multimodal data in cross-protocol adaptation.

[0004] In recent years, with the advancement of deep learning, particularly transfer learning and adversarial learning, cross-domain feature adaptation and distribution alignment have made significant progress in fields such as image classification and natural language processing. Some studies have begun to utilize deep neural networks to extract features from multimodal medical data and perform end-to-end protocol conversion, improving the adaptation of multimodal data across different protocols. However, these methods still have significant shortcomings. First, existing deep transfer learning methods mostly use general-purpose multimodal feature extraction networks and lack structured designs tailored to the protocol adaptation scenarios of the Medical Internet of Things (IoT), resulting in inadequate adaptation to contextual semantic differences between protocols. Second, existing conditional adversarial domain adaptation methods often overlook the dynamic incorporation of protocol context labels and are unable to dynamically adjust feature distribution alignment strategies based on differences in protocol labels. Third, traditional gradient reversal mechanisms use fixed adversarial coefficients and are unable to adapt to the adaptation progress and discriminant signal strength during training, limiting the effectiveness and convergence of adversarial training. Fourth, existing deep models generally suffer from high computational overhead and poor real-time performance when dealing with multi-protocol, multi-modal, and dynamically changing data in the Medical Internet of Things (IoT) environment, making them difficult to meet the real-time and reliability requirements of medical scenarios.

[0005] Therefore, how to combine advanced technologies such as deep transfer learning, conditional adversarial adaptation, and dynamic gradient reversal to build an efficient, robust, and adaptive protocol adaptation method for the complex characteristics of multi-protocol, multi-modal, and dynamic changes in medical Internet of Things protocol adaptation is an urgent problem to be solved in the current field of medical Internet of Things protocol adaptation technology. Summary of the Invention

[0006] One purpose of the present invention is to propose a medical Internet of Things protocol adaptation method based on deep transfer learning. The present invention integrates a multi-channel multimodal feature extraction network with a conditional adversarial adaptive algorithm. By introducing conditional information of the source protocol label and the target protocol label and combining it with a dynamic adaptive gradient reversal mechanism, the present invention realizes deep feature fusion and protocol distribution adaptive alignment of multimodal medical Internet of Things data. It has the advantages of strong protocol adaptation flexibility, good feature alignment robustness, adaptive adjustment of the training process, and high real-time reasoning. It is suitable for efficient interconnection and intelligent data interaction of multi-protocol devices in the medical Internet of Things environment.

[0007] A method for adapting a medical Internet of Things protocol based on deep transfer learning according to an embodiment of the present invention includes the following steps:

[0008] S1. Collect source protocol data and target protocol data in the medical Internet of Things environment, use a protocol parser to perform structured parsing, and obtain multimodal input data;

[0009] S2. Input the multimodal input data into the multi-channel multimodal adaptation network and extract the multimodal features through the multimodal feature encoder;

[0010] S3. In the multi-channel multimodal adaptation network, intra-modal fusion and inter-modal fusion are performed on the multimodal features to generate a fused feature representation;

[0011] S4. Input the fused feature representation into the conditional adversarial adaptive algorithm, generate adversarial samples based on the source protocol label and the target protocol label through the condition generator, receive the fused feature representation through the conditional discriminator and perform the protocol domain discrimination operation;

[0012] S5. In the conditional adversarial adaptive algorithm, gradient reversal is used to perform protocol domain adversarial training. The network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive module are updated through the back-propagation algorithm, and the updated fusion feature representation is output.

[0013] S6. Input the updated fusion feature representation into the protocol adapter module, perform protocol mapping and data format conversion according to the target protocol requirements, generate the target protocol data and transmit it to the medical Internet of Things gateway.

[0014] Optionally, the multimodal input data specifically includes text modality input data, numerical modality input data, and image modality input data;

[0015] Optionally, the multi-channel multimodal adaptation network specifically includes a multimodal feature encoder, an intra-modal fusion module, an inter-modal fusion module, a residual connection module, and a self-attention mechanism module:

[0016] The multimodal feature encoder receives multimodal input data, uses a convolutional neural network structure to perform feature extraction, and outputs multimodal features, wherein the multimodal features include text modal features, numerical modal features, and image modal features;

[0017] The intra-modality fusion module includes intra-modality fusion units corresponding to text modality, numerical modality and image modality. Each intra-modality fusion unit fuses features of the same modality through weighted fusion and outputs intra-modality fusion features.

[0018] The inter-modal fusion module includes a splicing unit and a weighted fusion unit. The splicing unit is used to splice the intra-modal fusion features. The weighted fusion unit is used to perform weighted fusion on the spliced ​​features according to a preset weighting coefficient to generate an inter-modal fusion feature.

[0019] The residual connection module includes a residual connection unit, which performs residual connection on the inter-modality fusion feature and the intra-modality fusion feature, and outputs a residual connection feature representation;

[0020] The self-attention mechanism module includes a self-attention calculation unit, which receives the residual connection feature representation, uses the self-attention mechanism to dynamically weight the correlation between feature dimensions, and finally generates a fused feature representation.

[0021] Optionally, the S3 specifically includes:

[0022] S31. Input the multimodal features into the corresponding intramodal fusion unit in the intramodal fusion module. Each intramodal fusion unit fuses the features of the same modality through weighted fusion and outputs the intramodal fusion features:

[0023]

[0024] in, is the intra-modal fusion feature of modality m, m∈{text,num,img} is the modality type, that is, is the fusion feature within the text modality, is the fusion feature within the numerical mode, is the image modality fusion feature, is the i-th feature channel in mode m, is the weighted coefficient of the i-th feature channel in mode m, Conv1D(·) is a one-dimensional convolution operation, is the weighted coefficient of the jth convolution channel in mode m, C m is the total number of characteristic channels of mode m, K m is the total number of convolution channels of modality m, and LayerNorm(·) is the layer normalization operation;

[0025] S32: Inputting the text modality intra-fusion features, the numerical modality intra-fusion features, and the image modality intra-fusion features into an inter-modality fusion module, wherein the inter-modality fusion module includes a splicing unit and a multi-head attention weighted fusion unit, and the splicing unit splices the intra-modality fusion features to generate modality splicing features;

[0026] S33. In the multi-head attention weighted fusion unit, the modal splicing features are received, and inter-modal weighted fusion is performed through the multi-head attention mechanism to generate an inter-modal fusion feature representation:

[0027]

[0028] in, is the inter-modal fusion feature representation, Concat(·) is the multi-head attention concatenation operation, Softmax(·) is the normalized weight assignment, Q is the query matrix, which is generated by the modal concatenation feature through the query weight matrix, K is the key matrix, which is generated by the modal concatenation feature through the key weight matrix, V is the value matrix, which is generated by the modal concatenation feature through the value weight matrix, and Wq 、W k 、W v is the learnable linear transformation weight matrix of the attention head, which is used to linearly transform the query matrix, key matrix and value matrix respectively. is the scaling factor;

[0029] S34. Output the inter-modal fusion feature representation as the final fusion feature representation.

[0030] Optionally, the conditional adversarial adaptive algorithm specifically includes a condition generator module, a condition feature generation module, a condition discriminator module, a gradient reversal module and a parameter update module:

[0031] The condition generator module receives a source protocol label and a target protocol label, wherein the condition embedding unit generates embedding vectors for the source protocol label and the target protocol label respectively, and splices them into a conditional embedding vector;

[0032] The conditional feature generation module receives the fused feature representation and the conditional embedding vector, performs conditional concatenation and one-dimensional convolution operations, and generates conditional adversarial features;

[0033] The conditional discriminator module includes a multi-layer convolution unit and a discrimination output unit, receives the conditional adversarial features, performs protocol domain discrimination, and generates a discrimination output;

[0034] The gradient reversal module introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the strength of the discriminant signal, and propagates the gradient of the discriminant output in an adaptive direction:

[0035]

[0036] in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, D out is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round;

[0037] The parameter updating module performs back propagation based on the dynamic inversion gradient, updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm, and outputs the updated fusion feature representation.

[0038] Optionally, the S4 specifically includes:

[0039] S41, the fusion feature representation is input into the conditional adversarial adaptive algorithm, the condition generator module receives the source protocol label and the target protocol label, and converts the source protocol label Label src and target protocol label tgt Input conditional embedding unit, the conditional embedding unit uses embedding function Embed(·) to transform vectors and generate conditional embedding vectors Embed(Label src ) and Embed(Label tgt );

[0040] S42, Embed(Label src ) and Embed(Label tgt ) Input the concatenation unit, concatenate according to the vector dimension, and generate a conditional embedding vector;

[0041] S43, inputting the conditional embedding vector and the fused feature representation into a splicing unit, performing a splicing operation on the feature dimension, and generating a conditional feature splicing input;

[0042] S44, inputting the conditional splicing input into a one-dimensional convolution unit in a conditional feature generation module, wherein the one-dimensional convolution unit performs a convolution operation and adds a bias term, and then performs nonlinear mapping through an activation function to generate a conditional adversarial feature;

[0043] S45. Input the conditional adversarial feature into a conditional discriminator, wherein the conditional discriminator includes a multi-layer convolution unit and a discriminant output unit. The conditional discriminator first performs feature transformation on the conditional adversarial feature through the multi-layer convolution unit to generate a discriminant feature, inputs the discriminant feature into the discriminant output unit, performs a convolution operation and adds a bias term, performs a protocol domain discrimination operation through an activation function, and generates a discriminant output;

[0044] Optionally, the S5 specifically includes:

[0045] S51. In the conditional adversarial adaptive algorithm, the gradient reversal module receives the discriminant output and introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the discriminant signal strength;

[0046] S52. Calculate the gradient of the discriminant output with respect to the network parameters, apply the dynamic weight coefficient to the gradient, and perform adaptive direction propagation on the gradient of the discriminant output:

[0047]

[0048] in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, Dout is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round;

[0049] S53. The parameter updating module updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm based on the dynamic inversion gradient through the back propagation algorithm, and outputs the updated fusion feature representation.

[0050] Optionally, the S6 specifically includes:

[0051] S61. The updated fused feature representation is passed as input to the protocol mapping unit of the protocol adapter module. The protocol mapping unit performs a field mapping operation based on the protocol field definition and protocol standard of the target protocol tag, maps the updated fused feature representation to the feature space of the target protocol, and generates a target protocol field feature representation.

[0052] S62: Input the target protocol field feature representation into the data format conversion unit of the protocol adapter module. The data format conversion unit performs field encapsulation, data rearrangement, and format conversion operations according to the data encapsulation standard of the target protocol to generate target protocol data that meets the requirements of the target protocol.

[0053] S63. The target protocol data is used as a protocol adaptation result and transmitted through the medical Internet of Things gateway.

[0054] The beneficial effects of the present invention are:

[0055] First, unlike the traditional medical Internet of Things protocol adaptation method based on rules or static mapping tables, the present invention proposes a multi-channel multimodal adaptation network based on deep transfer learning, which can fully utilize the feature expression power of multimodal data such as text, numerical values, and images. In the feature extraction stage, the intra-modal fusion module uses weighted fusion and convolution operations to fully model the local features within each modality. The inter-modal fusion module combines splicing operations with a multi-head attention mechanism to achieve deep coupling of multimodal global features. Through the residual connection module and the self-attention mechanism module, the network's expression robustness and global dependence on multimodal information are effectively enhanced, ensuring the stability and applicability of the fused feature representation in the protocol adaptation task.

[0056] Secondly, the present invention introduces a conditional adversarial adaptive algorithm during the protocol adaptation process. The source and target protocol labels are dynamically injected into the model in the form of conditional vectors. The conditional adversarial features generated by the conditional generator module are used in conjunction with the conditional discriminator module to discriminate the protocol domain, thereby constructing a deep adversarial adaptation framework based on conditional information. In particular, the dynamic adaptive gradient reversal mechanism proposed in the present invention combines training progress and discriminant signal strength to dynamically and smoothly adjust the adversarial weights. This effectively addresses the inability of traditional fixed gradient reversal coefficients to flexibly adapt to the training process, significantly improving the robustness and adaptability of feature distribution alignment.

[0057] During the protocol mapping and data format conversion phase, the present invention uses the protocol mapping unit and data format conversion unit of the protocol adapter module to efficiently convert the updated fusion features to the target protocol data according to the field definition and encapsulation standard of the target protocol, thereby ensuring the standardization and availability of the adaptation results. Through the deep fusion of multimodal features and conditional adversarial adaptive training, the present invention has higher protocol adaptation flexibility and multi-protocol adaptation capabilities, and can adapt to the complex scenario requirements of multiple devices, multiple protocols, and multiple data types in the medical Internet of Things environment.

[0058] Furthermore, the present invention utilizes an end-to-end differentiable network structure, combined with a lightweight model design, and can be deployed on edge computing devices equipped with neural network inference units, enabling low-latency, low-power, and real-time protocol adaptation. Compared to existing technologies, the present invention achieves significant breakthroughs in fusion feature expression, protocol distribution alignment, and adaptation robustness, demonstrating strong engineering capabilities and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0060] Figure 1 This is an overall flow chart of a medical Internet of Things protocol adaptation method based on deep transfer learning proposed by the present invention;

[0061] Figure 2 This is a schematic diagram of the structure of a multi-channel and multi-modal adaptation network for a medical Internet of Things protocol adaptation method based on deep transfer learning proposed in the present invention;

[0062] Figure 3 This is a structural diagram of the conditional adversarial adaptive algorithm of the medical Internet of Things protocol adaptation method based on deep transfer learning proposed in the present invention. DETAILED DESCRIPTION

[0063] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0064] refer to Figure 1-3 , a medical Internet of Things protocol adaptation method based on deep transfer learning, comprising the following steps:

[0065] S1. Collect source protocol data and target protocol data in the medical Internet of Things environment, use a protocol parser to perform structured parsing, and obtain multimodal input data;

[0066] S2. Input the multimodal input data into the multi-channel multimodal adaptation network and extract the multimodal features through the multimodal feature encoder;

[0067] S3. In the multi-channel multimodal adaptation network, intra-modal fusion and inter-modal fusion are performed on the multimodal features to generate a fused feature representation;

[0068] S4. Input the fused feature representation into the conditional adversarial adaptive algorithm, generate adversarial samples based on the source protocol label and the target protocol label through the condition generator, receive the fused feature representation through the conditional discriminator and perform the protocol domain discrimination operation;

[0069] S5. In the conditional adversarial adaptive algorithm, gradient reversal is used to perform protocol domain adversarial training. The network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive module are updated through the back-propagation algorithm, and the updated fusion feature representation is output.

[0070] S6. Input the updated fusion feature representation into the protocol adapter module, perform protocol mapping and data format conversion according to the target protocol requirements, generate the target protocol data and transmit it to the medical Internet of Things gateway.

[0071] Through a multi-channel multimodal adaptation network, the deep features of multimodal data such as text, numbers, and images are fully extracted. Combined with modules such as intra-modal fusion, inter-modal fusion, and self-attention mechanisms, the richness and accuracy of the fused feature expression are ensured. By introducing a conditional adversarial adaptive algorithm, the feature distributions between the source and target protocols are dynamically aligned, effectively overcoming the adaptation barriers caused by protocol differences and improving the generalization and robustness of protocol conversion. At the same time, combined with an adaptive dynamic gradient reversal mechanism, the stability and adaptability of adversarial training are enhanced, ensuring seamless alignment of feature distributions. This method is adaptable to the complex multi-protocol, multi-modal, and multi-variable scenarios in the medical Internet of Things (MIoT) environment. It boasts strong adaptability, high adaptation accuracy, and good real-time performance, meeting the data interaction needs of multi-device, high-concurrency scenarios in the MIoT.

[0072] In this embodiment, the multimodal input data specifically includes text modal input data, numerical modal input data and image modal input data;

[0073] By clearly segmenting multimodal input data into textual, numerical, and image modal input data, it is possible to comprehensively collect and utilize multi-source heterogeneous data in the medical Internet of Things environment. This segmentation enables more targeted processing of data from different modalities during the multimodal feature encoding stage, reduces interference between modalities, and improves the integrity and accuracy of feature extraction. For the diverse types of devices and complex data streams in medical scenarios, textual modal input can support textual information such as medical records, numerical modal input can cover vital sign parameters collected by sensors, and image modal input can meet the needs of image data processing, thereby comprehensively improving the applicability of protocol adaptation and the efficiency of multimodal data interaction.

[0074] In this embodiment, the multi-channel multimodal adaptation network specifically includes a multimodal feature encoder, an intra-modal fusion module, an inter-modal fusion module, a residual connection module and a self-attention mechanism module:

[0075] The multimodal feature encoder receives multimodal input data, uses a convolutional neural network structure to perform feature extraction, and outputs multimodal features, wherein the multimodal features include text modal features, numerical modal features, and image modal features;

[0076] The intra-modality fusion module includes intra-modality fusion units corresponding to text modality, numerical modality and image modality. Each intra-modality fusion unit fuses features of the same modality through weighted fusion and outputs intra-modality fusion features.

[0077] The inter-modal fusion module includes a splicing unit and a weighted fusion unit. The splicing unit is used to splice the intra-modal fusion features. The weighted fusion unit is used to perform weighted fusion on the spliced ​​features according to a preset weighting coefficient to generate an inter-modal fusion feature.

[0078] The residual connection module includes a residual connection unit, which performs residual connection on the inter-modality fusion feature and the intra-modality fusion feature, and outputs a residual connection feature representation;

[0079] The self-attention mechanism module includes a self-attention calculation unit, which receives the residual connection feature representation, uses the self-attention mechanism to dynamically weight the correlation between feature dimensions, and finally generates a fused feature representation.

[0080] By integrating a multimodal feature encoder, an intra-modal fusion module, an inter-modal fusion module, a residual connection module, and a self-attention mechanism module into a multi-channel multimodal adaptation network, a comprehensive multimodal deep feature fusion architecture is formed. This architecture not only extracts semantic information from multimodal data but also enables deep mutual enhancement and dynamic weighting between different modalities. The residual connection module further prevents information loss and enhances the network's feature retention capabilities, while the self-attention mechanism module adaptively captures global dependencies between key modalities, improving the expressive power of multimodal features. This deep fusion capability is crucial for complex medical IoT protocol data and can effectively improve the accuracy and flexibility of adaptation in actual protocol adaptation.

[0081] In this embodiment, S3 specifically includes:

[0082] S31. Input the multimodal features into the corresponding intramodal fusion unit in the intramodal fusion module. Each intramodal fusion unit fuses the features of the same modality through weighted fusion and outputs the intramodal fusion features:

[0083]

[0084] in, is the intra-modal fusion feature of modality m, m∈{text,num,img} is the modality type, that is, is the fusion feature within the text modality, is the fusion feature within the numerical mode, is the image modality fusion feature, is the i-th feature channel in mode m, is the weighted coefficient of the i-th feature channel in mode m, Conv1D(·) is a one-dimensional convolution operation, is the weighted coefficient of the jth convolution channel in mode m, C m is the total number of characteristic channels of mode m, K m is the total number of convolution channels of modality m, and LayerNorm(·) is the layer normalization operation;

[0085] S32: Inputting the text modality intra-fusion features, the numerical modality intra-fusion features, and the image modality intra-fusion features into an inter-modality fusion module, wherein the inter-modality fusion module includes a splicing unit and a multi-head attention weighted fusion unit, and the splicing unit splices the intra-modality fusion features to generate modality splicing features;

[0086] S33. In the multi-head attention weighted fusion unit, the modal splicing features are received, and inter-modal weighted fusion is performed through the multi-head attention mechanism to generate an inter-modal fusion feature representation:

[0087]

[0088] in, is the inter-modal fusion feature representation, Concat(·) is the multi-head attention concatenation operation, Softmax(·) is the normalized weight assignment, Q is the query matrix, which is generated by the modal concatenation feature through the query weight matrix, K is the key matrix, which is generated by the modal concatenation feature through the key weight matrix, V is the value matrix, which is generated by the modal concatenation feature through the value weight matrix, and W q 、W k 、W v is the learnable linear transformation weight matrix of the attention head, which is used to linearly transform the query matrix, key matrix and value matrix respectively. is the scaling factor;

[0089] S34. Output the inter-modal fusion feature representation as the final fusion feature representation.

[0090] By refining the intra-modal and inter-modal fusion processes of multimodal features, the feature fusion strategy is concretized. The intra-modal fusion portion, through a combination of weighted summation and one-dimensional convolution operations, can fully exploit the contextual feature information within each modality and enhance the expressive power of single-modal features. The inter-modal fusion portion, through a multi-head attention weighting mechanism, achieves dynamic weight distribution of multimodal information based on the splicing of intra-modal fusion features. The multi-view feature capture of the multi-head attention mechanism enables a more comprehensive characterization of protocol differences, enhancing the applicability and adaptability of the fused features to downstream protocol adaptation tasks and meeting the cross-protocol feature alignment requirements of the Medical Internet of Things.

[0091] In this embodiment, the conditional adversarial adaptive algorithm specifically includes a condition generator module, a condition feature generation module, a condition discriminator module, a gradient reversal module and a parameter update module:

[0092] The condition generator module receives a source protocol label and a target protocol label, wherein the condition embedding unit generates embedding vectors for the source protocol label and the target protocol label respectively, and splices them into a conditional embedding vector;

[0093] The conditional feature generation module receives the fused feature representation and the conditional embedding vector, performs conditional concatenation and one-dimensional convolution operations, and generates conditional adversarial features;

[0094] The conditional discriminator module includes a multi-layer convolution unit and a discrimination output unit, receives the conditional adversarial features, performs protocol domain discrimination, and generates a discrimination output;

[0095] The gradient reversal module introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the strength of the discriminant signal, and propagates the gradient of the discriminant output in an adaptive direction:

[0096]

[0097] in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, D out is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round;

[0098] The parameter updating module performs back propagation based on the dynamic inversion gradient, updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm, and outputs the updated fusion feature representation.

[0099] The conditional adversarial adaptive algorithm integrates a conditional generation module, a conditional discriminator module, and a dynamic weight-based adaptive gradient reversal module to achieve dynamic alignment of the distribution between the source and target protocols. The dynamic weight coefficient not only changes smoothly based on the training progress but also introduces adaptive adjustment of the discriminant signal strength, enabling flexible and dynamic adjustment of adversarial strength. This reduces training oscillations in the early stages of protocol adaptation while enhancing the model's adversarial capabilities in later stages. This further promotes the deep integration and distribution consistency of multi-protocol features, helping to continuously output more stable and reliable adaptation results in diverse medical IoT protocol scenarios.

[0100] In this embodiment, the S4 specifically includes:

[0101] S41, the fusion feature representation is input into the conditional adversarial adaptive algorithm, the condition generator module receives the source protocol label and the target protocol label, and converts the source protocol label Label src and target protocol label tgt Input conditional embedding unit, the conditional embedding unit uses embedding function Embed(·) to transform vectors and generate conditional embedding vectors Embed(Label src ) and Embed(Label tgt );

[0102] S42, Embed(Label src ) and Embed(Label tgt ) Input the concatenation unit, concatenate according to the vector dimension, and generate a conditional embedding vector;

[0103] S43, inputting the conditional embedding vector and the fused feature representation into a splicing unit, performing a splicing operation on the feature dimension, and generating a conditional feature splicing input;

[0104] S44, inputting the conditional splicing input into a one-dimensional convolution unit in a conditional feature generation module, wherein the one-dimensional convolution unit performs a convolution operation and adds a bias term, and then performs nonlinear mapping through an activation function to generate a conditional adversarial feature;

[0105] S45. Input the conditional adversarial feature into a conditional discriminator, wherein the conditional discriminator includes a multi-layer convolution unit and a discriminant output unit. The conditional discriminator first performs feature transformation on the conditional adversarial feature through the multi-layer convolution unit to generate a discriminant feature, inputs the discriminant feature into the discriminant output unit, performs a convolution operation and adds a bias term, performs a protocol domain discrimination operation through an activation function, and generates a discriminant output;

[0106] By defining the input, concatenation, and convolution operations of the conditional generator and conditional discriminator in detail, the implementation path of the conditional adversarial adaptive module is made more specific and executable. Through steps such as label embedding and feature concatenation, the conditional generator helps dynamically inject protocol label information into the protocol domain discrimination process, enabling flexible adaptation to the characteristics of the source and target domains. The conditional discriminator extracts conditional adversarial features and generates discriminant outputs through multi-layer convolution, ensuring the discriminator's efficient and flexible protocol domain classification capabilities, enhancing the robustness and adaptability of the entire system in multi-protocol scenarios.

[0107] In this embodiment, the S5 specifically includes:

[0108] S51. In the conditional adversarial adaptive algorithm, the gradient reversal module receives the discriminant output and introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the discriminant signal strength;

[0109] S52. Calculate the gradient of the discriminant output with respect to the network parameters, apply the dynamic weight coefficient to the gradient, and perform adaptive direction propagation on the gradient of the discriminant output:

[0110]

[0111] in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, D out is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round;

[0112] S53. The parameter updating module updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm based on the dynamic inversion gradient through the back propagation algorithm, and outputs the updated fusion feature representation.

[0113] The proposed dynamic adaptive gradient reversal mechanism adjusts the gradient direction and magnitude during backpropagation in real time based on training progress and discriminator signal strength. This dynamic adjustment avoids the drawbacks of a fixed gradient reversal coefficient, enabling the model to flexibly distribute the impact of adversarial learning at different training stages, promoting network convergence and adaptability. By introducing an adaptive adjustment term based on discriminator signal strength, the mechanism accurately matches the distribution differences between protocols, improving the domain alignment capabilities of the adversarial adaptive module and possessing greater practical value in multi-source adaptation scenarios for medical IoT protocols.

[0114] In this embodiment, S6 specifically includes:

[0115] S61. The updated fused feature representation is passed as input to the protocol mapping unit of the protocol adapter module. The protocol mapping unit performs a field mapping operation based on the protocol field definition and protocol standard of the target protocol tag, maps the updated fused feature representation to the feature space of the target protocol, and generates a target protocol field feature representation.

[0116] S62: Input the target protocol field feature representation into the data format conversion unit of the protocol adapter module. The data format conversion unit performs field encapsulation, data rearrangement, and format conversion operations according to the data encapsulation standard of the target protocol to generate target protocol data that meets the requirements of the target protocol.

[0117] S63. The target protocol data is used as a protocol adaptation result and transmitted through the medical Internet of Things gateway.

[0118] By inputting the updated fusion feature representation into the protocol mapping unit and data format conversion unit of the protocol adapter module, efficient conversion of fusion features to target protocol data is ensured. The protocol mapping unit dynamically maps features based on the target protocol label, ensuring that multimodal features are properly assigned to the target protocol fields to avoid information loss. The data format conversion unit automatically completes field encapsulation and data rearrangement according to the data format specification of the target protocol, allowing the adaptation results to seamlessly connect with the medical Internet of Things gateway. This modular and standardized data adaptation process helps achieve seamless integration of medical data in cross-protocol scenarios and enhances system flexibility and versatility.

[0119] Example 1:

[0120] In order to verify the actual application effect of the present invention in the adaptation of medical Internet of Things protocols, the research team deployed this method in a medical Internet of Things data adaptation test scenario in a certain hospital. The test scenario covers the data interaction process of medical equipment in various departments of the hospital, and the data types include text electronic medical records, numerical vital signs monitoring data and medical imaging data. The types of equipment involve multiple multi-parameter monitors, medical imaging workstations, intelligent infusion pumps and information management servers, and the protocols involve HL7, DICOM, FHIR and self-developed Internet of Things protocols. The test time is from December 2024 to March 2025, spanning 3 months, covering multiple scenarios such as daily outpatient, inpatient and emergency departments in hospitals. During the test, the system needs to connect to more than 60 different models of medical equipment, covering a variety of brands and manufacturers, to fully verify the cross-protocol and cross-device adaptation capabilities.

[0121] In this test, the system first collects and structurally parses multimodal input data, extracting electronic medical records (EMRs) from the textual input data, real-time monitoring data from the numerical input data, and medical imaging files from the imaging modality. A multi-channel multimodal adaptation network receives the parsed multimodal input data and extracts multimodal features using convolutional neural networks. It then integrates the features within the modality through intra-modal weighted fusion and convolutional fusion. It then dynamically fuses the features from different modalities through inter-modal concatenation and multi-head attention weighted fusion, ultimately generating a unified fused feature representation. In the conditional adversarial adaptive algorithm, the conditional generator dynamically incorporates source and target protocol label information to generate conditional adversarial features. The conditional discriminator performs discrimination, and the gradient reversal module performs adversarial training using a dynamic gradient adaptation strategy to ensure consistency and stability in feature distribution across different protocols. The updated fused feature representation is input to the protocol adapter module, which maps and formats the protocol fields. The output is the target protocol data, which is then transmitted back to the target device or platform via the hospital's IoT gateway.

[0122] The data collected and sorted during the experiment are shown in the following table:

[0123] Table 1 Multimodal input data adaptation performance of this embodiment

[0124]

[0125] Table 2 Comparative experimental results of the present invention and the traditional method

[0126]

[0127]

[0128] As shown in the table above, during the entire testing period, the system collected and processed over 18,000 pieces of text data, 25,000 pieces of numerical monitoring data, and 8,000 images. The adaptation task involved the automatic conversion of different data protocols. The results demonstrate that the proposed method achieves a protocol adaptation success rate of 99.2%, maintaining high adaptation consistency even when data distributions vary significantly between protocols. Compared to traditional rule-based adaptation methods, the average conversion latency of adapted data was reduced from 210ms to 98ms, and the average packet loss rate dropped from 0.8% to 0.1%. The quality of feature reconstruction during the fusion phase of multimodal data was significantly improved. For textual data, the accuracy of electronic medical record field mapping reached 99.5%. For numerical data, the transmission integrity of vital sign parameters in the target protocol was improved to 99.8%. For image data, the structural similarity (SSIM) of medical image files after cross-protocol adaptation and restoration increased to 0.91, and the edge information integrity increased to 0.88, ensuring the diagnostic usability and compatibility of medical imaging data.

[0129] At the same time, the system's dynamic gradient reversal strategy demonstrates significant advantages. The discriminator's signal strength adaptive adjustment factor automatically adjusts to around 0.7 in the later stages of training, reducing the convergence time of the adaptation network from 350 rounds in traditional methods to 280 rounds, and achieving a domain alignment accuracy of 96.7%. In complex scenarios with diverse protocol combinations, the system can adaptively assign multimodal feature weights, enabling targeted data protocol adaptation and real-time transmission, demonstrating engineering feasibility.

[0130] In summary, the examples demonstrate the present invention's efficient adaptation capabilities for multi-protocol, multi-modal data in a medical IoT environment, as well as its robust adaptability to protocol distribution differences. The system not only excels in protocol adaptation success rate, data integrity, and real-time performance, but also demonstrates the feasibility and stability of engineering implementation in terms of feature expression and dynamic alignment. It is particularly suitable for deployment in edge computing devices, enabling low-latency, highly reliable medical IoT protocol adaptation and data interaction.

[0131] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A medical Internet of Things protocol adaptation method based on deep transfer learning, characterized in that: The steps include: S1. Collect source protocol data and target protocol data in the medical Internet of Things environment, use a protocol parser to perform structured parsing, and obtain multimodal input data; S2. Input the multimodal input data into the multi-channel multimodal adaptation network and extract the multimodal features through the multimodal feature encoder; S3. In the multi-channel multimodal adaptation network, intra-modal fusion and inter-modal fusion are performed on the multimodal features to generate a fused feature representation; S4. Input the fused feature representation into the conditional adversarial adaptive algorithm, generate adversarial samples based on the source protocol label and the target protocol label through the condition generator, receive the fused feature representation through the conditional discriminator and perform the protocol domain discrimination operation; S5. In the conditional adversarial adaptive algorithm, gradient reversal is used to perform protocol domain adversarial training. The network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive module are updated through the back-propagation algorithm, and the updated fusion feature representation is output. S6. Input the updated fusion feature representation into the protocol adapter module, perform protocol mapping and data format conversion according to the target protocol requirements, generate the target protocol data and transmit it to the medical Internet of Things gateway.

2. A medical Internet of Things protocol adaptation method based on deep transfer learning according to claim 1, characterized in that: The multimodal input data specifically includes text modality input data, numerical modality input data and image modality input data.

3. The method for adapting a medical Internet of Things protocol based on deep transfer learning according to claim 1, characterized in that: The multi-channel multimodal adaptation network specifically includes a multimodal feature encoder, an intra-modal fusion module, an inter-modal fusion module, a residual connection module, and a self-attention mechanism module: The multimodal feature encoder receives multimodal input data, uses a convolutional neural network structure to perform feature extraction, and outputs multimodal features, wherein the multimodal features include text modal features, numerical modal features, and image modal features; The intra-modality fusion module includes intra-modality fusion units corresponding to text modality, numerical modality and image modality. Each intra-modality fusion unit fuses features of the same modality through weighted fusion and outputs intra-modality fusion features. The inter-modal fusion module includes a splicing unit and a weighted fusion unit. The splicing unit is used to splice the intra-modal fusion features. The weighted fusion unit is used to perform weighted fusion on the spliced ​​features according to a preset weighting coefficient to generate an inter-modal fusion feature. The residual connection module includes a residual connection unit, which performs residual connection on the inter-modality fusion feature and the intra-modality fusion feature, and outputs a residual connection feature representation; The self-attention mechanism module includes a self-attention calculation unit, which receives the residual connection feature representation, uses the self-attention mechanism to dynamically weight the correlation between feature dimensions, and finally generates a fused feature representation.

4. The method for adapting a medical Internet of Things protocol based on deep transfer learning according to claim 1, characterized in that: The S3 specifically includes: S31. Input the multimodal features into the corresponding intramodal fusion unit in the intramodal fusion module. Each intramodal fusion unit fuses the features of the same modality through weighted fusion and outputs the intramodal fusion features: in, is the intra-modal fusion feature of modality m, m∈{text,num,img} is the modality type, that is, is the fusion feature within the text modality, is the fusion feature within the numerical mode, is the image modality fusion feature, is the i-th feature channel in mode m, is the weighted coefficient of the i-th feature channel in mode m, Conv1D(·) is a one-dimensional convolution operation, is the weighted coefficient of the jth convolution channel in mode m, C m is the total number of characteristic channels of mode m, K m is the total number of convolution channels of modality m, and LayerNorm(·) is the layer normalization operation; S32: Inputting the text modality intra-fusion features, the numerical modality intra-fusion features, and the image modality intra-fusion features into an inter-modality fusion module, wherein the inter-modality fusion module includes a splicing unit and a multi-head attention weighted fusion unit, and the splicing unit splices the intra-modality fusion features to generate modality splicing features; S33. In the multi-head attention weighted fusion unit, the modal splicing features are received, and inter-modal weighted fusion is performed through the multi-head attention mechanism to generate an inter-modal fusion feature representation: in, is the inter-modal fusion feature representation, Concat(·) is the multi-head attention concatenation operation, Softmax(·) is the normalized weight assignment, Q is the query matrix, which is generated by the modal concatenation feature through the query weight matrix, K is the key matrix, which is generated by the modal concatenation feature through the key weight matrix, V is the value matrix, which is generated by the modal concatenation feature through the value weight matrix, and W q 、W k 、W v is the learnable linear transformation weight matrix of the attention head, which is used to linearly transform the query matrix, key matrix and value matrix respectively. is the scaling factor; S34. Output the inter-modal fusion feature representation as the final fusion feature representation.

5. The method for adapting a medical Internet of Things protocol based on deep transfer learning according to claim 1, characterized in that: The conditional adversarial adaptive algorithm specifically includes a condition generator module, a condition feature generation module, a condition discriminator module, a gradient reversal module and a parameter update module: The condition generator module receives a source protocol label and a target protocol label, wherein the condition embedding unit generates embedding vectors for the source protocol label and the target protocol label respectively, and splices them into a conditional embedding vector; The conditional feature generation module receives the fused feature representation and the conditional embedding vector, performs conditional splicing and one-dimensional convolution operations, and generates conditional adversarial features; The conditional discriminator module includes a multi-layer convolution unit and a discrimination output unit, receives the conditional adversarial features, performs protocol domain discrimination, and generates a discrimination output; The gradient reversal module introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the strength of the discriminant signal, and propagates the gradient of the discriminant output in an adaptive direction: in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, D out is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round; The parameter updating module performs back propagation based on the dynamic inversion gradient, updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm, and outputs the updated fusion feature representation.

6. The method for medical Internet of Things protocol adaptation based on deep transfer learning according to claim 1, characterized in that: The S4 specifically includes: S41, the fusion feature representation is input into the conditional adversarial adaptive algorithm, the condition generator module receives the source protocol label and the target protocol label, and converts the source protocol label Label src and target protocol label tgt Input conditional embedding unit, the conditional embedding unit uses embedding function Embed(·) to transform vectors and generate conditional embedding vectors Embed(Label src ) and Embed(Label tgt ); S42, Embed(Label src ) and Embed(Label tgt ) Input the concatenation unit, concatenate according to the vector dimension, and generate a conditional embedding vector; S43, inputting the conditional embedding vector and the fused feature representation into a splicing unit, performing a splicing operation on the feature dimension, and generating a conditional feature splicing input; S44, inputting the conditional splicing input into a one-dimensional convolution unit in a conditional feature generation module, wherein the one-dimensional convolution unit performs a convolution operation and adds a bias term, and then performs nonlinear mapping through an activation function to generate a conditional adversarial feature; S45. Input the conditional adversarial feature into the conditional discriminator, which includes a multi-layer convolution unit and a discriminant output unit. The conditional discriminator first performs feature transformation on the conditional adversarial feature through the multi-layer convolution unit to generate a discriminant feature, inputs the discriminant feature into the discriminant output unit, performs a convolution operation and adds a bias term, performs a protocol domain discrimination operation through an activation function, and generates a discriminant output.

7. The method for adapting a medical Internet of Things protocol based on deep transfer learning according to claim 1, characterized in that: The S5 specifically includes: S51. In the conditional adversarial adaptive algorithm, the gradient reversal module receives the discriminant output and introduces a dynamic weight coefficient that is dynamically adaptive based on the training progress and the discriminant signal strength; S52. Calculate the gradient of the discriminant output with respect to the network parameters, apply the dynamic weight coefficient to the gradient, and perform adaptive direction propagation on the gradient of the discriminant output: in, is the dynamic reversal gradient, μ is the smoothing adjustment coefficient, is the training progress coefficient, t is the current training round, T is the maximum training round, α is the adaptive adjustment factor, which controls the influence of the discrimination signal strength, D out is the conditional discrimination output result, distinguishing the source protocol domain from the target protocol domain, θ is the learnable parameter in the multi-channel multimodal adaptation network and the conditional adversarial adaptive module, D avg is the average value of the discriminator output in the current round; S53. The parameter updating module updates the network parameters of the multi-channel multimodal adaptation network and the conditional adversarial adaptive algorithm based on the dynamic inversion gradient through the back propagation algorithm, and outputs the updated fusion feature representation.

8. The method for medical Internet of Things protocol adaptation based on deep transfer learning according to claim 1, characterized in that: The S6 specifically includes: S61. The updated fused feature representation is passed as input to the protocol mapping unit of the protocol adapter module. The protocol mapping unit performs a field mapping operation based on the protocol field definition and protocol standard of the target protocol tag, maps the updated fused feature representation to the feature space of the target protocol, and generates a target protocol field feature representation. S62: Input the target protocol field feature representation into the data format conversion unit of the protocol adapter module. The data format conversion unit performs field encapsulation, data rearrangement, and format conversion operations according to the data encapsulation standard of the target protocol to generate target protocol data that meets the requirements of the target protocol. S63. The target protocol data is used as a protocol adaptation result and transmitted through the medical Internet of Things gateway.